Karinoya Learning Room

Qualifications · Cloud / AI / Python Success Lab

AI in Society: Law and Ethics

Read the questions and explanations in English. The lectures (explanatory articles) are available in Japanese only.

View the Japanese version (with lectures) →

Q1 | CRISP-DM

In CRISP-DM, the standard approach to a data analysis project, what is the content of the stage placed first?

  1. Examine the quality and distribution of the data at hand and grasp its tendencies
  2. Handle missing values and create features, preparing the data for training
  3. Clarify the business problem to be solved and the criteria for what counts as success
  4. Incorporate the created model into a business system and operate it
AnswerC. Clarify the business problem to be solved and the criteria for what counts as success

CRISP-DM consists of six stages: business understanding, data understanding, data preparation, modeling, evaluation, and deployment. The first is business understanding, which decides the business problem to be solved and the criteria for what counts as success. Examining the quality and distribution of the data is data understanding, handling missing values and creating features is data preparation, and incorporating it into a business system is the deployment stage. It is also worth keeping in mind that the six stages are not a one-way flow but are depicted as a cycle that iterates, going back to earlier stages.

Q2 | PoC

What is the most appropriate positioning of a PoC in an AI project?

  1. The process, before committing to full-scale development, of trying it out on a small scale to confirm feasibility and effectiveness
  2. The process of assigning correct-answer labels to training data, according to a set standard
  3. The process, after a completed model is placed into production, of continuing to monitor and retrain it
  4. The process of redesigning the business workflow itself, going back to its purpose, to raise productivity
AnswerA. The process, before committing to full-scale development, of trying it out on a small scale to confirm feasibility and effectiveness

PoC stands for proof of concept, and it is the process of building something small and trying it out before investing in full-scale development, to confirm whether that data achieves the needed accuracy and, if it does, whether the business can actually run on it. Since the goal is not to build a finished product, unless the conditions for moving forward and for withdrawing are decided beforehand, you end up repeating trials that produce no results. Running monitoring and retraining after production deployment is MLOps, assigning correct-answer labels is annotation, and redesigning the business workflow itself is BPR — all of which refer to processes separate from a PoC.

Q3 | Annotation

Which of the following is an appropriate description of annotation, performed in data preparation for machine learning?

  1. Removing missing values and outliers from collected data, and adjusting its distribution
  2. Gathering and organizing a large amount of collected text into a set of data for language processing
  3. Splitting collected data into a training set and a validation set, putting it in a form where performance can be measured
  4. Manually attaching correct-answer labels or region information, as training targets, to collected data
AnswerD. Manually attaching correct-answer labels or region information, as training targets, to collected data

Annotation is the work of manually attaching correct-answer labels or region information to raw data, and it is indispensable for supervised learning. If the standard for attaching labels varies from worker to worker, it sets the ceiling on the model's performance right there, so designing work instructions and a quality check is key. Gathering and organizing a large amount of text is a corpus; splitting into training and validation sets, and handling missing values and outliers, are separate tasks within data preparation — neither is the act of attaching labels itself.

Q4 | Leakage

What is the most appropriate description of data leakage, arising from how training data is handled?

  1. The phenomenon in which the validation data is too small, so the measured accuracy itself becomes unstable
  2. The phenomenon in which information that would not actually be available at prediction time gets mixed into training, making performance look unfairly high
  3. The phenomenon in which performance on unseen data drops as a result of fitting too closely to the training data
  4. The phenomenon in which the amount of data needed balloons sharply as feature dimensions increase too much
AnswerB. The phenomenon in which information that would not actually be available at prediction time gets mixed into training, making performance look unfairly high

Data leakage is a phenomenon in which information that would not actually be obtainable at the time of prediction gets mixed into the training data, making the accuracy measured during validation come out far higher than the true capability. It happens in forms such as including the date a cancellation request was received when predicting churn, computing the mean and standard deviation used in preprocessing from the whole dataset including the validation data, or randomly splitting time-series data while ignoring time order. Fitting too closely to the training data is overfitting, and its cause lies not in "using information you should not have used" but in "fitting too closely." The amount of needed data ballooning as dimensions increase is the curse of dimensionality — both have a different cause from leakage.

Q5 | MLOps

What is the most appropriate idea behind MLOps in the operation of an AI system?

  1. Continuing to monitor the data and accuracy after deployment, and running a continuous cycle of retraining and redeployment
  2. Borrowing the computing resources used for training from the cloud and keeping costs limited to what is used
  3. Bundling the development environment into a container so the same environment can be reproduced in production too
  4. Confirming once, before production deployment, that the model's accuracy meets the requirements
AnswerA. Continuing to monitor the data and accuracy after deployment, and running a continuous cycle of retraining and redeployment

MLOps is the idea of not ending things once a model is deployed, but monitoring the distribution of input data, the tendency of predictions, and even business metrics, and running as a system the cycle of updating the training data and rebuilding the model when a gap appears, then deploying it again. As the world changes, the tendency of inputs changes too, and accuracy quietly declines, so a single check is not enough. Borrowing computing resources is the use of the cloud, and bundling the runtime environment for reproducibility is the role of Docker — both are tools that support MLOps, but neither is itself the definition of MLOps.

Q6 | BPR

What is the content of BPR, discussed alongside AI adoption?

  1. Gathering the people in charge of the business and hearing what they require from AI
  2. Collecting business data and putting it into a form AI can learn from
  3. Redesigning the business workflow itself, going back to its purpose
  4. Replacing part of the work with AI while keeping the existing business procedure
AnswerC. Redesigning the business workflow itself, going back to its purpose

BPR stands for business process reengineering, and it refers to the effort of redesigning the workflow itself, going back to the purpose of the business, rather than partially improving the existing procedure. If you place AI on top while keeping the procedure as is, it can end up just adding a step where a person transcribes the AI's output, producing no effect; the success or failure of AI adoption is decided here more often than by accuracy alone. Hearing requirements is grasping stakeholders' needs, and shaping data into a learnable form is data preparation — neither is the redesign of the business process itself.

Q7 | Development approach

What is the most appropriate reason the waterfall model is considered hard to apply as is to AI development?

  1. Because there are many choices of language and library used in development, and standardization has not progressed
  2. Because the roles involved in development span many different job types, making it hard to decide how to divide the work
  3. Because you do not know what accuracy you can reach until you actually train the model, so the requirements cannot be fixed in advance
  4. Because procuring the computing resources needed for training takes time, and the schedule does not proceed as planned
AnswerC. Because you do not know what accuracy you can reach until you actually train the model, so the requirements cannot be fixed in advance

Waterfall is an approach that fixes each stage in order starting from requirements definition, on the premise of not going back. However, AI development has the nature that you do not know how far the accuracy will reach until you actually train the model, so the deliverable and the requirements cannot be fixed in advance. For that reason, the exploration and PoC stages take an agile form that proceeds through short iterations, and once the achievable accuracy comes into view, a combination is used where the main development stage is then fixed. This property also directly affects how the development contract is designed, that is, whether to use a contract for work (ukeoi) or a quasi-mandate contract (jun-inin).

Q8 | Docker

What is the main purpose of using Docker in AI development?

  1. Bundling the runtime environment, including libraries, so it runs in the same state in a different environment
  2. Publishing a trained model in a form that other systems can call over HTTP
  3. Keeping the progress and results of training, together with an explanation, in a single document
  4. Borrowing a large amount of computing resources for only as long as needed, and returning it when done
AnswerA. Bundling the runtime environment, including libraries, so it runs in the same state in a different environment

Docker is a technology that bundles the runtime environment, including the application and its libraries, as a container. Because the same container can be run on both the development machine and in production, mismatches such as "it worked on my machine" due to version differences are reduced. Keeping progress and results together with an explanation in a single document is Jupyter Notebook, making something callable over HTTP is a web API, and borrowing computing resources only for as long as needed is the use of the cloud — all of which refer to other tools used in AI development.

Q9 | Corpus

Which of the following correctly describes what a corpus means in the field of natural language processing?

  1. A collection of image or audio data widely published for research and development
  2. A collection of morphological analysis processes that split text into words and determine parts of speech
  3. A collection of a large amount of text data gathered and organized for language processing
  4. A collection of weight parameters trained to represent words as vectors
AnswerC. A collection of a large amount of text data gathered and organized for language processing

A corpus refers to a large amount of text data gathered and organized for natural language processing. Depending on its use, it may also have part-of-speech or syntactic information attached. A general collection of data published for research and development is an open dataset, which also includes images and audio and is not limited to text. A vector representation of words is a product of training, not a collection of data, and morphological analysis is a method for processing text — neither matches the definition of a corpus. Note that a corpus also has a license and terms of use, so whether commercial use or redistribution is allowed needs to be checked individually.

Q10 | Working with the outside

What is the most appropriate description of open innovation, as discussed in the context of AI development?

  1. Making use of published datasets to reduce the effort of collection
  2. Releasing a developed model and its source code for free, in a form anyone can use
  3. Gathering people from each in-house department to build a cross-department development structure
  4. Actively bringing in outside technology, knowledge, and people to create new value
AnswerD. Actively bringing in outside technology, knowledge, and people to create new value

Open innovation is the idea of not staying closed within one's own in-house research and development, but actively bringing in outside technology, knowledge, and people to create value. The syllabus lists industry-academia collaboration and collaboration with other companies or other industries under the same mid-level topic, and its means include joint research with universities and other research institutions, and teaming up with a partner that has data or a sales channel one does not have in-house. Releasing something for free is the idea of open source, which points in the opposite direction from bringing in outside resources. An in-house cross-department structure and the use of published datasets do not, either of them, refer to collaboration with an outside party.

Q11 | Types of processed information

Which is correct regarding the difference between anonymously processed information and pseudonymously processed information under the Act on the Protection of Personal Information?

  1. Both can be provided to a third party if the individual consents, and neither can be provided without consent
  2. Anonymously processed information cannot be provided to a third party without the individual's consent, but pseudonymously processed information can
  3. Both are uniformly prohibited from being provided to a third party, regardless of the degree of processing
  4. Anonymously processed information can be provided to a third party without the individual's consent, but pseudonymously processed information generally cannot
AnswerD. Anonymously processed information can be provided to a third party without the individual's consent, but pseudonymously processed information generally cannot

Anonymously processed information is processed so that a specific individual cannot be identified and so that the original personal information cannot be restored. Because it has been processed this far, it can, following the prescribed procedure, be provided to a third party without the individual's consent. Pseudonymously processed information, on the other hand, is processed to the degree that a specific individual cannot be identified unless cross-referenced with other information; the effort of processing is lighter, such as just removing the name, and it retains analytical value, but in exchange it generally cannot be provided to a third party, and it is intended for internal analytical use by the business. Remember it as: the lighter the processing, the less it can be taken outside — that helps avoid mixing the two up.

Q12 | Personal identification codes

Which is a correct description regarding personal identification codes under the Act on the Protection of Personal Information?

  1. A passport number, facial recognition data, and the like are treated as personal information on their own
  2. They are treated as identifying a specific individual only when combined with a name
  3. They are treated as personal information only when acquired with the individual's consent
  4. They are outside the scope at the time of acquisition, and become personal information once put into a database
AnswerA. A passport number, facial recognition data, and the like are treated as personal information on their own

Personal identification codes include physical-characteristic data such as fingerprint data, facial recognition data, and codes converted from DNA base sequences, as well as numbers assigned to an individual such as My Number, passport numbers, and driver's license numbers. These count as personal information on their own, so it is incorrect to think that without a name they are not personal information. Whether something counts as personal information is also not decided by whether consent was obtained at acquisition, and it does not become personal information only once put into a database. Note that information that has been systematically organized into a searchable state is called personal data, and additional obligations attach to it.

Q13 | GDPR

Which is correct regarding the positioning of GDPR?

  1. Guidance issued by an international organization, carrying no legal binding force on any country
  2. A regulation established to internationally extend Japan's Act on the Protection of Personal Information
  3. An EU regulation, applying only to businesses with a base within the EU
  4. An EU regulation that can also apply to businesses outside the EU in some cases
AnswerD. An EU regulation that can also apply to businesses outside the EU in some cases

GDPR is the EU's General Data Protection Regulation, a piece of EU legislation separate from Japan's Act on the Protection of Personal Information. Its distinguishing feature is extraterritorial application: if there are circumstances such as offering goods or services to people located within the EU, or monitoring the behavior of people within the EU, it can apply even to a business developing within Japan. So whether a business is subject to it cannot be judged solely by whether its base is in the EU. It also should be kept in mind that it is not an extension of Japanese law, and it is not non-binding guidance either.

Q14 | Article 30-4

In which case does Article 30-4 of the Copyright Act allow a copyrighted work to be used without the copyright holder's permission?

  1. When the purpose is not to enjoy the thoughts or feelings expressed in the work, and not to have another person enjoy them, either
  2. When the amount of the copyrighted work being used is only a small part of the whole
  3. When the copyrighted work being used is published for free on the internet
  4. When the purpose of use is not commercial and is limited to academic research
AnswerA. When the purpose is not to enjoy the thoughts or feelings expressed in the work, and not to have another person enjoy them, either

Article 30-4 of the Copyright Act provides that, when the purpose is not to enjoy the thoughts or feelings expressed in a copyrighted work oneself, nor to have another person enjoy them, the work may be used to the extent found necessary. "Enjoy" refers to appreciating the thoughts or feelings expressed in a work to obtain intellectual or emotional satisfaction, and the act of reading it in as data for machine learning is understood not to fall under this. Item 2 of the same article explicitly names information analysis, and the creation and use of training data falls under this. Being published for free, not being for a commercial purpose, and the amount used being small are, none of them, requirements set by this provision.

Q15 | The proviso

Which is a correct understanding of Article 30-4 of the Copyright Act, regarding the use of copyrighted works for AI training?

  1. Use for training is permitted only when the copyrighted work was purchased for a fee
  2. If the use is for training, the copyright holder's interests are never a concern, in any case
  3. This provision does not apply in cases that would unreasonably harm the copyright holder's interests
  4. If the training is lawful, then even if the generated output resembles an existing copyrighted work, it does not constitute infringement
AnswerC. This provision does not apply in cases that would unreasonably harm the copyright holder's interests

Article 30-4 carries a proviso stating that this does not apply in cases that, in light of the type and purpose of the copyrighted work and the manner of the use, would unreasonably harm the interests of the copyright holder. So it cannot be said that "anything goes for AI training." For example, if you copy a database sold for information-analysis purposes without purchasing it and use it for training, that is considered to fall under the proviso because it harms the market for that database. Also, even if the input stage of training is lawful, if the generated output resembles an existing copyrighted work, copyright infringement can separately become an issue at the output stage of generation or use. Input and output are evaluated separately. Whether the work was purchased for a fee is not the standard that decides whether this provision applies.

Q16 | Copyright in generated output

Which is closest to the currently established view on the copyrightability of AI-generated output?

  1. Copyright naturally belongs to the provider of the service that was used
  2. Because it is output by AI, it can never be recognized as a copyrighted work
  3. It is judged case by case, depending on whether a creative contribution by a human is recognized
  4. The person who gave the instruction for generation is always recognized as having copyright
AnswerC. It is judged case by case, depending on whether a creative contribution by a human is recognized

Because copyright is a system that protects a human's creative expression, something an AI merely output automatically does not become a copyrighted work. On the other hand, if a human has been involved to the degree that they can be said to have given it its essential creative characteristics in expression, there is room for copyrightability to be recognized. The key to the judgment is creative contribution, and things such as the specificity of the instructions and trial and error, and the degree of selection and revision of the generated output, are examined case by case. Therefore both the view that it never becomes a copyrighted work and the view that the person who gave the instruction always gains rights are incorrect. Note also that, separately from copyright law, the terms of use of the service used may set out how the generated output is handled, so both need to be checked.

Q17 | Employee inventions

Which is correct regarding the treatment of an employee invention under the Patent Act?

  1. It can be made to belong to the employer if provided for in work rules or similar, and the employee receives a reasonable benefit
  2. Even if it is an invention made as part of the employee's duties, there is no room to have the right belong to the employer
  3. The right always belongs to the employer, and a work rule cannot set a different arrangement
  4. Even when the right belongs to the employer, no payment of a benefit to the employee is required
AnswerA. It can be made to belong to the employer if provided for in work rules or similar, and the employee receives a reasonable benefit

An employee invention is an invention made by an employee as part of their duties. The right to obtain a patent arises in principle in the employee, the inventor, but if provided for in advance in work rules or the like, it can instead be made to belong to the employer from the moment it arises. In that case, the employee is recognized as having the right to receive a reasonable monetary or other economic benefit, so it is not a system where the right is taken without any compensation. The key point is also that it does not always belong to the employer, and there is room for it to belong to the employer, too — the treatment changes depending on whether such a provision exists.

Q18 | Trade secrets

Which combination represents the requirements for something to be protected as a trade secret under the Unfair Competition Prevention Act?

  1. Secret management, usefulness, inventive step
  2. Secret management, usefulness, non-public knowledge
  3. Usefulness, novelty, inventive step
  4. Secret management, novelty, non-public knowledge
AnswerB. Secret management, usefulness, non-public knowledge

The three requirements for a trade secret are that it is managed as a secret (secret management), that it is technical or business information useful for business activities (usefulness), and that it is not publicly known (non-public knowledge); protection is granted only once all three are satisfied. Novelty and inventive step are requirements for obtaining a patent, not requirements for a trade secret. Whereas a patent is a system that grants a monopoly in exchange for disclosing the content, a trade secret is a system that protects by not disclosing, so information one does not want to disclose, such as the weights of a trained model or ingenuity in preprocessing, is protected this way instead.

Q19 | Limited provision data

Which is correct regarding limited provision data under the Unfair Competition Prevention Act?

  1. It is a category covering technical information for which a patent was applied for but not granted
  2. It covers information that satisfies all three requirements of a trade secret and is additionally registered
  3. It is a category covering information that is not provided to anyone and is kept only within one's own company
  4. It covers electromagnetic records that are provided as a business to specific persons and that are accumulated in a substantial amount and managed
AnswerD. It covers electromagnetic records that are provided as a business to specific persons and that are accumulated in a substantial amount and managed

Limited provision data refers to information that is provided as a business to specific persons, accumulated in a substantial amount by electromagnetic means, and managed. This category was set up to protect transactions such as selling or sharing data, and its significance is that it can be protected without satisfying all three requirements of a trade secret. This is because, once it is provided to multiple parties, it does not necessarily strictly satisfy non-public knowledge. Information kept as a secret only within one's own company is protected on the trade secret side instead, and this category has no relation to whether a patent was applied for or granted.

Q20 | Development contracts

Which is correct regarding the difference between a contract for work (ukeoi) and a quasi-mandate contract (jun-inin) in the outsourcing of AI development?

  1. A contract for work promises the completion of the job, and a quasi-mandate contract carries out the work with the duty of care of a good manager
  2. A contract for work is a contract bearing a duty of due care, and a quasi-mandate contract is a contract promising completion
  3. Both are contracts promising completion of a deliverable, differing only in the level of accuracy guaranteed
  4. Neither promises completion, and they differ only in the timing of payment
AnswerA. A contract for work promises the completion of the job, and a quasi-mandate contract carries out the work with the duty of care of a good manager

A contract for work (ukeoi) is a contract that promises the completion of a job; if it is not completed, payment cannot be claimed, and liability is borne for a deliverable that does not conform to the contract. A quasi-mandate contract (jun-inin) is a contract for entrusting the handling of affairs, and what it promises is not completion but carrying out the work with the duty of care of a good manager. Because in AI development you do not know what accuracy you can reach until you actually train the model, a common split is to use a quasi-mandate contract for the exploration and PoC stages and a contract for work once the specification is fixed for the implementation stage. The Ministry of Economy, Trade and Industry's Contract Guidelines on the Utilization of AI and Data also present this pattern of splitting the contract by stage.

Q21 | Hard law

Which is the correct distinction between hard law and soft law?

  1. Hard law refers to legislation with binding force, and soft law refers to rules with no binding force
  2. Hard law refers to legislation with penalties, and soft law refers to legislation without penalties
  3. Hard law refers to international agreements, and soft law refers to domestic legislation
  4. Hard law refers to newly created legislation, and soft law refers to old customary practice
AnswerA. Hard law refers to legislation with binding force, and soft law refers to rules with no binding force

Hard law is legislation established by the state, whose violation carries binding force, and the Act on the Protection of Personal Information and the Copyright Act fall under this. Soft law is a set of rules with no legal binding force, such as guidelines, industry self-regulation, and corporate codes of conduct. Since soft law is not legislation, the distinction between legislation with and without penalties does not apply to it. It is also not a distinction between domestic and international, nor between old and new. In fields where technology changes quickly, an approach is taken of first indicating a direction with soft law and turning only the necessary parts into legislation, and Japan's AI policy is, in the main, centered on soft law.

Q22 | Risk response

What is the idea behind the risk-based approach, as discussed for AI regulation and in-house responses?

  1. Stop providing a service through a given use case entirely, as soon as a risk is confirmed
  2. Enumerate the anticipated risks and then impose the same level of measures on every use case
  3. Leave the risk assessment to the user, with the business only providing information
  4. Vary the strength of the response or regulation according to the magnitude of the risk posed by the use case
AnswerD. Vary the strength of the response or regulation according to the magnitude of the risk posed by the use case

A risk-based approach is the idea of varying the strength of regulation or response according to the magnitude of risk posed by an AI's use case. Even the same facial recognition technology has a completely different impact on people depending on whether it is used to unlock a phone screen versus for constant surveillance in a public space. Imposing the same strong regulation regardless of use case would stop even low-harm uses, while uniformly loosening it would leave high-harm uses unchecked. So strong obligations are assigned to high-impact uses and light treatment to low-impact ones. Businesses also use the same idea internally, sorting in-house AI projects by risk level and routing only the high-risk ones to strict review.

Q23 | Consideration at the design stage

What is the most appropriate description of the idea behind privacy by design?

  1. Limit the staff who handle personal information so that only authorized people can view it
  2. Build privacy protection mechanisms in from the planning and design stage
  3. Disclose the purposes for which collected data is used, upon request from users
  4. Investigate the cause once a problem occurs and add the necessary measures afterward
AnswerB. Build privacy protection mechanisms in from the planning and design stage

Privacy by design is the idea of building in privacy protection from the planning and design stage, rather than adding measures after a problem occurs. Decisions such as designing not to collect data in the first place, deleting it after the necessary period, or processing it in a form that cannot be identified, become more costly the later they are made. Disclosing the purpose of use and limiting viewing privileges are, in themselves, correct efforts, but they lack the viewpoint of when they are built in, so they do not describe this idea. It is a counterpart concept to security by design, which builds in security measures from the design stage.

Q24 | Proxy variables

Even after removing sensitive attributes from training data, unfair output can remain. What is the main reason for this?

  1. Because other features, such as postal codes, end up acting as substitutes for the removed attribute
  2. Because it gets retrained on other data that includes that attribute even after removal
  3. Because removing an attribute reduces the amount of data, lowering the model's accuracy itself
  4. Because the process of removing an attribute itself distorts the distribution of the data
AnswerA. Because other features, such as postal codes, end up acting as substitutes for the removed attribute

A feature that ends up acting as a substitute for a sensitive attribute is called a proxy variable. A postal code correlates strongly with race and income through the area of residence, and the history of one's school or club activities can correlate with gender. As long as such features remain, the model indirectly learns the same discrimination even if the gender or race column is removed from the input. Therefore, checking fairness is done not by removing input fields but by tabulating outputs by attribute and measuring the difference. Furthermore, there are multiple definitions of fairness — such as equalizing the acceptance rate across attributes, or equalizing the error rate — and in general it is not possible to satisfy all of them at the same time, so which definition to adopt needs to be chosen and explained according to the use case.

Q25 | Adversarial input

Which of the following correctly describes an Adversarial Attack (Adversarial Examples)?

  1. Adding a change to the input that is imperceptible to humans, causing an incorrect judgment
  2. Making a large number of queries and collecting the outputs, to build an equivalent model on hand
  3. Tampering with the parameters of a distributed model, embedding a specific behavior
  4. Mixing tampered data into the training data, distorting the training result
AnswerA. Adding a change to the input that is imperceptible to humans, causing an incorrect judgment

An Adversarial Attack is an attack that adds a tiny change to the input, one that is almost imperceptible to the human eye, to make a model in operation produce an incorrect judgment. It targets the inference stage; the model itself is not rewritten. Mixing tampering into the training data to distort the training result is data poisoning, which instead targets the training stage. Making a large number of queries to reproduce an equivalent model is model extraction, and tampering with a model's parameters or its distributed artifacts is model poisoning — each targets a different point in time and a different target. Defenses against any of these proceed based on the idea of security by design, which builds in security from the design stage.

Q26 | Distinguishing extraction

Among attacks on AI, which one is model extraction?

  1. Mixing tampered data into the training data, distorting the training result
  2. Using the outputs as a clue to infer and extract information that was in the training data
  3. Tampering with the parameters of a distributed model to change its behavior for a specific input
  4. Making a large number of queries and collecting the outputs, to build an equivalent model on hand
AnswerD. Making a large number of queries and collecting the outputs, to build an equivalent model on hand

Model extraction is an attack that feeds a large number of inputs into a target model, collects the outputs, and reproduces an equivalent model on hand. What is stolen is the model itself, and this can be pulled off even if the model is only offered as a public API. Data extraction, by contrast, is an attack that pulls out information that was in the training data by observing the outputs; what is stolen there is the data. Mixing tampering into the training data is data poisoning, and tampering with a model's parameters or its distributed artifacts is model poisoning — the former targets the training stage, and the latter the training or distribution stage. Organizing by what is stolen and at what point in time helps avoid mixing these up.

Q27 | An information bubble

Which phenomenon is said to arise from optimization by recommendation algorithms?

  1. The basis for a judgment becomes hidden internally, so people can no longer trace the reason for it
  2. Only information that suits one's preferences gets selected and shown, so other information stops reaching one
  3. People with the same opinion repeat broadcasting and agreeing with each other, reinforcing that opinion
  4. Generated fake video or audio comes to be widely received as genuine
AnswerB. Only information that suits one's preferences gets selected and shown, so other information stops reaching one

A filter bubble is a phenomenon in which a recommendation algorithm selects information to match a user's preferences, and as a result that person ends up wrapped in a bubble, seeing only information suited to their own tastes; the cause lies in optimization by the algorithm. Users are not even aware that this selection is happening. An echo chamber is when people with the same opinion repeat broadcasting and agreeing with each other in a closed space, coming to believe that opinion is the only correct one, and its cause lies in the human tendency to gather with similar people. Fake video or audio being received as genuine is deepfakes and fake news, and the basis for a judgment being untraceable is the black box problem — both belong to separate mid-level topics.

Q28 | Fake video

Which is the most appropriate description of a deepfake?

  1. Using deep learning to create a short piece of text that summarizes the key points of a large amount of text
  2. Using deep learning to visualize the region of an image that formed the basis for a judgment
  3. Using deep learning to create video or audio that makes it look as if a real person is speaking
  4. Using deep learning to newly generate and display the face photos of people who do not exist
AnswerC. Using deep learning to create video or audio that makes it look as if a real person is speaking

A deepfake refers to a technology, and its resulting product, that uses deep learning to create video or audio that makes it look as if a real person said things they never said and did things they never did. Impersonating a real person is at its core, so its aim differs from generating the face of a nonexistent person. It is used for defamation, fraud, and interference with elections, and it greatly increases the persuasive power of fake news. As a countermeasure, in addition to research into detection technology, efforts are being made to embed provenance information at the time of filming or generation, so that later alteration can be verified. Summarizing text and visualizing the basis for a judgment are, neither of them, about misuse.

Q29 | Recording provenance

Regarding ensuring transparency in AI, what is the most appropriate purpose of keeping a record of the provenance of data?

  1. To hold down the amount of computing resources consumed for training and reduce the resulting cost
  2. To be able to trace the source and the process by which it was processed, and reach the cause when a problem occurs
  3. To speed up the model's inference and shorten the time until a response is returned
  4. To increase the amount of data available for training and raise the model's accuracy
AnswerB. To be able to trace the source and the process by which it was processed, and reach the cause when a problem occurs

The provenance of data is a record of what data was obtained from where, and how it was processed before being used for training. Without this record, when a problem arises in the output, one cannot trace back to the data that caused it, making it impossible to fix or explain. Transparency does not mean only disclosing the model's internals; it also includes showing, according to the audience, what it is being used for, what data it was built with, and where its limitations lie. Reducing computing resources, speeding up inference, and increasing the amount of data are, none of them, the purpose of recording provenance. Along with this, it is also worth keeping track of traceability, which makes it possible to trace which data, through which model, produced which decision.

Q30 | Ethics review

Which of the following correctly describes an ethics assessment in AI governance?

  1. A document that puts one's own company's policy on how to handle AI into writing, and presents it externally
  2. A procedure of continuing to watch trends in output and complaints after operation has begun
  3. A procedure, at the planning stage, of identifying and evaluating anticipated impacts and risks
  4. A mechanism that, from a position independent of the department in charge, confirms whether operation follows policy
AnswerC. A procedure, at the planning stage, of identifying and evaluating anticipated impacts and risks

An ethics assessment is a procedure that, at the planning stage of an AI project, identifies and evaluates the anticipated impacts and risks, and the strictness of the subsequent review is adjusted according to the result. Continuing to watch trends in output and complaints after operation has begun is monitoring, confirming from an independent position whether operation follows policy is an AI audit, and putting one's own company's policy into writing and presenting it externally is an AI policy — each is a different component that makes up AI governance. Alongside these are keeping human involvement at key decision points, reproducibility, which makes it possible to rebuild a model later, and traceability, which makes it possible to trace the path of a decision.

Practice: answer the questions on this page

This practice tool asks questions in random order (it works when JavaScript is enabled). You can still read all the questions and explanations above without it.

* The explanations are information for study purposes. Exam scope and systems change from year to year, so always check the official announcements of the organization that administers the exam.

This page is a translation of the Japanese original. If the translation and the original differ, the Japanese version takes precedence. View the Japanese original