Karinoya Learning Room

Qualifications · Cloud / AI / Python Success Lab

Deep Learning Methods and Applications

Read the questions and explanations in English. The lectures (explanatory articles) are available in Japanese only.

View the Japanese version (with lectures) →

Q1 | Simple perceptron

Which of the following correctly describes the simple perceptron?

  1. It consists only of an input layer and an output layer, and can only solve linearly separable problems
  2. It excels at sliding a filter over an image to extract local features
  3. It has wiring that feeds its state back in the time direction, so it can handle sequential data in order
  4. It has multiple hidden layers and can perform nonlinear separation, so it can also represent exclusive OR
AnswerA. It consists only of an input layer and an output layer, and can only solve linearly separable problems

The simple perceptron is the simplest neural network, consisting only of an input layer and an output layer, and it can only solve linearly separable problems that can be divided by a single straight line. For exclusive OR (XOR), the points where the output is 1 and the points where it is 0 alternate diagonally, so no straight line can separate them, and the simple perceptron cannot represent it. A multilayer perceptron, which adds a hidden layer so the boundary can be bent, can represent XOR, and the first choice describes that multilayer perceptron. Feeding back the state from the previous time step to handle sequences is the description of a recurrent neural network, and sliding a filter to extract local features is the description of a convolutional layer.

Q2 | GPU suitability

What is the most appropriate reason GPUs are considered suitable for deep learning training?

  1. Because their storage capacity for holding training data is far larger than that of a CPU
  2. Because they have a dedicated circuit built in from the start for the chain rule used in backpropagation
  3. Because they can quickly perform complex processing with many branches in sequence using a small number of powerful cores
  4. Because they have a large number of small computing cores and can process the same calculation in parallel all at once
AnswerD. Because they have a large number of small computing cores and can process the same calculation in parallel all at once

Training in deep learning is a computation that repeats matrix and vector multiplication and addition an enormous number of times. A GPU has a large number of small cores that perform simple operations and is suited to running the same calculation in parallel all at once, so it is a good match for this form of computation. Quickly performing branch-heavy processing in sequence with a small number of powerful cores is the strength of a CPU, which does not have many cores. The advantage of a GPU is its degree of parallelism, not its storage capacity. There is no fact that it has a dedicated circuit built in for the chain rule, and it is the TPU that is designed specifically for the tensor operations of deep learning.

Q3 | Cause of vanishing gradients

Why does using the sigmoid function in a hidden layer tend to cause the vanishing gradient problem?

  1. Because the output stays within the range 0 to 1, so stacking layers makes the output itself approach 0
  2. Because the output is normalized so that its sum becomes 1, so the value gets smaller with each layer
  3. Because the maximum value of its derivative is 0.25, and this is multiplied together each time you go back a layer
  4. Because the output is always 0 for negative inputs, so that unit stops learning
AnswerC. Because the maximum value of its derivative is 0.25, and this is multiplied together each time you go back a layer

The vanishing gradient problem is a phenomenon in which the gradient approaches 0 during backpropagation, so the layers on the input side stop learning. Since the maximum value of the sigmoid function's derivative is only 0.25, a number of 0.25 or less is multiplied in each time you go back a layer, and in a deep network almost none of the gradient remains. What matters is not the output value itself but the magnitude of its derivative. The gradient becoming 0 for negative inputs, stopping learning, is a weakness of the ReLU function, and the Leaky ReLU function addresses this. Normalizing so the sum of the outputs becomes 1 is what the softmax function does.

Q4 | The ReLU function

Which of the following correctly describes the ReLU function?

  1. A function with an S-shaped curve that squeezes the input into the range 0 to 1
  2. A function that normalizes all outputs together so that their sum becomes 1
  3. A function that keeps the input within the range -1 to 1, taking a shape symmetric about the origin
  4. A function that returns the input value as is if it is positive, and returns 0 if it is negative
AnswerD. A function that returns the input value as is if it is positive, and returns 0 if it is negative

The ReLU function is a simple function that returns the input value as is if it is positive and returns 0 if it is negative. In the positive region the derivative is always 1, so the gradient does not shrink even when going back through layers, greatly easing the vanishing gradient problem. However, in the negative region the derivative is 0, so a unit that once falls into the negative side can stop learning, and the Leaky ReLU function, which gives a slight slope to the negative side, was created for this. The S-shaped curve that squeezes into 0 to 1 describes the sigmoid function, the shape within -1 to 1 symmetric about the origin describes the tanh function, and normalizing so the sum of the outputs becomes 1 describes the softmax function.

Q5 | Leaky ReLU

What point of the ReLU function is the Leaky ReLU function meant to address?

  1. The sum of the outputs does not become 1, so it cannot be interpreted as a probability
  2. The gradient becomes 0 for negative inputs, so that unit stops learning
  3. The computation includes an exponential function, making each pass heavier
  4. The gradient stays at 1 for positive inputs, so the update amount becomes too large
AnswerB. The gradient becomes 0 for negative inputs, so that unit stops learning

Because the ReLU function has both output and derivative equal to 0 in the negative region, if a certain unit's input stays negative the whole time, no gradient flows through it and that unit never learns again. The Leaky ReLU function gives a slight slope to the negative side too, avoiding this dead end. Having a gradient of 1 in the positive region is an advantage of the ReLU function, not a drawback, and it is precisely this property that prevents vanishing gradients. It is only natural for an activation function that the sum of the outputs does not become 1, and the softmax function is used for an output layer meant to be read as probabilities. It is the sigmoid function and the tanh function whose computation includes an exponential function; the ReLU function is, if anything, lightweight to compute.

Q6 | Softmax

Which of the following is a typical use case for the softmax function?

  1. Place it in the output layer of a regression problem and output continuous values as is
  2. Use it as the activation function of a hidden layer to give the layer nonlinearity
  3. Place it in the output layer of binary classification and output the probability of one class
  4. Place it in the output layer of multi-class classification and output the probability of each class
AnswerD. Place it in the output layer of multi-class classification and output the probability of each class

The softmax function normalizes the whole set so that the sum of the outputs is exactly 1, so it is placed in the output layer of multi-class classification, where there are three or more classes, and its outputs are read as the probability of belonging to each class. Combining it with cross-entropy as the loss function is the basic pattern. It is the sigmoid function that is placed in the output layer of binary classification to output a probability, and it is the ReLU function and the tanh function that are widely used as the activation function of hidden layers. In regression, the output layer outputs the value as is, and the mean squared error function is used as the loss function.

Q7 | Loss function for regression

For a regression problem that predicts a continuous value such as a house price, which is generally chosen as the loss function?

  1. Cross-entropy, which measures the gap between the predicted probability distribution and the true distribution
  2. Triplet Loss, which learns the distance relationship among a triplet
  3. Contrastive Loss, which learns whether two things are the same using distance
  4. The mean squared error function, which squares and averages the difference between the prediction and the true value
AnswerD. The mean squared error function, which squares and averages the difference between the prediction and the true value

Since regression is a problem of predicting a continuous value, the mean squared error function, which directly squares and averages the difference between the prediction and the true value, is the natural choice. Cross-entropy is a function that measures the gap between probability distributions and is used for classification problems; for multi-class classification, combining it with the softmax function in the output layer is the basic form. Triplet Loss learns the distance relationship among a triplet of an anchor, a positive example, and a negative example, and Contrastive Loss learns whether two pieces of data are the same using distance; both are loss functions for obtaining a representation that measures similarity. Choosing the loss function according to the task is the goal of this mid-level topic.

Q8 | Loss over a triplet

Which of the following correctly describes the loss function Triplet Loss?

  1. It squares and averages the difference between the prediction and the true value, directly measuring the gap in magnitude
  2. It measures the gap between the predicted probability distribution and the true distribution and is widely used for training classification
  3. It uses a triplet of an anchor, a positive example, and a negative example so that the positive example ends up closer than the negative example
  4. It measures the gap between two probability distributions and works to bring the distributions closer together
AnswerC. It uses a triplet of an anchor, a positive example, and a negative example so that the positive example ends up closer than the negative example

Triplet Loss is a loss function that uses a triplet of an anchor data point, data of the same kind as it (a positive example), and data of a different kind (a negative example), and trains the model so that the distance between the anchor and the positive example becomes smaller than the distance between the anchor and the negative example. It is used to build a space that represents similarity through distance. Measuring the gap between two probability distributions is Kullback-Leibler divergence (KL divergence); squaring and averaging the difference between the prediction and the true value is the mean squared error function; and measuring the gap between probability distributions and being widely used for classification is cross-entropy. Note that Contrastive Loss, which trains on a pair of two data points for whether they are the same, is included under the same mid-level topic.

Q9 | Backpropagation

Which of the following is the most appropriate description of backpropagation?

  1. The procedure of directly rewriting the weights in order from the output layer — the update itself
  2. The procedure of moving the weights little by little with random numbers and adopting the direction in which the error decreased
  3. The procedure of proceeding the computation in order from the input side to the output side to obtain the output
  4. The procedure of using the chain rule to work back from the output side to the input side to find the gradient
AnswerD. The procedure of using the chain rule to work back from the output side to the input side to find the gradient

Backpropagation is a method that uses the chain rule, which decomposes the derivative of a composite function into a product of the derivatives of its parts, to efficiently find the gradient of each weight while working back from the output side to the input side. What is found here is strictly the gradient; it is the job of gradient descent to actually move the weights using that gradient. It would be inaccurate to state flatly that backpropagation updates the weights, so care is needed here. Moving weights with random numbers and adopting a good direction is the idea behind search methods that do not use gradients, and proceeding the computation from the input side to the output side is the procedure of forward propagation, that is, inference.

Q10 | The credit assignment problem

What is it called when, for the final error, you assign how much each unit of each layer is responsible?

  1. The credit assignment problem
  2. The exploding gradient problem
  3. The vanishing gradient problem
  4. The curse of dimensionality
AnswerA. The credit assignment problem

The credit assignment problem is the question of how much of the error appearing at the output should be attributed to which unit of which layer along the way. Backpropagation is positioned as having made this assignment computable through the chain rule. The vanishing gradient problem is a phenomenon in which the gradient approaches 0 while working backward and the layers on the input side stop learning; the exploding gradient problem is the opposite phenomenon in which the gradient becomes too large and diverges — both are problems that arise when applying backpropagation. The curse of dimensionality is the phenomenon in which the amount of data needed increases sharply as the number of feature dimensions grows, and in the syllabus it is placed under the mid-level topic of machine learning.

Q11 | L1 regularization

Which of the following is a characteristic of L1 regularization?

  1. It uses the sum of squares of the weights as a penalty, keeping the values small and uniform but rarely driving them to exactly 0
  2. It uses the sum of the absolute values of the weights as a penalty, and the values tend to become exactly 0
  3. It randomly disables units at each training step to prevent bias toward a particular dependency
  4. It stops training at the point where the validation error starts to worsen
AnswerB. It uses the sum of the absolute values of the weights as a penalty, and the values tend to become exactly 0

L1 regularization is a form of regularization that adds the sum of the absolute values of the weights as a penalty to the loss function, and it tends to push weights all the way to exactly 0. As a result, the coefficients of unused features become 0 and are automatically dropped, producing a feature-selection effect. This state is called sparse. Using the sum of squares of the weights as a penalty is L2 regularization, which keeps values small and uniform but rarely drives them to exactly 0. Randomly disabling units during training is dropout, and stopping training when the validation error starts to worsen is early stopping — both address overfitting, but through different mechanisms.

Q12 | Ridge regression

Which regularization is used in ridge regression?

  1. L0 regularization, which uses the number of nonzero weights as a penalty
  2. L2 regularization, which adds the sum of squares of the weights as a penalty
  3. L1 regularization, which uses the sum of the absolute values of the weights as a penalty
  4. Dropout, which randomly disables units during training
AnswerB. L2 regularization, which adds the sum of squares of the weights as a penalty

Ridge regression is a method that combines linear regression with L2 regularization. Because the sum of squares of the weights is added as a penalty, the coefficients as a whole are kept small, which prevents overfitting. Linear regression combined with L1 regularization instead is lasso regression, which tends to drive coefficients to exactly 0 and yields a sparse solution. The correspondence between L1 and lasso, and between L2 and ridge, is easy to mix up, so be sure to memorize them as pairs. L0 regularization is the idea of using the number of nonzero weights itself as a penalty; since a count cannot be differentiated, it is hard to incorporate into gradient-based training. Dropout is a regularization method for neural networks.

Q13 | Dropout

Which of the following correctly describes dropout?

  1. It flips or crops images to pad out the training data itself
  2. It randomly disables a fixed proportion of units during training to suppress overfitting
  3. It stops training at the point where the validation error starts to rise, to prevent overfitting
  4. It adjusts the output of a middle layer within a mini-batch to have mean 0 and variance 1 to stabilize training
AnswerB. It randomly disables a fixed proportion of units during training to suppress overfitting

Dropout is a method that randomly disables a fixed proportion of the units of a hidden layer at each training step. Because training is done on a slightly different network shape each time, the model can no longer rely entirely on a particular combination of units, producing an effect similar to averaging many models. Units are disabled only during training; at inference time all units are used. Despite the similar name, in the syllabus dropout is classified under regularization item 14, not under normalization layers item 19. Adjusting the distribution within a mini-batch is batch normalization, padding out the data is data augmentation, and stopping on worsening error is early stopping.

Q14 | Mini-batch training

Which of the following correctly describes mini-batch training?

  1. An approach that does not update the weights and keeps using a predetermined value
  2. An approach that computes the gradient over the entire training set and updates the weights just once
  3. An approach that updates the weights for each chunk of a few dozen to a few hundred examples
  4. An approach that updates the weights every time a single training example is used
AnswerC. An approach that updates the weights for each chunk of a few dozen to a few hundred examples

Mini-batch training is an approach that divides the training data into chunks of a few dozen to a few hundred examples and computes the gradient and updates the weights for each chunk. It has an intermediate property: its per-step gradient is not as noisy as online learning, which updates one example at a time, and it is not as computationally heavy as batch training, which updates just once using the entire dataset, so in practice it accounts for most training. Stochastic gradient descent (SGD) computes the gradient from only part of the data, and this fluctuation has the advantage of making it easier to escape saddle points and shallow local optima. Not updating the weights is not training at all, but plain inference.

Q15 | Epoch

Which of the following correctly describes the relationship between epoch and iteration?

  1. An epoch refers to the size of the mini-batch, and an iteration refers to the size of the learning rate
  2. An epoch means going through the data once, and an iteration means updating the weights once
  3. An epoch means updating the weights once, and an iteration means going through the data once
  4. Both epoch and iteration mean going through the data once
AnswerB. An epoch means going through the data once, and an iteration means updating the weights once

An epoch means going through the entire training set once, and an iteration means updating the weights once. The two are linked through the mini-batch size: the number of iterations per epoch equals the number of data examples divided by the mini-batch size. For instance, running 3,000 examples with a batch size of 50 makes one epoch equal to 60 iterations. Making the mini-batch size larger stabilizes each gradient but reduces the number of updates per epoch, while making it smaller increases the number of updates but increases the fluctuation. Both the mini-batch size and the learning rate are hyperparameters decided in advance by a person, and are separate concepts from epoch and iteration.

Q16 | Saddle point

Which of the following correctly describes a saddle point?

  1. A point that is a local minimum when viewed from one direction and a local maximum when viewed from another
  2. A point that has the smallest error in its nearby range, but there is a better point overall
  3. A region where the error value is constant over the entire range, with no slope in any direction
  4. A point where the error is at its smallest over the entire feasible range
AnswerA. A point that is a local minimum when viewed from one direction and a local maximum when viewed from another

A saddle point is a point that looks like the bottom of a valley from one direction but looks like the top of a mountain from another direction, and it is named for its resemblance to the shape of a horse's saddle. In deep learning, where the number of parameters is very large, it is considered more likely to get stuck at a saddle point than to reach a local optimum that is simultaneously the bottom of a valley in every direction. The point where the error is smallest over the entire range is the global optimum, and the point that is smallest only in its nearby range is a local optimum. Because stochastic gradient descent (SGD) has fluctuation in its gradient, it can more easily escape from such a stall.

Q17 | Adam

Which ideas does the Adam optimizer combine?

  1. It combines two kinds of penalty: the sum of the absolute values of the weights and the sum of their squares
  2. It combines the idea of using momentum with the idea of adapting the learning rate
  3. It alternates between two approaches, batch training and online learning
  4. It combines two kinds of search: grid search and random search
AnswerB. It combines the idea of using momentum with the idea of adapting the learning rate

Adam is an optimization method that combines the idea of momentum, which adds in the previous update amount as inertia, with the idea of RMSprop, which automatically adjusts the learning rate for each parameter, and it is widely used as the default choice in practice. The lineage of adapting the learning rate starts with AdaGrad; AdaGrad lowers the learning rate more for parameters that had larger updates, but it keeps lowering it indefinitely, so RMSprop eased this with an exponential moving average. The same lineage also includes AdaDelta, AdaBound, and AMSBound. Penalties on the weights are a matter of regularization, and grid search and random search are matters of hyperparameter search — neither is about the update rule itself.

Q18 | No free lunch

Which of the following does the No Free Lunch theorem state?

  1. That there is no universal algorithm that outperforms all others across every problem
  2. That the amount of training data needed increases sharply as the number of feature dimensions grows
  3. That continuing to lower the learning rate is guaranteed to reach the global optimum
  4. That making a model larger causes the error to worsen once before dropping again
AnswerA. That there is no universal algorithm that outperforms all others across every problem

The No Free Lunch theorem states that, averaged over every conceivable problem, no universal algorithm outperforms all the others. Consequently, the choice of method has to be considered case by case, and there is no single answer that always works. Making a model larger causing the error to worsen once before dropping again is double descent, which is also included under the same mid-level topic. The amount of data needed increasing sharply as feature dimensions grow is the curse of dimensionality. There is no guarantee that continuing to lower the learning rate will reach the global optimum, since it can get stuck at a local optimum or a saddle point.

Q19 | Random search

Which of the following correctly describes random search, a method for searching hyperparameters?

  1. Arrange candidate values in a grid and try every combination in turn
  2. Randomly disable units at each training step to obtain average behavior
  3. Stop training at the point where the validation error starts to worsen
  4. Randomly pick points from within the search range and try that combination
AnswerD. Randomly pick points from within the search range and try that combination

Random search is a method that randomly picks points from within the hyperparameter search range and tries them. Unlike grid search, which arranges candidates in a grid and tries every combination exhaustively, it is less likely to waste trials on hyperparameters that have little effect, and can explore the axes that matter more finely with the same number of trials. Values decided in advance by a person, such as the learning rate, the mini-batch size, and the number of layers, are hyperparameters. Stopping training when the validation error worsens is early stopping, and randomly disabling units during training is dropout — neither is a search method.

Q20 | Double descent

What is it called when, as you keep increasing a model's size or the amount of training, the validation error worsens once and then decreases again?

  1. The exploding gradient problem
  2. Acquisition of invariance
  3. Double descent
  4. The credit assignment problem
AnswerC. Double descent

Double descent is a phenomenon in which, as you keep increasing a model's size or amount of training, the validation error worsens once and then, if you keep increasing further, starts decreasing again. It is named for the error curve descending twice. It runs counter to the naive intuition that a larger model overfits more and gets worse, but it is included in the syllabus's optimization methods as one of the frameworks that explain the recent experience that larger models can sometimes give better results. The credit assignment problem is the question of which unit should be attributed responsibility for the error, the exploding gradient problem is the phenomenon in which the gradient diverges during backpropagation, and acquisition of invariance is the role of pooling layers.

Q21 | Convolutional layers

Compared with a fully connected layer, which is an appropriate characteristic of a convolutional layer?

  1. It connects to every unit of the previous layer, so the weights increase as the input grows larger
  2. It reuses the same filter while shifting its position, so it can extract features with few weights
  3. It feeds the previous time step's state back into the input, handling sequential data that has an order
  4. It has no weights to learn at all, and reduces the output by taking a representative value of a region
AnswerB. It reuses the same filter while shifting its position, so it can extract features with few weights

A convolutional layer extracts local features by sliding a small filter over the input while reusing the same weights. Because of two things — that it only connects to nearby elements, and that it uses the same weights regardless of position — it can handle images with far fewer parameters than a fully connected layer. As a side effect of this reuse, the subject can be picked up as the same feature no matter where it appears in the image. Connecting to every unit of the previous layer describes a fully connected layer, having no weights to learn and taking a representative value describes a pooling layer, and feeding back the previous time step's state describes a recurrent layer. The result of passing one filter through is called a feature map, the width by which it is shifted is the stride, and filling the surrounding area is padding.

Q22 | Batch normalization

Which of the following correctly describes batch normalization?

  1. It adds the sum of squares of the weights as a penalty to the loss function, keeping the weights small and uniform
  2. It disables a fixed proportion of units at each training step, removing bias toward a particular dependency
  3. It adjusts the values across a whole layer together within a single sample, so it is not affected by the mini-batch size
  4. It adjusts the values of the same channel together within a mini-batch to have mean 0 and variance 1
AnswerD. It adjusts the values of the same channel together within a mini-batch to have mean 0 and variance 1

Batch normalization is a normalization layer that adjusts the values of the same channel together within a mini-batch to have mean 0 and variance 1, stabilizing and speeding up training. Its weakness is that when the mini-batch size is small, the statistics become unreliable. Adjusting the values across a whole layer together within a single sample is layer normalization, which is not affected by the mini-batch size and so is widely used in models that handle sequences. There are also instance normalization and group normalization, and these four are the ones listed in the syllabus's normalization layers. Note that this mid-level topic's name changed from regularization layers to normalization layers in version 1.1. L2 regularization, which uses the sum of squares of the weights as a penalty, and dropout, which disables units, are classified under regularization, not normalization layers.

Q23 | Pooling layers

What is the most appropriate role a pooling layer plays?

  1. It skips over layers to add the input into a later layer, creating a path the gradient can travel
  2. It aligns the distribution of a middle layer's outputs, stabilizing and speeding up training
  3. It reduces a region to a representative value, making the model robust to slight shifts in position
  4. It compresses the input to a low dimension and then makes it possible to reconstruct the original input
AnswerC. It reduces a region to a representative value, making the model robust to slight shifts in position

A pooling layer is a layer that reduces a fixed-size region to a single representative value; there is max pooling, which takes the maximum value of a region, and average pooling, which takes the average. It not only lowers the resolution and lightens the computation, but also makes the representative value less likely to change even if the subject shifts by a few pixels, making the model robust to positional shifts. This is called acquisition of invariance. Global average pooling (GAP), which averages a whole feature map into a single value, can greatly reduce the number of parameters when used instead of stacking a fully connected layer at the end. A pooling layer having no weights to learn is also a difference from a convolutional layer. Skipping over layers to add in is a skip connection, aligning the distribution is a normalization layer, and compressing and then reconstructing is an autoencoder.

Q24 | Skip connections

Why does adopting a skip connection allow training to proceed even in a very deep network?

  1. Because a path that skips over layers is created, letting the gradient reach the input side by a shallow route
  2. Because units are randomly disabled, so it no longer depends on a particular path
  3. Because the output distribution of each layer is aligned, keeping the value fluctuation suppressed layer by layer
  4. Because the number of weights decreases, making it less prone to overfitting even with limited data
AnswerA. Because a path that skips over layers is created, letting the gradient reach the input side by a shallow route

A skip connection is wiring that adds the input, unchanged, into the output of a layer several layers ahead, skipping over the layers in between. During backpropagation, the gradient reaches the input side not only through the long path passing through many layers but also through the short skipping path, so the gradient is less likely to vanish even when stacked deeply. The representative example that adopted this idea and made depths beyond 100 layers a reality is ResNet, and in the syllabus the only keyword placed under the mid-level topic of skip connections is ResNet. Randomly disabling units is dropout, and aligning the distribution of each layer is a normalization layer — both are separate mechanisms from a skip connection. A skip connection is also not a way of reducing the number of weights.

Q25 | LSTM and GRU

Which is the correct combination regarding the gates that LSTM and GRU have?

  1. Both LSTM and GRU commonly have the same two gates: reset and update
  2. Both LSTM and GRU commonly have the same three gates: input, output, and forget
  3. LSTM has three, input, output, and forget; GRU has two, reset and update
  4. LSTM has two, reset and update; GRU has three, input, output, and forget
AnswerC. LSTM has three, input, output, and forget; GRU has two, reset and update

LSTM has a memory pathway called the CEC that stores information, and controls what goes in and out of it with three gates: the input gate, which decides what to write in; the forget gate, which decides what to discard; and the output gate, which decides what to send out. This mechanism makes it possible to preserve long-range dependencies, and it greatly eased the vanishing gradient problem of recurrent neural networks. GRU is a simplification of LSTM, with only two gates, the reset gate and the update gate. It has fewer parameters and is lighter to compute, but it is not always higher-performing than LSTM. The fact that the gate counts are 3 and 2, and that CEC is a term on the LSTM side, is the distinction that is repeatedly tested.

Q26 | Jordan network

Which of the following is an appropriate description of the structure of a Jordan network?

  1. It feeds the output layer's output back into the hidden layer at the next time step, as its input
  2. It feeds the hidden layer's output back into the hidden layer at the next time step, as its input
  3. It feeds the output layer's output back into the input layer at the same time step, as its input
  4. It skips the input layer's value straight through to the output layer and adds it in
AnswerA. It feeds the output layer's output back into the hidden layer at the next time step, as its input

A Jordan network is a type of recurrent neural network that feeds the output layer's output back into the hidden layer at the next time step, as its input. By contrast, feeding the hidden layer's output back into the hidden layer at the next time step is an Elman network, and what is fed back is the difference between the two. Both are structures for handling time-series data in order, and their training uses BPTT, which unrolls the network in the time direction and then applies backpropagation. For long sequences, the vanishing gradient problem and the exploding gradient problem show up strongly, so LSTM and GRU, which have gating mechanisms, are used. Skipping over layers to add in is a skip connection, which is not about the time direction.

Q27 | Attention

Which of the following correctly describes Source-Target Attention?

  1. A mechanism that computes the strength of the relationship between elements within the same sequence
  2. A mechanism that computes the strength of the relationship between two separate sequences
  3. A mechanism that runs multiple ways of paying attention in parallel and combines the results
  4. A mechanism that adds a signal unique to each position to give word-order information
AnswerB. A mechanism that computes the strength of the relationship between two separate sequences

Attention is a mechanism that computes, as weights, where in the input attention should be paid, and it is expressed using a query, a key, and a value. Among Attention mechanisms, Source-Target Attention computes the strength of the relationship between two separate sequences, in a setting such as translation where the input sequence and the output sequence are distinct. Computing the relationship between elements within the same sequence is instead Self-Attention, which is the central component of the Transformer. Adding a signal unique to each position to give word order is positional encoding, and running multiple ways of paying attention in parallel and combining the results is Multi-Head Attention — both are separate components that make up a Transformer.

Q28 | Positional encoding

Which of the following correctly describes the Transformer?

  1. A model that compresses the input to a low dimension and trains to reconstruct the original input from it
  2. A model that alternately stacks convolution and pooling to gather local features step by step
  3. A model built with Attention instead of recurrent connections, giving word order through positional encoding
  4. A model with a structure of stacked recurrent connections that hands the previous time step's state along in order
AnswerC. A model built with Attention instead of recurrent connections, giving word order through positional encoding

The Transformer is a sequence model built mainly with Attention, without recurrent connections. Because it does not need to wait for the previous time step's computation, it can process an entire sequence in parallel, and it can directly capture relationships between distant words in a single computation. However, processing in parallel erases word-order information, so positional encoding, which adds a signal unique to each position, is used to give word order. The point that catches people out is that the Transformer is not a type of recurrent neural network. Stacking recurrent connections to hand along the previous time step's state describes an RNN, stacking convolution and pooling describes a CNN, and compressing and then reconstructing describes an autoencoder.

Q29 | VAE

What is the point on which a variational autoencoder (VAE) differs from an ordinary autoencoder?

  1. That it compresses the input to a low dimension once and then reconstructs the original input
  2. That it learns the compressed representation as a probability distribution rather than a single point, allowing new data to be created
  3. That it can proceed with training without preparing correct-answer labels
  4. That it trains each layer one at a time in sequence to create the initial values of a deep network
AnswerB. That it learns the compressed representation as a probability distribution rather than a single point, allowing new data to be created

A variational autoencoder (VAE) learns the compressed representation not as a single point but as a probability distribution. Sampling a value from that distribution and reconstructing it lets you create new data that was not in the training data, so it is treated as a generative model. Compressing to a low dimension and then reconstructing, and being able to train without correct-answer labels, are properties shared with an ordinary autoencoder too, not features unique to VAE. Training each layer one at a time in sequence to create the initial values of a deep network is pretraining via a stacked autoencoder. The syllabus also lists VQ-VAE, info VAE, and β-VAE under the same mid-level topic.

Q30 | Mixup

Among image data augmentation methods, which operation does Mixup perform?

  1. Crop part of the image and enlarge it to make a new example
  2. Overlay two images to create one new example
  3. Flip the image left and right to create examples facing a different direction
  4. Paint over part of the image with a solid rectangle to hide the information in that area
AnswerB. Overlay two images to create one new example

Mixup is a data augmentation method that overlays two images at a certain ratio and mixes their correct-answer labels at the same ratio to create new training data. The similarly named CutMix does not overlay images but instead cuts and pastes parts of two images together into one. Painting over part of the image with a solid rectangle is Cutout or Random Erasing, flipping left and right is Random Flip, and cropping part of it is Crop. There are also Rotate, which rotates the image, Brightness, which changes the brightness, Contrast, which changes the contrast, and RandAugment, which automatically decides which transformations to apply and how much, and for text, paraphrasing and noising are listed. In every case, it is important to avoid choosing a transformation that destroys the meaning.

Q31 | GoogLeNet

Which of the following is a characteristic of GoogLeNet?

  1. It used an Inception module that applies filters of different sizes in parallel
  2. It won the 2012 ILSVRC and used the ReLU function and GPU training
  3. It introduced skip connections that jump over layers, achieving a depth of over 100 layers
  4. It took a deep configuration by stacking only small 3-by-3 convolutions
AnswerA. It used an Inception module that applies filters of different sizes in parallel

GoogLeNet is a model that used an Inception module, which applies filters of different sizes in parallel and combines the results, achieving a deep configuration while keeping the number of parameters down. VGG took a simple, deep configuration by stacking only small 3-by-3 convolutions, and it placed in the top ranks of the same 2014 ILSVRC as GoogLeNet. ResNet made depths beyond 100 layers possible with skip connections that jump over layers, and AlexNet won the 2012 ILSVRC, combining the ReLU function, dropout, and GPU training. After that, this lineage expanded to Wide ResNet, DenseNet, SENet, EfficientNet, MobileNet, and others, and the Vision Transformer, which brings in the Transformer idea instead of convolution, is also included.

Q32 | One-stage detection

Which is a characteristic that YOLO and SSD have in common?

  1. They detect objects through a two-stage procedure of proposing candidate regions and then classifying them
  2. They assign a class to each individual pixel, painting the regions of the image
  3. They perform candidate-region extraction and classification together in a single pass of inference, emphasizing speed
  4. They aim to estimate the positions of human body joints and determine posture
AnswerC. They perform candidate-region extraction and classification together in a single pass of inference, emphasizing speed

YOLO and SSD are one-stage object detection models that perform candidate-region extraction and class determination together in a single pass of inference, and their strength is speed. By contrast, the lineage that starts with R-CNN and developed through Fast R-CNN and Faster R-CNN is a two-stage approach that proposes candidate regions first and then classifies them; it is slower but tends to achieve higher accuracy. Assigning a class to each pixel is semantic segmentation, for which FCN, SegNet, U-Net, PSPNet, and DeepLab are representative examples. Instance segmentation, which separates objects of the same class down to individual instances, is handled by Mask R-CNN, and pose estimation, which estimates human body joint positions, is represented by OpenPose.

Q33 | BERT

Which of the following correctly describes BERT?

  1. It maps images and their captions into the same vector space to connect the two
  2. It works backward through a process of adding noise, generating data little by little
  3. It uses the Transformer's encoder and looks at context from both directions, before and after
  4. It stacks recurrent layers and reads words one at a time in order to build context
AnswerC. It uses the Transformer's encoder and looks at context from both directions, before and after

BERT is a language model that uses the Transformer's encoder, and its distinguishing feature is that it is pretrained looking at context from both directions, before and after, in a sentence. It popularized the approach of pretraining on a large amount of text and then adapting to individual tasks. Stacking recurrent layers and reading words in order in sequence is a model that uses a recurrent neural network; the Transformer has no recurrent connections. Mapping images and captions into the same vector space is CLIP, which in the syllabus is placed under multimodal. Working backward through a noise-adding process to generate data is a Diffusion Model, which falls under the mid-level topic of data generation.

Q34 | CBOW

In which direction does CBOW of word2vec train?

  1. It trains to predict, from the words that appear around it, the central word
  2. It trains to predict, from the central word, the words that appear around it
  3. It trains to predict, from the whole sentence, whether that sentence is positive or negative
  4. It trains to predict, from the first half of a sentence, whether the following second half comes next
AnswerA. It trains to predict, from the words that appear around it, the central word

CBOW trains to predict the central word from the words that appear around it. The reverse — training to predict the surrounding words from the central word — is skip-gram, and since the direction is reversed, be careful not to mix the two up. Both are word2vec methods, and their purpose is to obtain a distributed representation, which expresses words as dense vectors. A one-hot vector, which assigns one dimension per word, has as many dimensions as the vocabulary size and cannot express closeness between words, whereas in a distributed representation, words with similar meaning are placed close together in the vector space. Predicting whether a sentence is positive or negative is the applied task of sentiment analysis, not a matter of how word representations are trained.

Q35 | RLHF

Which method builds a reward model from human feedback and then uses reinforcement learning to adjust the model?

  1. DQN
  2. A3C
  3. PPO
  4. RLHF
AnswerD. RLHF

RLHF is a method that builds a reward model from preference ratings given by humans and then uses that as the reward to adjust the model through reinforcement learning. It is well known for adjusting language models, but a surprising point worth remembering is that in the syllabus it is placed under the mid-level topic of deep reinforcement learning, not natural language processing. DQN is a representative deep reinforcement learning method that approximates the action-value function with a neural network, PPO is a policy-training method that stabilizes learning by limiting the size of policy updates, and A3C is a method that runs multiple environments in parallel to advance training. Double DQN, dueling networks, Rainbow, Ape-X, and Agent57 are also listed under the same mid-level topic.

Q36 | Diffusion Model

Which of the following correctly describes a Diffusion Model?

  1. It learns the three-dimensional appearance from photos taken from multiple viewpoints and creates images from new viewpoints
  2. It learns the compressed representation as a probability distribution and reconstructs by sampling from it
  3. It generates realistic-looking data by making two networks compete against each other
  4. It works backward through a process of gradually adding noise, generating data
AnswerD. It works backward through a process of gradually adding noise, generating data

A Diffusion Model is a generative model that works backward through a process of gradually adding noise to data, creating the target data starting from noise. Its mechanism is entirely different from that of a generative adversarial network (GAN), which is trained by making two networks compete against each other, so it is not a type of GAN. DCGAN, CycleGAN, and Pix2Pix are derived from GAN. Learning the three-dimensional appearance from photos taken from multiple viewpoints is NeRF, and learning the compressed representation as a probability distribution is the variational autoencoder (VAE). It is worth keeping in mind that GAN, VAE, and the Diffusion Model are separate lineages placed side by side.

Q37 | Transfer learning

Which of the following correctly describes the difference between transfer learning and fine-tuning?

  1. Transfer learning retrains the entire set of weights, and fine-tuning trains only the output layer
  2. Both transfer learning and fine-tuning train the weights from scratch without pretraining
  3. Transfer learning mainly trains around the output layer, and fine-tuning retrains a broad range
  4. Transfer learning can only be used for images, and fine-tuning only for language
AnswerC. Transfer learning mainly trains around the output layer, and fine-tuning retrains a broad range

Transfer learning is a method that keeps the feature-extracting part of a pretrained model fixed and trains mainly around the output layer alone on the new task. Fine-tuning retrains the entire set of weights, or a broad range of them, on the new data; it requires more data and computation, but it tends to do better when the target domain differs greatly from the pretraining data. With fine-tuning, catastrophic forgetting can occur, in which the original model's capabilities are lost. Both methods presuppose a pretrained model; neither trains from scratch. Nor is the choice between them determined by the type of data being handled.

Q38 | Foundation models

Under which mid-level topic in the syllabus is the term foundation model placed?

  1. Under transfer learning and fine-tuning, alongside few-shot
  2. Under multimodal, alongside CLIP and DALL-E
  3. Under deep reinforcement learning, alongside DQN and RLHF
  4. Under natural language processing, alongside large language models (LLMs)
AnswerB. Under multimodal, alongside CLIP and DALL-E

A foundation model refers to a large model that is pretrained on a large amount of data and can be repurposed for a variety of downstream tasks; in the syllabus it is placed under item 32, multimodal, alongside CLIP, DALL-E, Flamingo, Unified-IO, and zero-shot. Terms around generative AI are scattered in their placement: the large language model (LLM) is included under item 27, natural language processing, and RLHF under item 29, deep reinforcement learning. What is listed alongside item 31, transfer learning and fine-tuning, is few-shot, one-shot, self-supervised learning, pretrained models, catastrophic forgetting, and the like. Keeping track of which mid-level topic each term is placed under makes it easier to grasp the overall shape of the scope.

Q39 | SHAP

Which of the following is the correct idea behind SHAP, a method for explainable AI (XAI)?

  1. It deliberately swaps the values of a feature and measures importance from how much accuracy drops
  2. It shows where in the image the model looked to make its decision, in a form like a heat map
  3. It fits a simple model near the prediction and explains using that model's coefficients
  4. Using ideas from cooperative game theory, it shows the contribution of each feature by allocating it
AnswerD. Using ideas from cooperative game theory, it shows the contribution of each feature by allocating it

SHAP is a method that borrows ideas used in cooperative game theory to fairly allocate and show the contribution of each feature to a prediction. Showing where in the image the model looked to make its decision, in a form like a heat map, is CAM or Grad-CAM; deliberately swapping the values of a feature and measuring importance from how much accuracy drops is Permutation Importance; and fitting a simple model near an individual prediction and explaining using its coefficients is LIME. All of these are included under item 33, model interpretability, in the syllabus, and the effort to present the basis for a decision in a form people can understand is collectively called explainable AI (XAI).

Q40 | The lottery ticket hypothesis

What is the content of the lottery ticket hypothesis regarding model lightweighting?

  1. The idea that having a small model learn from the output of a large teacher model can preserve performance
  2. The idea that a large network contains a part that, trained on its own, reaches performance equal to the whole
  3. The idea that making a model larger requires the amount of training data needed to increase exponentially
  4. The idea that lowering the precision of the numbers representing the weights causes only a very small drop in accuracy
AnswerB. The idea that a large network contains a part that, trained on its own, reaches performance equal to the whole

The lottery ticket hypothesis is the hypothesis that a large network contains a winning sub-network that, if taken out and trained on its own, reaches performance equal to the whole. It is cited as background explaining why pruning, which removes connections and layers with small contributions, works well. Having a small model learn from the output of a large teacher model is distillation, and lowering the precision of the numbers representing the weights is quantization; together with pruning, these three make up the syllabus's lightweighting methods, collectively called model compression. The amount of training data needed increasing sharply is a matter of the curse of dimensionality. A typical case where lightweighting is needed is edge AI, which runs inference on-device.

Practice: answer the questions on this page

This practice tool asks questions in random order (it works when JavaScript is enabled). You can still read all the questions and explanations above without it.

* The explanations are information for study purposes. Exam scope and systems change from year to year, so always check the official announcements of the organization that administers the exam.

This page is a translation of the Japanese original. If the translation and the original differ, the Japanese version takes precedence. View the Japanese original