It is used to improve how actions are chosen, guided by reward
As its name suggests, it is used for regression problems, outputting a continuous value directly
Despite its name, it is used for classification problems, outputting the probability of belonging to a class
It is used without training data, to find clusters within the data
AnswerC. Despite its name, it is used for classification problems, outputting the probability of belonging to a class
Although logistic regression has "regression" in its name, it is a classification method that converts the result of a linear computation into a probability form to determine a class. Predicting a continuous value directly is the job of methods for regression problems, such as linear regression, and being drawn to that answer by the name is a classic mistake. Finding clusters without training data is clustering, a topic in unsupervised learning, and improving action choices guided by reward is reinforcement learning — both placed in separate subsections.
Q2 | Regression and classification
What is the difference between a regression problem and a classification problem in supervised learning?
A regression problem uses training labels, and a classification problem learns using only reward as a guide
A regression problem predicts a discrete class, and a classification problem predicts a continuous value
A regression problem predicts a continuous value, and a classification problem predicts a discrete class
A regression problem uses only features, and a classification problem uses features and training labels
AnswerC. A regression problem predicts a continuous value, and a classification problem predicts a discrete class
A regression problem predicts a continuous value, such as sales or temperature, and a classification problem predicts a discrete class, such as whether an email is spam. Swapping these two is the most common mistake. Supervised learning requires a pair of features and a training label for either type of problem, so it is not the case that only one of them needs just features. Learning guided only by reward is reinforcement learning, and does not describe supervised learning.
Q3 | Simple vs. multiple regression
Which correctly describes the difference between simple regression analysis and multiple regression analysis?
Simple regression analysis does not need training labels, and multiple regression analysis does
Simple regression analysis has several explanatory variables, and multiple regression analysis has one
Simple regression analysis has one explanatory variable, and multiple regression analysis has several
Simple regression analysis is used for classification, and multiple regression analysis is used for predicting continuous values
AnswerC. Simple regression analysis has one explanatory variable, and multiple regression analysis has several
Simple regression analysis has one explanatory variable, and multiple regression analysis has two or more. Since the characters for "simple" and "multiple" directly indicate the number of explanatory variables, an explanation that swaps them is clearly wrong. Both are regression methods that predict a continuous value, so it is also incorrect to say one of them is for classification. Furthermore, regression analysis is supervised learning, requiring training labels regardless of the number of variables.
Q4 | Margin maximization
What is the idea behind how a support vector machine (SVM) decides a boundary?
It repeatedly redraws the boundary while increasing the weight of misclassified data
It draws the boundary through the midpoint of the line segment connecting the centroids of the classes
It draws the boundary at the position that maximizes the margin between it and the nearest data point
It draws the boundary aligned with the direction in which the variance of the explanatory variables is maximized
AnswerC. It draws the boundary at the position that maximizes the margin between it and the nearest data point
SVM performs margin maximization, drawing the boundary at the position that maximizes the margin between it and the nearest data point. These closest points are called support vectors. Repeatedly adding learners while increasing the weight of misclassified data is the idea behind boosting, and finding the direction of maximum variance is the idea behind principal component analysis (PCA) — neither describes SVM. Drawing the boundary through the midpoint between centroids does not consider margin at all.
Q5 | The kernel trick
What does the kernel trick make possible?
Performing nonlinear separation without actually computing the coordinates after mapping to a higher dimension
Computing high-dimensional features one by one and then separating them linearly
Automatically splitting data into multiple clusters without using training labels
Greatly reducing the number of data records used for training to make it faster
AnswerA. Performing nonlinear separation without actually computing the coordinates after mapping to a higher dimension
The kernel trick makes nonlinear separation possible without actually computing the coordinates after mapping to a higher-dimensional space, by directly computing the inner product after that mapping using a kernel function. The reason it is called a "trick" is precisely that the coordinates do not need to be computed one by one, so an explanation saying it computes them first and then separates has the direction reversed. Splitting data without training labels is a clustering topic, and it is not a mechanism for reducing the number of training records.
Q6 | Parallel and sequential
What is the difference between bagging and boosting in ensemble learning?
Bagging can only be used with decision trees, and boosting only with linear regression
Bagging trains weak learners sequentially, and boosting trains them in parallel
Bagging trains weak learners in parallel, and boosting trains them sequentially
Bagging does not use training labels, and boosting does
AnswerC. Bagging trains weak learners in parallel, and boosting trains them sequentially
Bagging builds learners in parallel and independently of one another from training data created by bootstrap sampling, then takes a majority vote or average. Boosting adds learners sequentially, with each new one weighting more heavily the examples the previous one got wrong, so it depends on order and cannot be built in parallel. Swapping parallel and sequential is the most common mistake here. Both use training labels as supervised learning, and the types of weak learner that can be used are not limited to decision trees or linear regression.
Q7 | Random forest
Which correctly describes how a random forest is trained?
Growing a single decision tree as deep as possible and equalizing the number of branches at the end
Adding decision trees one by one, each weighting more heavily the examples the previous tree got wrong
Training many decision trees in parallel, built from data created by bootstrap sampling
Using no decision trees at all, and averaging the predictions of multiple linear regressions instead
AnswerC. Training many decision trees in parallel, built from data created by bootstrap sampling
A random forest is a representative example of bagging: it creates slightly different training data through bootstrap sampling, which is sampling with replacement, and grows many decision trees in parallel and independently from that data, then takes a majority vote or average. Adding trees one by one while weighting more heavily the previous tree's errors is the description of the boosting side, such as gradient boosting, and swapping this is a classic mistake. It uses not a single tree but a large number of them, and what is combined is decision trees, not linear regressions.
Q8 | A weakness of decision trees
Which correctly describes a property of decision trees?
The deeper the branching, the worse the fit to the training data necessarily becomes
It is a method that requires no training labels, so it can be used for neither classification nor regression
The branching conditions cannot be read by a person, but overfitting is hard to induce
A person can read the branching conditions, but a single tree is prone to overfitting
AnswerD. A person can read the branching conditions, but a single tree is prone to overfitting
A decision tree is a method that arranges branching conditions on features into a tree shape, and its strength is that a person can read which condition led down which branch. On the other hand, the deeper the tree is grown, the more it memorizes fine idiosyncrasies of the training data, making a single tree prone to overfitting. Growing it deeper actually improves the fit to the training data, and the problem is that performance drops on new data anyway, so saying the fit necessarily worsens is incorrect. A decision tree is supervised learning and can be used for both classification and regression.
Q9 | AR and VAR
What is the relationship between the autoregressive model (AR model) and the vector autoregressive model (VAR model)?
AR extracts local features of an image, and VAR extracts frequency components of audio
AR handles multiple time series simultaneously, and VAR explains only one time series using its past values
AR requires no training labels, and VAR is a method that requires training labels
AR explains one time series by its own past values, and VAR explains multiple time series by each other's past values
AnswerD. AR explains one time series by its own past values, and VAR explains multiple time series by each other's past values
An AR model explains the current value of a single time series using that series's own past values. A VAR model extends this to multiple time series, having them explain each other using each other's past values. The difference is one time series versus multiple, so an explanation that swaps this is wrong. Both are models for handling time series, not tools for extracting local image features or audio frequency components, nor is the distinction one of needing training labels or not.
Q10 | Multiclass classification
Which of the following is an example of multiclass classification?
Assigning an image of a handwritten digit to one of 0 through 9
Choosing the next move based on reward from the state of the board
Finding clusters of similar customers from purchase history
Predicting tomorrow's high temperature value from temperature and humidity
AnswerA. Assigning an image of a handwritten digit to one of 0 through 9
Multiclass classification is a classification problem that determines which of three or more classes something belongs to, and assigning a handwritten digit to one of 0 through 9 is a typical example. Predicting a continuous value, such as a high temperature value, is a regression problem, not classification. Finding clusters of similar customers is clustering, a form of unsupervised learning, and choosing a move based on reward is reinforcement learning — neither is a classification problem in supervised learning to begin with.
Q11 | Unsupervised input
Which correctly describes unsupervised learning?
Given only features, it finds the structure or regularity that the data itself has
Given pairs of features and training labels, it reduces the difference from the correct answer
Using only a small amount of labeled data, it discards the rest and proceeds with training
Guided by reward received from the environment, it improves how actions are chosen
AnswerA. Given only features, it finds the structure or regularity that the data itself has
The syllabus lists as a goal understanding that unsupervised learning requires only features. Without correct-answer labels, unsupervised learning finds the structure inherent in the data itself through clustering or dimensionality reduction. Presupposing pairs of features and training labels is supervised learning, and being guided by reward is reinforcement learning — each placed in a separate subsection. Discarding remaining data is also not a description of unsupervised learning.
Q12 | The k-means algorithm
Which correctly describes how the k-means algorithm proceeds?
Decide the number of clusters in advance, then repeat updating the centroids and reassigning points
Given correct-answer labels, maximize the margin of the boundary
Without deciding the number of clusters, repeatedly merge the closest items in order
Estimate the proportion of topics latent in a document from how words appear
AnswerA. Decide the number of clusters in advance, then repeat updating the centroids and reassigning points
The k-means algorithm is a non-hierarchical clustering method that fixes the number of clusters k in advance and repeats the operations of assigning each point to the nearest centroid and recomputing the centroids. Repeatedly merging the closest items without deciding the number of clusters is hierarchical clustering, such as Ward's method, which can produce a dendrogram (tree diagram). Maximizing the margin is SVM, a form of supervised learning, and estimating topics from how words appear is a topic model such as latent Dirichlet allocation (LDA).
Q13 | Ward's method
What is the relationship between Ward's method and a dendrogram (tree diagram)?
It is a dimensionality reduction method, drawing the two reduced axes as a dendrogram
It is a hierarchical clustering method, and the process of merging can be drawn as a dendrogram
It is a supervised learning method, drawing the classification boundary as a dendrogram
It is a non-hierarchical clustering method, and a dendrogram cannot in principle be drawn
AnswerB. It is a hierarchical clustering method, and the process of merging can be drawn as a dendrogram
Ward's method is hierarchical clustering that repeatedly merges the closest items in order, and the process of merging can be drawn as a dendrogram (tree diagram). Its advantage is that the number of clusters can be decided afterward by looking at where to cut. It is k-means that is non-hierarchical and fixes the number of clusters in advance, and it does not produce a dendrogram. Ward's method is neither a dimensionality reduction method nor a supervised learning method, and it does not use correct-answer labels.
Q14 | PCA and t-SNE
Which correctly describes the difference between principal component analysis (PCA) and t-SNE?
PCA is linear dimensionality reduction, and t-SNE is nonlinear and suited to visualization
PCA requires training labels, and t-SNE is a method that runs on features alone
PCA is nonlinear dimensionality reduction, and t-SNE is linear and suited to visualization
PCA estimates the topics of a document, and t-SNE is a method for decomposing matrices
AnswerA. PCA is linear dimensionality reduction, and t-SNE is nonlinear and suited to visualization
PCA is linear dimensionality reduction that re-orients the axes toward the direction of maximum variance, while t-SNE is a nonlinear method that maps to a lower dimension while preserving the closeness of nearby points, and is often used for visualization in two or three dimensions. Swapping linear and nonlinear here is a classic mistake. Estimating the topics of a document is a topic model such as latent Dirichlet allocation (LDA), and decomposing a matrix is singular value decomposition (SVD). Both PCA and t-SNE are unsupervised learning, so neither requires training labels.
Q15 | Cold start
What is the cold start problem in collaborative filtering?
That a recommended product is out of stock and cannot be delivered to the user
That relying only on a product's description text biases the recommendations
That for a user with too much history, the recommendation computation never finishes
That good recommendations cannot be produced for a user or product with no history yet
AnswerD. That good recommendations cannot be produced for a user or product with no history yet
Because collaborative filtering recommends based on the similarity of ratings between users or between products, a new user or new product with no rating history yet provides no clue to go on, making it hard to produce recommendations. This is the cold start problem. Content-based filtering, which relies on a product's description text or attributes themselves, is discussed alongside it as an approach that compensates for this weakness. It is not about computation time or inventory.
Q16 | The signal in reinforcement learning
What does reinforcement learning use as its guide for learning?
The reward signal received from the environment
The count of word occurrences contained in a document
A correct-answer label given by a person for each instance
Only the closeness of distance between features
AnswerA. The reward signal received from the environment
Reinforcement learning interacts with an environment through trial and error, learning a policy that maximizes the cumulative reward received. The guide is reward, not a correct-answer label for each move, so it is not a form of supervised learning. Using only closeness of distance as a guide is the idea behind clustering and other unsupervised learning, and word occurrence counts are a feature-representation topic in natural language processing — neither is the learning signal of reinforcement learning. Evaluating the future with a discount rate, in case reward is not returned immediately, is also a characteristic of this framework.
Q17 | Value and policy
Which correctly describes the two representative approaches in reinforcement learning?
There is only the method of learning a value function, and the policy is uniquely determined from it
There is a method that learns a value function and a method that learns a policy directly, and Actor-Critic combines both
There is only the method of learning a policy directly, and a value function cannot be used partway through learning
Both the value function and the policy are designed by hand by a person and are never a target of learning
AnswerB. There is a method that learns a value function and a method that learns a policy directly, and Actor-Critic combines both
The syllabus lists as a goal understanding the two representative approaches of learning a value function and learning a policy. The former estimates a state-value function or an action-value function and then chooses an action based on it, with Q-learning and SARSA as representative examples. The latter represents the policy itself with parameters and updates it directly, with policy gradient methods and REINFORCE as representative examples. Actor-Critic combines an Actor responsible for the policy and a Critic responsible for the value, having both together. An explanation saying only one of the two exists, or that neither can be a target of learning, is incorrect.
Q18 | Q-learning and SARSA
Which correctly describes the difference between Q-learning and SARSA?
Q-learning uses the value of the action actually chosen, and SARSA uses the maximum action value at the next state
Q-learning requires training labels, and SARSA updates without using training labels
Q-learning does not use an action-value function, and only SARSA updates one
Q-learning uses the maximum action value at the next state, and SARSA uses the value of the action actually chosen
AnswerD. Q-learning uses the maximum action value at the next state, and SARSA uses the value of the action actually chosen
Both are reinforcement learning methods that update an action-value function, differing in what is used for the update. Q-learning uses whichever available action at the next state has the maximum value, while SARSA updates using the value of the next action actually chosen. Swapping the two is a classic mistake. Q-learning also updates an action-value function, so saying only one of them uses one does not hold. Both are frameworks that learn from reward, so it is also wrong to say training labels are required.
Q19 | ε-greedy
Consider a 4-armed bandit problem where, with probability ε, one of the 4 arms is chosen uniformly at random, and with probability 1−ε, the arm with the highest estimated value is chosen, following an ε-greedy policy. When ε is 0.1, what is the probability that a single selection results in the arm with the highest estimated value being chosen?
0.900
0.100
0.250
0.925
AnswerD. 0.925
The arm with the highest estimated value is chosen in two ways: through the greedy choice, or by chance when the random choice happens to land on that same arm. The former is 1 minus 0.1, giving 0.900, and the latter is 0.1 divided by 4, giving 0.025, for a total of 0.925. 0.900 is the value obtained by overlooking the chance of landing on it via the random choice; 0.250 is the probability of choosing uniformly from the 4 arms; and 0.100 is just ε itself. The larger ε is, the more exploration occurs, and this probability falls.
Q20 | Computing a discounted sum
From a certain point in time, you receive a reward of 4 one step later, a reward of 12 two steps later, and a reward of 8 three steps later. If you multiply the reward one step later by 0.5, the reward two steps later by 0.5 squared, and the reward three steps later by 0.5 cubed, and sum the results, what is that value?
24.0
6.0
7.5
12.0
AnswerB. 6.0
Multiplying 4 by 0.5 gives 2, multiplying 12 by 0.5 squared, i.e., 0.25, gives 3, and multiplying 8 by 0.5 cubed, i.e., 0.125, gives 1, for a total of 6.0. 24.0 is the value obtained by adding 4, 12, and 8 with no discounting applied at all; 12.0 is that 24 multiplied by 0.5 just once; and 7.5 is obtained by reversing the order of multiplication, applying 0.125 to 4, 0.25 to 12, and 0.5 to 8 — all common mix-ups. Since the role of the discount rate is to value more distant future rewards less, the nearer ones carry the larger weight.
Q21 | Computing precision
A model's predictions were: 40 true positives, 10 false positives, 20 false negatives, and 130 true negatives. What is the precision?
0.85
0.80
0.73
0.67
AnswerB. 0.80
Precision takes as its denominator the items predicted positive, so dividing the 40 true positives by 50, the sum of 40 true positives and 10 false positives, gives 0.80. 0.67 is recall, which takes as its denominator the items that are actually positive, dividing 40 by 60, the sum of 40 true positives and 20 false negatives. 0.85 is accuracy, the proportion of the 170 correctly predicted out of all 200 cases, and 0.73 is the F-score, the harmonic mean of the precision 0.80 and the recall 0.67 — all values obtained by computing a different metric.
Q22 | The difference in denominator
Which correctly describes the difference in denominator between precision and recall?
Precision takes the number actually positive as its denominator, and recall takes the number predicted positive
Precision takes only the number of true negatives as its denominator, and recall takes only the number of false negatives
Both precision and recall take the total number of data records directly as their denominator
Precision takes the number predicted positive as its denominator, and recall takes the number actually positive
AnswerD. Precision takes the number predicted positive as its denominator, and recall takes the number actually positive
Precision is "of those predicted positive, the proportion that were actually positive," so its denominator is on the prediction side. Recall is "of those actually positive, the proportion that were caught," so its denominator is on the actual side. Swapping the two is the most common mistake, so it helps to recall the difference as prediction versus actual. Taking the total data count as the denominator is accuracy, and a metric that takes only true negatives or only false negatives as its denominator is neither of these two. When you want to reduce false detections, weight precision more; when you want to reduce missed detections, weight recall more.
Q23 | Computing the F-score
If precision is 0.6 and recall is 0.9, what is the F-score?
1.50
0.54
0.75
0.72
AnswerD. 0.72
The F-score is the harmonic mean of precision and recall, so multiplying 2 by 0.54, the product of 0.6 and 0.9, gives 1.08, and dividing that by 1.5, the sum of 0.6 and 0.9, gives 0.72. 0.75 is the value obtained by computing an arithmetic mean instead of a harmonic mean; 0.54 is simply the product of the two; and 1.50 is simply their sum. Worth confirming too is that a harmonic mean is pulled toward the smaller value, so it comes out smaller than the arithmetic mean of 0.75.
Q24 | Imbalanced data
You built a model that predicts everything as negative, for data where only 1% of the whole is positive. Which correctly describes this situation?
Accuracy comes out to 50%, showing that its performance is the same as guessing at random
Accuracy is high at 99%, but recall is 0%, meaning not a single positive case is caught
Accuracy is as low as 1%, and recall is also 0%, so the problem shows up in either metric
Accuracy is high at 99%, and recall is also high at 99%, so it can be considered fit for practical use
AnswerB. Accuracy is high at 99%, but recall is 0%, meaning not a single positive case is caught
Since only 1% is positive, simply answering negative for everything still gets 99% right, so accuracy is 99%. However, because not a single actual positive is caught, recall is 0%, making this model useless. Looking only at accuracy when class counts are imbalanced leads to a misjudgment, so recall, F-score, and the ROC curve and AUC also need to be checked together. In this setting, accuracy never comes out to 1% or 50%.
Q25 | The axes of the ROC curve
Which correctly pairs the vertical and horizontal axes of the ROC curve?
Vertical axis is the size of the error, and the horizontal axis is the number of times training was run
Vertical axis is accuracy, and the horizontal axis is the number of data records used for training
Vertical axis is the true positive rate, horizontal axis is the false positive rate, and the area beneath it is the AUC
Vertical axis is precision, horizontal axis is recall, and the area beneath it is the AUC
AnswerC. Vertical axis is the true positive rate, horizontal axis is the false positive rate, and the area beneath it is the AUC
The ROC curve is drawn by varying the decision threshold and plotting the true positive rate on the vertical axis against the false positive rate on the horizontal axis, and the area beneath it is the AUC. The closer the AUC is to 1, the better, while 0.5 means a level equivalent to random judging. Putting precision on the vertical axis and recall on the horizontal axis produces a different curve altogether, which is where the classic mix-up occurs. Accuracy paired with data count, or error paired with iteration count, are different graphs for looking at how training is progressing.
Q26 | AIC and BIC
How are the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC) used?
Look at the balance between fit and number of parameters, and choose the model with the larger value
Look only at the number of parameters, and choose the model with the most parameters
Look only at the fit to the training data, and choose the model with the smallest error
Look at the balance between fit and number of parameters, and choose the model with the smaller value
AnswerD. Look at the balance between fit and number of parameters, and choose the model with the smaller value
Both AIC and BIC are metrics that add together how well a model fits the data and a penalty for having more parameters, and a smaller value is considered a better model. Treating a larger value as better is wrong, and this is the point most often targeted. Choosing based only on fit to training data would end up always selecting models with more parameters, tending toward overfitting, which is why a penalty is included to balance it out. This is aligned with the idea of Occam's razor: do not make something more complex than necessary.
Q27 | The count in 5-fold
When 200 records of data undergo 5-fold cross-validation, what is the correct pairing of the number of validation records per round and the number of times training and evaluation are repeated?
160 for validation, with training and evaluation repeated 5 times
40 for validation, with training and evaluation repeated 5 times
40 for validation, with training and evaluation performed only once
100 for validation, with training and evaluation repeated 2 times
AnswerB. 40 for validation, with training and evaluation repeated 5 times
K-fold cross-validation splits the data into k parts, and repeats the procedure of using one part for validation and the rest for training, k times, averaging the results. Splitting 200 records into 5 gives 40 records per block, so validation uses 40 records each time, training uses the remaining 160, and this is repeated 5 times. Splitting only once and finishing is holdout validation, which is computationally lighter but more sensitive to how the split happens to fall. Using 160 records for validation confuses training and validation.
Q28 | Computing RMSE
For 4 predictions, the errors — actual value minus predicted value — were, in order, 2, 4, -4, and 0. What is the root mean squared error (RMSE)?
6.0
9.0
2.5
3.0
AnswerD. 3.0
The squared errors are 4, 16, 16, and 0, summing to 36. Dividing this by the count of 4 gives 9, the mean squared error (MSE), and its square root, 3.0, is the RMSE. 9.0 is the answer left as MSE without taking the square root; 2.5 is the mean absolute error (MAE), the average of the absolute errors 2, 4, 4, and 0; and 6.0 is obtained by forgetting to divide by the count and taking the square root of the sum of squares, 36, directly. RMSE is easy to read because it shares the same unit as the original error.
Q29 | Mean and median
For 7 data records — 2, 4, 4, 6, 8, 10, and 50 — what is the correct combination of mean, median, and mode?
Mean 12, median 4, mode 6
Mean 12, median 6, mode 4
Mean 6, median 12, mode 4
Mean 6, median 4, mode 10
AnswerB. Mean 12, median 6, mode 4
The sum, 84, divided by the count, 7, gives a mean of 12. The 4th value when sorted in ascending order is 6, so the median is 6, and 4, which appears twice, is the mode. The key point is that only the mean gets pulled substantially by the outlier 50, while the median and mode barely move — an example illustrating why you should not judge based on a single summary statistic alone. Swapping mean and median, or confusing median and mode, are common mistakes.
Q30 | Spurious correlation
Which correctly describes spurious correlation?
That if the correlation coefficient is 0, it can be concluded with certainty that there is no relationship between the two
That a rule requiring outliers to always be removed when computing a correlation coefficient
That correlation appears between two things with no causal relationship, because of a common underlying factor
That if the correlation coefficient is positive, one must necessarily be the cause of the other
AnswerC. That correlation appears between two things with no causal relationship, because of a common underlying factor
Spurious correlation is the phenomenon in which correlation appears between two things that have no direct causal relationship, because of a common underlying factor — as when ice cream sales and the number of drowning accidents move together because of a shared factor, temperature. The lesson here is that correlation does not necessarily imply causation, so a positive correlation coefficient does not mean one thing is the cause of the other. Also, even a correlation coefficient of 0 can hide a non-linear relationship, so it cannot be concluded with certainty that no relationship exists. How outliers are handled is a separate issue from this.
Practice: answer the questions on this page
This practice tool asks questions in random order (it works when JavaScript is enabled). You can still read all the questions and explanations above without it.
* The explanations are information for study purposes. Exam scope and systems change from year to year, so always check the official announcements of the organization that administers the exam.
This page is a translation of the Japanese original. If the translation and the original differ, the Japanese version takes precedence. View the Japanese original