Karinoya Learning Room

High School · High School Math: Regular Test Lab

Data analysis (Mathematics I)

Read the questions and explanations in English. The lectures (explanatory articles) are available in Japanese only.

View the Japanese version (with lectures) →

Q1 | Mean

Which is the mean of the data 2, 4, 6, 8, 10?

  1. 6
  2. 7
  3. 8
  4. 5
AnswerA. 6

The mean is the total divided by the count: (2+4+6+8+10)/5=30/5=6. For equally spaced data the middle value coincides with the mean. Do not forget to divide by the count, 5.

Q2 | Median with an odd count

Which is the median of the data 2, 9, 4, 7, 3?

  1. 5
  2. 3
  3. 4
  4. 7
AnswerC. 4

In increasing order the data are 2, 3, 4, 7, 9. With 5 values, an odd number, the median is the third one, namely 4. The mean is 25/5=5, and in general the median and the mean do not agree. Do not forget to sort first.

Q3 | Median with an even count

Which is the median of the data 2, 5, 3, 8, 9, 7?

  1. 7
  2. 5
  3. 5.5
  4. 6
AnswerD. 6

In increasing order the data are 2, 3, 5, 7, 8, 9. With 6 values, an even number, the median is the mean of the middle two, the third and fourth: (5+7)/2=6. Taking only one of the middle values and answering 5 or 7 is wrong.

Q4 | Mode

Which is the mode of the data 5, 7, 7, 8, 9, 9, 9, 10?

  1. 8
  2. 9
  3. 8.5
  4. 7
AnswerB. 9

The mode is the value that occurs most often. The value 9 occurs 3 times and every other value at most twice, so the mode is 9. The value 8 is the mean and 8.5 is the median; these three measures of centre are different things.

Q5 | Range

Which is the range of the data 3, 8, 2, 10, 7?

  1. 7
  2. 10
  3. 5
  4. 8
AnswerD. 8

The range is maximum minus minimum: 10−2=8. Do not answer with the maximum 10 itself, and do not confuse the range with the interquartile range. A single outlier can change the range a great deal.

Q6 | Interquartile range

Which is the interquartile range of the data 1, 2, 3, 4, 5, 6, 7?

  1. 6
  2. 2
  3. 3
  4. 4
AnswerD. 4

The median is 4. The lower half 1, 2, 3 has median Q1=2, and the upper half 5, 6, 7 has median Q3=6. The interquartile range is Q3−Q1=6−2=4. Do not confuse it with the range, which is 7−1=6.

Q7 | Third quartile

Which is the third quartile Q3 of the data 1, 3, 4, 5, 6, 8, 9, 10, listed in increasing order with 8 values?

  1. 5.5
  2. 9
  3. 8
  4. 8.5
AnswerD. 8.5

With 8 values the lower half is 1, 3, 4, 5 and the upper half is 6, 8, 9, 10. Q3 is the median of the upper half, (8+9)/2=8.5. The value 5.5 is the overall median Q2. Do not forget to average the middle two values of the upper half.

Q8 | Box plot

What does the length of the box in a box plot represent?

  1. the standard deviation
  2. the deviation from the mean
  3. the range
  4. the interquartile range
AnswerD. the interquartile range

The left end of the box is Q1 and the right end is Q3, so the length of the box is Q3−Q1, the interquartile range. The range runs from the end of one whisker to the end of the other, that is maximum minus minimum. The standard deviation cannot be read off a box plot.

Q9 | Outliers

For a set of data the first quartile is 10 and the third quartile is 16. Using the rule that any value greater than Q3+1.5×IQR is an outlier, which value is an outlier?

  1. 24
  2. 20
  3. 12
  4. 26
AnswerD. 26

IQR=16−10=6, so the boundary is 16+1.5×6=16+9=25. The value 26 exceeds it and is therefore an outlier. The values 24 and 20 are at most the boundary 25, so they are not outliers. Remember to add 1.5×IQR to Q3.

Q10 | Computing a mean

Which is the mean of the data 98, 102, 101, 99, 100?

  1. 100.5
  2. 100
  3. 99
  4. 101
AnswerB. 100

The total is 98+102+101+99+100=500, and 500/5=100. You can also see it from the differences from the working mean 100, namely −2, +2, +1, −1, 0, whose average is 0, so the mean is 100. Since the total is 500, the mean cannot be 100.5.

Q11 | Corrected mean

Five data values had mean 10, but one of them, recorded as 8, was wrong and should have been 13. Which is the correct mean?

  1. 11.5
  2. 12
  3. 10.5
  4. 11
AnswerD. 11

The original total was 10×5=50. Correcting the error makes the total 50−8+13=55, so the mean is 55/5=11. You can also divide the increase of 5 by the count 5 and see that the mean rises by 1.

Q12 | Definition of variance

Which is the variance of the data 1, 2, 3, 4, 5?

  1. 10
  2. 2
  3. 4
  4. √2
AnswerB. 2

The mean is 3. The deviations are −2, −1, 0, 1, 2, so the variance is (4+1+0+1+4)/5=10/5=2. Do not forget to divide the sum of squared deviations, 10, by the count. The value √2 is the standard deviation.

Q13 | Standard deviation

Which is the standard deviation of the data 2, 4, 6, 8?

  1. 5
  2. √5
  3. √10
  4. 25
AnswerB. √5

The mean is 5. The variance is (9+1+1+9)/4=20/4=5, and the standard deviation is its positive square root, √5. Do not report the variance 5 as the standard deviation.

Q14 | Computational formula

For the data 1, 3, 5, 7, which is the variance obtained using variance = mean of the squares − square of the mean?

  1. √5
  2. 16
  3. 5
  4. 21
AnswerC. 5

The mean is 4 and the mean of the squares is (1+9+25+49)/4=84/4=21. The variance is 21−4²=21−16=5. The value 21 stops at the mean of the squares and 16 at the square of the mean, both forgetting the subtraction. The defining formula also gives 5.

Q15 | From variance to standard deviation

If a set of data has variance 9, which is its standard deviation?

  1. 9
  2. 81
  3. √3
  4. 3
AnswerD. 3

The standard deviation is √(variance)=√9=3. The value 81 squares the variance again, and √3 takes one square root too many. Keep the relation exact: the standard deviation is the positive square root of the variance.

Q16 | Linear change and standard deviation

If the variable x has standard deviation 4, which is the standard deviation of the variable y defined by y=2x+3?

  1. 8
  2. 16
  3. 11
  4. 4
AnswerA. 8

For y=bx+a the standard deviation is multiplied by |b|, so it is 2×4=8. The +3 merely shifts everything and does not change the spread. The value 11 adds the constant as well, computing 2×4+3, and 16 confuses this with the factor b²=4 for the variance.

Q17 | Linear change and variance

If the variable x has variance 2, which is the variance of the variable y defined by y=3x−2?

  1. 18
  2. still 2
  3. 6
  4. 4
AnswerA. 18

For y=bx+a the variance is multiplied by b², so it is 3²×2=18. The value 6 multiplies only by 3, which is the factor for the standard deviation. The −2 has no effect on the variance. Keep the two apart: the standard deviation scales by |b| and the variance by b².

Q18 | Combined mean

Group A has 4 members with mean score 5, and group B has 6 members with mean score 10. Which is the mean score of all 10 people?

  1. 7.5
  2. 8.5
  3. 7
  4. 8
AnswerD. 8

The overall total is 4×5+6×10=20+60=80 points, so the mean is 80/10=8 points. The typical error is to ignore the group sizes and compute (5+10)/2=7.5. Weight by the sizes and work from the total.

Q19 | A negative coefficient

If the variable x has standard deviation 2, which is the standard deviation of the variable y defined by y=−2x+1?

  1. 4
  2. −4
  3. 8
  4. 2
AnswerA. 4

The standard deviation is multiplied by |b|, so it is |−2|×2=4. A standard deviation measures spread and can never be negative, so −4 is wrong. Changing the sign leaves the width of the spread the same. The value 8 confuses this with the factor b²=4 for the variance.

Q20 | Adding the same amount to everyone

If 5 points are added to everyone's test score, what happens to the mean and to the variance?

  1. the mean is unchanged and the variance increases by 5
  2. both the mean and the variance increase by 5
  3. the mean increases by 5 and the variance is unchanged
  4. neither the mean nor the variance changes
AnswerC. the mean increases by 5 and the variance is unchanged

Adding the same value to every item raises the mean by that amount, but each deviation from the mean is unchanged, so the variance and standard deviation stay the same. It helps to picture the whole shape of the spread simply sliding along.

Q21 | Sum of squares

Five data values have mean 4 and variance 6. Which is the sum of their squares, x₁²+x₂²+…+x₅²?

  1. 30
  2. 50
  3. 80
  4. 110
AnswerD. 110

From variance = mean of the squares − square of the mean, the mean of the squares is 6+4²=22. Hence the sum of the squares is 22×5=110. The value 50 comes from computing (6+4)×5. Use the formula the other way round: mean of the squares = variance + square of the mean.

Q22 | Positive correlation

If the points on a scatter plot slope upward to the right overall, what can be said to exist between the two variables?

  1. a causal relationship
  2. no correlation
  3. negative correlation
  4. positive correlation
AnswerD. positive correlation

Sloping upward to the right, meaning y tends to increase as x increases, is positive correlation, while sloping downward to the right is negative correlation. Note that the shape of a scatter plot alone tells you nothing about causation.

Q23 | Range of r

Which range must the correlation coefficient r always lie in?

  1. −1 < r < 1
  2. −1 ≦ r ≦ 1
  3. r ≧ 0
  4. 0 ≦ r ≦ 1
AnswerB. −1 ≦ r ≦ 1

The correlation coefficient is the covariance divided by the product of the standard deviations, and it always lies in −1≦r≦1. The endpoints ±1 do occur, when all the points lie on a single straight line, so the equalities are included. Negative values also occur.

Q24 | r=−0.9

Two variables have correlation coefficient −0.9. Which best describes the scatter plot?

  1. the points cluster around a line sloping downward to the right
  2. the points are spread evenly all over
  3. the points cluster around a line sloping upward to the right
  4. the points lie along a curve such as y=x²
AnswerA. the points cluster around a line sloping downward to the right

Since r is close to −1 the correlation is strongly negative, so the points gather near a line sloping downward to the right. If it sloped upward, r would be positive. Because r measures a linear relationship, a parabolic arrangement would not give a value such as −0.9.

Q25 | Points on a line

All the points of a scatter plot lie on a single line sloping downward to the right. Which is the correlation coefficient?

  1. 0
  2. 1
  3. −1
  4. −0.5
AnswerC. −1

When all points lie on one straight line the correlation coefficient is ±1, and for a downward slope r=−1. The value r=−1 is the strongest possible negative correlation, 0 means no correlation, and 1 corresponds to points on an upward-sloping line.

Q26 | Computing r

For 5 pairs of data x: 1, 2, 3, 4, 5 and y: 3, 5, 4, 7, 6, the variance of x is 2, the variance of y is 2 and the covariance is 1.6. Which is the correlation coefficient?

  1. 1.6
  2. 0.4
  3. −0.8
  4. 0.8
AnswerD. 0.8

r=sxy/(sx·sy)=1.6/(√2×√2)=1.6/2=0.8. Dividing by the product of the variances, 2, instead of the standard deviations, √2, gives the wrong answer 0.4. Do not simply report the covariance 1.6.

Q27 | Correlation and causation

A strong positive correlation was found between ice cream sales and the number of people taken to hospital for heatstroke. Which is the most appropriate interpretation?

  1. restricting ice cream sales would reduce the number of people taken to hospital
  2. the correlation is strong, so one must be the cause of the other
  3. ice cream is the cause of heatstroke
  4. a common factor, namely temperature, may be affecting both
AnswerD. a common factor, namely temperature, may be affecting both

Correlation is not evidence of causation. In this example it is natural to think that a third variable, temperature, is raising both. The strength of a correlation, the size of r, is a separate matter from whether there is causation.

Q28 | Example of negative correlation

Which pair of variables is most likely to show a negative correlation?

  1. the population of a city and the number of shops
  2. temperature and sales of hot coffee
  3. height and weight
  4. study time and test score
AnswerB. temperature and sales of hot coffee

Sales of hot coffee would be expected to fall as the temperature rises, which is a downward slope, that is negative correlation. Height and weight, study time and score, and population and number of shops all suggest positive correlation, where one rises as the other rises.

Q29 | Hypothesis testing

A coin was tossed 8 times and came up heads all 8 times. Under the hypothesis that the coin is fair, with probability 1/2 of heads, the probability of 8 heads in a row is (1/2)⁸=1/256, about 0.4%. With a threshold of 5%, which judgement follows the idea of hypothesis testing?

  1. accept the hypothesis and judge that the coin is fair
  2. since the probability fell below 50%, judge that the coin is not fair
  3. since the probability is not 0, judge that the coin is fair
  4. reject the hypothesis and judge that the coin favours heads
AnswerD. reject the hypothesis and judge that the coin favours heads

Under the assumption of fairness, the probability of the observed outcome, 1/256≒0.4%, is smaller than the threshold of 5%. Rather than believing that something very rare has happened, it is more natural to doubt the assumption, so we reject the hypothesis and judge that heads is favoured. The comparison is with the threshold of 5%, not with 50%.

Q30 | Failing to reject

In a hypothesis test, the probability of the observed outcome under the hypothesis was computed as 6%. With a threshold of 5%, which judgement is most appropriate?

  1. the hypothesis has been proved correct
  2. reject the hypothesis
  3. the hypothesis cannot be rejected, so withhold judgement
  4. the probability is close to the threshold, so change the threshold to 10%
AnswerC. the hypothesis cannot be rejected, so withhold judgement

A probability of 6% is at or above the threshold of 5%, so the outcome cannot be called rare and the hypothesis cannot be rejected. Failing to reject, however, is not a proof that the hypothesis is correct. Moving the threshold after seeing the result is a mistaken attitude.

Practice: answer the questions on this page

This practice tool asks questions in random order (it works when JavaScript is enabled). You can still read all the questions and explanations above without it.

* This is for reviewing high school study content. If notation or treatment differs from your textbook or your school's teaching, follow your textbook and your teacher's explanations.

This page is a translation of the Japanese original. If the translation and the original differ, the Japanese version takes precedence. View the Japanese original