CognitiveDrill

Cognitive Reflection Test

Three short puzzles. Each has an answer that arrives before you have finished reading, and it is wrong.

Settings

Changing one restarts the attempt, and scores set under different settings are not comparable.

Quick start

  1. 1Press Play now. One puzzle appears, with a box under it.
  2. 2Type your answer as a whole number and press Answer.
  3. 3Say whether you had met that puzzle before. It changes nothing about your score.
  4. 4Read the working, then press Next puzzle.
  5. 5Answer all three. Your score is how many you got right.

Loading test

How you compare

Puzzles you got right (of 3). Higher is better. Reference distribution, not CognitiveDrill data.

Puzzles you got right (of 3)Reference

The median for this test is 1 of 3. Take a run and your score appears on the bar.

Reference distribution, not CognitiveDrill data. It is shaped to match published results for this task and is replaced by our own norms once a cohort reaches n = 1,000. Shape based on: Stieger S, Reips UD; Frederick S.

Try next

About this test

Why the wrong answer gets there first

None of these three puzzles is hard. Each one is a single sum, and every one of them can be done in your head.

What they have in common is that a number is available before the sum is. A bat and a ball cost £1.10, and the bat costs a pound more. Ten pence turns up on its own, before anybody has worked anything out.

Shane Frederick published the three in 2005 and called what they measure reflection. Not arithmetic, and not care. The specific thing is noticing that your first answer is a first answer, and looking at it again before you say it.

That is why the people who get these right are not simply the people who are quicker at sums. The sum was never the obstacle.

Three puzzles, and why there is not a fourth

Your score can only be 0, 1, 2 or 3. That is coarse, and it is the price of the one thing this page has that its neighbours do not.

These three items have been put to tens of thousands of people since 2005, in the same words. The spread of the score is published, so a number from this page can be placed against something real rather than against a guess.

Add a fourth question of our own and that comparison is gone. Whatever the new set measured, nobody would have measured it before. The page would be back where the rest of this hub is, with a number and nothing under it. A score that can be read against published figures is worth a great deal. The same argument sits behind what counts as a good digit span.

One run is still one run. Three items is a short measurement, and a second run cannot tell you much, because by then you know the answers.

Nearly half the people who take this have met it before

This is the largest problem the test has, and it is bigger than anything about the puzzles themselves.

Stieger and Reips put the three items to 2,137 people online and asked a plain question afterwards: had you done this before? Almost half said yes, and the ones who said yes scored 1.65 out of three against 1.21 for everybody else.

So this page asks the same question, after every puzzle, and it asks before it marks you. Asking afterwards would collect nothing useful, because everybody remembers a puzzle once they have been shown the answer.

The count sits on the result card next to your score, and the comparison underneath is built from the people who said no.

If you already knew the answers, the settings panel has the same three puzzles with the numbers moved. Those are ours rather than a published form. Thomson and Oppenheimer built a proper second version in 2016, for exactly this reason, and this is not it.

The tempting answer is counted apart from the other wrong ones

Each puzzle has one wrong answer it was built to produce. Ten pence, a hundred minutes, twenty four days.

Any other wrong answer is a different event. Somebody who types six pence was working the problem and slipped. Somebody who types ten pence did not get as far as working it.

So the card reports them separately. Two right with one tempting answer is a different run from two right with none, and the card shows which you had.

The tag belongs to the number rather than to the person who typed it. Nothing on this page knows what you were thinking, and it does not guess.

  • Bat and ball. The answer is 5 pence. Ten pence is the pull.
  • Machines and widgets. The answer is 5 minutes. A hundred is the pull.
  • Lily pads. The answer is 47 days. Twenty four is the pull.

Where the curve under your score comes from

The reference distribution is the naive half of the Stieger and Reips sample, which is 1,174 people who had not met the test.

Their scores ran 33 in 100 with none right, 26 with one, 26 with two and 14 with all three. That is a median of one, and it is why a single correct answer is close to the middle rather than close to the bottom.

Each score is placed at the middle of its own band rather than the top of it. Two correct beats about 73 in 100, rather than the 86 in 100 who scored two or fewer. The other people who scored two are not people you beat.

It is somebody else's sample and not ours, and the page says so wherever the bar appears. It was recruited online, it was German speaking, and the education mix of one study is never the education mix of the internet.

Our own norm goes up when a cohort here reaches a thousand runs, and not one day before that.

The claim this page will not make

Frederick reported that scores on these three items lined up with how his participants answered questions about risk and about waiting for money. That is a finding about a sample. It is not a statement about you, and nothing on this page has been checked against anything.

It is not an intelligence test. It is three questions, and no score here has been compared with any measure of intelligence.

It screens for nothing, it rules out nothing, and a score of zero is not a diagnosis. It is three puzzles on a Tuesday.

Getting better at these three is getting better at these three. That is what researchers mean by near transfer, and whether brain training games are worth it covers what happened when the wider claims were tested.

If something has genuinely changed in how you think, that belongs with a doctor rather than with a web page.

Questions

What is the answer to the bat and ball question?

Five pence. The bat is £1.05, which is a pound more than the ball, and the two together come to £1.10. Ten pence is the answer that arrives first, and it makes the pair £1.20.

What is a good score on the cognitive reflection test?

Two out of three sits above about 73 in 100 people who had not seen the test before. One is roughly the middle. All three is around the top one in seven.

Why is the test only three questions long?

Because those are the three that were published, and their spread is the only reason this page can rank a score at all. A fourth question of ours would throw the comparison away.

Does knowing the answers already ruin my score?

It inflates it. People who had met the test scored about a third of a point higher. The settings panel offers the same three puzzles with the numbers moved if you already know them.

Is the cognitive reflection test an IQ test?

No. It is three word puzzles with a published score distribution. No score on this page has been checked against any measure of intelligence.

Who wrote the cognitive reflection test?

Shane Frederick, in a 2005 paper in the Journal of Economic Perspectives. The three items are quoted here in his order, with the money in pounds and pence.

What does the tempting answers figure mean?

It counts the puzzles where your answer was the specific wrong one that puzzle was built to produce. Any other wrong number is not counted there.

Is there a time limit?

No, and there is no clock anywhere on the canvas. The card does report the longest you sat with a single puzzle, because that is the part of the run worth knowing about.

Sources

  • Frederick S (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25-42. Link
  • Stieger S, Reips UD (2016). A limitation of the Cognitive Reflection Test: familiarity. PeerJ, 4, e2395. Link
  • Thomson KS, Oppenheimer DM (2016). Investigating an alternate form of the cognitive reflection test. Judgment and Decision Making, 11(1), 99-113. Link