N-Back Test
Say match when the square returns to where it was n steps ago.
Settings
Changing one restarts the attempt, and scores set under different settings are not comparable.
Quick start
- 1Pick your n. Two is standard, and it is what the reference distribution describes.
- 2Click Start. A square lights somewhere on the grid, about two seconds each.
- 3Press M when the square is where it was n steps back.
- 4Do nothing when it does not. Staying silent counts as an answer, and it is scored.
- 5At the end you get your accuracy, your catches and your false alarms.
Press Match when the square returns to where it was 2 back
20 squares on a three by three grid, about two seconds each. Press M or tap Match. Doing nothing is also a response, and it is scored.
How you compare
2-back accuracy (%). Higher is better. Reference distribution, not CognitiveDrill data.
The median for this test is 80 %. Take a run and your score appears on the bar.
Reference distribution, not CognitiveDrill data. It is shaped to match published results for this task and is replaced by our own norms once a cohort reaches n = 1,000. Shape based on: Kirchner WK; Jaeggi SM, Buschkuehl M, Jonides J, Perrig WJ.
Try next
About this test
Where the task comes from
Wayne Kirchner invented this task in 1958. He wanted to know how well younger and older adults kept up with information that would not sit still. His volunteers watched a row of lights and reported the one lit a few steps back. The result that made it famous was how sharply people failed as the gap grew.
The task stayed quiet for decades, then became one of the most used tools in memory research. The reason is what it asks of you. Most memory tests let you learn a list and then recall it. This one never lets you finish. Every new letter asks you a question and paints over what you were holding at the same time.
Why accuracy is scored the way it is
There are two ways to be right and two ways to be wrong. Press on a match and you are right. Stay quiet on a non-match and you are right too. Press on a non-match and you are wrong. Miss a match and you are also wrong.
Count only the matches you caught and you could score perfectly by pressing every single time. So we do not count it that way. We score every trial, presses and silences alike, and show your catches and your false alarms separately underneath.
If your accuracy looks high and your false alarms are high too, you are guessing freely and the headline number is flattering you. About one trial in four is a match, so pressing on everything scores near 25 per cent, not 50.
Single n-back, and why we did not build the dual version
The famous training studies used dual n-back. It runs a picture stream and a sound stream at once, and you answer each one separately. It is much harder, much more irritating, and much harder to read, because the two streams interfere with each other.
Single n-back on a grid is the cleaner measurement, and the version most people came here for. Nine positions, one lit at a time. There are no words in it, which is the point. A letter stream can be held in the silent voice you use to repeat a phone number, so it partly measures how well you rehearse. A position has nothing to say.
About one square in four is a real match. We also plant near misses one step either side of the target. So a position that feels almost familiar is often exactly the trap it looks like.
The training claim, stated honestly
In 2008 Jaeggi and colleagues reported that training on dual n-back raised scores on reasoning tests. That result launched an industry.
In 2013 Redick and colleagues ran a bigger study. Their comparison group did a real task rather than nothing, which is the harder and fairer test. They found no such benefit. Reviews since have been consistently unimpressive. Training makes you better at n-back, and at tasks that look like n-back. That is mostly where it stops.
We will not tell you this makes you smarter. It will push your n-back score up, quickly at first, and that is a real and satisfying thing to chase. For what working memory is, and why span tests and n-back measure different things, see how working memory works. For the storage side, take the digit span test and read what counts as a good digit span.
Questions
What is a good n-back score?
On 2-back, 80 per cent accuracy sits near the middle of the reference distribution and above 90 per cent is strong. Scores are not comparable across different values of n or different trial counts.
Is n-back the same as a memory span test?
No. A span test asks how much you can hold. N-back asks how well you can keep updating while holding. People often score very differently on the two, which is the interesting part.
Does n-back training raise IQ?
The best-controlled studies say no. Training improves n-back performance and closely related tasks, with little evidence of transfer to fluid intelligence once expectation effects are controlled.
Why do I get worse when I try harder?
Because most people cope by holding a rough picture of the last few positions, and effort tends to break the rhythm of that. A steady pace usually beats straining.
What is a lure?
A square that matches the position at n plus one or n minus one steps back rather than exactly n. It feels familiar and produces most false alarms, which is precisely why it is included.
Should I use 3-back or 4-back?
Only if 2-back has stopped being difficult. Above 3-back most people fall to near-guessing, and a score in that region carries almost no information.
Sources
- Kirchner WK (1958). Age differences in short-term retention of rapidly changing information. Journal of Experimental Psychology, 55(4), 352-358. Link
- Jaeggi SM, Buschkuehl M, Jonides J, Perrig WJ (2008). Improving fluid intelligence with training on working memory. PNAS, 105(19), 6829-6833. Link
- Redick TS, Shipstead Z, Harrison TL, et al. (2013). No evidence of intelligence improvement after working memory training: a randomized, placebo-controlled study. Journal of Experimental Psychology: General, 142(2), 359-379. Link
- Melby-Lervåg M, Hulme C (2013). Is working memory training effective? A meta-analytic review. Developmental Psychology, 49(2), 270-291. Link