CognitiveDrill

The Stroop Effect Explained

Ninety years of a very robust finding, and the claims it will not support.

The Stroop effect is the delay you feel when you name the ink colour of a word that spells a different colour. It is one of the most reliable findings in psychology. It has been repeated in hundreds of forms across nine decades. And it is routinely described as measuring things it does not measure.

This page covers what Stroop actually did and the best current explanation of why it happens. It also covers what the score is worth and how to read your own. The test itself is below, because the effect is more convincing when it happens to you.

Try it, then read on

Thirty trials. Half have the word and the ink agreeing, half have them clashing. Your score is the gap between the two in milliseconds, which is what the effect actually is.

Settings

Changing one restarts the attempt, and scores set under different settings are not comparable.

Quick start

  1. 1Click Start. A colour word appears, usually printed in a different colour.
  2. 2Answer with the colour of the ink. Never the word.
  3. 3Use the four keys shown under the canvas, or click the labelled buttons.
  4. 4Half the trials match and half clash, in random order.
  5. 5Your score is the gap between the two, in milliseconds.

Name the ink colour, not the word

30 trials, half of them in conflict. Press 1 to 4, or click the buttons. Speed matters, but so does getting it right.

How you compare

Stroop interference (ms). Lower is faster. Reference distribution, not CognitiveDrill data.

Stroop interference (ms)Reference

The median for this test is 75 ms. Take a run and your score appears on the bar.

Reference distribution, not CognitiveDrill data. It is shaped to match published results for this task and is replaced by our own norms once a cohort reaches n = 1,000. Shape based on: Stroop JR; MacLeod CM.

Full page, settings and the detail behind the scoring: Stroop Test.

What Stroop did in 1935

John Ridley Stroop's doctoral research ran three experiments.

The first asked people to read colour words printed in clashing ink. Reading barely slowed. The second asked them to name the ink colour of those same words. That slowed a lot, compared with naming plain patches of colour. The third looked at practice.

The gap between the first two is the finding. Reading gets in the way of naming colours. Naming colours does not much get in the way of reading. Any explanation has to account for that one-way street, not just for the delay.

Why reading wins

The old explanation is habit. Adults read so much that reading runs without being asked, so the word gets processed whether you want it or not. The answer it throws up then has to be pushed down.

The better modern account is about speed. Reading is a faster and better practised route to an answer than naming a colour is. By the time the colour is ready to drive your response, the word has already put a rival answer on the table.

Cohen and colleagues modelled this in 1990 as a race between two routes of different strengths. The model reproduces the one-way street without needing reading to be entirely beyond your control.

One piece of supporting evidence is neat. The effect is small in people learning to read and grows as reading gets fluent. In a second language its size follows how good you are. It tracks practice, exactly as a race between routes would predict.

What your score measures, and what it does not

It measures the cost, in milliseconds, of settling a fight between two pieces of information pointing at different answers. That is a narrow and real thing. It moves with sleep, alcohol and sudden stress.

It does not measure attention span, concentration, willpower, mental toughness or intelligence. It is not a screening test for any condition. Stroop tasks appear in clinical assessments as one part among many, read by someone who can see the rest of your history. Those versions are usually spoken and card-based rather than typed.

There is also a measurement problem worth knowing. A score made by subtracting one number from another is always shakier than either number alone, because the noise in both carries through.

That is why a single 30-trial run can land tens of milliseconds away from your own true value. It is also why a negative score is almost always noise rather than a discovery.

The keyboard version is a different animal

The effect is biggest when you answer out loud. The thing fighting you is a word, your answer is a word, and the clash is direct. Pressing a key adds a step, because you have to turn a colour into a key position. Typed versions give smaller and shakier numbers than spoken ones.

Browsers cannot time your voice reliably, so every online Stroop test, ours included, is a typed version. Compare your score only against other typed runs, ideally your own. The Stroop test reports your matching and clashing times separately alongside the gap, so you can see what drove it.

The variants, and what each one was built to isolate

Because the effect is so reliable, the task became a workbench.

The reverse version asks you to read the word and ignore the ink. It produces almost no delay. That is the one-way street again, and it is what any explanation has to survive.

The number version sets a digit against its own size, such as a large 2 beside a small 8. It shows the conflict is not about words at all. Two features of the same object are simply pulling towards different answers.

The emotional version swaps colour words for charged ones, and measures the delay caused by words tied to whatever is preying on a person's mind. It looks like the classic task but behaves quite differently. The delay often turns up on the trial after the charged word rather than on the word itself. People have argued about what it means for decades.

Spatial versions create conflict with no reading at all. In the Simon task a signal appears on the left but calls for a right-hand answer. In the flanker task a target arrow sits among arrows pointing the other way.

Why there is no single focus score

Those spatial versions matter here for one reason. They show that conflict costs are not one ability.

Someone with a large Stroop effect will not reliably show a large flanker effect. These tasks track each other only loosely. That is the main reason a single number labelled attention or focus is worth so little.

The same holds across this site. Holding back a habit is a different job, measured by the go/no-go test in errors rather than milliseconds. Swapping between two rules is different again, measured by the trail making test as the gap between its two parts.

People who look excellent on one of these are often ordinary on another. That is a finding, not a flaw. For a fourth angle, average reaction time by age shows how much any reaction-based measure moves for reasons that have nothing to do with control.

Related

Questions

What does the Stroop test measure?

The time cost of overriding an automatic reading response in order to report the ink colour instead. That cost is the interference effect, measured in milliseconds.

Why does the Stroop effect happen?

Reading is faster and better practised than naming a colour. The word puts a rival answer on the table before the colour can, and settling that fight takes time.

Is a small Stroop effect good?

Not necessarily. Responding cautiously slows congruent trials and shrinks the difference, so a small effect can reflect strategy rather than superior control.

Can the Stroop test detect ADHD or dementia?

No. It appears in clinical batteries as one component interpreted alongside a full assessment. On its own, and in a browser, it detects nothing.

Does the effect disappear with practice?

It shrinks with practice on a specific set of colours and mappings, but it is remarkably persistent and does not vanish.

Does it work in a second language?

Yes, and its size tracks reading fluency in that language, which is good evidence that the effect depends on how strongly practised reading is.

Sources

  • Stroop JR (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643-662. Link
  • MacLeod CM (1991). Half a century of research on the Stroop effect: an integrative review. Psychological Bulletin, 109(2), 163-203. Link