the evidence

Fifty years of arguing about how people pick up languages

Immersion is the most-studied idea in second language acquisition, which means it also has one of the messiest literatures. Here are the core hypotheses, the people who objected to them, and an honest read on what's settled.


01 — the foundation

Krashen's five hypotheses

Stephen Krashen published these across the late 1970s and bundled them into Principles and Practice in Second Language Acquisition in 1982. Nearly every immersion method you'll encounter is either built on this model or built in reaction to it, so it's worth knowing what it actually claims.

HYPOTHESIS 1

Acquisition and learning are different things

Adults build competence two separate ways. Acquisition is subconscious, the way children get their first language: you end up with a feel for what sounds right without being able to say why. Learning is conscious rule knowledge, the stuff you can talk about. Krashen's strong claim is that the two never merge. Knowing a rule doesn't turn into having the rule.

HYPOTHESIS 2

Grammar arrives in a fixed order

Learners acquire structures in a predictable sequence that barely moves between individuals, whatever their first language. English -ing and plurals land early; third-person -s lands late, roughly a year later. The teaching implication is uncomfortable for textbooks: you acquire a structure when you're ready for it, not when the syllabus schedules it.

HYPOTHESIS 3

Conscious rules only work as an editor

Learned grammar gets one job: editing output after the acquired system has already produced it. And the editor only runs when three conditions all hold at once. You have time. You're thinking about form rather than meaning. You actually know the rule. Normal conversation satisfies almost none of these, which is why people who can diagram a sentence still can't speak one.

HYPOTHESIS 4

The input hypothesis

The one that matters. We acquire in exactly one way: by understanding messages that contain structure slightly beyond our current level, written as i+1. Comprehension comes first, carried by context and general knowledge, and the structure gets absorbed as a by-product. Speaking can't be taught directly under this model. It emerges after enough input has accumulated, which is why beginners go quiet for months and that's normal rather than broken.

HYPOTHESIS 5

The affective filter

Motivation, confidence and anxiety don't change the mechanism. They gate the supply line. A stressed or bored learner can sit in a room full of comprehensible input and acquire almost nothing, because the input never reaches the part of the brain that does the acquiring. This is the research case for learning through shows you actually like. Interest isn't a nice extra, it's what drops the filter.

iWHERE YOU ARE
what you already understand without effort
+
1HALF A STEP ON
new structure, decoded through context
=
i+1COMPREHENSIBLE INPUT
understood messages become acquired competence
!

Where Krashen gets argued with. The acquisition/learning split has weak experimental support, i+1 was never defined precisely enough to test, and his claim that error correction is useless drew decades of pushback. But his harshest critics kept the load-bearing finding: comprehension of large amounts of meaningful input is necessary for acquisition. Everything below modifies that claim. Nothing removes it.


02 — the objection that stuck

Swain: comprehension alone leaves holes

Merrill Swain's challenge came from inside the immersion world. She studied Canadian French immersion students who had years of massive comprehensible input behind them, and found they understood French at near-native levels while their speaking and writing kept the same grammatical gaps year after year. Input turned out to be necessary but not sufficient. Her 1985 Output Hypothesis lists three things producing language does that comprehension never forces.

Noticing the gap

The moment you try to say something, you find out what you can't say. Swain's phrasing: producing the language "may be the trigger that forces the learner to pay attention to the means of expression needed in order to successfully convey their own intended meaning." You can understand past tense for years while cheerfully saying "yesterday I go", because comprehension never made you need the form.

Testing hypotheses

Speech is a trial run. You verbalise a guess about how the language works and the world pushes back, and wrong guesses get revised. Izumi's 2002 study found that producing output containing a target form correlated strongly with noticing it, while merely receiving modified input barely moved the needle.

Talking about the language

Discussing the language itself, with a tutor or a partner or out loud alone, turns it into an object you can think about. Lantolf found learning in 80% of the metalinguistic episodes he tracked during meaning-focused tasks. The discussion happens between two people first and inside your head later.

The mechanism underneath all three is a shift from semantic to syntactic processing. Comprehension lets you cheat: keywords, context and anticipation get you the gist without parsing anything properly. Output removes the cheat. That's why the immersion students plateaued, and it's why every serious method adds speaking practice a few months in, including the ones that started life as pure-input cults.

"It is possible to comprehend input, to get the message, without a syntactic analysis of that input."Merrill Swain, 1985

03 — the middle position

Long: conversation is where input gets fixed

Michael Long's interaction hypothesis (1981, revised 1996) argues the active ingredient isn't raw input but negotiated input. When communication breaks down, speakers repair it: they simplify, repeat, rephrase, slow down. Those repairs convert input that was too hard into input that's exactly at your edge, in real time, tuned to you specifically. No textbook can do that.

What the negotiation does

Long's revised version folded in attention: negotiation doesn't just deliver input, it highlights form while meaning is still on the table. He called this focus on form, a brief glance at structure inside real communication, as opposed to decontextualised grammar lessons. Rod Ellis's 1991 critique accepted the core claim (input is necessary, interaction makes it comprehensible) while demanding a tighter account of how modified input becomes acquisition, which is where noticing, comparison and integration came in.

Why it matters at a desk

This is the argument for tutors, exchange partners and voice-chat practice. A thirty minute lesson generates more negotiation of meaning per minute than three hours of podcasts, because every misunderstanding gets repaired on the spot. It also explains why background listening is weaker per minute: nobody is negotiating anything with you. Passive input still builds the ear, but the repair loops live in conversation.


04 — the mechanism

Ellis: your brain is running statistics on everything you hear

Nick Ellis's usage-based account explains why massive input works at the cognitive level. Learners unconsciously estimate the language from the sample they've heard, the way you'd estimate a population from a survey. Every construction you know is a form-meaning pattern extracted from the tokens you've encountered, and the extraction is driven by a handful of variables.

Read that way, immersion stops being a philosophy and becomes sampling theory: acquisition quality scales with the size, variety and attention of your input sample. One podcast and one genre gives you a biased estimate of the language. Different speakers, registers and contexts is what builds the feel for correctness Krashen talks about. Ellis also leaves room for explicit instruction, but in a specific role: briefly pointing out low-salience features so learners notice them in the flood. That's exactly what a good tutor and a well-built flashcard deck do.


05 — the numbers

Nation: what the reading actually costs

Paul Nation put arithmetic on the input side. Vocabulary learning depends on how many times you meet a word and how much attention you pay each time, and when the two are compared directly, quality of attention wins. His findings turn "read more" into a budget.

The 98% rule

Texts where about 2% of running words are unknown are the sweet spot for extensive reading: enough known context for guessing to work, enough new words to keep learning. Below that coverage, reading stops being reading and becomes decoding, which is slow, miserable and low-yield. This is the empirical version of comprehensible input.

Twelve meetings per word

To learn roughly a thousand word families a year, Nation's models assume each word gets met about twelve times across varied contexts. Meeting a word once in a lesson is not learning. Meeting it a dozen times across different stories is.

Vocabulary levelWords to read per year (12 meetings each)Weekly reading at 150 wpmPer day, five days a week
2nd 1,000 word family200,00033 min7 min
3rd 1,000300,00050 min10 min
4th 1,000500,0001 h 23 min17 min
5th 1,0001,000,0002 h 47 min33 min
6th 1,0001,500,0004 h 10 min50 min
7th 1,0002,000,0005 h 33 min1 h 07 min
8th 1,0002,500,0006 h 57 min1 h 23 min

From Nation (2015), based on his 2014 volume estimates. Note the shape of the curve: the first few thousand words come cheap, then the bill arrives. That's the intermediate plateau, in tabular form.

Nation's other findings are all free technique. Guessing from context works, but confirming the guess with a dictionary multiplies retention, and electronic lookup is now fast enough that it doesn't break reading flow. Narrow reading, staying with one topic or author, cuts the number of distinct new words roughly in half because the same vocabulary keeps recurring. Re-reading a book within a few weeks adds retrieval practice for nothing. And his Four Strands framework says a balanced course splits time roughly evenly between meaning-focused input, meaning-focused output, deliberate language study, and fluency work on easy material. Extensive reading should be about three sixteenths of total study time, two thirds of it on just-hard-enough texts and a third on deliberately easy ones.


06 — the retention half

Spacing: the most reliable effect in learning science

Immersion supplies the meetings. Spaced repetition makes sure they land at useful times. Cepeda and colleagues' 2006 meta-analysis covered 839 assessments across 317 experiments in 184 articles, and two of its findings matter here.

Spread beats cram, reliably

The same total study time produces much better retention when the episodes are spread out. The effect has been replicating since Ebbinghaus in the 1880s. It is about as settled as psychology gets.

The right gap depends on how long you want to keep it

The optimal interval between study episodes grows with the retention interval you're aiming for. Cramming the night before a test wins the test and loses the year. A word you want to keep for life wants reviews spaced over weeks and months, which is what an SRS scheduler automates.

This is the licence for sentence mining. Pull sentences out of what you watched, let the scheduler space the reviews, and the words you met in last night's episode get their twelve Nation-style meetings over the following months instead of evaporating by Friday.


07 — the uncomfortable question

Age, finally measured instead of guessed

"Am I too old?" had no good answer until 2018, when Hartshorne, Tenenbaum and Pinker published the largest language-learning dataset anyone has assembled: 669,498 people who took a viral grammar quiz and reported their age, when they started English, and how they learned it.

What that means for an adult doing this at home: native-like grammar is statistically out of reach, and it was never a sensible target. What the study actually says is that your brain keeps learning grammar well into adulthood, and that of all the variables in the dataset, the one an adult in a suburb can move is immersive exposure time. That's the one the whole method is built on.


08 — the natural experiment

Canada ran this at population scale for thirty years

Canadian French immersion began in 1965: English-speaking children take at least half their entire school curriculum through French, from maths to science. Jim Cummins's 1998 review of that research is the definitive summary, and it reads like a controlled trial of everything on this page.

What worked

Students gained fluency and literacy in French at no cost to their English. After a brief catch-up period, mostly spelling, resolved by grade five, immersion kids matched or beat English-stream peers on English academic measures. No long-term lag in subject matter taught through French. By grade six, French comprehension and reading sat close to native norms. Bilingualism brought measurable cognitive and metalinguistic advantages.

What didn't

Speaking and writing accuracy lagged native norms significantly even after years of immersion, grammar worst of all (Harley, Allen, Cummins and Swain, 1991). The diagnosis: classrooms were transmission-oriented. The teacher talked in French, students mostly listened, and there was little creative production, little authentic French children's literature, and no francophone peers to talk to because they attended a separate school system. Attrition was brutal too: Alberta data showed 43 to 68 percent of students gone by grade six and up to 97 percent by grade twelve.

Cummins's conclusion is the bridge between Krashen and Swain. Comprehension-based immersion is extremely good at building comprehension. Building production requires deliberately manufacturing reasons to produce: project work, creative writing, genuine two-way communication. Sixty years of national-scale data, and the lesson for someone learning at a desk is the same one. Input gets you understanding. Only pushed output gets you accuracy.


09 — where everyone lands

The parts nobody argues about anymore

Strip out the turf wars and the field converges. Krashen's input, Swain's output, Long's interaction, Ellis's frequency statistics and Nation's four strands are different projections of one object.

PrincipleWho says itWhat it looks like at home
Understanding messages drives acquisitionKrashen, input hypothesisWatch, listen and read things you mostly understand and actually enjoy
Volume and variety of exposure is the mechanismEllis; Hartshorne et al.Hours per day beat hours per week; mix speakers, genres and registers
Production builds accuracy that input can'tSwain; the Canadian dataTutors, exchange partners, journaling, from about month three to six
Negotiation sharpens bothLong, interaction hypothesisPrefer two-way conversation over one-way media whenever you can get it
Deliberate review locks vocabulary inNation; Cepeda et al.Anki sentence mining, fifteen to thirty minutes a day, spaced rather than crammed
Anxiety blocks intakeKrashen, affective filterCompelling content, low-stakes practice, no shame about mistakes

The method pages are this table turned into a schedule.