Immersion is the most-studied idea in second language acquisition, which means it also has one of the messiest literatures. Here are the core hypotheses, the people who objected to them, and an honest read on what's settled.
Stephen Krashen published these across the late 1970s and bundled them into Principles and Practice in Second Language Acquisition in 1982. Nearly every immersion method you'll encounter is either built on this model or built in reaction to it, so it's worth knowing what it actually claims.
Adults build competence two separate ways. Acquisition is subconscious, the way children get their first language: you end up with a feel for what sounds right without being able to say why. Learning is conscious rule knowledge, the stuff you can talk about. Krashen's strong claim is that the two never merge. Knowing a rule doesn't turn into having the rule.
Learners acquire structures in a predictable sequence that barely moves between individuals, whatever their first language. English -ing and plurals land early; third-person -s lands late, roughly a year later. The teaching implication is uncomfortable for textbooks: you acquire a structure when you're ready for it, not when the syllabus schedules it.
Learned grammar gets one job: editing output after the acquired system has already produced it. And the editor only runs when three conditions all hold at once. You have time. You're thinking about form rather than meaning. You actually know the rule. Normal conversation satisfies almost none of these, which is why people who can diagram a sentence still can't speak one.
The one that matters. We acquire in exactly one way: by understanding messages that contain structure slightly beyond our current level, written as i+1. Comprehension comes first, carried by context and general knowledge, and the structure gets absorbed as a by-product. Speaking can't be taught directly under this model. It emerges after enough input has accumulated, which is why beginners go quiet for months and that's normal rather than broken.
Motivation, confidence and anxiety don't change the mechanism. They gate the supply line. A stressed or bored learner can sit in a room full of comprehensible input and acquire almost nothing, because the input never reaches the part of the brain that does the acquiring. This is the research case for learning through shows you actually like. Interest isn't a nice extra, it's what drops the filter.
Where Krashen gets argued with. The acquisition/learning split has weak experimental support, i+1 was never defined precisely enough to test, and his claim that error correction is useless drew decades of pushback. But his harshest critics kept the load-bearing finding: comprehension of large amounts of meaningful input is necessary for acquisition. Everything below modifies that claim. Nothing removes it.
Merrill Swain's challenge came from inside the immersion world. She studied Canadian French immersion students who had years of massive comprehensible input behind them, and found they understood French at near-native levels while their speaking and writing kept the same grammatical gaps year after year. Input turned out to be necessary but not sufficient. Her 1985 Output Hypothesis lists three things producing language does that comprehension never forces.
The moment you try to say something, you find out what you can't say. Swain's phrasing: producing the language "may be the trigger that forces the learner to pay attention to the means of expression needed in order to successfully convey their own intended meaning." You can understand past tense for years while cheerfully saying "yesterday I go", because comprehension never made you need the form.
Speech is a trial run. You verbalise a guess about how the language works and the world pushes back, and wrong guesses get revised. Izumi's 2002 study found that producing output containing a target form correlated strongly with noticing it, while merely receiving modified input barely moved the needle.
Discussing the language itself, with a tutor or a partner or out loud alone, turns it into an object you can think about. Lantolf found learning in 80% of the metalinguistic episodes he tracked during meaning-focused tasks. The discussion happens between two people first and inside your head later.
The mechanism underneath all three is a shift from semantic to syntactic processing. Comprehension lets you cheat: keywords, context and anticipation get you the gist without parsing anything properly. Output removes the cheat. That's why the immersion students plateaued, and it's why every serious method adds speaking practice a few months in, including the ones that started life as pure-input cults.
"It is possible to comprehend input, to get the message, without a syntactic analysis of that input."Merrill Swain, 1985
Michael Long's interaction hypothesis (1981, revised 1996) argues the active ingredient isn't raw input but negotiated input. When communication breaks down, speakers repair it: they simplify, repeat, rephrase, slow down. Those repairs convert input that was too hard into input that's exactly at your edge, in real time, tuned to you specifically. No textbook can do that.
Long's revised version folded in attention: negotiation doesn't just deliver input, it highlights form while meaning is still on the table. He called this focus on form, a brief glance at structure inside real communication, as opposed to decontextualised grammar lessons. Rod Ellis's 1991 critique accepted the core claim (input is necessary, interaction makes it comprehensible) while demanding a tighter account of how modified input becomes acquisition, which is where noticing, comparison and integration came in.
This is the argument for tutors, exchange partners and voice-chat practice. A thirty minute lesson generates more negotiation of meaning per minute than three hours of podcasts, because every misunderstanding gets repaired on the spot. It also explains why background listening is weaker per minute: nobody is negotiating anything with you. Passive input still builds the ear, but the repair loops live in conversation.
Nick Ellis's usage-based account explains why massive input works at the cognitive level. Learners unconsciously estimate the language from the sample they've heard, the way you'd estimate a population from a survey. Every construction you know is a form-meaning pattern extracted from the tokens you've encountered, and the extraction is driven by a handful of variables.
Read that way, immersion stops being a philosophy and becomes sampling theory: acquisition quality scales with the size, variety and attention of your input sample. One podcast and one genre gives you a biased estimate of the language. Different speakers, registers and contexts is what builds the feel for correctness Krashen talks about. Ellis also leaves room for explicit instruction, but in a specific role: briefly pointing out low-salience features so learners notice them in the flood. That's exactly what a good tutor and a well-built flashcard deck do.
Paul Nation put arithmetic on the input side. Vocabulary learning depends on how many times you meet a word and how much attention you pay each time, and when the two are compared directly, quality of attention wins. His findings turn "read more" into a budget.
Texts where about 2% of running words are unknown are the sweet spot for extensive reading: enough known context for guessing to work, enough new words to keep learning. Below that coverage, reading stops being reading and becomes decoding, which is slow, miserable and low-yield. This is the empirical version of comprehensible input.
To learn roughly a thousand word families a year, Nation's models assume each word gets met about twelve times across varied contexts. Meeting a word once in a lesson is not learning. Meeting it a dozen times across different stories is.
| Vocabulary level | Words to read per year (12 meetings each) | Weekly reading at 150 wpm | Per day, five days a week |
|---|---|---|---|
| 2nd 1,000 word family | 200,000 | 33 min | 7 min |
| 3rd 1,000 | 300,000 | 50 min | 10 min |
| 4th 1,000 | 500,000 | 1 h 23 min | 17 min |
| 5th 1,000 | 1,000,000 | 2 h 47 min | 33 min |
| 6th 1,000 | 1,500,000 | 4 h 10 min | 50 min |
| 7th 1,000 | 2,000,000 | 5 h 33 min | 1 h 07 min |
| 8th 1,000 | 2,500,000 | 6 h 57 min | 1 h 23 min |
From Nation (2015), based on his 2014 volume estimates. Note the shape of the curve: the first few thousand words come cheap, then the bill arrives. That's the intermediate plateau, in tabular form.
Nation's other findings are all free technique. Guessing from context works, but confirming the guess with a dictionary multiplies retention, and electronic lookup is now fast enough that it doesn't break reading flow. Narrow reading, staying with one topic or author, cuts the number of distinct new words roughly in half because the same vocabulary keeps recurring. Re-reading a book within a few weeks adds retrieval practice for nothing. And his Four Strands framework says a balanced course splits time roughly evenly between meaning-focused input, meaning-focused output, deliberate language study, and fluency work on easy material. Extensive reading should be about three sixteenths of total study time, two thirds of it on just-hard-enough texts and a third on deliberately easy ones.
Immersion supplies the meetings. Spaced repetition makes sure they land at useful times. Cepeda and colleagues' 2006 meta-analysis covered 839 assessments across 317 experiments in 184 articles, and two of its findings matter here.
The same total study time produces much better retention when the episodes are spread out. The effect has been replicating since Ebbinghaus in the 1880s. It is about as settled as psychology gets.
The optimal interval between study episodes grows with the retention interval you're aiming for. Cramming the night before a test wins the test and loses the year. A word you want to keep for life wants reviews spaced over weeks and months, which is what an SRS scheduler automates.
This is the licence for sentence mining. Pull sentences out of what you watched, let the scheduler space the reviews, and the words you met in last night's episode get their twelve Nation-style meetings over the following months instead of evaporating by Friday.
"Am I too old?" had no good answer until 2018, when Hartshorne, Tenenbaum and Pinker published the largest language-learning dataset anyone has assembled: 669,498 people who took a viral grammar quiz and reported their age, when they started English, and how they learned it.
What that means for an adult doing this at home: native-like grammar is statistically out of reach, and it was never a sensible target. What the study actually says is that your brain keeps learning grammar well into adulthood, and that of all the variables in the dataset, the one an adult in a suburb can move is immersive exposure time. That's the one the whole method is built on.
Canadian French immersion began in 1965: English-speaking children take at least half their entire school curriculum through French, from maths to science. Jim Cummins's 1998 review of that research is the definitive summary, and it reads like a controlled trial of everything on this page.
Students gained fluency and literacy in French at no cost to their English. After a brief catch-up period, mostly spelling, resolved by grade five, immersion kids matched or beat English-stream peers on English academic measures. No long-term lag in subject matter taught through French. By grade six, French comprehension and reading sat close to native norms. Bilingualism brought measurable cognitive and metalinguistic advantages.
Speaking and writing accuracy lagged native norms significantly even after years of immersion, grammar worst of all (Harley, Allen, Cummins and Swain, 1991). The diagnosis: classrooms were transmission-oriented. The teacher talked in French, students mostly listened, and there was little creative production, little authentic French children's literature, and no francophone peers to talk to because they attended a separate school system. Attrition was brutal too: Alberta data showed 43 to 68 percent of students gone by grade six and up to 97 percent by grade twelve.
Cummins's conclusion is the bridge between Krashen and Swain. Comprehension-based immersion is extremely good at building comprehension. Building production requires deliberately manufacturing reasons to produce: project work, creative writing, genuine two-way communication. Sixty years of national-scale data, and the lesson for someone learning at a desk is the same one. Input gets you understanding. Only pushed output gets you accuracy.
Strip out the turf wars and the field converges. Krashen's input, Swain's output, Long's interaction, Ellis's frequency statistics and Nation's four strands are different projections of one object.
| Principle | Who says it | What it looks like at home |
|---|---|---|
| Understanding messages drives acquisition | Krashen, input hypothesis | Watch, listen and read things you mostly understand and actually enjoy |
| Volume and variety of exposure is the mechanism | Ellis; Hartshorne et al. | Hours per day beat hours per week; mix speakers, genres and registers |
| Production builds accuracy that input can't | Swain; the Canadian data | Tutors, exchange partners, journaling, from about month three to six |
| Negotiation sharpens both | Long, interaction hypothesis | Prefer two-way conversation over one-way media whenever you can get it |
| Deliberate review locks vocabulary in | Nation; Cepeda et al. | Anki sentence mining, fifteen to thirty minutes a day, spaced rather than crammed |
| Anxiety blocks intake | Krashen, affective filter | Compelling content, low-stakes practice, no shame about mistakes |
The method pages are this table turned into a schedule.