AJATT's founding idea was "bring Japan to you." The internet made that cheap. Five layers, built in order. Each one costs a bit of setup once and pays exposure back forever.
These are settings you change once and never think about again, and each one converts time you were already spending into input. Ellis's sampling argument is the reason: every flip makes the sample bigger without costing you a minute.
Switch the OS language on both. You already know where everything lives by muscle memory, so the first annoying week costs you almost nothing in function, and after that you get hundreds of incidental word-meetings a day from menus and dialogs you'd have skimmed in English anyway. Best exposure-per-effort ratio available anywhere.
Sticky notes on objects, in characters plus pinyin. The classic beginner move, and it does work for a while, mostly because it forces the question "how do I say this?" fifty times a week. Retire it after a few months. Labels only teach nouns you already point at every day.
Inner monologue in the target language while making coffee, driving, in the shower. No equipment, no cost, works anywhere. When you hit a gap, and you'll hit one constantly, note the phrase and look it up later. This is Swain's noticing function running without a conversation partner.
Make target-language audio the default for every slot that currently holds silence or English: driving, dishes, gym, walking the dog. Early on you won't understand most of it, and that's fine. You're training the ear to find word boundaries in the stream, which everything else depends on.
Google, YouTube, Wikipedia, recipe sites. Search in the target language for things you were going to look up anyway. Recipes, tech fixes, game guides and hobby content are ideal because you already know the domain, so context does the comprehension work Krashen needs.
Make a separate YouTube or streaming profile used only for target-language content, and only interact with that content there. Within about a week the recommendation engine becomes a self-renewing comprehensible-input machine. For most people this is the highest-leverage item on the list, because it hijacks time you were already spending.
The failure mode of home immersion isn't laziness. It's buying a native crime drama in week two, understanding twenty percent of it, and quietly deciding you're bad at languages. Nation's 98% rule from reading research transfers almost directly to listening: below roughly ninety percent comprehension you're not acquiring, you're enduring.
| Rung | Setup | What it's for | How long to stay on it |
|---|---|---|---|
| 0 | Native audio, English subs | Getting used to the sound with zero comprehension demand | Weeks zero to four, then get off. With English subs on, your eyes read English and your ears tune the audio out. |
| 1 | Native audio, target-language subs | Binding sound to written form | The workhorse. Most learners spend most of their first year here. |
| 2 | Dual subs | Unsticking yourself when rung 1 stalls | Sparingly. The English line competes for your attention and usually wins. |
| 3 | No subs | Pure listening; the ear does all the work | Add gradually once rung 1 feels easy. Rewatching known content is the gentle way in. |
| 4 | Re-watch, no subs | Known plot means near-total comprehensibility | Any time. A second viewing of a known episode beats a first viewing of a new one, per minute. |
Full attention on the content. Watching with subs, looking words up, mining sentences, reading with a dictionary open. This is where most acquisition happens, especially early.
high yield per minutecosts real attentionthe daily core
Audio in the background while your hands are busy: dishes, commute, gym, work. Comprehension is partial and intermittent, so the yield per minute is much lower, but the minutes are essentially free.
free minutesbuilds the earweak on its own
The ratio that works: thirty to sixty minutes of active immersion as the daily core, then as much passive as your life allows. And re-run yesterday's episode as today's background audio. Passive listening to content you've already watched actively is dramatically better than passive listening to new material, because you're consolidating known input instead of filtering noise.
What passive immersion is genuinely good for: prosody and rhythm, which is make-or-break for tonal languages; getting used to the speed and blur of real speech; and phonological segmentation, hearing where words begin and end, which adult beginners find nearly impossible at first. What it can't do: teach you vocabulary you've never met, or fix grammar you don't notice. It's a supplement with real value, not a substitute.
Listening builds the ear and the gut feel. Reading builds the vocabulary and the precision, and Nation's numbers make it the highest-yield vocabulary source available, provided you read at the right level.
The retention engine from the AJATT and Refold world. While consuming media, pull out sentences containing one unknown item and put them in Anki with native audio.
Nation's supporting techniques, all free: guess then confirm (guess a word from context, then check the dictionary; confirming a guess beats looking it up cold), narrow reading (one author or topic at a time, which halves the distinct new words), re-reading within a few weeks (free retrieval practice), and easy-reading sessions (about a third of your reading time on material with almost nothing unknown, purely for speed and fluency).
Swain's research and the Canadian data agree that comprehension without production plateaus. You need to be pushed to say things, and you need someone on the other end negotiating meaning with you. All of it now happens over a video call from your desk.
italki and Preply. Community tutors run roughly ten to eighteen Australian dollars an hour for Mandarin, professional teachers more. Book two or three thirty-minute sessions a week and ask them to correct you. That correction is the pushed output and the negative evidence Long's research says drives form learning.
HelloTalk, Tandem, Discord servers. Free, and the reciprocity (you help them with English) is what makes the arrangement last. Less structured than a tutor, more social, and social is what keeps it going at month eight when the novelty has worn off.
Voice-mode LLM conversation: infinite patience, zero social risk, available at eleven at night. It targets Krashen's affective filter directly, because no judgment means no anxiety means no blocked intake. Good for drilling the twenty conversation scripts you actually need.
Repeat audio a beat behind the speaker, matching rhythm, intonation and, for Mandarin, tones exactly. Builds the motor patterns for pronunciation without a partner, and forces the syntactic processing Swain says comprehension alone skips.
Three to five sentences a day about what you did. Writing gives you time to run the Monitor properly, since Krashen's three conditions all hold when you write, so accuracy improves faster here than in speech and then transfers.
Once a week, record two minutes of free speech and listen back. It's unpleasant and it's the fastest way to hear your own gaps. You'll catch fossilised errors instantly that a tutor has been correcting for months.
Timing. Don't force output on day one. Krashen's silent period is real, and speaking before you can hear the difference between tones just practises errors. Rough guide: two to three months of mostly input to build the ear, then start speaking and never stop. For tonal languages, listen longer than feels necessary before you open your mouth.
| Stage | Focus | Media | Mining and SRS | Output |
|---|---|---|---|---|
| Absolute beginner weeks 0–8 | Sound system first: tones and pinyin, or the alphabet. Then the highest-frequency few hundred words. | Pronunciation videos, beginner comprehensible-input channels, kids' songs. Passive audio from day one. | A beginner Anki deck plus pronunciation pairs. | None. Quiet shadowing only. |
| Beginner months 2–6 | Listening stamina and the core thousand words. | Learner content, graded readers, easy kids' animation. Rungs zero and one. | Mine ten to fifteen sentences a day from what you watch. | Self-talk, journaling, first tutor session around month four. |
| Intermediate months 6–18 | The long slog: one to five thousand words, native-speed listening. Where most people stall. | Native content you enjoy, re-watches of known shows, web novels, podcasts. | Mining continues; the review queue becomes the main workload. | Tutor two or three times a week. This stage needs it most. |
| Advanced year 2+ | Depth, register, domain vocabulary, accent. | News, books, film, technical material in your field, unscripted native media. | Mining tapers. Most new words now come from volume alone. | Regular conversation and writing, no scaffolding. |