What TOPIK listening actually tests: 24 past papers, measured
You can read a TOPIK II passage and follow most of it. Then the listening audio starts and it slides past you. If you’ve been treating that as a grammar problem, as a sign there’s a stack of intermediate forms you haven’t met yet, the past papers say otherwise.
We took every released TOPIK listening booklet that prints its script, tagged all 1,561 spoken lines in them against a fixed inventory of 88 grammar points, and compared the two levels. The corpus is 24 papers across 12 exam rounds, 2014 through 2025, from the official archive at topik.go.kr. Every number below comes out of that tagging run.
What grammar do you need for TOPIK II listening?
Almost exactly the grammar you already needed for TOPIK I. Nineteen forms appear in at least nine of the twelve TOPIK I papers, and all nineteen are just as reliable in TOPIK II. Sixty-nine percent of every grammar hit in the TOPIK II scripts comes from those same nineteen forms.
What changes is how much Korean arrives per question.
| TOPIK I | TOPIK II | ratio | |
|---|---|---|---|
| Script characters per paper (mean) | 1,266 | 3,183 | 2.5× |
| Characters per spoken line (mean) | 19.6 | 48.6 | 2.5× |
| Characters per spoken line (median) | 16 | 29 | 1.8× |
| Characters per spoken line (90th percentile) | 32 | 104 | 3.3× |
| Distinct grammar points seen | 56 | 78 | 1.4× |
| Grammar points per 1,000 characters | 94.6 | 74.7 | 0.8× |
| Share of hits that are first-year forms | 88.9% | 73.4% | — |
Look at the last two rows, because they’re the ones that undercut the intuition. Grammar density goes down at TOPIK II — fewer tagged forms per thousand characters, not more. And nearly three quarters of what’s there is still first-year material. The sentences aren’t denser with grammar. They’re just longer, and there are more of them.
The volume ratio holds in every round we have: the TOPIK II script runs between 1.82× and 3.68× the TOPIK I script of the same sitting, median 2.53×. It isn’t an artifact of one unusual year.
The 19 forms that carry both levels
These are the forms appearing in at least nine of the twelve TOPIK I papers. The last column is how often the same form turns up in the TOPIK II scripts — which is the whole point: none of them go away.
| Form | What it does | TOPIK I papers | Hits in TOPIK I | Hits in TOPIK II |
|---|---|---|---|---|
| -아/어요 | polite informal ending | 12/12 | 394 | 162 |
| -(으)세요 | polite command, honorific | 12/12 | 133 | 43 |
| -(스)ㅂ니다 | formal polite declarative | 12/12 | 103 | 415 |
| -는데 / -(으)ㄴ데 | background, contrast | 12/12 | 96 | 268 |
| -아/어서 | so, and then | 12/12 | 78 | 110 |
| -(으)ㄹ 수 있다 / 없다 | can, cannot | 11/12 | 49 | 119 |
| -네요 | noticing, mild surprise | 12/12 | 47 | 34 |
| -(으)ㄹ 것이다 / 거예요 | will, probably | 10/12 | 37 | 36 |
| -(으)니까 | because (spoken) | 9/12 | 33 | 60 |
| -(으)ㄹ까요 | shall we, I wonder | 12/12 | 33 | 57 |
| -고 싶다 | want to | 11/12 | 31 | 24 |
| -는 것 / -(으)ㄴ 것 | nominalized clause | 10/12 | 28 | 164 |
| -는데요 / -(으)ㄴ데요 | softened statement | 10/12 | 28 | 134 |
| -(으)ㄴ 후에 / -기 전에 | after, before | 10/12 | 27 | 23 |
| -지 않다 | negation (long form) | 10/12 | 25 | 78 |
| -아/어 주다 | do something for someone | 12/12 | 25 | 60 |
| -아/어 보다 | try doing | 11/12 | 25 | 53 |
| -지요 / -죠 | seeking agreement | 10/12 | 24 | 100 |
| -(으)ㄹ게요 | speaker promise | 10/12 | 18 | 28 |
Notice which ones grow rather than shrink. -(스)ㅂ니다 goes from 103 hits to 415, and -는 것 / -(으)ㄴ 것 from 28 to 164. That’s TOPIK II leaning into lectures, announcements and interviews, where the register is formal and clauses get packed inside other clauses. Same forms, heavier use.
So what does TOPIK II actually add?
Seventeen more forms reach the same reliability bar, in at least nine of the twelve TOPIK II papers, that don’t reach it in TOPIK I. This is the honest answer to “what should I study for the jump”:
| Form | What it does | TOPIK II papers | Hits in TOPIK II | Hits in TOPIK I |
|---|---|---|---|---|
| -고 있다 | progressive | 12/12 | 83 | 15 |
| -지만 | but | 12/12 | 51 | 10 |
| -에 대해 / -에 대한 | about, regarding | 10/12 | 45 | 2 |
| -(으)ㄹ 때 | when | 12/12 | 42 | 18 |
| -아/어지다 | become, get | 11/12 | 40 | 2 |
| -(으)면서 | while (simultaneous) | 12/12 | 38 | 6 |
| -나요 / -(으)ㄴ가요 | softened question | 12/12 | 36 | 5 |
| -(으)ㄴ/는 것 같다 | seems (observed) | 12/12 | 35 | 6 |
| -잖아요 | as you know | 11/12 | 35 | 4 |
| -(으)ㄹ 것 같다 | seems, probably will | 10/12 | 30 | 11 |
| -기 위해(서) | in order to | 12/12 | 29 | 0 |
| -는데도 / -아/어도 | even though, even if | 10/12 | 28 | 5 |
| -게 되다 | come to, end up | 11/12 | 26 | 5 |
| -아/어야 하다 / 되다 | must | 10/12 | 24 | 5 |
| -다고 하다 | reported statement | 9/12 | 16 | 0 |
| -거나 | or | 9/12 | 15 | 7 |
| -(으)라고 하다 | reported command | 9/12 | 13 | 0 |
Seventeen forms. Most of them you’ve almost certainly already studied.
Three of them are genuinely new territory in the sense that they never appear in a single TOPIK I script in this corpus: -기 위해(서), in all twelve TOPIK II papers, and the two reported-speech forms -다고 하다 and -(으)라고 하다, in nine each. Reported speech is worth flagging on its own — TOPIK II listening frequently asks you to track who said what to whom, and that grammar is how the script marks it.
Twenty-four points in total appear in TOPIK II and never in TOPIK I. Those three are the only ones that recur reliably; the other twenty-one are thin — mostly a handful of hits scattered over a few papers. -(으)ㄹ 뿐만 아니라, -(으)로 인해, -와/과 달리, -더라도, -(으)ㄹ수록. Real, but not what the exam rests on.
The forms that never showed up
Eight points from the inventory appear in none of the 24 papers:
-자마자, -았/었더니, -길래, -는 바람에, -기는커녕, -(으)ㄹ 뻔하다, -(으)ㄴ/는 셈이다, -조차 / -마저.
Several of those get drilled hard in intermediate textbooks. That’s not an argument against learning them, they’re actually common in conversation and in writing, and this is a listening corpus, not the whole language. It’s an argument about sequencing. If the exam is what you’re preparing for in the next few months, they’re not where the next hour goes.
And two forms run the other way, appearing only in TOPIK I: -(으)십시오 (three papers) and -기로 하다 (two). Both thin enough to be sampling noise rather than a pattern.
What this changes about how you practice
If the grammar inventory is nearly the same at both levels, then “study more grammar” is the wrong lever for the TOPIK II listening jump. The lever is holding a longer stretch of familiar Korean in your head while it’s still arriving.
Concretely, from the numbers above: a typical TOPIK I line is 16 characters. The TOPIK II line at the 90th percentile is 104. That’s the reach, not a harder sentence, a longer one, built from the same nineteen forms you already know, and it doesn’t wait for you.
So the practice that matches the data is length, not novelty. Take material built from grammar you’ve already met and push the duration you can follow without a reset — a sixty-second stretch before a thirty-second one, monologue before dialogue, formal register before casual, since -(스)ㅂ니다 is where TOPIK II lives. That’s a different session from a grammar drill, and it’s the one the past papers argue for.
The catch is that “grammar you’ve already met” is different for every learner, which is why nobody can publish that material for you. That’s the problem up103 works on: we read your study notes and write Korean texts and dialogues from what’s actually in them, and generate incredible audio. And the difficulty comes from length and speed rather than from words you’ve never seen. Founding Learners get 6 months free, then 50% off for life, at app.up103.com.
Method
Corpus. All released TOPIK listening booklets that print the listening script — the 듣기 통합 and 듣기 대본 papers — from the official past-paper archive at topik.go.kr (학습하기 > 학습 자료실). That is 24 papers: rounds 35, 36, 37, 41, 47, 52, 60, 64, 83, 91, 96 and 102, one TOPIK I and one TOPIK II booklet for each, spanning 2014 to 2025. TOPIK I covers levels 1–2 and TOPIK II levels 3–6, so “level” here means the two exam papers, not the six certified levels. Question-only booklets and answer keys from the same rounds were excluded, since they don’t print the audio script.
Extraction. The booklets are 300-dpi page scans with no text layer, so each was OCR’d (ocrmypdf --skip-text -l kor+eng) and extracted with pdftotext -layout. Lines carrying a speaker label (남자, 여자, 가, 나) were kept as spoken utterances; printed answer choices were not. That yielded 1,561 utterances and 53,391 characters of script. Counting speaker labels in the raw OCR text as a denominator, the extractor recovered 90.0% of spoken lines (87.9% at TOPIK I, 92.3% at TOPIK II) — the rest were lost to mangled labels. Because recall is similar at both levels, the ratios above are more trustworthy than the absolute totals, which are lower bounds.
Tagging. A fixed inventory of 88 grammar points, written before the counting was run, matched by regular expression. Korean OCR inserts spurious spaces inside words and the original word spacing can’t be recovered, so matching runs on a whitespace-free form of each line, with per-pattern stop-lists to block collisions that only exist because the spaces are gone. Two very common connectives — bare -고 and bare -(으)면 — were deliberately left out of the inventory: without spacing they collide with too many ordinary nouns to count honestly. Their distinctive composite forms (-고 있다, -고 싶다, -(으)면서, -(으)면 되다) are counted. The totals are an undercount by design.
Model-assisted step, and what it was. The tagger is deterministic, not a model. What was model-assisted is its validation: An LLM read 120 sampled hits with their surrounding context and judged, one by one, whether each tagged form really is that grammar point in that sentence. On the final frozen pattern set the agreement rate was 94.2% (113 of 120) — and 98.5% on the ten most frequent points, which account for 59.5% of all hits. Every judgement, including the seven errors and why each is wrong, is recorded alongside the analysis. Earlier rounds of the same check drove three pattern fixes before the patterns were frozen; the reported rate is measured on a sample the patterns were not tuned against.
What is and isn’t new here. Measuring TOPIK listening is not new ground. Korean-language academic work on the listening section exists, theses and journal articles on 한국어능력시험 듣기 텍스트 난이도 분석 (listening-text difficulty, with vocabulary and grammar difficulty among the factors) and on 듣기 평가 문항의 선택지 제시 방법 (how the answer options are presented), indexed on KCI, DBpia and KISS. If you read Korean and want the scholarship, that’s where it is. What this post is instead is a free, open, per-level coverage table you can read, quote and check the method of without an institutional login — an artifact, not a first measurement.
Limitations. Only released papers, not the full exam history, and NIIED has not released every round. OCR errors are a real error source — the scans mis-read circled answer digits and occasional syllables, and one paper (round 96, TOPIK I) extracted at only 78.8% recall. Tagging against a fixed inventory undercounts anything outside it. Per-point counts below roughly ten hits should be read as indicative, not exact, since that’s where the tagger’s precision is weakest. And the whole study measures the script, not the audio: speech rate, pausing and speaker overlap are a separate question, and a real part of listening difficulty that a transcript can’t see.
— 민지, up103
Founding Learners get 6 months free, then 50% off for life · 50 spots available