IELTS Speaking: all three parts and how to prepare
The three parts are not three difficulty levels of the same task. They test different things — and the answer length that earns marks in Part 1 will cost you in Part 3.
The three parts are not three difficulty levels of the same task. They test different things — and the answer length that earns marks in Part 1 will cost you in Part 3.
The commonest misunderstanding about Speaking is treating the three parts as three rising difficulty levels of one task.
They are not. The three parts measure three different things, and an answer length that suits one will lose marks in another. A Part 1 answer given at Part 3 length reads as rambling; a Part 3 answer given at Part 1 length reads as having nothing to say.
The whole test runs 11 to 14 minutes, face to face with an examiner.
What it tests: talking about yourself and familiar things fluently and naturally. Home, work, study, hobbies, weather, food.
The right length: two or three sentences. No more.
The two failure modes: a one-word answer gives the examiner nothing to mark; a forty-second answer cuts down how many questions you get and starts to sound rehearsed.
A safe shape: answer directly, then add one reason or one example. Do you like cooking? → Not really, no. I find it stressful after a long day, so I usually just reheat something. That is enough.
How to practise: take twenty familiar topics and answer each in exactly three sentences, recording yourself. The aim is not brilliance but consistency — every answer arriving immediately, with no silence while you hunt for content.
The shape: you get a cue card, one minute to prepare with paper and pencil, then speak for up to two minutes without interruption.
What it tests: sustaining a long turn alone and staying coherent — something ordinary conversation never demands of you.
The real problem: running out of things to say at the forty-second mark.
The fix: do not try to generate more points — go deeper into one, through four layers: the event (what happened) → the feeling (what you felt, and why) → comparison (before and after, against other people, against expectations) → hypothetical (what if it had gone differently). One small experience taken through those four fills two minutes easily and sounds far more natural than a list.
How to use the preparation minute: do not write sentences. Write four keywords, one per layer. Writing full sentences leads to reading them aloud, and examiners hear that instantly.
How to practise: time two minutes, record, then listen back and mark where you stalled. That is exactly where a layer is missing.
What it tests: discussing abstract issues — analysing, comparing, speculating. The topic follows from Part 2 but widens to the societal level.
The right length: considerably longer than Part 1. Four to six sentences is reasonable.
You do not need the right opinion — you need one that is developed. The shape: position → reason → concrete example → concession.
That concession is what lifts an answer from band 6 to band 7, because it shows you can hold more than one side: Having said that, I suppose some people would argue…
When you have no ready opinion: it is entirely acceptable to say That's an interesting question — I've never really thought about it. I suppose… It buys you two seconds and sounds completely natural.
Fluency and Coherence — keeping going, with a thread. Hesitating to find an idea costs less than hesitating to find a word.
Lexical Resource — natural collocation, and the ability to paraphrase around a word you do not have. Talking your way round a gap is a strength, not a weakness.
Grammatical Range and Accuracy — both range and accuracy together.
Pronunciation — and here something needs saying plainly.
This is the most expensive misunderstanding in Speaking preparation. The band 7 Pronunciation descriptor reads: is easy to understand throughout; L1 accent has minimal effect on intelligibility.
What is assessed is intelligibility, not native-likeness. Concretely, three things: word and sentence stress (stressing the words that carry meaning), intonation (rising and falling with the sense rather than staying flat), and chunking (pausing between meaning groups, not inside them).
A clearly Vietnamese accent with stress and pauses in the right places can reach Pronunciation band 7 or above. Time spent trying to erase a first-language accent is better spent on those three.
Examiners identify a memorised answer by its recitation intonation, its unnaturally even pace, and its collapse the moment the question steps off the script. Examiners are trained to step off the script, and memorised answers are marked down.
What is worth preparing is not sentences but idea frames — angles of approach for each topic group — and functional phrases: It's a bit of a double-edged sword because…, It depends largely on…, I'd say the main reason is… Flexible raw material, not a rigid script.
The two are marked as separate criteria, so do not sacrifice one entirely. But in the moment, when you must choose: keep talking. Stopping to self-correct every small slip damages Fluency more than the slip damages Grammar.
Safe fillers: well, let me think, that's an interesting question — in moderation. Penalised fillers: drawn-out umm and err, and meaningless repetition.
This is worth stating plainly, because the misunderstanding is common enough to shape how people prepare: if you book IELTS on computer at a test centre, Speaking is still face-to-face with a human examiner. You do not talk to a machine, and there is no version of the centre-based test in which you do. The official description is unambiguous — the Speaking section is completed during a face-to-face session with a trained IELTS examiner.
What the computer format changes is only the other three sections and the logistics around them. Speaking may be scheduled on the same day as the rest of the test or within a few days of it, before or after; the interview itself is the same three parts, the same eleven to fourteen minutes, the same criteria, the same examiner behaviour described throughout this article.
There is one genuine variant, and it is a different product rather than a different format. IELTS Online — available for the Academic test in selected countries, taken from home — does conduct Speaking as a live video call with a trained examiner. Still a person, still the same three parts, still the same marking; only the room is different. Two things are worth knowing before choosing it: the questions and marking criteria are identical to the centre-based test, but each receiving organisation decides for itself whether it accepts IELTS Online results, so confirm with your university or immigration authority first rather than after.
If you do take the online version, rehearse the conditions and not just the answers. Speaking to a camera changes what you do with your eyes and how long you tolerate a silence; a two-second connection lag will make you talk over the examiner unless you have felt it before. Practise a full mock over video, with headphones, once — that is enough to remove the strangeness.
One scheduling note that does affect planning: results for the computer format are released in one to five days, against thirteen days for paper. Speaking is graded no differently in either case, but if a deadline is near, the wait is the thing to weigh.
Four recording sessions, ten minutes each: two on Part 1 with twenty short questions, one timed Part 2, one on Part 3 with three abstract questions.
Listening back to yourself is the least pleasant and most useful part. Mark three things: where you stalled for an idea, where you stalled for a word, and which words you keep repeating.
One pronunciation session focused on sentence stress rather than individual sounds.
One conversation with a real person each week if you can manage it. Alone, you cannot manufacture the pressure of having to answer immediately — and that pressure is precisely what the test measures.
The same article, told in plain words — for younger readers, or for anyone who wants the point quickly.
Three parts, about fourteen minutes. Here is what each actually tests, and the mistakes that cost the most.
A cue card, one minute to prepare, then up to two minutes speaking without interruption. It tests something ordinary conversation never asks of you: sustaining a long turn alone.
The real problem is running out of things to say at the forty-second mark.
The fix is not to find more points. It is to go deeper into one, through four layers:
One small experience taken through those four fills two minutes easily, and sounds far more natural than a list.
Use the preparation minute properly: do not write sentences. Write four keywords, one per layer. Writing full sentences leads to reading them aloud, and examiners hear that instantly.
Abstract questions, widening from your Part 2 topic to society. Answers should be considerably longer than in Part 1 — four to six sentences is reasonable.
You do not need the right opinion. You need a developed one: position → reason → concrete example → concession.
That last step is what lifts an answer from band 6 to band 7, because it shows you can hold more than one side.
This is the most expensive misunderstanding in the whole preparation. The band 7 pronunciation descriptor says: easy to understand throughout; L1 accent has minimal effect on intelligibility.
What is assessed is being understood, not sounding native. Concretely, three things: stressing the words that carry meaning, letting your pitch rise and fall with the sense instead of staying flat, and pausing between meaning groups rather than inside them.
A clearly Vietnamese accent, with the stress and pauses in the right places, can reach band 7 or above. Time spent trying to erase your accent is better spent on those three.
Examiners identify a memorised answer by its recitation rhythm, its unnaturally even pace, and its collapse the moment the question steps off the script. Examiners are trained to step off the script.
What is worth preparing is not sentences but idea frames — angles of approach for each topic group — and useful phrases: It's a bit of a double-edged sword because…, It depends largely on… Flexible raw material, not a rigid script.
They are marked separately, so do not sacrifice one entirely. But in the moment, when you must choose: keep talking. Stopping to correct every small slip damages Fluency more than the slip damages Grammar.
These articles summarize well-established research in learning science and linguistics. Key sources and further reading:
The recording plays exactly once. That single fact governs the whole paper, and every technique worth learning is a cons...
Task 2 counts twice as much as Task 1. So the most consequential decision you make is not a good phrase — it is stopping...
Unlike Listening, Reading gives you no transfer time. Sixty minutes is everything you get, and that single fact shapes e...