The six components
Drawing on Canale and Swain's communicative competence framework (1980) and the extensions by Bachman and Palmer (1996), "speaking English" decomposes into:
1. Language production
The ability to produce well-formed English at speed. Grammar, vocabulary, pronunciation, fluency of delivery. This is what most ESL textbooks foreground and what AI tutors mostly target.
It is the most measurable component and the easiest to drill. It is also only one of six.
2. Comprehension
The ability to understand spoken English at speed, including variations of accent, pace, and register. Some accent-and-register variety we covered in the accents post and the slang and idioms post.
AI tutors deliver a narrow band of comprehension input. Real classrooms deliver broad bands automatically. The breadth is what real-world comprehension requires.
3. Repair
The ability to recover when communication breaks down. Asking for clarification, rephrasing, offering examples, acknowledging confusion, signalling that you're not following. We covered this in detail in the predictability trap post.
This is the central skill of conversational competence and one of the weakest areas of AI tutor practice. AI tutors smooth over breakdowns; real partners require students to repair them.
4. Register
The ability to match the formality and style of your language to the context. Knowing when to say "would you mind" versus "can you" versus "gimme". Knowing when to use idioms and when to drop them. Knowing how academic English differs from casual English differs from professional English.
AI tutors default to a narrow neutral register. Real classrooms expose students to multiple registers and require them to switch between them.
5. Nonverbal cues
The visual, prosodic, and gestural channels that carry meaning alongside the words. Eye contact, facial expression, pitch contour, pause patterns, head movement. The literature estimates that something like 30-40% of communicative meaning in face-to-face interaction is carried in these channels.
AI tutors handle approximately zero of this. The voice-only ones miss everything visual; the video-avatar ones have schematic facial expressions that don't carry real meaning. Real classrooms deliver the full nonverbal channel for free.
6. Intercultural sensitivity
The awareness that communication norms vary across cultures. Politeness norms, turn-taking conventions, topic appropriateness, conversational distance, the meaning of silence. A student who is grammatically perfect but interculturally tone-deaf will misfire constantly in international communication.
AI tutors are not culturally calibrated. Real classrooms with international students naturally expose students to intercultural variation. The diversity of the class itself is the training ground.
How the bundle interacts in real use
The six components are not separable in real conversation. Every utterance involves several of them simultaneously. A student asking a question in a job interview is:
- Producing language (component 1)
- Comprehending the interviewer's previous turn (component 2)
- Repairing the moment they realise they misunderstood the original question (component 3)
- Matching register to the professional context (component 4)
- Maintaining appropriate eye contact and tone (component 5)
- Calibrating the directness of the question to the interviewer's cultural background (component 6)
A student who's only practised production sounds grammatically fine and behaves interculturally wrong. The interviewer hears the grammar but reacts to the cultural misfire. The student doesn't get the job.
This is why "I can speak English" is not a sufficient description of communicative competence. The student probably means component 1. They have not necessarily got 2-6.