What AI does well with discussion questions
Speed and volume. No argument here. ChatGPT, Claude, or Gemini can produce 20 discussion questions on any topic in under a minute. This is dramatically faster than teacher writing time.
Topic specificity. AI handles niche, specific, or unusual topics better than pre-written question banks. "Generate 10 discussion questions about deep-sea mining" or "questions about the role of the grandmother in Japanese family life" produces workable questions on topics most banks don't cover.
Level adaptation. With a clear prompt specifying
CEFR level and what that means, AI produces reasonably level-appropriate questions. The adaptation is not always perfect, but with prompt refinement it improves.
Cultural customisation. AI can be instructed to generate questions relevant to a specific cultural context or student background. "Generate questions about technology for adults in Vietnam where smartphone use is near-universal" produces different questions from a generic prompt.
Varied formats. AI readily produces questions in specific formats:
debate motions, this-or-that choices, ranking tasks, role play scenarios - whatever the format, AI can generate it on request.
What AI does poorly
Consistent CEFR calibration. AI frequently misjudges what B1 looks like versus B2, or C1 versus B2. Questions labelled "B2" are often either too simple or too abstract. Without manual checking, you may be giving students questions that are pitched wrong.
Avoiding repetition across sessions. AI has no memory of what your class has discussed before. Every session, you're generating fresh questions without any filter for what's been covered. Students may end up discussing similar themes repeatedly without you noticing.
Classroom safety filtering. AI doesn't know your specific students. A question that's fine for adult professionals may be completely inappropriate for teenagers, or vice versa. Age-group filtering requires deliberate prompting and then checking.
The "obvious correct answer" problem. A surprising proportion of AI-generated discussion questions have an implied correct answer that most people would give. Questions with obvious right answers produce short, shallow conversations. A question bank designed for pair speaking is engineered to avoid this; AI is not.
Grammatical idiosyncrasy. AI occasionally produces questions with slightly unnatural phrasing that experienced teachers recognise immediately but which students find confusing. This is rare but worth checking.