Search and Filter

Repository Feedback

Your feedback helps us improve the repository's content relevance and usability. Please share your thoughts to help us better serve researchers and practitioners.

Submit feedback

Submit a research study

Contribute to the repository:

Add a paper

Beyond Direct Answering: Aligning Educational LLMs As Socratic Guides Via Heuristic Reinforcement Learning

Authors
Xiaokun Wang,
Siyu Song,
Wentao Liu,
Xiaodong Zou
Date
Publisher
arXiv
Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children's science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.
What is the application?
Who is the user?