Breadcrumb
- Home
- AI Hub For Education
- Research Study Repository
- Outcomes – Numeracy
Search and Filter
Repository Feedback
Your feedback helps us improve the repository's content relevance and usability. Please share your thoughts to help us better serve researchers and practitioners.
Outcomes – Numeracy
Research synthesis is AI-generated, human reviewed. Updated 05/2026.
Displaying 1 - 30 of 239
Generative AI Can Harm Teaching
Alp Sungu, Benjamin Lira, Angela L. Duckworth. (07/2026). SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7007339
Beyond Helpfulness: A Teaching-Over-Solving Diagnostic For Measuring Educational Impact In LLM Tutors
Junyi Yao, Zihao Zheng, Baichuan Li. (06/2026). arXiv. https://arxiv.org/abs/2606.16206v1
Faster Completion, Less Learning: Generative AI Reduced Study Time On Math Problems And The Knowledge They Build
Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn, Eyad Kurd-Misto. (06/2026). arXiv. https://arxiv.org/abs/2605.21629v2
Context-Aware Prediction Of Student Quiz Performance With Multimodal Textbook Features
Samin Khan. (05/2026). arXiv. https://arxiv.org/abs/2606.24770v1
Facet: Multi-Agent AI Supporting Teachers In Scaling Differentiated Learning For Diverse Students
Jana Gonnermann-Muller, Jennifer Haase, Nicolas Leins, Moritz Igel, Konstantin Fackeldey and Sebastian Pokutta. (05/2026). arXiv. https://arxiv.org/abs/2601.22788v4
"Would You Want an AI Tutor?" Understanding Stakeholder Perceptions of LLM-based Systems in the Classroom
Caterina Fuligni, Daniel Dominguez Figaredo, Armanda Lewis, Julia Stoyanovich. (05/2026). arXiv. https://arxiv.org/abs/2503.02885v3
Llama Lima: A Living Meta-Analysis On The Effects Of Generative AI On Learning Mathematics
Anselm Strohmaier, Samira Bödefeld, Oliver Straser, Frank Reinhold. (05/2026). arXiv. https://arxiv.org/abs/2601.18685v3
A Scoping Review Of Large Language Model-Based Pedagogical Agents
Shan Li, Juan Zheng. (05/2026). arXiv. https://arxiv.org/abs/2604.12253v2
Addressing The Reality Gap: A Three-Tension Framework For Agentic AI Adoption
Jason Fournier, Kacper Łodzikowski. (05/2026). arXiv. https://arxiv.org/abs/2604.27245v2
Promptdecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions
Miina Koyama, Ruiwei Xiao, John Stamper. (05/2026). arXiv. https://arxiv.org/abs/2605.16605v1
Teaching with Gemini: Measuring the impact of Guided Learning on student mathematics progress in Sierra Leone
LearnLM Team, Google & Fab AI. (05/2026). Google. https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26…
Little Impact Of ChatGPT Availability On High School Student Test Score Performance
Nick Huntington-Klein. (05/2026). arXiv. https://arxiv.org/abs/2605.08812v2
When Should Teachers Control AI Generation For Mathematics Visuals?
Zhengxu Li, Junling Wang, April Yi Wang. (05/2026). arXiv. https://arxiv.org/abs/2605.10672v1
Math Education Digital Shadows For Facilitating Learning With LLMs: Math Performance, Anxiety And Confidence In Simulated Students And Ais
Naomi Esposito, Anthony Tricarico, Luisa Porzio, Ali Aghazadeh Ardebili, and Massimo Stella. (04/2026). arXiv. https://arxiv.org/abs/2604.27618v1
Human-In-The-Loop Benchmarking Of Heterogeneous LLMs For Automated Competency Assessment In Secondary Level Mathematics
Jatin Bhusal, Nancy Mahatha, Aayush Acharya, Raunak Regmi. (04/2026). arXiv. https://arxiv.org/abs/2604.26607v1
From Test-Taking To Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark For LLMs On English Standardized Tests
Luoxi Tang, Tharunya Sundar, Yuqiao Meng, Shuai Yang, Ankita Patra, Lakshmi Manohar Chippada, Jiqian Zhao, Yi Li, Weicheng Ma, Zhaohan Xi. (04/2026). arXiv. https://arxiv.org/abs/2505.17056v2
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
Michael Hardy, Yunsung Kim. (04/2026). arXiv. https://arxiv.org/abs/2603.00883
Beyond The AI Tutor: Social Learning With LLM Agents
Harsh Kumar, Jonathan Vincentius, Zi Kang (Jace) Mu, Ashton Anderson. (04/2026). arXiv. https://arxiv.org/abs/2604.02677v1
How Motivation Relates To Generative AI Use: A Large-Scale Survey Of Mexican High School Students
Echo Zexuan Pan, Danny Glick, Ying Xu. (04/2026). arXiv. https://arxiv.org/abs/2603.19263v2
Evaluating Vision-Language And Large Language Models For Automated Student Assessment In Indonesian Classrooms
Nurul Aisyah, Muhammad Dehan Al Kautsar, Arif Hidayat, Raqib Chowdhury, and Fajri Koto. (04/2026). arXiv. https://arxiv.org/abs/2506.04822v3
Evaluating A Data-Driven Redesign Process For Intelligent Tutoring Systems
Qianru Lyu, Conrad Borchers, Meng Xia, Karen Xiao, Paulo F. Carvalho, Kenneth R. Koedinger, and Vincent Aleven. (03/2026). arXiv. https://arxiv.org/abs/2603.29094v1
Practitioner Voices Summit: How Teachers Evaluate AI Tools Through Deliberative Sensemaking
Dorottya Demszky, Christopher Mah, Helen Higgins. (03/2026). arXiv. https://arxiv.org/abs/2603.22588v3
Exploring Student Perception On Gen AI Adoption In Higher Education: A Descriptive Study
Harpreet Singh, Jaspreet Singh, Satwant Singh, Rupinder Singh, Shamim Ibne Shahid, Mohammad Hassan Tayarani Najaran. (03/2026). arXiv. https://arxiv.org/abs/2603.27777v1
Artificial Intelligence In Secondary Education: Educational Affordances And Constraints Of ChatGPT-4O Use
Tryfon Sivenas, Panagiota Maragkaki. (03/2026). arXiv. https://arxiv.org/abs/2602.13717v2
Facet: Teacher-Centred LLM-Based Multi-Agent Systems- Towards Personalized Educational Worksheets
Jana Gonnermann-Muller, Jennifer Haase, Konstantin Fackeldey, Sebastian Pokutta. (03/2026). arXiv. https://arxiv.org/abs/2508.11401v4
Heal: Hindsight Entropy-Assisted Learning For Reasoning Distillation
Wenjing Zhang, Jiangze Yan, Jieyun Huang, Yi Shen, Shuming Shi, Ping Chen, Ning Wang, Zhaoxiang Liu, Kai Wang, Shiguo Lian. (03/2026). arXiv. https://arxiv.org/abs/2603.10359v1
Conversational Learning Diagnosis Via Reasoning Multi-Turn Interactive Learning
Fangzhou Yao, Sheng Chang, Weibo Gao, Qi Liu. (03/2026). arXiv. https://arxiv.org/pdf/2603.03236v1
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform Llms
Prarthana Bhattacharyya, Joshua Mitton, Simon Woodhead, Ralph Abboud. (03/2026). arXiv. https://arxiv.org/pdf/2603.02830v1
When Shallow Wins: Silent Failures And The Depth-Accuracy Paradox In Latent Reasoning
Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary. (03/2026). arXiv. https://arxiv.org/pdf/2603.03475v1

