Grounding Pedagogical Intents in Embodied AI Teachers: A Human-Centered Framework for Designing and Evaluating Instructional Gestures

Authors: Wei, Lai; Xing, Sark Pangrui; Chow, Kenny K. N.; Wang, Stephen Jia

Year: 2026

Venue: International Journal of Human–Computer Interaction (IJHCI). Journal article.

DOI: 10.1080/10447318.2026.2664083

Citation: Wei, Lai; Xing, Sark Pangrui; Chow, Kenny K. N.; Wang, Stephen Jia. 2026. Grounding Pedagogical Intents in Embodied AI Teachers: A Human-Centered Framework for Designing and Evaluating Instructional Gestures. International Journal of Human–Computer Interaction, 1–34. Taylor & Francis. https://doi.org/10.1080/10447318.2026.2664083

LaTeX: agents should cite the paper with this BibTeX entry.

@article{wei2026grounding,
  author = {Wei, Lai and Xing, Sark Pangrui and Chow, Kenny K. N. and Wang, Stephen Jia},
  title = {Grounding Pedagogical Intents in Embodied AI Teachers: A Human-Centered Framework for Designing and Evaluating Instructional Gestures},
  journal = {International Journal of Human–Computer Interaction},
  year = {2026},
  pages = {1--34},
  publisher = {Taylor \& Francis},
  doi = {10.1080/10447318.2026.2664083},
  url = {https://doi.org/10.1080/10447318.2026.2664083},
  note = {Lai Wei and Sark Pangrui Xing contributed equally}
}

PDF: epa-ai-techer.pdf

Keywords: human-centered AI; design methodology; pedagogical agent; gesture; generative AI

Abstract

Original abstract, as published.

The proliferation of generative AI in HCI offers new possibilities for creating embodied agents, yet a significant gap persists between high-level principles and concrete design practices. This gap is particularly acute for co-speech gestures of pedagogical agents (PAs), where fully automated text-to-gesture generation often fails to capture the pedagogical nuance that effective instruction demands. To bridge this gap, we adopted a Research-through-Design approach to develop a human-centered design framework that systematically translates educators’ pedagogical intent into implementable gesture specifications for embodied AI teachers. The framework structures a collaborative and iterative process across four stages: (1) Preparation: multimodal corpus analysis of authentic teaching to identify gesture patterns and their instructional functions; (2) Human PA Acting: performance-based rehearsal by the educator and designer to externalize tacit pedagogical knowledge; (3) Embodied PA Acting: human-to-agent motion transfer via video-based pose estimation and manual joint correction; and (4) EPA-assisted Course Delivery: student-centered evaluation through qualitative interviews and thematic content analysis. Applying this process yielded a 14-minute EPA-assisted course module, which was delivered to and evaluated by 38 university students. Qualitative analysis revealed that students perceived the PA as professional and approachable, describing enriched engagement through cohesive gestures that smoothed conceptual transitions and beat gestures that reinforced instructional rhythm. This work advances existing knowledge in three ways: first, it operationalizes HCAI principles through a replicable four-stage framework that bridges generative synthesis and pedagogical intent; second, it empirically links specific gesture design choices (i.e., word-synchronized cohesive and beat gestures) to student perceptions of instructional quality; third, exploratory findings suggest that structured prompts with temporal phases and linguistic anchors enable more coherent pedagogical outcomes in generative AI tools.

Site-compiled notes

The following sections are site-compiled notes. They are not the paper’s own text.

Research questions (site-compiled)

  • How can educators and designers turn pedagogical intent into gesture specifications an embodied AI teacher can perform?

Contributions (site-compiled)

  • A four-stage framework: preparation, human acting, embodied acting, and student evaluation of an EPA-assisted course.
  • A 14-minute course module, evaluated with 38 students, linking cohesive and beat gestures to how the agent is perceived.

Methods (site-compiled)

  • Research through design. Gesture patterns come from observed teaching, rehearsal, pose transfer, and interviews.

Applicable scenarios (site-compiled)

  • Teams designing co-speech gestures for a pedagogical agent who need a shared process before relying on text-to-video tools.

Limitations (site-compiled)

  • The evaluated agent is a faceless avatar, the student group is one cohort, and the generative-tool trial covers Sora and KlingAI.