Doctoral thesis · doctoral thesis
Scalable Hybrid Approach of Co-speech Text-to-Gesture Generation for Interactive Digital Humans
2023 · University of Science and Technology, KIST School
The evolution of text-driven gesture systems for interactive digital humans.
An external manuscript archive link will be added when the public deposit is available.

From the paper
Author abstract
This thesis chronicles developing and optimizing co-speech text-to-gesture generation for digital humans, underlining a significant leap towards enhancing user engagement and interaction in wearable mixed reality environments. The research narrative begins with the conception of a multimodal interaction framework for intelligent virtual agents and concludes with the advanced co-speech gesture generation system, ConGRets, capable of zero-shot speaker style adaptation. The research began with creating an intelligent virtual agent system for mixedreality environments. This system integrated features such as speech recognition, domainspecific chatbot capabilities, and manual mapping of body animations to speech content, highlighting the potential of digital humans as interactive tools. However, it also emphasized the need for a more automated and expressive body gesture generation system to enhance communication and engagement further. Motivated by this requirement, the research introduced an automated, rule-based co-speech gesture mapping system. This system, leveraging large publicly available video datasets and word embeddings, generated a diverse and accurate set of gestures that bypassed the limitations of traditional human-created mappings. A significant advancement was marked by the introduction of GestureCLR, a contrastive learning model that drastically improved the matching of 2D poses from videos to 3D gesture units, outperforming the previous mean cosine similarity method. The integration of GestureCLR and K-Means clustering on gesture units derived from English monologue videos and Korean speaker motion capture sequences led to the development of a cross-lingual method. This method effectively generated realistic, humanlike gestures for digital humans, augmenting their non-verbal communication capabiliii ities across different languages without compromising user experience. The final phase of the research introduced ConGRets, a unique framework optimizing real-time performance and resource usage. Addressing the limitations of previous methods, such as slow query speeds and averaged gesture generation, ConGRets incorporated zero-shot speaker style adaptation. This novel feature allowed the system to generate gestures in alignment with the text input and the individual speaker ’s motion style. The thesis narrates the journey of this groundbreaking research, which is now permeating diverse applications, including medical screening kiosks, the virtual therapy system, the actor-focused pre-visualization system ASAP , the ChatGPT frontend, and the virtual presenter. This comprehensive exploration of the evolution of co-speech text-to-gesture generation for digital humans underscores a significant stride towards enriching their non-verbal communication abilities and enhancing user interaction and engagement in various environments.
Author-written abstract from the author manuscript.
In plain language
What this work does
The dissertation connects wearable mixed-reality agents, automatic rule mining, GestureCLR pose matching, multilingual gesture retrieval, and ConGRets speaker-style adaptation. It documents the research lineage and should be cited as a dissertation rather than as a journal article.
- 01Interactive MR agent
- 02Automatic maps + GestureCLR
- 03Style-aware gesture retrieval
At a glance
Method, evidence, and scope
| Input | Text, video-derived pose, and captured gesture data across the included systems |
|---|---|
| Output | Text-driven co-speech gestures for interactive digital humans |
| Method | Multimodal agents, automated rule maps, contrastive matching, and style-aware retrieval |
| Data and scope | Research conducted for the 2023 AI-Robotics dissertation |
| Evaluation | System evaluations and user studies across dissertation chapters |
| Limitations | Chapter results reflect their respective systems and datasets; a dissertation chapter is not evidence of a later journal publication. |
Implementation and artifacts
Code availability
Original project implementations are held by the institute. Separate educational repositories will be handled paper by paper.
Reference this work
Citation
Ghazanfar Ali. Scalable Hybrid Approach of Co-speech Text-to-Gesture Generation for Interactive Digital Humans. University of Science and Technology, KIST School, 2023.
@phdthesis{doctoralthesis2023,
title = {{Scalable Hybrid Approach of Co-speech Text-to-Gesture Generation for Interactive Digital Humans}},
author = {Ali, Ghazanfar},
year = {2023},
school = {University of Science and Technology, KIST School},
url = {https://ghazanfarali.com/research/doctoral-thesis/}
}