GAGhazanfar Ali / Research

Published · journal article

Expanding Multilingual Co-Speech Interaction: The Impact of Enhanced Gesture Units in Text-to-Gesture Synthesis for Digital Humans

Ghazanfar Ali · Woojoo Kim · Muhammad Shahid Anwar · Jae-In Hwang · Ahyoung Choi

2025 · IEEE Access

First author

More gesture variety supports multilingual digital-human interaction.

Research illustration for Multilingual GestureCLR
Research illustration

From the paper

Author abstract

In this study, we explore the effects of co-speech gesture generation on user experience in 3D digital human interaction by testing two key hypotheses. The first hypothesis posits that increasing the number of gestures enhances the user experience across criteria such as naturalness, human-likeness, temporal consistency, semantic consistency, and social presence. The second hypothesis suggests that language translation does not degrade the user experience across these criteria. To explore these hypotheses, we investigated three conditions using a digital human: voice only with no gestures, limited(56 gestures) cospeech gestures, and full system functionality with over 2000 unique gestures. For the second hypothesis, we used language translation to provide multilingual support, retrieving gestures from an English rule base. We obtained text and pose from English videos and matched the pose with gesture units derived from Korean speakers’ motion-capture sequences, enhancing a comprehensive rule base that we used for gesture retrieval for given text input. We used translation of non-English input language to English for text matching. Our novel method utilizes an improved pipeline to extract text, 2D pose data, and 3D gesture units. Incorporating a cutting-edge gesture-pose matching model with deep contrastive learning, we retrieved gestures from a comprehensive rule base containing 210,000 rules. This approach optimizes alignment and generates realistic, semantically consistent co-speech gestures adaptable to various languages. A comprehensive user study evaluated our hypotheses. The results underscored the positive impact of diverse gestures, supporting the first hypothesis. Additionally, multilingual capabilities did not degrade the user experience, confirming the second hypothesis. Highlighting the scalability and flexibility of our method, this study provides valuable insights into cross-lingual data and expert systems for gesture generation, contributing significantly to more engaging and immersive digital human interactions and the broader field of human-computer interaction.

Author-written abstract from the author manuscript.

In plain language

What this work does

The system extracts text and 2D pose from English monologue videos, matches those poses to captured 3D gesture units with GestureCLR, and builds a text-to-gesture rule base. Non-English input is translated into English before gesture retrieval. The study compares no gestures, a small library, and the expanded library.

  1. 01Video + captured motion
  2. 02GestureCLR rule-map
  3. 03Text-driven gesture retrieval

At a glance

Method, evidence, and scope

Graphical abstract: GestureCLR rule construction and translation-based multilingual gesture retrieval
Graphical abstract diagram. GestureCLR expands a clustered gesture library for translation-based multilingual retrieval. View full size
Method and evidence for Multilingual GestureCLR
InputText; non-English text translated into English
OutputRetrieved 3D co-speech gesture units for a digital human
MethodContrastive 2D-to-3D matching offline; rule-based retrieval at runtime
Data and scopeEnglish monologue videos; Korean-speaker motion capture; 2,035 units and 210,000 rules
Evaluation51-participant study of gesture diversity and translation; reported high-noise matching improvement
LimitationsMultilingual support uses translation and an English rule base. The study does not establish equivalence across all languages or cultures.

Implementation and artifacts

Code and setup

Independent implementation of the paper’s core ideas, with setup instructions and data preparation documented in the repository README. The institute’s original source, datasets and trained models are not distributed.

Browse code and setup guide

Reference this work

Citation

Ghazanfar Ali, Woojoo Kim, Muhammad Shahid Anwar, Jae-In Hwang, Ahyoung Choi. Expanding Multilingual Co-Speech Interaction: The Impact of Enhanced Gesture Units in Text-to-Gesture Synthesis for Digital Humans. IEEE Access, 2025. Volume 13. Pages 145144-145157. DOI: 10.1109/access.2025.3596328.

Download BibTeX
@article{multilingualgesture2025,
  title = {{Expanding Multilingual Co-Speech Interaction: The Impact of Enhanced Gesture Units in Text-to-Gesture Synthesis for Digital Humans}},
  author = {Ali, Ghazanfar and Kim, Woojoo and Anwar, Muhammad Shahid and Hwang, Jae-In and Choi, Ahyoung},
  year = {2025},
  journal = {IEEE Access},
  volume = {13},
  pages = {145144-145157},
  doi = {10.1109/access.2025.3596328},
  url = {https://ghazanfarali.com/research/multilingual-gesture/}
}