Published · journal article
RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation
2025 · Computer Animation and Virtual Worlds
First author
High-confidence rules and learned similarity retrieve semantically aligned gestures.

From the paper
Author abstract
Co-speech gestures are essential for natural human communication, yet existing synthesis methods fall short in delivering semantically aligned and contextually appropriate motions. In this paper, we present RIDGE, a hybrid system that combines rule-based and deep learning approaches to generate realistic gestures for virtual avatars and human-computer interaction. RIDGE employs a high-fidelity rule base generated from motion capture data with the assistance of large language models, to select reliable gesture mappings. When a high-confidence match is not available, a contrastively trained deep learning model steps in to produce semantically appropriate gestures. Evaluated using a novel Gesture Cluster Affinity (GCA) metric, our system outperforms existing baselines, achieving a GCA score of 0.73 compared to rule-based baseline 0.6 and end-toend: 0.52, while ground truth score was 0.90. Detailed analyses of system architecture, data preprocessing, and evaluation methodologies demonstrate RIDGE’s potential to enhance gesture synthesis. Project Url: https://www. mrlab.co.kr/research/ridge
Author-written abstract from the author manuscript.
In plain language
What this work does
RIDGE first searches a motion-derived rule base enriched with language-model assistance. When a rule does not meet the confidence threshold, a contrastively trained text–motion embedding retrieves an appropriate recorded gesture. Both paths use existing animation segments; direct decoding into new motion frames is described as future work.
- 01Text query
- 02Rule match or learned fallback
- 03Recorded gesture clip
At a glance
Method, evidence, and scope

| Input | Text |
|---|---|
| Output | Retrieved recorded gesture clips |
| Method | Rule retrieval with confidence gating; contrastive text–motion retrieval fallback |
| Data and scope | BEAT co-speech motion and in-the-wild video data |
| Evaluation | Gesture Cluster Affinity: RIDGE 0.73, rule baseline 0.60, end-to-end baseline 0.52, ground truth 0.90 |
| Limitations | Direct latent-to-motion decoding is outside the study. Retrieved motion inherits source quality, including finger artifacts. |
Implementation and artifacts
Code and setup
Independent implementation of the paper’s core ideas, with setup instructions and data preparation documented in the repository README. The institute’s original source, datasets and trained models are not distributed.
Browse code and setup guideReference this work
Citation
Ghazanfar Ali, HwangYoun Kim, Jae‐In Hwang. RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation. Computer Animation and Virtual Worlds, 2025. Volume 36. Issue 4. Article e70034. DOI: 10.1002/cav.70034.
@article{ridge2025,
title = {{RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation}},
author = {Ali, Ghazanfar and Kim, HwangYoun and Hwang, Jae‐In},
year = {2025},
journal = {Computer Animation and Virtual Worlds},
volume = {36},
number = {4},
eid = {e70034},
doi = {10.1002/cav.70034},
url = {https://ghazanfarali.com/research/ridge/}
}