GAGhazanfar Ali / Research

Published · journal article

RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation

Ghazanfar Ali · HwangYoun Kim · Jae‐In Hwang

2025 · Computer Animation and Virtual Worlds

First author

High-confidence rules and learned similarity retrieve semantically aligned gestures.

Research illustration for RIDGE
Research illustration

From the paper

Author abstract

Co-speech gestures are essential for natural human communication, yet existing synthesis methods fall short in delivering semantically aligned and contextually appropriate motions. In this paper, we present RIDGE, a hybrid system that combines rule-based and deep learning approaches to generate realistic gestures for virtual avatars and human-computer interaction. RIDGE employs a high-fidelity rule base generated from motion capture data with the assistance of large language models, to select reliable gesture mappings. When a high-confidence match is not available, a contrastively trained deep learning model steps in to produce semantically appropriate gestures. Evaluated using a novel Gesture Cluster Affinity (GCA) metric, our system outperforms existing baselines, achieving a GCA score of 0.73 compared to rule-based baseline 0.6 and end-toend: 0.52, while ground truth score was 0.90. Detailed analyses of system architecture, data preprocessing, and evaluation methodologies demonstrate RIDGE’s potential to enhance gesture synthesis. Project Url: https://www. mrlab.co.kr/research/ridge

Author-written abstract from the author manuscript.

In plain language

What this work does

RIDGE first searches a motion-derived rule base enriched with language-model assistance. When a rule does not meet the confidence threshold, a contrastively trained text–motion embedding retrieves an appropriate recorded gesture. Both paths use existing animation segments; direct decoding into new motion frames is described as future work.

  1. 01Text query
  2. 02Rule match or learned fallback
  3. 03Recorded gesture clip

At a glance

Method, evidence, and scope

RIDGE system architecture: rule-base construction, contrastive representation learning, and threshold-gated hybrid gesture retrieval
Graphical abstract diagram. Confidence-gated rules and learned similarity retrieve recorded gesture clips. View full size
Method and evidence for RIDGE
InputText
OutputRetrieved recorded gesture clips
MethodRule retrieval with confidence gating; contrastive text–motion retrieval fallback
Data and scopeBEAT co-speech motion and in-the-wild video data
EvaluationGesture Cluster Affinity: RIDGE 0.73, rule baseline 0.60, end-to-end baseline 0.52, ground truth 0.90
LimitationsDirect latent-to-motion decoding is outside the study. Retrieved motion inherits source quality, including finger artifacts.

Implementation and artifacts

Code and setup

Independent implementation of the paper’s core ideas, with setup instructions and data preparation documented in the repository README. The institute’s original source, datasets and trained models are not distributed.

Browse code and setup guide

Reference this work

Citation

Ghazanfar Ali, HwangYoun Kim, Jae‐In Hwang. RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation. Computer Animation and Virtual Worlds, 2025. Volume 36. Issue 4. Article e70034. DOI: 10.1002/cav.70034.

Download BibTeX
@article{ridge2025,
  title = {{RIDGE: Rule‐Infused Deep Learning for Realistic Co‐Speech Gesture Generation}},
  author = {Ali, Ghazanfar and Kim, HwangYoun and Hwang, Jae‐In},
  year = {2025},
  journal = {Computer Animation and Virtual Worlds},
  volume = {36},
  number = {4},
  eid = {e70034},
  doi = {10.1002/cav.70034},
  url = {https://ghazanfarali.com/research/ridge/}
}