Assistant Professor · AI Researcher · End-to-End Builder

I make machines behave.

Language is only one channel. I build intelligent systems that can speak, move, remember, adapt—and survive contact with real people.

Abstract winged signal creature representing Buraq orchestrating language, speech, face, and body motion across independent clients
System 01 · LiveBuraqA multimodal runtime with a name—and a nervous system.
01 / 26
17scholarly
outputs
8national R&D projects
participated in
10live research
demos
600+hmotion audited
in OmniMo
Descend

00 / Thesis

Intelligence is not the model.
It is the whole encounter.

I work from corpus to client: motion data, representations, retrieval and generation, APIs, real-time animation, embodied interfaces, and the studies that decide whether any of it matters.

The through-line is simple: make intelligent behavior scalable, controllable, grounded, and measurable.

01 / Now

The live research ecosystem.

Working platforms, active evidence programs, manuscripts, studies, and intentionally labeled research directions. No proposal is dressed up as a finished result.

Buraq visualized as parallel language, speech, facial, and motion signals converging through a live runtime A.01Working system

Live multimodal agent runtime

Buraq

ASR, LLM, speech, gesture, sessions and instrumentation—overlapped behind one documented contract.
  • 314 server tests
  • Studio · Unity · Web
  • Provider-driven
Different motion datasets passing through a crystalline calibration prism into one shared joint spaceA.02Working platform

Motion data infrastructure

OmniMo

Many skeleton dialects in. One auditable Shared Joint Space out.
  • 6+ corpora
  • 600+ hours
  • Continuously expanding
A five-stage MotionBridge pipeline from running video through pose and volumetric motion into tokens and a complete retargetable 3D skeletonA.03Research program

Video → retargetable 4D motion

MotionBridge

A consumer-accessible path from monocular video to clean, complete motion and a frozen token vocabulary.
  • 4-stage program
  • Hands + root
  • Economics-aware
A synthetic performer surrounded by cameras generating video, exact joints, masks, occlusions, and camera geometryA.04Specification

Motion → video + exact truth

MotionStage

A deliberately controllable renderer where ground truth—not photorealism—is the product.
  • Exact 2D/3D
  • Occlusion + masks
  • Free path
An hourglass measuring compute, labor, failures, and usable human motion framesA.05Evidence in progress

Motion-corpus economics review

What Does a Motion-Hour Cost?

A field-wide audit of the compute, yield, labor, licensing and reporting practices hidden behind motion datasets.
  • 375+ verified works
  • 343+ paper notes
  • Living evidence base
Three motion conditions flow through a glass evaluation arena with seven diagnostic dimensions and a temporal waveformA.06Benchmark active

Evaluation that audits itself

Motion EvalSuite + GQNet

From perceptual quality and semantic fit to foot sliding, latency, VRAM and whether a metric agrees with people.
  • 7 evaluation axes
  • ≈35 metrics planned
  • 6/7 baselines measured
A vast spiral library of curated gesture capsules searched by language and a motion-derived speaker-style signalA.07Manuscript · under review

Low-latency zero-shot style

ConGRets

Contrastive text-and-style retrieval from curated gesture units—speaker personalization without audio or frame-by-frame decoding.
  • <500 ms / 1,000 CPU queries*
  • <100 MB*
  • Motion-only style*
*Under-review manuscript results.
An abstract digital human surrounded by nine semantic controls for handedness, emotion, amplitude, speed, height, and prosodyA.08Design v1

Text + authored visual persona

ControllableGesture

Nine interpretable control signals at inference, without requiring audio.
  • 9 control axes
  • Text-only runtime
  • Study pending
Four behavioral reflections of one identical avatar with one subtle channel inconsistency under human observationA.09Proposal · v0.4

Identity–behavior incongruence

Behavioral Imposter Test

A signal-detection and survival-analysis instrument for when a digital persona stops feeling like itself.
  • Live interrogation
  • 4 behavior channels
  • No data collected yet
A continuous gesture ribbon crystallizing into variable-length motion tokens at meaningful movement phasesA.10Planned · gated

Token boundaries that mean something

Phase-Aligned Tokenizer

Test whether gesture-phase boundaries concentrate semantics better than arbitrary fixed windows.
  • 4 matched tokenizers
  • Variable rate
  • Segmenter gate first
Co-speech and action-language streams balanced through the same model to isolate semantic transferA.11Planned · dependent

What auxiliary data actually buys

Multi-Task Transfer

A matched-compute study of whether action language improves co-speech gesture meaning.
  • 5 controlled arms
  • Same backbone
  • Tokenizer-dependent
Text producing predicted duration, pitch, and energy signals before gesture, without synthesizing audio firstA.12Planned · can start

Middle-Out TTS

Prosody-Lite Conditioning

Price the perceptual value of duration, pitch and energy before audio exists.
  • 4 conditioning levels
  • Latency-value curve
  • Corpus phase first
A figure selecting future motion frames from semantic, kinematic, and persona signals while different skeletons retain one signatureA.13Research direction

On-manifold, interruption-ready motion

Semantic Motion Matching

Frame-level retrieval steered by kinematics, meaning and persona—plus skeleton-invariant matching.
  • ≈100 ms cadence proposed
  • No decoder at runtime
  • Not yet a completed system

02 / Published

Seventeen outputs.
Thirteen visual ideas.

Journal papers, conference papers, adjunct items, posters, and a live demonstration are grouped by the system they actually describe. Every bibliographic output remains visible below.

Translucent eyes and brows composed from speech, emotion, and liveness signals

2026 · IEEE Access + ISMAR-Adjunct 2025

Speech-conditioned upper-face animation

Emotion and liveness composed into a lightweight real-time signal.
Noisy multilingual pose and language signals aligning with a clean large 3D gesture library

2025 · IEEE Access · First author

Multilingual motion retrieval + GestureCLR

High-noise 2D pose aligned to 2,035 clean gesture units; cross-lingual interaction evaluated with 51 people.
A hybrid rule path and learned motion manifold converging into one semantically aligned gesture

2025 · CAVW · First author

RIDGE

High-confidence captured rules when possible; contrastive learning when necessary.
A long historical archive tunnel retrieving sourced evidence for an embodied guide

2025 · CAVW + ISMAR-Adjunct

Joseon Dynasty RAG agent

Dynamic article chunking, date-aware retrieval, source grounding, and an embodied voice-and-motion interface.
A screenplay ribbon unfolding into storyboard, 3D previz, actors, cameras, and an immersive scene

2025 journal · 2021 adjunct · 2022 live

ASAP: screenplay → many visual worlds

Storyboards, animated 3D previz, and immersive scenes generated from a screenplay.
A warm virtual physician using expression and gesture to explain a surgical model to a patient

2024 · CAVW · CASA Best Paper

Virtual physician

Grounded medical answers carried by carefully designed expression, gesture, and social presence.
A virtual person convincingly interacting with the silhouette geometry of real furniture through mobile AR

2021 · Applied Sciences

Silhouettes make mobile AR believable

Lightweight real-object geometry gives virtual humans occlusion and physical context.
Public video and spoken semantics growing into an automatically mined text-to-gesture rule garden

2020 · CAVW · First author

Automatic text-to-gesture rules

Mine useful behavior from 106 hours of public video instead of asking experts to author every map.
A plain scene transformed through a broad orbit of artist-level style examples rather than one copied painting

2026 · ECCV

Through Van Gogh’s Eyes

Global artist-level diffusion style transfer designed to learn a distribution, not repeat one iconic exemplar.
Noisy wild 2D pose traces snapping through a contrastive lens onto clean clustered 3D gestures

2022 · SIGGRAPH Asia Posters · First author

Wild pose matching with GestureCLR

Robust 2D-to-3D matching expands the gesture bank to 2,035 units and 210,000 rules.
Simple conversation blocks flowing into automatically generated digital-human speech, face, and gesture

2022 · SIGGRAPH Asia Posters

Flow Human

No-code conversation flows become verbal and nonverbal digital-human behavior.
One virtual human carrying out a sequence of grounded room actions inferred from natural language context

2021 · IEEE VR Workshops

Context → virtual-human action

Joint sentence and entity understanding turns conversation into grounded room behavior.
An intelligent virtual guide in a spatially mapped botanical garden connecting gaze, speech, objects, emotion, and gesture

2019 · CASA · First author

Wearable mixed-reality agent

The foundational system: place-aware, multimodal, embodied, and designed for real cultural spaces.

03 / Record

The complete output ledger.

Eight journal articles and nine conference, adjunct, poster, or live-demonstration items. Labels matter; not every item is presented as a full paper.

  1. Journal · IEEE AccessLightweight Speech-Conditioned Upper-Face Animation for Virtual Agents via Emotion–Liveness CompositionHwang Youn Kim, Ghazanfar Ali, Jeongha Lee, Jae-In Hwang
    PDF
  2. Conference · ECCVThrough Van Gogh’s Eyes: Global Style Transfer with Diffusion ModelJeongha Lee, Yujin Kim, Ghazanfar Ali, Suhyun Kim, Jae-In Hwang
    PDF
  3. Journal · IEEE Access · First authorExpanding Multilingual Co-Speech Interaction: The Impact of Enhanced Gesture Units in Text-to-Gesture Synthesis for Digital HumansGhazanfar Ali, Woojoo Kim, Muhammad Shahid Anwar, Jae-In Hwang, Ahyoung Choi
    PDF
  4. Journal · Computer Animation and Virtual Worlds · First authorRIDGE: Rule-Infused Deep Learning for Realistic Co-Speech Gesture GenerationGhazanfar Ali, Hwang Youn Kim, Jae-In Hwang
    PDF
  5. Journal · Computer Animation and Virtual WorldsA Retrieval-Augmented Generation System for Accurate and Contextual Historical Analysis: AI-Agent for the Annals of the Joseon DynastyJeong Ha Lee, Ghazanfar Ali, Jae-In Hwang
    PDF
  6. Journal · Multimedia Tools and Applications · Equal contributionASAP for Multi-Outputs: Auto-generating Storyboard and Pre-visualization with Virtual Actors Based on ScreenplayHanseob Kim, Ghazanfar Ali, Bin Han, Hwang Youn Kim, Jieun Kim, Hyemin Shin, Gerard Jounghyun Kim, Jae-In Hwang
    PDF
  7. Adjunct · IEEE ISMARRAG-based AI-Agent for Contextualized Analysis of High-Density Historical Records: Application to the Annals of the Joseon DynastyJeongha Lee, Ghazanfar Ali, Jae-In Hwang
    PDF
  8. Adjunct · IEEE ISMARLUFA: Lightweight Upper-Face Animation for VR/MR AvatarsHwang Youn Kim, Ghazanfar Ali, Jae-In Hwang
    VIEW
  9. Journal · Computer Animation and Virtual Worlds · CASA Best PaperEnhancing Doctor-Patient Communication in Surgical Explanations: Designing Effective Facial Expressions and Gestures for Animated Physician CharactersHwang Youn Kim, Ghazanfar Ali, Jae-In Hwang
    PDF
  10. Poster · SIGGRAPH Asia · First authorImproving Co-Speech Gesture Rule-Map Generation via Wild Pose Matching with Gesture UnitsGhazanfar Ali, Jae-In Hwang
    PDF
  11. Real-Time Live! · SIGGRAPH AsiaASAP: Auto-generating Storyboard and PrevizHanseob Kim, Ghazanfar Ali, Bin Han, Hwang Youn Kim, Jieun Kim, Jae-In Hwang
    VIEW
  12. Poster · SIGGRAPH AsiaNo-Code Digital Human for Conversational BehaviorHanseob Kim, Jieun Kim, Ghazanfar Ali, Jae-In Hwang
    PDF
  13. Journal · Applied SciencesSilhouettes from Real Objects Enable Realistic Interactions with a Virtual Human in Mobile Augmented RealityHanseob Kim, Ghazanfar Ali, Andreas Pastor, Myungho Lee, Gerard J. Kim, Jae-In Hwang
    PDF
  14. Adjunct · IEEE ISMARASAP: Auto-generating Storyboard and Previz with Virtual HumansHanseob Kim, Ghazanfar Ali, Jae-In Hwang
    VIEW
  15. Abstracts & Workshops · IEEE VRAuto-generating Virtual Human Behavior by Understanding User ContextsHanseob Kim, Ghazanfar Ali, Seungwon Kim, Gerard J. Kim, Jae-In Hwang
    PDF
  16. Journal · Computer Animation and Virtual Worlds · First authorAutomatic Text-to-Gesture Rule Generation for Embodied Conversational AgentsGhazanfar Ali, Myungho Lee, Jae-In Hwang
    PDF
  17. Conference · CASA · First authorDesign of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality EnvironmentGhazanfar Ali, Hong-Quan Le, Junho Kim, Seung-Won Hwang, Jae-In Hwang
    PDF

04 / Trajectory

Formation and
practice.

Four stages of education. Seven professional chapters. One continuous path from engineering and software delivery to embodied intelligence and research leadership.

EDU

Formation

Education

04
South Korea

Ph.D. in AI-Robotics

KIST School, University of Science and Technology

Scalable Hybrid Approach of Co-speech Text-to-Gesture Generation for Interactive Digital Humans.
Pakistan

B.E. in Computer Engineering

National University of Sciences and Technology

Final-year project: driving simulator for autonomous vehicles.
Pakistan

College · Pre-Engineering

Military College Jhelum

Pakistan

School Education

Army Public School & College, Kharian Cantt

EXP

Practice

Employment

07
Gachon University · South Korea

Assistant Professor

Department of AI

Leads multimodal-agent, human-motion, and persona research.
KAIST · South Korea

InnoCORE Postdoctoral Researcher

LLM Center

Developed the major components and control pipeline for ControllableGesture.
KIST · South Korea

Postdoctoral Researcher

Korea Institute of Science and Technology

Led personalized gesture, embodied RAG, digital-human systems, and evaluation research.
KIST · South Korea

Research Assistant

Korea Institute of Science and Technology

Built the research lineage from wearable agents to GestureCLR and multilingual interaction.
Penumbra Digital · Remote

Web Development Advisor & Team Lead

Technical strategy and full-cycle web delivery

Confidential B2B, corporate, NGO, and technology work across multiple regions.
University of Central Punjab · Pakistan

Lab Engineer

Computer-programming laboratories

Planned and supervised practical work supporting approximately 600 students.
Xetecx Solutions · Pakistan

Co-Founder & CTO

Created and managed software solutions and products.

Portrait of Ghazanfar AliSeoul · Republic of Korea

05 / The person in the loop

Ghazanfar Ali,
Ph.D.

Assistant Professor in the Department of AI at Gachon University. Research originator, system architect, and builder across multimodal agents, human motion, computer vision, grounded intelligence, and HCI.

His path runs from wearable mixed-reality agents and automatically mined gesture rules to contrastive motion representations, real-time personalized gesture, historical RAG, medical agents, and an increasingly rigorous motion-data and evaluation stack.

He builds the pieces that usually fall between papers: datasets, model interfaces, servers, protocol contracts, rendering clients, latency instruments, study logging, and the tests that keep the whole thing honest.

2023Ph.D. · UST, KIST School2024CASA Best Paper Award2026Assistant Professor · Gachon AI

06 / Quick answers

For humans and machines.

Direct answers for collaborators, search engines, and AI systems trying to understand the work without flattening it.

What does Ghazanfar Ali research?+

Multimodal intelligent agents, human motion and co-speech gesture, grounded LLM and RAG systems, real-time AI infrastructure, computer vision, and human-centered evaluation.

What is Buraq?+

Buraq is the live agent runtime formerly described generically as LiveAgent. It coordinates recognition, language models, speech, gesture, streaming, sessions, logging, and multiple independent clients behind explicit interfaces.

Are all active projects completed systems?+

No. Each card carries an explicit status. Buraq and OmniMo are working systems; several others are active research programs, evidence bases, study specifications, or deliberately early research directions.

What is the unifying research idea?+

Behavior is a systems problem. Data, representation, runtime, embodiment, and human evaluation have to work together—and the cost and failure modes of that full chain should be measured rather than hidden.

07 / Contact

Let’s build something
that can answer back.

[email protected]
Research specimen