Published · adjunct paper
RAG based AI-Agent for Contextualized Analysis of High-Density Historical Records: Application to the Annals of the Joseon Dynasty
2025 · IEEE ISMAR-Adjunct
An embodied agent makes historical source retrieval conversational.

From the paper
Author abstract
The field of digital heritage has increasingly focused on digitizing and reinterpreting historical materials to preserve and transmit cultural assets. Accurately conveying the content of historical records is critical for both academic research and education. Traditional digital archiving and retrieval systems have commonly relied on keyword-based searches or simple text matching, resulting in limited contextual understanding, insufficient source citation, and inadequate alignment with user intent. To address these challenges, this paper proposes a novel methodology that integrates a large language model (LLM) with a Retrieval-Augmented Generation (RAG) framework for high-density historical records such as the Annals of the Joseon Dynasty (49,646,667 characters, 13921910).Our system retrieves the most relevant historical sources, and delivers both objective facts and contextual analysis. Ultimately, we integrate this system with a Unity-based 3D AI-Agent, providing answers through synchronized voice output and body motion for an immersive interactive experience. By enabling natural, embodied interaction, this integration makes historical knowledge more accessible and engaging for users. Experimental results on a set of 30 benchmark questions demonstrate that our model outperforms both ChatGPT-4o and AI-Assistant v2 in Factual Accuracy, Reliability, and Reasonableness. This research expands the potential applications of digital heritage and lays the groundwork for broader integration with diverse historical data sources.
Author-written abstract from the author manuscript.
In plain language
What this work does
This adjunct paper integrates a retrieval-augmented language model with a Unity-based 3D agent. Questions produce grounded historical answers delivered through synchronized voice and body motion. It is a separate publication from the related journal article.
- 01User question
- 02Historical RAG
- 03Embodied answer
At a glance
Method, evidence, and scope

| Input | Voice or text questions about the Annals |
|---|---|
| Output | Historical answers with voice and contextual body animation |
| Method | Historical RAG pipeline integrated with a Unity agent |
| Data and scope | Annals corpus of 49,646,667 characters; 30 benchmark questions |
| Evaluation | Comparison on factual accuracy, reliability, and reasonableness against the paper’s GPT-4o and AI-Assistant v2 baselines |
| Limitations | A 30-question historical benchmark is limited in scope; the journal’s measurements should not be treated as measurements of this adjunct paper. |
Implementation and artifacts
Code and setup
Independent implementation of the paper’s core ideas, with setup instructions and data preparation documented in the repository README. The institute’s original source, datasets and trained models are not distributed.
Browse code and setup guideReference this work
Citation
Jeongha Lee, Ghazanfar Ali, Jae-In Hwang. RAG based AI-Agent for Contextualized Analysis of High-Density Historical Records: Application to the Annals of the Joseon Dynasty. IEEE ISMAR-Adjunct, 2025. Pages 893-894. DOI: 10.1109/ismar-adjunct68609.2025.00243.
@inproceedings{joseonragismar2025,
title = {{RAG based AI-Agent for Contextualized Analysis of High-Density Historical Records: Application to the Annals of the Joseon Dynasty}},
author = {Lee, Jeongha and Ali, Ghazanfar and Hwang, Jae-In},
year = {2025},
booktitle = {IEEE ISMAR-Adjunct},
pages = {893-894},
doi = {10.1109/ismar-adjunct68609.2025.00243},
url = {https://ghazanfarali.com/research/joseon-rag-ismar/}
}