ELION Lab
Language Intelligence & Representation
Our research group explores the frontiers of natural language processing and AI systems, guided by the vision of Elucidating Language Intelligence & RepresentatiON (ELION), to understand how language models represent knowledge, reason, and generate meaning toward interpretable and controllable language intelligence for real-world interaction.
최신 뉴스 (Latest News)
Stay updated with our recent achievements and announcements.
- Sep 2026 1 paper is accepted at TKDD (ACM Transactions on Knowledge Discovery from Data).
- Sep 2026 1 paper is accepted at AACL-IJCNLP 2026.
- Sep 2026 한글 및 한국어 정보처리 학술대회(HCLT 2026)에 9편의 논문이 게재 승인되었습니다 (구두 발표 7편, 포스터 발표 2편).
- Aug 2026 5 papers are accepted at EMNLP 2026 (2 Main, 3 Findings).
- Jul 2026 Selected for AI StarFellowship 2026 (AI 최고급신진연구자지원), a 6-year project with total funding of KRW 11.0 billion (KRW 5.0 billion to Konkuk University). Serving as Project 1 Leader, in collaboration with Jeju National University (lead institution), NC AI, AIVIS, and Metsakuur Company.
- Jul 2026 Selected for the KETI AI Research Computing Support Project (H100 × 8).
- May 2026 Our paper "SERA: Self-referential Assessment Framework for Bidirectional Generative Commonsense Reasoning" has been accepted to Knowledge-Based Systems (to appear, 28 June 2026).
- May 2026 Our paper "No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand" has been selected for the ACL 2026 Best Paper Award Nomination (top 1%).
- Apr 2026 한국콘텐츠진흥원(KOCCA) 신규 과제에 선정되었습니다 — 「개인 맞춤형 국어생활종합상담 서비스를 위한 한국어 지식 연계 시스템 개발」 (2026~2028, 총 41억 원).
- Apr 2026 3 papers are accepted at ACL 2026.
- Mar 2026 2 papers are accepted at CVPR 2026.
- Mar 2026 Established the ELION Lab at Konkuk University
- Aug 2025 5 papers are accepted at EMNLP 2025.
모집 중! (NOW HIRING!)
Ongoing Projects
ELION Lab은 언어 모델의 한계를 넘어, 세상을 이해하고 행동하는 차세대 AI를 구축하기 위해 다음의 핵심 과제들을 수행하고 있어요.
1. Continual Representation Learning
언어 모델이 새로운 정보를 배울 때 기존 지식을 잊지 않으면서 계속 업데이트할 수 있는 방법을 연구하고 있어요. 나아가 Life2Vec 스타일로 확장하여, 다양한 유형의 데이터를 활용해 인간과 모델의 생애 전반에 걸친 사건과 리스크를 예측하는 지속학습 시스템을 만들고 있어요.
* Status: CVPR 2026 Highlight, EMNLP 2025, NAACL 2025 논문 게재
2. Reducing Hallucination & RAG
LLM이 사실과 다른 내용을 생성하는 환각(Hallucination) 문제를 해결하기 위한 연구를 하고 있어요. Calibration과 Uncertainty Estimation으로 모델이 얼마나 확신하는지 측정하고, 검색 증강(RAG) 기반의 Grounding 기술을 개발해 생성 결과의 신뢰성을 높이고 있어요.
* Project: IITP 과제 수행 중 (2024-2026, 생성형 AI 성과물의 신뢰성 및 일관성 연구)
3. Multi-Agent Systems & Orchestration
하나의 모델로는 어려운 문제를, 여러 AI 에이전트가 함께 풀 수 있는 멀티에이전트 시스템을 만들고 있어요. 효율적인 오케스트레이션으로 Planning과 온톨로지 탐색을 연결하고, 작은 모델로도 복잡한 작업을 해낼 수 있는 에이전트 시스템을 목표로 해요.
* Project: 한국콘텐츠진흥원(KOCCA) - 개인 맞춤형 국어생활종합상담 서비스를 위한 한국어 지식 연계 시스템 개발 과제 수행 중 (2026 ~ 2028)
4. Multimodal LM & World Understanding
텍스트뿐 아니라 시각 등 여러 모달리티를 함께 다루며, 현실 세계의 맥락을 이해하는 연구를 하고 있어요. Vision-Language 모델의 Dependency Parsing과 Semantic Chunking 성능을 끌어올려 멀티모달 문서 검색과 이해를 개선하고, AI가 실제 세상의 상식까지 추론할 수 있도록 해요.
* Collaboration: 고려대학교 Document AI & Multimodal 팀과 공동 연구 수행 중
5. World-Interactive Data Augmentation
웹 데이터가 점점 고갈되는 시대에, 현실 환경과의 상호작용을 통해 데이터를 늘리는 방법을 연구해요. 환경과 주고받는 피드백으로 고차원 추론 데이터와 다양한 언어 데이터를 스스로 만들어내는 데이터 엔진 기술을 개발하고 있어요.
* Collaboration: 싱가포르 A*STAR Research, Microsoft Research Asia (MSRA), IIT Delhi 협업 연구 수행 중 (ACL 2026, EMNLP 2026)
ELION Lab is dedicated to pushing the boundaries of language models to build next-generation AI that interacts with the real world.
1. Continual Representation Learning
We investigate methods for language models to continuously update and refine their internal knowledge without experiencing "catastrophic forgetting." Expanding this into Life2Vec-style research, we aim to build lifelong learning systems that utilize AnyType data to predict events and risks across both human and model lifecycles.
* Status: Published at CVPR 2026 (Highlight), EMNLP 2025, and NAACL 2025.
2. Reducing Hallucination & RAG
We explore cutting-edge methodologies to detect and mitigate the hallucination phenomenon in LLMs. Internally, we focus on measuring model confidence through calibration and uncertainty estimation. Externally, we strive to maximize the reliability and consistency of generated outputs by developing advanced grounding techniques within precise Retrieval-Augmented Generation (RAG) pipelines.
* Project: Supported by IITP (2024–2026, Research on the Reliability and Coherence of Generative AI Outcomes).
3. Multi-Agent Systems & Orchestration
Moving beyond the limitations of single models, we aim to develop practical and precise multi-agent systems. Through efficient orchestration, we integrate planning and ontology exploration to achieve high performance even with lightweight models, enabling them to execute complex, multi-step tasks.
* Project: Supported by KOCCA (Korea Creative Content Agency) — "Development of a Korean Knowledge-Linked System for Personalized Korean Language Consulting Services" (2026–2028).
4. Multimodal LM & World Understanding
We conduct research on integrating diverse modalities, such as vision, to understand physical and social contexts beyond text. By enhancing dependency parsing and semantic chunking in Vision-Language Models (VLMs), we innovate multimodal document retrieval and empower AI to reason with real-world common sense.
* Collaboration: Ongoing joint research with the Korea University Document AI & Multimodal Team.
5. World-Interactive Data Augmentation
To address the era of "data exhaustion" as online data becomes saturated, we explore world-interactive data augmentation. We research innovative data engines that autonomously generate high-level reasoning and multi-dimensional language data through direct feedback and interaction with real-world environments.
* Collaboration: Ongoing international collaboration with Singapore A*STAR Research, Microsoft Research Asia (MSRA), and IIT Delhi (ACL 2026, EMNLP 2026).
ELION Lab에서 인공지능 연구에 열정을 지닌 인턴, 석사, 박사과정 학생을 모집합니다
(We are now looking for talented M.S/Ph.D students, and research interns.)
주요 논문 (Featured Publications)
Selected recent papers from our research group.
🌟 Large Language Models Create Hallucinations in Response to Negated Text
ACM Transactions on Knowledge Discovery from Data (TKDD), 2026
🌟 Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations
International Joint Conference on Natural Language Processing and Asia-Pacific Chapter of the Association for Computational Linguistics (AACL-IJCNLP), 2026
🌟 The Geometry of Hallucination Detection: Why Simple Probes Are Sufficient
Empirical Methods in Natural Language Processing (EMNLP) Main, 2026
🌟 PILAR: A Page-Grounded Unified Evidence Representation via an Entity-Linked Assertion Graph for Open-Domain QA Agents over Multimodal Document Corpora
Empirical Methods in Natural Language Processing (EMNLP) Findings, 2026
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Empirical Methods in Natural Language Processing (EMNLP) Main, 2026
🌟 SERA: Self-referential Assessment Framework for Bidirectional Generative Commonsense Reasoning
Knowledge-Based Systems, Volume 345, 116152, 2026
🌟 No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
Annual Meeting of the Association for Computational Linguistics (ACL), 2026
🌟 HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
Annual Meeting of the Association for Computational Linguistics (ACL), 2026
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
Annual Meeting of the Association for Computational Linguistics (ACL), 2026
🌟 Evidential Transformation Network: Turning Pretrained Models into Evidential Models for Uncertainty Estimation
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
🌟 M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
🌟 The Impact of Negated Text on Hallucination with Large Language Models
Empirical Methods in Natural Language Processing (EMNLP), 2025
🌟 KoLEG: On-the-Fly Korean Legal Knowledge Editing with Continuous Retrieval
Empirical Methods in Natural Language Processing (EMNLP) Findings, 2025
🌟 MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
Empirical Methods in Natural Language Processing (EMNLP), 2025
🌟 Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
Empirical Methods in Natural Language Processing (EMNLP), 2025
🌟 K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language Models
International Conference on Learning Representations (ICLR), 2025
🌟 KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models
Annual Meeting of the Association for Computational Linguistics (ACL) Findings, 2024
🌟 CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients
Empirical Methods in Natural Language Processing (EMNLP), 2023
🌟 PU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledge
Knowledge-Based Systems, 2022
🌟 Plain Template Insertion: Korean-Prompt-Based Engineering for Few-Shot Learners
IEEE Access, 2022
🌟 A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation
North American Chapter of the ACL (NAACL) Findings, 2022
🌟 Dense-to-Question and Sparse-to-Answer: Hybrid Retriever System for Industrial Frequently Asked Questions
Mathematics, 2022
🌟 번역이 아니라 검색이 문제다: 쿼리-코퍼스 언어 불일치로 보는 검색 에이전트 격차
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 한국어 다중 에이전트의 KV 캐시 재사용: 엔트로피 게이트 판정의 언어 간 차이
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 코드 생성 언어모델의 지침 준수 실패: 어텐션 배분과 Value 경로의 분리
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 조사를 몇 개 떼어야 하는가? 한국어 서브워드 토큰화에서 부분 분리의 토큰 비용과 어간 일관성
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 검증 가능한 신호 기반 DPO를 통한 아첨 현상 완화의 한계
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 K-RLC: 다듬은 말의 장르별 사용과 거대 언어모델의 명칭 선택을 연결한 벤치마크
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 같은 페이지, 다른 근거: 시각 문서 질의응답에서 검색기 근거와 리더 어텐션 간의 비대칭성
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 다영역 글쓰기 자동 채점에서 점수 이산화가 모델 평가에 미치는 영향
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 효율적 한국어 정보검색을 위한 점수 마진 기반 선택적 재순위화
한글 및 한국어 정보처리 학술대회 (HCLT), 2026
🌟 Post-negation Text Induce New Hallucinations in Large Language Models
Annual Conference on Human and Cognitive Language Technology (HCLT), 2024
🌟 KommonGen: A Dataset for Korean Generative Commonsense Reasoning Evaluation
Annual Conference on Human and Cognitive Language Technology (HCLT), 2021