Non-profit · Open source · Est. 2023

Preserving and advancing
Korean NLP.

HAE-RAE is a research community working on the evaluation and interpretability of Korean and multilingual language models. We build the benchmarks the field lacks, and we release them openly.

Our name is the 해례 — the commentary bound with the Hunminjeongeum, which walks through the new Korean script letter by letter so that people can actually use it. The alphabet is the invention; the commentary is what makes it usable. Models are the invention. We write the commentary.

We believe in

01

Open source

Research should not remain inside a few labs. Open models, data, benchmarks, and tools allow more people to reproduce, challenge, and extend the work.

02

Diverse backgrounds

AI research should not be reserved for a narrow group of computer scientists. People from chemistry, law, medicine, and beyond bring questions and perspectives the field would otherwise miss.

03

Young talent

Starting young gives a researcher a longer horizon. Early access to real problems, tools, and responsibility creates more time to explore, take risks, and make lasting contributions.

What we work on

Drudgery for the community.

Someone has to build the benchmarks, run the evaluations, and check whether the numbers mean anything. We do that work, and we give it away.

01

Evaluation & benchmarks

HAE-RAE Bench, KMMLU, KMMLU-Pro, KMMMU, KAIO. What Korean language models know, how they reason, and whether they hold on to cultural grounding — measured task by task.

02

Multilingual reasoning

Why reasoning degrades outside English, and what closes the gap: language-mixed chain-of-thought, understand-solve-translate, language confusion control, and evaluation that reaches 100+ languages.

03

Multimodal & agents

Vision-language models facing under-specified queries, text-only VLM training, and browsing agents tested against real Korean contexts rather than translated ones.

04

LLM-as-a-judge

What automated evaluators and reward models can and cannot do — including the failures that stay invisible as long as you only ever check them in English.

Selected work

Recent papers

All publications →