Is micro domain-adaptive pre-training for LLMs effective in real-world operations? Insights from a multi-step evaluation

Introduction
Large language models (LLMs) are increasingly being adopted in real-world enterprise operations. Many real-world applications require LLMs to work with proprietary knowledge in highly specialized and narrow fields — so-called micro domains.[1]
At Hitachi, we deliver solutions for mission-critical societal infrastructure, where accuracy, reliability and trustworthiness are essential. In such environments, standard LLMs often struggle to adapt to domain-specific knowledge, even as consistently high performance is required.
Domain-adaptive pre-training (DAPT) [2] is one approach to incorporate domain knowledge into LLMs. While DAPT has shown strong results in large domains such as healthcare and finance, its effectiveness in micro domains — where only limited training data is available — remains unclear.
In our recent paper [8], accepted for the EACL 2026 Industry Track, we address this question by examining not only final answer accuracy, but also where and why LLMs succeed or fail during the answer generation process.

