DetRefiner: 変化する実世界環境におけるオープンボキャブラリ物体検知の信頼性向上

Why detection confidence matters in real-world AI
What happens when an AI system correctly detects an object but assigns it a confidence score too low to trust—or incorrectly gives a high score to a false detection? In industrial environments, such errors can lead to missed inspections, unnecessary interventions, or reduced confidence in AI-assisted operations. This blog discusses research that addresses a practical challenge in deploying visual AI: how to improve recognition reliability without rebuilding or retraining the underlying detection system.
Real-world environments add another layer of complexity. New, rare, or previously unseen objects can appear long after an AI system has been deployed. Object detection, a visual AI technology that identifies and localizes objects in images, is typically trained using a fixed set of categories. As a result, conventional detectors may struggle when encountering such objects in practice.
To address this limitation, researchers have developed open-vocabulary object detection (OVOD), which allows object categories to be specified using natural-language names rather than a fixed set of predefined classes. This enables AI systems to recognize a much wider range of objects, including categories not seen during training. Recent open-vocabulary and grounding-based object detection models, such as GLIP [2], Grounding DINO [3], MM-Grounding DINO [4], and LLMDet [5] have significantly expanded the ability to detect categories beyond those seen during training. However, an important challenge remains: the confidence score—how confident the model is that a detection is correct—is not always reliable. For example, a correct rare object may receive a low score, while a similar-looking false detection may receive a high one.
To address this issue, we propose DetRefiner [1], a lightweight, model-agnostic module, that improves these confidence scores after detection. An important aspect of this approach is that existing detection systems can remain unchanged. DetRefiner works as a post-processing module, refining confidence scores without modifying the detector itself. It does this by evaluating each detection result from two perspectives: whether the category makes sense in the overall scene, and whether the object inside the bounding box visually matches that category.

