← 論文一覧に戻る

検索器が見えないもの:ESG報告書検索における視覚表現ギャップの定量化

What the Retriever Cannot See: Quantifying the Visual Representation Gap in ESG-Report Retrieval (原題)

Ivan Gentile, Motaz Saad, Kianna Kazemi, Antonella Longo

Preprints.orgプレプリント2026-09-22#AI×ESG経営インパクト: 資金調達対象セクター: finance
DOI: 10.20944/preprints202609.1870.v1
原典: https://doi.org/10.20944/preprints202609.1870.v1

🤖 gxceed AI 要約

日本語

標準的なPDF→Markdown変換は埋め込み図表を捨て、RAGに「表現ギャップ」を生む。ESG報告書274件・画像48,896枚のうち21.7%がデータ保有画像でありながらテキスト検索では不可視だった。視覚言語モデルによる拡張で数値トークン+16.0%、語数+17.8%を回復し、図表依存クエリの検索完全性が+7.0ポイント改善した。

English

Standard PDF-to-Markdown pipelines silently discard embedded figures, creating a representation gap in RAG. Across 274 ESG reports, 21.7% of 48,896 images are data-bearing yet invisible to text-only retrieval. A vision-language augmentation stage recovers +16.0% numeric tokens and +17.8% words, improving retrieval completeness by +7.0/100 on figure-dependent queries.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

SSBJ基準・有価証券報告書のサステナビリティ開示が本格化する中、日本企業の統合報告書やESGデータ集は図表に重要KPIを多く含む。本手法は国内開示文書の検索性・比較可能性を高め、投資家のESG評価インフラ整備に直結する。

In the global GX context

As ISSB/CSRD disclosure scales, ESG reports remain figure-heavy and machine-unreadable, undermining automated analysis and assurance. This work reframes corpus-side enrichment—not generator choice—as the binding constraint on disclosure retrieval, offering a measurable ceiling for global ESG data pipelines.

👥 読者別の含意

🔬研究者:RAG評価において生成器選択よりコーパス側の視覚情報欠落が検索上限を規定することを示す実証的ベースラインを提供する。

🏢実務担当者:自社ESG報告書の図表が自動分析・ESG評価ツールで正しく読まれているかを点検し、開示データ設計を見直す根拠になる。

🏛政策担当者:XBRL等の構造化開示義務が図表内KPIを捕捉できているか、デジタル開示インフラの盲点として検討すべき。

📄 Abstract(原文)

Standard PDF-to-Markdown pipelines silently discard embedded figures, creating a representation gap in RAG that prior work has overlooked by focusing on generator choice. We quantify this gap in the ESG domain: on 274 reports, 21.7% of 48,896 images are data-bearing yet invisible to text-only retrieval. A vision–language augmentation stage (Step3-VL-10B) recovers this content, adding 116,673 numeric tokens (+16.0%) and 1.3M words (+17.8%) across all GRI topic-standard series with 0% truncation over 8,112 descriptions. LLM-as-judge evaluation achieves 99.8% KPI-domain coverage, with 77.9% of descriptions hitting two or more KPI buckets (vs. 44.9% keyword floor). A retrieval A/B on 83 reports shows +7.0/100 completeness improvement on figure-dependent queries. These results establish the representation gap as a measurable retrieval ceiling that corpus-side enrichment must first close.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。