同じテキスト、異なる数値:LLMベース指標の乖離
Same Text, Different Numbers: The Divergence of LLM-Based Measures (原題)
Hamid Boustanifar, Sasan Mansouri
🤖 gxceed AI 要約
日本語
7社のLLMでS&P500企業の決算説明会テキストを13指標(センチメント、経営陣の明確さ、不確実性、気候・政治リスク等)で採点し、モデル間の順位相関は平均0.52にとどまることを示した。モデル間不一致はアナリストや市場の不一致を予測せず、モデル固有の要素が大きい。モデル選択は回帰係数の符号・有意性を左右し、LLM生成変数はモデル依存的測定として複数プロバイダーで検証すべきと結論づける。
English
Seven LLMs from different providers scored S&P 500 earnings call transcripts on thirteen textual measures including climate and political risk. Cross-model rank correlations averaged only 0.52, and disagreement did not predict analyst or market disagreement, indicating a large model-specific component. Model choice altered coefficient signs and significance, so LLM-derived variables should be treated as model-contingent and validated across providers.
Unofficial AI-generated summary based on the public title and abstract. Not an official translation.
📝 gxceed 編集解説 — Why this matters
日本のGX文脈において
SSBJ・有報・統合報告書のテキスト分析にLLMを活用する動きが日本でも進む中、モデル選択が気候リスク開示の評価結果を大きく変えうることを示す。日本企業の開示評価やESGスコアリングにAIを用いる際の頑健性検証の必要性を裏付ける。
In the global GX context
As TCFD/ISSB/CSRD disclosure analysis increasingly relies on LLMs, this paper warns that model choice materially shifts measured climate-risk and sentiment variables. It sets a validation standard for AI-driven disclosure research and ESG scoring pipelines globally.
👥 読者別の含意
🔬研究者:LLMで構築したテキスト変数はモデル依存的であり、単一モデルでの推定結果の頑健性を必ず複数モデルで確認すべき。
🏢実務担当者:AIによる開示・ESG評価を導入する際は、使用モデルによって評価が変わりうるため、ベンダー選定と結果の検証体制が重要。
🏛政策担当者:AIを用いた開示モニタリングやESG格付規制を検討する際、モデル間の測定不一致を前提とした検証・透明性要件が必要。
📄 Abstract(原文)
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.
🔗 Provenance — このレコードを発見したソース
- semanticscholar https://www.semanticscholar.org/paper/4518d48f1ca525f4ef929ff168e79bd563a5341afirst seen 2026-09-29 05:38:22
🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。
gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。