← 論文一覧に戻る

FinTextSim: a domain-specific sentence-transformer for extracting predictive latent topics from financial disclosures

FinTextSim:財務開示から予測的な潜在トピックを抽出するためのドメイン固有文変換モデル (AI 翻訳)

Simon Jehnen, Javier Villalba-Díez, Joaquín B. Ordieres Meré

Frontiers Artif. Intell.📚 査読済 / ジャーナル2026-03-02#その他Origin: US
DOI: 10.3389/frai.2026.1752103
原典: https://doi.org/10.3389/frai.2026.1752103

🤖 gxceed AI 要約

日本語

本研究は、財務開示テキストから予測的なトピックを抽出するためのドメイン特化型文変換モデルFinTextSimを提案する。S&P500企業の10-K報告書(2016-2023)を用いて評価した結果、BERTopicとFinTextSimの組み合わせが最も明確で財務的に関連性の高いトピッククラスタを生成し、予測性能も向上した。

English

This study proposes FinTextSim, a domain-specific sentence-transformer for extracting predictive latent topics from financial disclosures. Using 10-K filings of S&P 500 companies (2016-2023), BERTopic with FinTextSim yields clearer, more coherent financial topic clusters and improves corporate performance prediction by 2 percentage points in ROC-AUC and F1-score over a purely financial baseline.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

日本企業の有報テキスト分析にも応用可能な手法だが、本論文はESG・気候関連開示を直接対象としていない。ただし、非構造化テキストを構造化表現に変換する枠組みは、SSBJ開示の分析に転用できる可能性がある。

In the global GX context

While not directly addressing climate or ESG disclosure, FinTextSim offers a methodology for converting unstructured financial text into structured representations that could be adapted for analyzing TCFD/ISSB reports. The domain-specific embedding approach demonstrates the value of fine-tuning on financial text, which may inform similar work in climate disclosure analysis.

👥 読者別の含意

🔬研究者:Researchers in NLP for finance can adopt FinTextSim for more coherent topic modeling and predictive feature extraction from disclosures.

🏢実務担当者:Corporate sustainability teams may find the methodology useful for extracting actionable insights from their own disclosures, but direct GX application requires further validation.

📄 抄録(日本語訳)

情報入手の進歩と計算能力の向上により、年次報告書の分析は変革を遂げ、従来の財務指標とテキストデータからの洞察を統合するようになった。この豊富なテキストデータから実用的な洞察を引き出すには、トピックモデリングなどの自動化されたレビュープロセスが不可欠である。本研究では、古典的なアプローチを現代のニューラル技術と比較評価し、財務テキスト用にファインチューニングされた文トランスフォーマーであるFinTextSimを紹介する。S&P 500企業(2016年~2023年)の10-K提出書類のItem 7およびItem 7Aを用いて、これらのモデルを質的および量的に系統的に評価する。BERTopicとFinTextSimの組み合わせは、他のすべての代替手法を一貫して上回り、著明に明確で、首尾一貫性があり、財務的に関連性の高いトピッククラスターを生成する。最も広く使用されている標準的な埋め込みモデルおよび財務ベースラインと比較して、FinTextSimはトピック内類似度を最大71%向上させ、トピック間類似度を108%以上低減し、ドメイン固有の埋め込みの重要性を浮き彫りにしている。決定的に重要なのは、これらの質的な利得が量的な予測上の利点に変換されることである。企業業績予測のためのロジスティック回帰フレームワークにFinTextSim由来のトピック特徴量を組み込むと、純粋な財務ベースラインと比較して、ROC-AUCおよびF1スコアの両方において統計的に有意な2パーセントポイントの向上がもたらされる。対照的に、既製の文トランスフォーマーおよび古典的なトピックモデルは、予測性能を低下させるノイズを導入する。非線形分類器では、いくつかのテキスト表現がわずかな利得をもたらし、これはノイズの多い特徴量を吸収するより大きな能力を反映している。しかし、FinTextSimは、線形および非線形の両方の設定において、最も安定して一貫した強力なパフォーマンスを発揮し続けている。全体として、FinTextSimはドメイン適応型情報フィルターとして機能し、構造化されていない財務テキストを、人間のアナリストや汎用モデルが見落としがちな、構造化された意味的に豊かな表現に変換する。解釈可能性と予測的有用性を橋渡しすることで、企業のナラティブから経済的に関連する情報の抽出を可能にし、より効果的な意思決定、リソース配分、および企業業績予測を支援する。

AI 翻訳(deepseek-v4-flash)。 正確を期す場合は下の原文を参照してください。

📄 Abstract(原文)

Recent advancements in information availability and computational capabilities have transformed the analysis of annual reports, integrating traditional financial metrics with insights from textual data. To extract actionable insights from this wealth of textual data, automated review processes, such as topic modeling, are essential. This study benchmarks classical approaches against contemporary neural techniques and introduces FinTextSim, a sentence-transformer finetuned for financial text. Using Item 7 and Item 7A of 10-K filings from S&P 500 companies (2016–2023), we systematically evaluate these models qualitatively and quantitatively. BERTopic in combination with FinTextSim consistently outperforms all alternatives, producing notably clearer, more coherent and financially relevant topic clusters. Compared to the most widely used standard embedding models and financial baselines, FinTextSim improves intratopic similarity by up to 71% and reduces intertopic similarity by more than 108%, highlighting the importance of domain-specific embeddings. Crucially, these qualitative gains translate into quantitative predictive benefits: incorporating FinTextSim-derived topic features into a logistic regression framework for corporate performance prediction leads to a statistically significant two-percentage-point increase in both ROC-AUC and F1-score over a purely financial baseline. In contrast, off-the-shelf sentence-transformers and classical topic models introduce noise that degrades predictive performance. For non-linear classifiers, several textual representations yield modest gains, reflecting their greater capacity to absorb noisier features. However, FinTextSim remains the most stable and consistently strong performer across both linear and non-linear settings. Overall, FinTextSim acts as a domain-adapted information filter, translating unstructured financial text into structured, semantically rich representations that human analysts and generic models often overlook. By bridging interpretability and predictive utility, it enables the extraction of economically relevant information from corporate narratives and supports more effective decision-making, resource allocation, and corporate performance forecasting.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。