Benchmarking Generalization in Financial Statement Fraud Detection: Robust Evaluation and Novel Tasks
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
Abstract
We address financial statement fraud detection (FSFD) by proposing a robust evaluation framework leveraging LLMs to integrate structured financial data and unstructured textual information (MD&A). We introduce the Company-Isolated FSFD (CI-FSFD) benchmark task for more realistic evaluation, constructing and releasing a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and AAER-derived fraud labels. We demonstrate that Fino1-8B with SMD&A text data achieves best performance (AUC 0.74) on CI-FSFD. Our results show that classic random splitting inflates performance (up to 0.96 AUC), while company-isolated evaluation drops dramatically (~0.70-0.74 AUC), revealing that previous work significantly overestimates generalization.
Key Results
- Fino1-8B with SMD&A text achieves best CI-FSFD AUC of 0.74
- Random split inflates AUC to 0.96 — company-isolated evaluation is essential
- Text-only SMD&A outperforms combined FIN+SMD&A (noise bottleneck)
- Zero-shot LLMs perform at chance (~0.50 AUC)
Citation
@inproceedings{waffo2026benchmarking,
title={Benchmarking Generalization in Financial Statement Fraud Detection: Robust Evaluation and Novel Tasks},
author={Waffo Dzuyo, Guy Stephane and Guibon, Ga{"{e}}l and Cerisara, Christophe and Belmar-Letelier, Luis},
booktitle={IJCAI-ECAI 2026, FINLLM Workshop},
year={2026}
}