Back to publications
IJCAI-ECAI 2026, FINLLM Workshop2026

Benchmarking Generalization in Financial Statement Fraud Detection: Robust Evaluation and Novel Tasks

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier

Abstract

We address financial statement fraud detection (FSFD) by proposing a robust evaluation framework leveraging LLMs to integrate structured financial data and unstructured textual information (MD&A). We introduce the Company-Isolated FSFD (CI-FSFD) benchmark task for more realistic evaluation, constructing and releasing a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and AAER-derived fraud labels. We demonstrate that Fino1-8B with SMD&A text data achieves best performance (AUC 0.74) on CI-FSFD. Our results show that classic random splitting inflates performance (up to 0.96 AUC), while company-isolated evaluation drops dramatically (~0.70-0.74 AUC), revealing that previous work significantly overestimates generalization.

Key Results

  • Fino1-8B with SMD&A text achieves best CI-FSFD AUC of 0.74
  • Random split inflates AUC to 0.96 — company-isolated evaluation is essential
  • Text-only SMD&A outperforms combined FIN+SMD&A (noise bottleneck)
  • Zero-shot LLMs perform at chance (~0.50 AUC)

Citation

@inproceedings{waffo2026benchmarking,
  title={Benchmarking Generalization in Financial Statement Fraud Detection: Robust Evaluation and Novel Tasks},
  author={Waffo Dzuyo, Guy Stephane and Guibon, Ga{"{e}}l and Cerisara, Christophe and Belmar-Letelier, Luis},
  booktitle={IJCAI-ECAI 2026, FINLLM Workshop},
  year={2026}
}