Publications

Research papers in AI for auditing, financial NLP, and fraud detection.

IJCAI-ECAI 2026, FINLLM Workshop2026

Benchmarking Generalization in Financial Statement Fraud Detection: Robust Evaluation and Novel Tasks

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier

We address financial statement fraud detection (FSFD) by proposing a robust evaluation framework leveraging LLMs to integrate structured financial data and unstructured textual information (MD&A). We introduce the Company-Isolated FSFD (CI-FSFD) benchmark task for more realistic evaluation, constructing and releasing a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and AAER-derived fraud labels. We demonstrate that Fino1-8B with SMD&A text data achieves best performance (AUC 0.74) on CI-FSFD. Our results show that classic random splitting inflates performance (up to 0.96 AUC), while company-isolated evaluation drops dramatically (~0.70-0.74 AUC), revealing that previous work significantly overestimates generalization.

EMNLP 2026 Industry Track2026

ForensicBench: Evaluating Agentic LLMs on Journal-Entry Fraud Detection

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier

Accounting fraud costs organizations billions annually, yet auditors rarely confront it as a single suspicious posting. Instead, they investigate coordinated fraud schemes such as fictitious vendor payments, ghost-employee payroll, and inflated revenues that span many journal entries, each individually routine, and only become visible when understood against the underlying accounting process. Existing AI systems either score isolated entries as anomalies or mine company-level disclosures; none evaluate whether a large language model (LLM) agent can query a live ledger, recover coordinated fraud patterns, and assign the correct scheme type. We release ForensicBench: the first benchmark for evaluating agentic LLMs on scheme-level journal-entry fraud detection. The dataset, Forensic Ledger, extends DataSynth with five injected fraud scheme types and scheme-level labels linking journal entries to coordinated instances. We evaluate twelve open-weight models under a single reference agent scaffold on a public leaderboard: among them, the best reaches only 34.7% Entry-F1, and Type-F1 is up to 13 points lower. We publicly release the dataset, evaluation harness, and a reference agentic baseline.

AAAI 20252025

Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company Classification

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier

We explore the potential of machine learning algorithms and language models to analyze the relationship between industry sector categories and companies' financial statements. We propose a supervised company classification methodology analyzing several types of representations for financial statements. We show that textual information in financial records can be leveraged by language models to match decision tree-based classifier performance while providing better explainability. Our proposed Text-Numeric Transformer — a fusion of tag embeddings with amounts via a gating mechanism — achieves the best MCC of 0.71. LLM-gen (generative classification) with FinLLaMA3 achieves MCC 0.66 and provides explainable predictions, while LightGBM establishes strong baselines with MCC 0.69.