CI-FSFD Benchmark Explorer
Interactive exploration of the benchmark from our IJCAI 2026 FINLLM paper.
What is CI-FSFD?
Company-Isolated Financial Statement Fraud Detection prevents data leakage by ensuring all data from the same company appears in either training or test set — never both. This reveals that prior random-split evaluations (up to 0.96 AUC) drastically overestimate real generalization (~0.70–0.74 AUC).
End-to-End Data Pipeline
Data Collection
SEC filings (10-K/10-Q) via SEC-API, AAER enforcement releases, XBRL financial data
MDA Extraction
Extract Management Discussion & Analysis sections (Item 7 for 10-K, Part 1 Item 2 for 10-Q)
XBRL Financials
Parse US-GAAP 2024 taxonomy, extract 122 financial features per company/quarter
Fraud Labeling
Link AAER enforcement actions to quarterly filings via CIK + fiscal quarter matching
MDA Summarization
Qwen3-32B extracts key insights per MD&A section (~3,800 tokens avg)
Feature Engineering
Aggregate, diff, ratio features; Beneish M-score; Dechow accruals
Cross-Validation
5-fold stratified split by SIC industry + time period to prevent data leakage
Model Training
LoRA fine-tuning with softmax classifier, per-epoch threshold optimization