CI-FSFD Benchmark Explorer

Interactive exploration of the benchmark from our IJCAI 2026 FINLLM paper.

What is CI-FSFD?

Company-Isolated Financial Statement Fraud Detection prevents data leakage by ensuring all data from the same company appears in either training or test set — never both. This reveals that prior random-split evaluations (up to 0.96 AUC) drastically overestimate real generalization (~0.70–0.74 AUC).

End-to-End Data Pipeline

📥Step 1

Data Collection

SEC filings (10-K/10-Q) via SEC-API, AAER enforcement releases, XBRL financial data

📝Step 2

MDA Extraction

Extract Management Discussion & Analysis sections (Item 7 for 10-K, Part 1 Item 2 for 10-Q)

📊Step 3

XBRL Financials

Parse US-GAAP 2024 taxonomy, extract 122 financial features per company/quarter

🏷️Step 4

Fraud Labeling

Link AAER enforcement actions to quarterly filings via CIK + fiscal quarter matching

🤖Step 5

MDA Summarization

Qwen3-32B extracts key insights per MD&A section (~3,800 tokens avg)

⚙️Step 6

Feature Engineering

Aggregate, diff, ratio features; Beneish M-score; Dechow accruals

🔄Step 7

Cross-Validation

5-fold stratified split by SIC industry + time period to prevent data leakage

🎯Step 8

Model Training

LoRA fine-tuning with softmax classifier, per-epoch threshold optimization

Pipeline Configuration

Taxonomy: US-GAAP 2024
Features extracted: 122 per quarter
Misstatement types: 11 (excl. Marketable Securities)
Fraud categories: 4 (Financial, Regulatory, Ethical, Market)
Summarizer: Qwen3-32B (open-source)
Max insights: 100 per section