Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks Guy Stephane Waffo Dzuyo1,2 , Gaël Guibon2,3 , Christophe Cerisara2 and Luis Belmar-Letelier1 1 Forvis Mazars 2 LORIA, CNRS, Université de Lorraine 3 Université Sorbonne Paris Nord, CNRS, Laboratoire d’Informatique de Paris Nord, LIPN, F-93430 Villetaneuse, France {guy.stephane.waffo, luis.belmar-letelier}@forvismazars.com, [email protected], [email protected]
arXiv:2607.19259v1 [cs.LG] 21 Jul 2026
Abstract Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and underutilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.
1
Introduction
The prevalence of financial statement fraud compromises the transparency and the integrity of financial markets and results in significant economic losses for stakeholders [Rezaee, 2005]. Traditional deterministic and statistical models [Altman, 1968; Beneish, 1999; Costa and Soares, 2022] are often insufficient against today’s complex and sophisticated fraudulent schemes, driving the need for more advanced detection approaches to identify them. Machine learning techniques, including tree-based models and deep learning, have shown promise in this domain by leveraging structured financial data [Craja et al., 2020; Ali et al., 2022]. However, financial reports contain rich, unstructured textual information, such as Management Discussion and Analysis (MD&A) sections, which often contain qualitative signals and narratives that complement numerical data, and can be indicative of fraudulent intent
or misrepresentation [Kirkos et al., 2024]. Recent advancements in Large Language Models (LLMs) have demonstrated their capacity to process, understand, and reason over complex information across various domains [Liu et al., 2025; Xu and Ding, 2025]. This capability presents a significant opportunity to leverage the textual components of financial reports more effectively [Wang and Brorsson, 2025] for tasks like fraud detection. However, applying LLMs to financial data, especially for fraud detection through anomaly detection, faces unique challenges [Li et al., 2023]. A primary limitation in current Financial Statement Fraud Detection (FSFD) research is the evaluation methodology itself. Many studies simply rely on random data splitting, which can inflate performance metrics by allowing models to learn company-specific patterns or exploit temporal dependencies present in the training data, thus failing to generalize to unseen companies or future periods [Wang et al., 2023], which is the main purpose of the FSFD task. This overestimation of predictive capabilities highlights the need for more robust and realistic evaluation frameworks that better reflect realworld deployment scenarios. To address these limitations and advance towards a more robust FSFD, we propose a novel framework grounded in two key hypotheses. Firstly, we hypothesize that (HYP1) generalization on companies in financial statement fraud detection is mandatory, which goes beyond the commonly used random train-test splitting. Indeed, realistic evaluation requires isolating models from specific company identities, which differs from what is currently standard practice. Secondly, we hypothesize that (HYP2) leveraging the rich textual information within financial reports, alongside traditional structured financial data, can improve fraud detection performance, particularly under these more challenging isolation conditions from HYP1. Driven by these hypotheses, we investigate three core research questions: (RQ1) How does evaluating FSFD performance under company isolation impact the model’s performance? (RQ2) To what extent does textual data from financial reports contribute to fraud detection, especially when structured data alone is less informative? (RQ3) Are LLMs capable of effectively detecting financial fraud over multimodal and multiformat financial data for robust fraud detection?
In this paper, we contribute as follows: A Novel Task for Realistic FSFD Evaluation. We introduce a novel task called Company-Isolated FSFD (CI-FSFD) based on professional expertise. This novel task provides a more realistic assessment of model generalization capabilities compared to the current standard practice limited to traditional random splitting. Multimodal Financial Dataset. We construct and make publicly available a comprehensive dataset of U.S. companies by integrating structured financial statement data with unstructured textual information from the Management Discussions and Analysis sections (MD&A). Crucially, we link these to fraud labels derived Accounting and Auditing Enforcement Releases (AAERs) from the U.S. Securities and Exchange Commission’s (SEC) through a two-stage, high-fidelity process. We detail our data extraction pipeline, including LLM-based summarization of MD&A and our manually-audited temporal linking of AAERs labels1 , which ensures the accuracy of our ground-truth labels. We also detail our preparation pipeline and temporal linking of AAERs. LLM-based Framework. We propose and implement a novel LLM-based framework capable of effectively processing and combining structured numerical features and Summarized text data from MD&A (SMD&A), for binary fraud classification. Benchmark-leading Metrics. We demonstrate that our LLM-based approach achieves the top performance on the novel CI-FSFD task, validating our hypotheses and highlighting the important role of both textual information and robust evaluation in FSFD. By introducing these challenging evaluation and benchmarks we demonstrate the power of LLMs on multimodal and multiformat financial data. Our work lays a foundation for developing more reliable and generalizable financial fraud detection systems. We believe it will help the community to tackle financial statement analysis2 .
2
Related Work
Financial fraud involves the intentional misrepresentation of financial information to deceive stakeholders or gain an unfair advantage [Rezaee, 2005]. This illicit activity can manifest in various forms, including revenue misstatement, asset misappropriation, and expense misstatement. The consequences of financial fraud are severe, leading to significant financial losses, legal repercussions, and reputational damage for companies and investors. While fraud can encompass issues such as disclosure violations, breaches of market regulations, bribery, and earning manipulation, the latter is most likely to be revealed through financial statements. Therefore, in this work, we define fraud as earning manipulations and the false reporting of any accounting information 1
https://www.sec.gov/enforcement-litigation/ accounting-auditing-enforcement-releases 2 https://github.com/WaguyMz/Financial-Statements-Fraud -Detection
intended to mislead investors, regulators, customers, or other parties [Healy and Palepu, 2003; Mishkin, 2011]. Early Statistical Techniques. Early attempts to detect financial fraud relied on statistical techniques based on financial ratios. They provide an estimate of the probability of fraud by analyzing various financial ratios and identifying patterns indicative of fraudulent behavior. The Altman ZScore focuses on bankruptcy prediction [Altman, 1968], the Beneish M-Score on earnings manipulation [Beneish, 1999], and the Jones model focuses on accruals [Costa and Soares, 2022]. In 2011, Dechow et al. [2011] proposed an efficient approach based on a logistic regression model and 7 features to predict earning mistatement. Machine Learning Approaches. The rise of machine learning in the 2000s spurred research into more advanced techniques for FSFD, such as Deep Neural Networks [Krizhevsky et al., 2012] and Random Forests [Breiman, 2001]. Later in the 2010s, the increased accessibility of powerful Natural Language Processing methods like LSTMs [Hochreiter and Schmidhuber, 1997] enabled researchers to leverage textual information within financial reports. For example, Craja et al. [2020] used the Management Discussion and Analysis (MD&A) sections of Form-10K3 , where executives explain financial performance, to enhance the performance of their classification model. In 2023, Wang et al. [2023] tackled financial statement fraud detection with a novel model, RCMA, emphasizing the importance of attentive mechanisms for distinguishing between modalities and coordinate financial ratios with textual data from financial reports. By addressing fusion ambiguity, their approach achieved strong fraud detection performance on CSMARD4 . The Emergence of Large Language Models (LLMs). LLMs [Radford et al., 2019; Touvron et al., 2023] offer new avenues for complex tasks like FSFD. Initial studies, such as Kirkos et al. [2024] using ChatGPT-4 on CEO letters and Kim et al. [2024] on general financial statement analysis, have highlighted LLMs’ potential in understanding financial narratives. However, these often rely on closed-source models, posing reproducibility challenges. Bhattacharya and Mickovic [2024] fine-tuned a BERT model using truncated MD&A sections, potentially missing key information. This underscores the need for LLM-based FSFD approaches that can utilize the extensive textual data in financial reports, a gap our work addresses. Frameworks for Financial Statement Fraud Detection (FSFD). Prevailing FSFD evaluation using random data splitting often yields overoptimistic performance, as models may learn company-specific or time-bound artifacts rather than generalizable fraud indicators, failing to reflect realworld deployment challenges. To address this, we introduce 3
Form 10-K is the comprehensive annual report that public companies file with the SEC. Form 10-Q is a quarterly report about the company’s financial performance during the quarter. 4 The China Stock Market & Accounting Research Database (CSMARD) offers data on the China stock markets and the financial statements of China’s listed companies.
more realistic evaluation via a novel task: Company-Isolated FSFD (CI-FSFD), evaluating generalization to unseen companies. To our knowledge, this is the first work to establish dedicated benchmark for that specific setting, aiming for more reliable assessments of FSFD systems. Open Data and Datasets. A significant challenge in FSFD research is the scarcity of readily available, open datasets. While U.S. AAERs [U.S. Securities and Exchange Commission, 2025] provide public fraud instances and China’s CSMARD [CSMAR Database, 2025] offers extensive data for Chinese markets, integrating these primary sources with structured financial statements (often in XBRL format5 ) and textual MD&A sections (both also available from the SEC) requires intensive, non-trivial preprocessing and accurate temporal linking. Curated datasets that perform this integration, such as those available through commercial providers like the Compustat database [S&P Global Market Intelligence, 2025], often come at a significant cost, limiting accessibility for widespread research. We address this gap by constructing and making publicly available a comprehensive FSFD dataset for U.S. companies, which combines financial statements, temporally linked AAERs, and processed MD&A text, and we detail our data collection and preprocessing pipeline.
3
Data Collection and Preprocessing
Our FSFD dataset integrates financial data, textual information from Form 10-Q’s MD&A sections, and fraud labels. This involved extracting, cleaning, and structuring these components from various sources, detailed below.
3.1
Financial Data
Quarterly financial data (Forms 10-Q) from 2009-2024 were sourced from the SEC website. Using the US-GAAP taxonomy, we mapped items to core accounts and, following Waffo Dzuyo et al. [2025], processed XBRL data to extract raw metrics (e.g., Total Revenue) and impute missing values. From these, we engineered 122 financial indicators (raw figures, change-based, ratios). For quality, reports with less than 25% of these features present were removed, resulting in 268,936 firm-quarter reports from 13,332 companies. Appendix A details the extraction and list all features.
3.2
Text Data
We collected 195,023 quarterly MD&A sections (Forms 10Q, 2009-2024) via a paid SEC-API6 . These raw HTML sections were lengthy and variable (1k-150k tokens, avg. 14k), making direct LLM processing computationally challenging. To make this text tractable, we summarized each section using the pretrained and open-source Qwen3 32B [Yang et al., 2025]. This step filters out non-material boilerplate legal language. Because forensic accounting anomalies (e.g., transaction misstatements) are embedded within hard factual disclosures rather than subtle linguistic style shifts, utilizing 5 XBRL (eXtensible Business Reporting Language): https:// www.xbrl.org/the-standard/what 6 https://sec-api.io/
Qwen3 32B distills core factual triggers while minimizing context distraction for the classifier. This yielded concise summaries averaging 3,800 tokens, forming our Summarized MD&A dataset, referred as SMD&A.
3.3
Fraud Dataset Preprocessing
Sourcing and aligning fraud labels is a critical step in constructing a robust FSFD dataset. Our fraud labels are derived from 3,300 AAERs obtained via the SEC-API. A significant challenge arises because the machine-readable JSON summaries for these releases lack the specific fiscal years and quarters of the violations, preventing a direct link to our quarterly financial data. To overcome this, we implemented a twostage pipeline to guarantee the accuracy of our ground-truth labels. Stage 1: Automated Extraction. First, we scraped the full, detailed legal documents linked within each AAER summary. We then leveraged the long-context capabilities of the Qwen3 32B model as a powerful parsing assistant. Using a structured prompt, we tasked the LLM with identifying and extracting a preliminary set of key information from each document: the fraudulent company or companies involved, a description of the fraudulent scheme, a list of fine-grained fraud categories, based on the 11 earning misstatement types proposed by Dechow et al. [2011], which we augmented with an additional Assets misstatement label (details in Appendix C). Stage 2: Manual Audit and Verification. Each of the 249 AAERs processed by the LLM was individually reviewed by one human expert in both Machine Learning and Auditing, who cross-referenced the extracted company, fiscal quarter(s), and fraud categories against the original legal source documents. This meticulous verification process confirmed that every extracted data point was correct, resulting in perfect accuracy for our fraud labels. This process yielded a set of 1,451 firm-quarter reports identified as fraudulent between 2000 and 2022. These verified instances form our core binary fraud labels which will serve to further construct the dataset. Additionally, our fraud dataset preprocessing involves extracting 12 fine-grained fraud labels. Although these labels could serve as valuable features for an advanced multi-label classification task, the scope of the current work is limited to binary classification to demonstrate the robustness of the novel CI-FSFD task.
3.4
Final Dataset Construction
The final dataset construction involved 3 key steps: merging the datasets, handling class imbalance, and ensuring temporal and company consistency. Merging Datasets. We merged the financial data, text data, and fraud labels. The financial and text data were aligned using company identifiers and fiscal quarters, creating distinct firm-quarter instances. The fraud labels, derived from AAERs, were linked to the financial and text data based on the extracted fiscal quarters. Critical Class Imbalance Handling. Financial fraud is an inherently rare event, leading to extreme class imbalance; in
our raw dataset, fraud cases are only about 0.03% of firmquarter observations. Training directly on such severe imbalance biases models towards the majority (non-fraud) class. Following rare-event ML paradigms, we target a stable 5% distribution. This preserves a realistic, severe class imbalance while ensuring gradient stability during training. Posthoc threshold calibration via validation F1-maximization ensures the model remains optimized for precision under these imbalanced constraints. This initial downsampling is done by preserving original industry and time distributions. Final Dataset Statistics. After merging and downsampling, the final dataset consists of 10,159 samples (511 fraud cases and 9,648 non-fraud cases). The distribution of samples across industries and time periods was maintained to ensure generalization and realistic evaluation.
4
Tasks Definition
Classic FSFD. In the common setting of binary fraud detection, the dataset usually consists of sets of firm-quarters observations either labelled as fraud or not. The dataset is split randomly into train and test sets, with the goal of predicting whether a given firm-quarter observation is fraudulent or not. CI-FSFD: Company Isolated FSFD. In the classic FSFD, random splitting of the dataset can lead to overfitting, as the model may learn to recognize specific patterns of individual companies. In contrast, our CI-FSFD task requires the model to generalize across different companies, ensuring that it can accurately identify frauds in firms it has never encountered before. This novel task is particularly relevant in real-world scenarios where models must be deployed to detect fraud in new companies.
5
Fraud Supervised Classification
Input Data. To train and evaluate our models, we explored three feature sets derived from our processed data. First, we used Financial Data Only (FIN), which comprises the 122 engineered financial indicators detailed in Appendix A. These indicators cover a range of metrics including raw figures, changebased values, financial ratios, and Beneish M-Score components. Second, we employed Text Data Only (SMD&A), consisting solely of the summarized quarterly MD&A sections. Finally, we utilized Combined Financial and Text Data (FIN+SMD&A) to leverage information from both sources. For this combined input, we serialized the 122 structured financial indicators into a key-value string (e.g., ”Total Revenue: 123456, Net Income: 7890, ...”), which was then directly concatenated with the SMD&A text. This straightforward fusion approach unified both modalities into a single text sequence for the model prompt (shown in Appendix F). Network Design. We propose a Large Language Model (LLM) based framework for financial fraud detection. Our primary model employs a pretrained LLM. The input to the LLM is a carefully constructed prompt that defines the binary fraud detection task. This prompt includes the relevant financial (FIN)
Figure 1: Financial Statement Fraud Classification.
and/or textual (SMD&A) data for a given firm-quarter, and is structured to elicit a classification response. The LLM is not fine-tuned on the autoregressive language modeling objective (i.e., predicting every subsequent token in the input sequence), but rather on predicting the final target token in the sequence, which represents the classification decision: either ”YES” (indicating fraud) or ”NO” (indicating non-fraud). Figure 1 shows an overview of our classification approach. The specifics of the fine-tuning methodology are detailed in section 6. Class Imbalance. Financial fraud is an inherently rare event, leading to highly imbalanced datasets. To mitigate the risk of models becoming biased towards the majority (non-fraud) class, we implement epoch-level undersampling during training. At each training epoch, we dynamically undersample the non-fraud cases from the training set to match the number of fraud instances. This prevents the model from overfitting to the majority class while still utilizing the full diversity of the majority class samples across different epochs.
6
Experiments and Results
This section details the experimental setup, baseline models, evaluation metrics, and the results obtained for the different FSFD tasks (Classic FSFD and CI-FSFD). We also present results for zero-shot performance of pretrained LLM on FSFD for further comparison.
Data Splitting. For both the Classic FSFD and CI-FSFD tasks, we employ a 5-fold cross-validation strategy on the 10,159 firm-quarter observations. In Classic FSFD setting, folds are created by randomly splitting these observations. For CI-FSFD, company isolation is enforced: all firm-quarter data from a specific company belong exclusively to either the training or test set within a fold. This company-based splitting also maintains the dataset’s original industry sector distribution and approximates a 5% fraud ratio across folds. Detailed information on these data splitting methodologies is provided in Appendix D.
6.1
Baseline Models
We compare our LLM-based approach against several established and contemporary baselines: Dechow Model. We include the logistic regression model proposed by Dechow et al. [2011], hereafter referred to as LR-DECHOW. It is reference logistic regression model, built on 7 financial features and widely recognized benchmark in prediction of earning misstatements. Its features are calculated from our 122 financial features. Appendix A.8 provides details on these features. Multi-Layer Perceptron (MLP). The MLP serves as a strong baseline for structured financial data. It is trained on the full set of 122 engineered financial indicators (FIN). Hyperparameters, including the number of layers and neurons per layer, are optimized using a Bayesian optimization approach via Hyperopt [Bergstra et al., 2013] to maximize the average AUC over the 5 validation sets. Tree-Based Ensemble Models. We benchmark our approach against tree-based ensemble models: Random Forest [Breiman, 2001], LightGBM [Ke et al., 2017], and XGBoost [Chen and Guestrin, 2016]. These methods are widely recognized for their robustness and efficacy in classification tasks, including financial fraud detection [Ashtiani and Raahemi, 2022]. Random Forest aggregates multiple decision trees to improve stability, while LightGBM and XGBoost, advanced gradient boosting frameworks, are known to offer optimized performance and scalability. We train them all using the 122 engineered financial indicators (FIN), with their respective hyperparameters tuned through Hyperopt as above. RCMA-adapted. We developed and benchmarked an adapted implementation of the Ratio-Chapter-ModalityAware (RCMA) model [Wang et al., 2023]. This adaptation was necessary due to two primary factors: the original model’s text subnetwork relies on legacy methods (Doc2Vec and LSTMs), and its source code and hyperparameters are not publicly available. Our principal modification was to replace those legacy methods by Jina Embedding V2Small [Nussbaum et al., 2025], a modern, open-source SentenceBERT model with a long-context architecture [Reimers and Gurevych, 2019], to align the model with current best practices. We fine-tuned this new component with LoRA [Hu et al., 2021]. For all other hyperparameters, we performed a grid search optimization to find the best configuration.
Crucially, the remainder of the RCMA architecture was replicated as faithfully as possible to Wang et al. [2023]’s description, especially its core modality-aware attention mechanisms. This ensures that our benchmark is a fair and up-todate representation of the RCMA design. More details are provided in Appendix E.
6.2
Experimental Setup
We fine-tune two foundation large language models: the general-purpose Llama-3.1 8B [Touvron et al., 2023] and the domain-specific Fino1-8B [Qian et al., 2025]. To ensure computational tractability, base models were loaded using 4bit quantization [Frantar et al., 2023; Zheng et al., 2024]. We utilized Low-Rank Adaptation (LoRA) [Hu et al., 2021] for parameter-efficient fine-tuning, applying adapters to all linear layers. Experiments were run on a single NVIDIA H100 GPU, requiring approximately 4 hours per fold. Hyperparameters are detailed in Appendix E.
6.3
Evaluation Methodology
Our primary evaluation metric is the ROC AUC score, a widely recognized standard for tasks with significant class imbalance [Fawcett, 2006]. We supplement this with standard classification metrics: precision, recall, and F1-score. Our model selection and calibration process follows a twostage approach on a validation set (10% of the training data). First, we select the model checkpoint that achieves the highest ROC AUC. Second, using this chosen model, we determine an optimal decision threshold by maximizing the F1-score on the same validation data. This final model is then used to generate predictions on the held-out test set.
6.4
Classic FSFD Results
The results for the Classic FSFD task, detailed in Table 1, show high performance across most models. The Llama3.1 8B configuration achieved the best AUC of 0.96, and nearly all approaches surpassed an AUC of 0.89, with the LRDECHOW model being the only exception. However, we argue that this high performance is more indicative of a methodological artifact than true generalization capability. The random splitting protocol results in significant data leakage, where the same companies appear in both training and evaluation sets. Our analysis confirms this issue: on average, each of the 321 fraudulent firms is present in 3.35 folds. This setup incites models to memorize companyspecific patterns instead of learning robust fraud signals. The strong results reported here and in the literature [Wang et al., 2023; Li et al., 2016] should be interpreted with caution, as they likely reflect this evaluation flaw. This observation responds to our research question (RQ1).
6.5
Company-Isolated FSFD Results
The company-isolated evaluation, presented in Table 2, provides a more rigorous test of model generalization by preventing data leakage. The dramatic drop in performance for all models validates our hypothesis (HYP1) that the Classic FSFD task is prone to optimistic bias. Against this challenging backdrop, a clear pattern emerges. The Fino1-8B model, when leveraging only narrative
Model
Input
AUC ± stdev
F1 ± stdev
Precision ± stdev
Recall ± stdev
LR-DECHOW MLP LightGBM XgBoost Random Forest RCMA-adapted
FIN FIN FIN FIN FIN FIN+SMD&A
0.68 ± 0.0272 0.89 ± 0.0098 0.95 ± 0.0098 0.96 ± 0.0108 0.92 ± 0.0119 0.89 ± 0.0081
0.15 ± 0.0135 0.40 ± 0.0827 0.74 ± 0.0355 0.76 ± 0.0930 0.54 ± 0.0245 0.39 ± 0.0598
0.09 ± 0.0121 0.30 ± 0.1080 0.84 ± 0.0541 0.84 ± 0.0604 0.52 ± 0.0443 0.28 ± 0.0705
0.53 ± 0.1530 0.70 ± 0.0800 0.66 ± 0.0451 0.69 ± 0.1240 0.57 ± 0.0366 0.71 ± 0.0572
Fino1 8B Fino1 8B Fino1-8B
FIN SMD&A FIN+SMD&A
0.90 ± 0.0195 0.95 ± 0.0178 0.94 ± 0.0094
0.46 ± 0.1032 0.66 ± 0.0818 0.60 ± 0.0920
0.38 ± 0.1564 0.56 ± 0.1049 0.50 ± 0.1169
0.70 ± 0.1162 0.83 ± 0.0584 0.80 ± 0.0573
Llama-3.1 8B Llama-3.1 8B Llama-3.1 8B
FIN SMD&A FIN+SMD&A
0.93 ± 0.0104 0.96 ± 0.0184 0.95 ± 0.0123
0.44 ± 0.0819 0.76 ± 0.0842 0.71 ± 0.0824
0.33 ± 0.0968 0.71 ± 0.1354 0.69 ± 0.1665
0.79 ± 0.0973 0.84 ± 0.0430 0.77 ± 0.0605
Table 1: Performance on the Classic FSFD task over 5 folds with standard deviation (stdev)
Model
Input
AUC
F1
Precision
Recall
∆AUC (p-value)
LR-DECHOW MLP LightGBM XgBoost Random Forest RCMA-adapted
FIN FIN FIN FIN FIN FIN+SMD&A
0.67 ± 0.04 0.69 ± 0.06 0.68 ± 0.01 0.66 ± 0.04 0.70 ± 0.03 0.65 ± 0.00
0.13 ± 0.02 0.14 ± 0.06 0.15 ± 0.03 0.13 ± 0.03 0.16 ± 0.03 0.14 ± 0.01
0.07 ± 0.01 0.10 ± 0.04 0.11 ± 0.03 0.08 ± 0.02 0.10 ± 0.02 0.08 ± 0.01
0.68 ± 0.08 0.39 ± 0.29 0.25 ± 0.08 0.38 ± 0.18 0.47 ± 0.18 0.71 ± 0.08
-0.074 (p=0.000) -0.058 (p=0.000) -0.085 (p=0.000) -0.080 (p=0.000) -0.042 (p=0.000) -0.134 (p=0.000)
Fino1 8B Fino1 8B Fino1 8B
FIN 0.69 ± 0.04 SMD&A 0.74 ± 0.03 FIN+SMD&A 0.72 ± 0.01
0.14 ± 0.03 0.18 ± 0.04 0.17 ± 0.03
0.10 ± 0.03 0.16 ± 0.03 0.12 ± 0.05
0.49 ± 0.29 0.23 ± 0.08 0.46 ± 0.18
-0.049 (p=0.000) (Reference) -0.026 (p=0.002)
Llama-3.1 8B Llama-3.1 8B Llama-3.1 8B
FIN 0.68 ± 0.04 SMD&A 0.68 ± 0.04 FIN+SMD&A 0.68 ± 0.04
0.12 ± 0.07 0.14 ± 0.01 0.13 ± 0.04
0.15 ± 0.08 0.09 ± 0.01 0.08 ± 0.01
0.24 ± 0.18 0.43 ± 0.08 0.35 ± 0.20
-0.067 (p=0.000) -0.066 (p=0.000) -0.067 (p=0.000)
Table 2: Performance on the Company-Isolated FSFD (CI-FSFD) task over 5 folds. Metrics are reported as mean ± standard deviation. The final column displays the results of a paired bootstrap test comparing each model against the top performer (Fino1 8B on SMD&A, in bold). This test reports the mean difference in AUC (∆ AUC) and the corresponding empirical p-value, calculated from 5,000 bootstrap iterations.
SMD&A data, significantly outperforms all other configurations, achieving a leading AUC of 0.74 and an F1-score of 0.18. This result also highlights the value of domain specialization, as Fino1-8B consistently surpassed the generalpurpose Llama-3.1 8B across all data modalities. Interestingly, this text-only model is more effective than the same LLM using financial data (AUC 0.69) or even the combination of both data types (AUC 0.72). The superiority of this approach over the best-performing classical model, Random Forest (AUC 0.70), further highlights the unique advantage of LLMs in this context.
6.6
Discussion of Results
Our LLM-based FSFD framework, particularly with summarized textual (SMD&A) data, yields significant insights, emphasizing the need for robust evaluation and thus validating our first hypothesis (HYP1). The substantial performance drop observed in Company-Isolated (CI-FSFD) scenarios vividly demonstrates how traditional random splitting inflates real-world generalization estimates. Notably, SMD&A text proved highly valuable, consistently boost-
Model
Input
AUC ± stdev
F1 ± stdev
Fino1 8B Fino1 8B Fino1 8B
FIN SMD&A FIN+SMD&A
0.48 ± 0.07 0.52 ± 0.05 0.51 ± 0.04
0.00 0.00 0.00
Llama-3.1 8B Llama-3.1 8B Llama-3.1 8B
FIN SMD&A FIN+SMD&A
0.49 ± 0.04 0.49 ± 0.04 0.52 ± 0.05
0.10 ± 0.01 0.10 ± 0.01 0.09 ± 0.00
Qwen3 32B Qwen3 32B Qwen3 32B
FIN SMD&A FIN+SMD&A
0.47 ± 0.04 0.48 ± 0.05 0.47 ± 0.04
0.02 ± 0.01 0.00 ± 0.01 0.02 ± 0.01
Table 3: Zero-shot FSFD performance (mean over 5 folds, with standard deviation, stdev). All models perform extremely poorly, needing finetuning. Threshold for computing F1 was set to 0.5.
ing model discrimination (higher AUC). Fino-1’s specialized financial fine-tuning enabled it to outperform the generalpurpose Llama-3.1 8B model across all input types in the challenging Company-Isolated Financial Statement Fraud Detection (CI-FSFD) task. Interestingly, the combined in-
Figure 3: Attn-LRP sentence-level relevancy. Red highlights mean positive contribution to Fraud prediction and blue ones mean negative contribution, with according intensity. Figure 2: Detection Performance per Misstatement type on the CIFSFD task (Fino1 8B with SMD&A input).The average AUC per mistatement is also reported above the bars.
put (FIN+SMD&A) underperformed compared to text alone, rejecting HYP2. We intentionally used simple serialization to establish a clean baseline; these results reveal a textual ”noise bottleneck,” proving that naive concatenation distracts the LLM and highlighting the need for future non-linear crossmodal gating structures. Further, our statistical analysis, using a bootstrap test with 1,000 iterations per fold, confirmed that this model significantly outperforms all others in terms of AUC score, with all observed empirical p-values being zero, except for one (Table 2). Finally, the consistently poor zeroshot LLM performance (AUC 0.50) confirms that LLMs’ pretrained alone are inadequate for fraud detection, highlighting the need for specialized FSFD fine-tuning. Fine-Grained Labels Analysis. To gain a deeper understanding of our model’s performance on different types of financial misstatements, we conducted a post-training analysis using the 12 fine-grained fraud categories extracted from the AAERs (details in Appendix C). Utilizing the predictions from our best-performing model (Fino1 8B with SMD&A input) on the CI-FSFD task, we computed the average AUC for each of these misstatement types. Figure 2 shows detection performance by category. ”Other Expense/Shareholder Equity Account” (AUC 0.64) and ”Revenue” (AUC 0.71) are the most frequent and relatively well-detected misstatement types. In contrast, ”Assets Valuation” recorded the lowest AUC (0.54), indicating it is particularly challenging to consider for fraud detection.
6.7
Explainability
In order to explain the LLM classification, we employ AttnLRP [Achtibat et al., 2024], a technique that calculates the relevancy of each input token to the LLM predictions using a gradient-perturbation of the input signal. For each sentence of the SMD&A document, we aggregate the relevance scores of all constituent tokens to derive a sentence-level relevance score. This approach allows us to identify and highlight key sentences that significantly influence the model’s decisions as shown in Figure 3. We acknowledge that those scores do not directly elicit explanations, but they can serve as clues, helping human experts to understand the model’s prediction.
7
Limitations
Our study, though it advances financial statement fraud detection (FSFD), has some limitations. First, The low F1-score (0.18) reflects the severe difficulty of cross-company generalization without identity leakage. However, this baseline is valuable to help human experts narrow down audit spaces, rather than acting as an automated judge. Second, while CI-FSFD eliminates company identity leakage, it does not strictly enforce chronological sequencing (e.g., historical-to-future splits). Merging company isolation with explicit rolling time windows is a crucial next trajectory for this benchmark to entirely prevent forward-looking bias. Third, our experiments are confined to the U.S. SEC dataset. To establish global generalizability, future work must extend this empirical evaluation to cross-country and multijurisdictional contexts using external databases such as CSMARD [CSMAR Database, 2025]. Finally, our approach assumes fraud signals reside primarily within factual disclosures, which may filter out subtle stylistic or linguistic anomalies. Future work should explore hybrid architectures that ingest original MD&A texts to capture a broader spectrum of behavioral fraud indicators.
8
Conclusion
Financial statement fraud detection is essential for market integrity, yet it faces significant challenges due to the sophistication of fraudulent schemes and subpar evaluation methods that often overestimate real-world performance. To address these issues, we introduced the novel Company-Isolated Financial Statement Fraud Detection (CI-FSFD) task that better evaluates models’ ability to generalize to unseen companies compared to standard practice. We created and publicly released a comprehensive dataset integrating structured financial data, summarized Management Discussion and Analysis (SMD&A) texts, and fraud labels derived from SEC AAERs. Our experiments showed that fine-tuned LLMs, particularly the specialized Fino-1 8B model using SMD&A data, outperformed other models. These results highlight the critical value of both structured and textual data in fraud detection and underscore the importance of robust evaluation frameworks and the poor zero-shot performance of LLMs emphasizes the necessity of task-specific fine-tuning. This work provides crucial benchmarks and resources, paving the way for more reliable fraud detection systems and evaluations.
References [Achtibat et al., 2024] Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. AttnLRP: Attention-aware layer-wise relevance propagation for transformers. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 135–168. PMLR, 21–27 Jul 2024. [Ali et al., 2022] Abdulalem Ali, Shukor Abd Razak, Siti Hajar Othman, Taiseer Abdalla Elfadil Eisa, Arafat Al-Dhaqm, Maged Nasser, Tusneem Elhassan, Hashim Elshafie, and Abdu Saif. Financial fraud detection based on machine learning: A systematic literature review. Applied Sciences, 12(19):9637, 2022. [Altman, 1968] Edward I. Altman. Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. The Journal of Finance, 23(4):589–609, 1968. [Ashtiani and Raahemi, 2022] Matin N. Ashtiani and Bijan Raahemi. Intelligent fraud detection in financial statements using machine learning and data mining: A systematic literature review. IEEE Access, 10:72504–72525, 2022. [Beneish, 1999] Messod D. Beneish. The detection of earnings manipulation. Financial Analysts Journal, 55(5):24– 36, 1999. [Bergstra et al., 2013] James Bergstra, Daniel Yamins, and David D Cox. Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures. In Proc. of the 30th International Conference on Machine Learning (ICML 2013), pages I–115– I–23, June 2013. [Bhattacharya and Mickovic, 2024] Indranil Bhattacharya and Ana Mickovic. Accounting fraud detection using contextual language learning. International Journal of Accounting Information Systems, 53:100682, 2024. [Breiman, 2001] Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001. [Chen and Guestrin, 2016] Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. CoRR, abs/1603.02754, 2016. [Costa and Soares, 2022] Cristiano Machado Costa and José Mauro Madeiros Velôso Soares. Standard jones and modified jones: An earnings management tutorial. Revista de Administração Contemporânea, 26(2):e200305, 2022. [Craja et al., 2020] Patricia Craja, Alisa Kim, and Stefan Lessmann. Deep learning for detecting financial statement fraud. Decision Support Systems, 139:113421, 2020. [CSMAR Database, 2025] CSMAR Database. China stock market & accounting research (csmar) database. https:// www.csmar.com/en/, 2025. Accessed: 2025-05-16.
[Dechow et al., 2011] Patricia M. Dechow, Weili Ge, Chad R. Larson, and Richard G. Sloan. Predicting material accounting misstatements*: Predicting material accounting misstatements. Contemporary Accounting Research, 28(1):17–82, 2011. [Fawcett, 2006] Tom Fawcett. An introduction to roc analysis. Pattern Recognition Letters, 27(8):861–874, 2006. [Frantar et al., 2023] Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023. [Healy and Palepu, 2003] Paul M. Healy and Krishna G. Palepu. The fall of enron. Journal of Economic Perspectives, 17(2):3–26, June 2003. [Hochreiter and Schmidhuber, 1997] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9:1735–80, 12 1997. [Hu et al., 2021] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. CoRR, abs/2106.09685, 2021. [Ke et al., 2017] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. Lightgbm: A highly efficient gradient boosting decision tree. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. [Kim et al., 2024] Alex G. Kim, Maximilian Muhn, and Valeri V. Nikolaev. Financial statement analysis with large language models, 2024. [Kirkos et al., 2024] Efstathios Kirkos, Georgia Boskou, Evrikleia Chatzipetrou, Eleftherios Tiakas, and Charalampos Spathis. Exploring the boundaries of financial statement fraud detection with large language models. SSRN Electronic Journal, 2024. [Krizhevsky et al., 2012] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, volume 25, 2012. [Li et al., 2016] Bin Li, Julia Yu, Jie Zhang, and Bin Ke. Detecting accounting frauds in publicly traded u.s. firms: A machine learning approach. In Geoffrey Holmes and TieYan Liu, editors, Asian Conference on Machine Learning, volume 45 of Proceedings of Machine Learning Research, pages 173–188, Hong Kong, 20–22 Nov 2016. PMLR. [Li et al., 2023] Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceedings of the Fourth ACM International Conference on AI in Finance, ICAIF ’23, page 374–382, New York, NY, USA, 2023. Association for Computing Machinery. [Liu et al., 2025] Shu Liu, Shangqing Zhao, Chenghao Jia, Xinlin Zhuang, Zhaoguang Long, Jie Zhou, Aimin Zhou,
Man Lan, and Yang Chong. FinDABench: Benchmarking financial data analysis ability of large language models. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 710–725, Abu Dhabi, UAE, January 2025. Association for Computational Linguistics. [Mishkin, 2011] Frederic S. Mishkin. Over the cliff: From the subprime to the global financial crisis. Journal of Economic Perspectives, 25(1):49–70, March 2011. [Nussbaum et al., 2025] Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. Nomic embed: Training a reproducible long context text embedder, 2025. [Qian et al., 2025] Lingfei Qian, Weipeng Zhou, Yan Wang, Xueqing Peng, Han Yi, Jimin Huang, Qianqian Xie, and Jianyun Nie. Fino1: On the transferability of reasoning enhanced llms to finance, 2025. [Radford et al., 2019] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. https://cdn.openai.com/better-language-models/ language models are unsupervised multitask learners. pdf, 2019. [Reimers and Gurevych, 2019] Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Conference on Empirical Methods in Natural Language Processing, 2019. [Rezaee, 2005] Zabihollah Rezaee. Causes, consequences, and deterence of financial statement fraud. Critical Perspectives on Accounting, 16(3):277–298, 2005. [S&P Global Market Intelligence, 2025] S&P Global Market Intelligence. Compustat via WRDS. https://wrds. wharton.upenn.edu/, 2025. Accessed: 2025-05-18. [Touvron et al., 2023] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. Arxiv 2302.13971, 2023. [U.S. Securities and Exchange Commission, 2025] U.S. Securities and Exchange Commission. Accounting and auditing enforcement releases, 2025. Accessed: 2025-05-16. [Waffo Dzuyo et al., 2025] Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, and Luis BelmarLetelier. Linking industry sectors and financial statements: A hybrid approach for company classification. Proceedings of the AAAI Conference on Artificial Intelligence, 39(16):16444–16452, Apr. 2025. [Wang and Brorsson, 2025] Xinlin Wang and Mats Brorsson. Can large language model analyze financial statements well? In Chung-Chi Chen, Antonio MorenoSandoval, Jimin Huang, Qianqian Xie, Sophia Ananiadou, and Hsin-Hsi Chen, editors, Proceedings of the Joint
Workshop of the 9th Financial Technology and Natural Language Processing (FinNLP), the 6th Financial Narrative Processing (FNP), and the 1st Workshop on Large Language Models for Finance and Legal (LLMFinLegal), pages 196–206, Abu Dhabi, UAE, January 2025. Association for Computational Linguistics. [Wang et al., 2023] Gang Wang, Jingling Ma, and Gang Chen. Attentive statement fraud detection: Distinguishing multimodal financial data with fine-grained attention. Decision Support Systems, 167:113913, 2023. [Xu and Ding, 2025] Ruiyao Xu and Kaize Ding. Large language models for anomaly and out-of-distribution detection: A survey. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics: NAACL 2025, pages 5992–6012, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. [Yang et al., 2025] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report, 2025. [Zheng et al., 2024] Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. Llamafactory: Unified efficient finetuning of 100+ language models, 2024.
A
Appendix A : Details on Financial Features
The financial data used in this study are derived from quarterly reports (Forms 10-Q and 10-K) sourced from the SEC, covering the period from 2009 to 2024. The process involved meticulous extraction, imputation, feature engineering, and quality control to construct a robust set of financial indicators for fraud detection.
A.1
Data Preparation Overview
Raw financial metrics were initially extracted by mapping reported items from company filings to the standardized US-GAAP (Generally Accepted Accounting Principles) taxonomy. This taxonomy provides a hierarchical structure for financial reporting elements. The quarterly financial datas are downloaded freely form the SEC -Website : https://www.sec.gov/data-research/ sec-markets-data/financial-statement-data-sets Taxonomy-based Data Imputation A significant challenge in processing financial statements is handling missing data. The hierarchical nature of the US-GAAP taxonomy was leveraged to impute missing values. For instance, if a parent account (e.g., Total Assets) is reported but some of its constituent child accounts are missing, their values can sometimes be inferred based on the reported parent value and other reported sibling accounts. This imputation helps in creating a more complete financial picture for each report. Figure 4 provides a simplified overview of the US-GAAP taxonomy structure. US GAAP Tree - Balance sheet tags Assets(1) Current Assets(2) Cash (3)
Non Current Assets(5)
Short-Term Investments(4)
Liabilities and Stockholder’s equity(6) Stockholder’s equity(7)
Liabilities(8)
Current Liabilities(9)
Non Current Liabilities(11)
Figure 4: Simplified overview of the US-GAAP taxonomy tree structure. The full taxonomy is extensive and can be explored via the FASB website (https://xbrlview.fasb.org/yeti/resources/yeti-gwt/Yeti.jsp). Only a few top-level balance sheet tags are presented for illustration.
Feature Engineering and Quality Control Following imputation, a comprehensive set of 122 financial indicators was engineered. These indicators are designed to capture a wide array of financial signals relevant to fraud detection. To ensure data quality, a cutoff filtering process was applied: reports with excessive missing information, specifically those where less than 25% of the 122 engineered features were present (i.e., non-zero and not NaN), were removed from the final dataset. The 122 engineered features are categorized into five groups as detailed below. Notation and Conventions: In the formulas presented, the subscript t denotes the current fiscal quarter, and t − 1 denotes the previous fiscal quarter. • ∆X = Xt − Xt−1 represents the change in feature X from the previous quarter to the current quarter. • Avg(X) = (Xt + Xt−1 )/2 represents the average value of feature X over the current and previous quarters. • Values for financial tags that are missing in a report are treated as 0. • Safe Division (safe divide(num, den)): If the denominator den is 0 or NaN, or if the numerator num is NaN, the result is 0. Otherwise, it is num / den. • Safe Summation (safe sum(args...)): If any of the arguments args is NaN or 0, the result is 0. Otherwise, it is the sum of the arguments. This specific behavior is adopted for consistency in calculations.
A.2
Basic Financial Numbers (43 features)
These features are core financial metrics extracted directly from financial statements, standardized according to the US-GAAP taxonomy. The 43 basic financial numbers include: 1. AccountsPayableCurrentAndNoncurrent 2. AccountsReceivableNetCurrent 3. AccountsReceivableNetNoncurrent 4. AccumulatedOtherComprehensiveIncomeLossNetOfTax 5. AdditionalPaidInCapital
6. AmortizationOfIntangibleAssets 7. Assets 8. AssetsCurrent 9. CashCashEquivalentsAndShortTermInvestments 10. CommonStockHeldBySubsidiary 11. CommonStockValue 12. CostOfRevenue 13. DebtCurrent (Short-Term Debt) 14. DeferredTaxAssetsDeferredIncome 15. DeferredTaxLiabilitiesDeferredExpense 16. DeferredTaxLiabilitiesTaxDeferredIncome 17. DepreciationAndAmortization 18. Goodwill 19. GrossProfit 20. IncomeLossFromContinuingOperations 21. IntangibleAssetsNetIncludingGoodwill 22. InterestAndDebtExpense 23. InventoryNet 24. Liabilities (Total Liabilities) 25. LiabilitiesCurrent 26. LongTermDebtCurrent 27. LongTermDebtNoncurrent 28. MinorityInterest 29. NetCashProvidedByUsedInFinancingActivities 30. NetCashProvidedByUsedInInvestingActivities 31. NetCashProvidedByUsedInOperatingActivities 32. NetIncomeLoss 33. OperatingExpenses 34. OperatingIncomeLoss 35. PreferredStockValue 36. PropertyPlantAndEquipmentNet 37. ReceivableFromShareholdersOrAffiliatesForIssuanceOfCapitalStock 38. RetainedEarningsAccumulatedDeficit 39. Revenues 40. SellingGeneralAndAdministrativeExpense 41. TemporaryEquityCarryingAmountIncludingPortionAttributableToNoncontrollingInterests 42. TreasuryStockValue 43. UnearnedESOPShares
A.3
Aggregated Measures (9 features)
These features are composite values derived by summing related basic financial numbers to represent broader financial concepts. 1. agg ACCOUNT RECEIVABLES: AccountsReceivableNetCurrentt + AccountsReceivableNetNoncurrentt 2. agg LONG TERM DEBT: LongTermDebtCurrentt + LongTermDebtNoncurrentt 3. agg EQUITY: safe sum of: |CommonStockValuet |, |PreferredStockValuet |, AdditionalPaidInCapitalt , RetainedEarningsAccumulatedDeficitt , AccumulatedOtherComprehensiveIncomeLossNetOfTaxt , −TreasuryStockValuet , −TemporaryEquityCarryingAmount...t , −ReceivableFromShareholders...t , −MinorityInterestt , UnearnedESOPSharest , CommonStockHeldBySubsidiaryt 4. agg TOTAL DEBT: DebtCurrentt + agg LONG TERM DEBT t 5. agg DEF TAX EXPENSE: DeferredTaxLiabilitiesDeferredExpenset − DeferredTaxAssetsDeferredIncomet 6. agg ACCRUALS: NetIncomeLosst − NetCashProvidedByUsedInOperatingActivitiest 7. agg EBIT: Revenuest − CostOfRevenuet − OperatingExpensest 8. agg EBITDA: agg EBIT t + DepreciationAndAmortizationt 9. agg NET CASH FLOW: NetCashProvidedByUsedInOperatingActivitiest +NetCashProvidedByUsedInFinancingActivitiest + NetCashProvidedByUsedInInvestingActivitiest
A.4
Change-based Measures (16 features)
This category includes 16 features designed to capture temporal changes in financial accounts and performance. These features are: 1. diff WC Accruals: ∆AssetsCurrent − ∆LiabilitiesCurrent − ∆CashCashEquivalentsAndShortTermInvestments 2. diff Inventories: safe divide(∆InventoryNet, Avg(Assets)) 3. diff Receivables: safe divide(∆agg ACCOUNT RECEIVABLES, Avg(Assets)) 4. diff CashSales: (safe divide(Revenuest , Avg(InventoryNet)) + safe divide(Revenuest−1 , Avg(InventoryNet))) /2 ∆agg ACCOUNT RECEIVABLES 5. diff CashMargin: safe divide(safe sum(CostOfRevenuet , −∆InventoryNet, ∆agg ACCOUNT RECEIVABLES), diff CashSalest ) 6. diff DefTaxExpense: safe divide(∆agg DEF TAX EXPENSE, Assetst−1 ) 7. diff Earnings: safe divide(∆NetIncomeLoss, Avg(Assets)) 8. diff AverageAssets: Avg(Assets) 9. diff Revenues: ∆Revenues 10. diff Cash: ∆CashCashEquivalentsAndShortTermInvestments
−
11. diff EBIT: ∆agg EBIT 12. diff EBITDA: ∆agg EBITDA 13. diff NetCashFlow: ∆agg NET CASH FLOW 14. diff Depreciation: ∆DepreciationAndAmortization 15. diff Assets: ∆Assets 16. diff Equity: ∆agg EQUITY
A.5
Ratio-based Measures (45 features)
These features are financial ratios calculated to assess profitability, liquidity, solvency, efficiency, and market valuation. Table 4 lists these ratios and their formulas. Table 4: List of 45 Ratio-based Measures
No.
Ratio Name (Feature ID)
Formula
1 2 3 4 5 6
ratio GrossProfitMargin ratio OperatingMargin ratio NetProfitMargin ratio EBITMargin ratio EBITDAMargin ratio CashFlowMargin
7 8 9 10 11
ratio ReturnOnAssets ratio ReturnOnEquity ratio CurrentRatio ratio QuickRatio ratio CashRatio
12 13 14 15 16 17 18 19 20 21 22 23 24 25
ratio WorkingCapitalToTotalAssets ratio DebtToAssetsRatio ratio DebtToEquityRatio ratio InterestCoverageRatio ratio TotalLiabilitiesToAssets ratio AssetTurnover ratio FixedAssetTurnover ratio ReceivablesTurnover ratio InventoryTurnover ratio SalesTurnover ratio EquityMultiplier ratio SGARatio ratio GoodwilltoAssets ratio CashFlowToDebtRatio
26
ratio CashFlowFinancingActivities
27
ratio CashFlowOperatingActivities
28 29
ratio EquityRatio ratio CashFlowToCurrentLiabilities
safe divide(GrossProfitt , Revenuest ) safe divide(OperatingIncomeLosst , Revenuest ) safe divide(NetIncomeLosst , Revenuest ) safe divide(agg EBIT t , Revenuest ) safe divide(agg EBITDAt , Revenuest ) safe divide(NetCashProvidedByUsedInOperatingActivitiest , Revenuest ) safe divide(NetIncomeLosst , Assetst ) safe divide(NetIncomeLosst , agg EQUITY t ) safe divide(AssetsCurrentt , LiabilitiesCurrentt ) safe divide(AssetsCurrentt −InventoryNett , LiabilitiesCurrentt ) safe divide(CashCashEquivalentsAndShortTermInvestmentst , LiabilitiesCurrentt ) safe divide(AssetsCurrentt − LiabilitiesCurrentt , Assetst ) safe divide(agg TOTAL DEBT t , Assetst ) safe divide(agg TOTAL DEBT t , agg EQUITY t ) safe divide(OperatingIncomeLosst , InterestAndDebtExpenset ) safe divide(Liabilitiest , Assetst ) safe divide(Revenuest , Avg(Assets)) safe divide(Revenuest , Avg(PropertyPlantAndEquipmentNet)) safe divide(Revenuest , agg ACCOUNT RECEIVABLESt ) safe divide(CostOfRevenuet , Avg(InventoryNet)) safe divide(Revenuest , Avg(InventoryNet)) safe divide(Assetst , agg EQUITY t ) safe divide(SellingGeneralAndAdministrativeExpenset , Revenuest ) safe divide(Goodwillt , Assetst ) safe divide(NetCashProvidedByUsedInOperatingActivitiest , agg TOTAL DEBT t ) safe divide(NetCashProvidedByUsedInFinancingActivitiest , agg NET CASH FLOW t ) safe divide(NetCashProvidedByUsedInOperatingActivitiest , agg NET CASH FLOW t ) safe divide(agg EQUITY t , Assetst ) safe divide(NetCashProvidedByUsedInOperatingActivitiest , LiabilitiesCurrentt )
Table 4: List of 45 Ratio-based Measures (Continued)
No.
Ratio Name (Feature ID)
Formula
safe divide(NetCashProvidedByUsedInOperatingActivitiest , Revenuest ) 31 ratio CashFlowCoverageRatio safe divide(NetCashProvidedByUsedInOperatingActivitiest , agg TOTAL DEBT t ) 32 ratio NetWorkingCapital AssetsCurrentt − LiabilitiesCurrentt safe divide(agg LONG TERM DEBT t , agg EQUITY t ) 33 ratio LongTermDebtToEquity 34 ratio DegreeOfFinancialLeverage safe divide(Revenuest , NetIncomeLosst ) 35 ratio InvestedCapitalRatio safe divide(PropertyPlantAndEquipmentNett + InventoryNett , Assetst ) 36 ratio CashToTotalAsset safe divide(CashCashEquivalentsAndShortTermInvestmentst , Assetst ) 37 ratio DebtServiceCoverage safe divide(OperatingIncomeLosst , agg TOTAL DEBT t ) 38 ratio FinancialLeverageIndex safe divide(OperatingIncomeLosst , Assetst ) 39 ratio TimesInterestEarnedRatio safe divide(NetIncomeLosst + InterestAndDebtExpenset , InterestAndDebtExpenset ) 40 ratio CurrentAssetToRevenues safe divide(AssetsCurrentt , Revenuest ) 41 ratio CurrentLiabilitiesToRevenues safe divide(LiabilitiesCurrentt , Revenuest ) 42 ratio ShortTermDebtToRevenue safe divide(DebtCurrentt , Revenuest ) 43 ratio IntangibleAssetToRevenue safe divide(IntangibleAssetsNetIncludingGoodwillt , Revenuest ) 44 ratio LongtermLeverage safe divide(agg LONG TERM DEBT t , Assetst ) safe divide(safe divide 45 ratio CFF (NetCashProvidedByUsedInFinancingActivitiest , agg NET CASH FLOW t ), Avg(Assets)) 30
ratio CashFlowToRevenue
A.6
Beneish M-Score Indicators (9 features)
These features are components of the Beneish M-Score model, designed to detect earnings manipulation. The individual indicators are: 1. Beneish DSRI (Days’ Sales in Receivables Index): agg ACCOUNT RECEIVABLESt agg ACCOUNT RECEIVABLESt−1 safe divide , Revenuest Revenuest−1 2. Beneish GMI (Gross Margin Index): Let GMt = safe divide(Revenuest − CostOfRevenuet , Revenuest ). safe divide(GMt−1 , GMt ).
Then Beneish GMI
=
3. Beneish AQI (Asset Quality Index): Let N CAt = Assetst − AssetsCurrentt − PropertyPlantAndEquipmentNett . Let AQt = safe divide(N CAt , Assetst ). Then Beneish AQI = safe divide(AQt , AQt−1 ). 4. Beneish SGI (Sales Growth Index): safe divide(Revenuest , Revenuest−1 ) 5. Beneish DEPI (Depreciation Index): Let DepRatet = safe divide(DepreciationAndAmortizationt , DepreciationAndAmortizationt + PropertyPlantAndEquipmentNett ). Then Beneish DEPI = safe divide(DepRatet−1 , DepRatet ). 6. Beneish SGAI (SG&A Index): t SellingGeneralAndAdministrativeExpenset−1 safe divide SellingGeneralAndAdministrativeExpense , Revenuest Revenuest−1 7. Beneish ACCRUALS (Total Accruals to Total Assets): safe divide(agg ACCRUALSt , Assetst ) 8. Beneish LVGI (Leverage Index): Let Levt = safe divide(agg LONG TERM DEBT t + DebtCurrentt , Assetst ). safe divide(Levt , Levt−1 ).
Then Beneish LVGI
9. Beneish PROBM (Beneish M-Score): This is the M-Score itself, calculated using the indicators above.
=
The Beneish PROBM (M-Score) is calculated as: M-Score = −4.84 + 0.920 × DSRI + 0.528 × GMI + 0.404 × AQI + 0.892 × SGI + 0.155 × DEPI − 0.172 × SGAI + 4.679 × ACCRUALS val − 0.327 × LVGI val Where DSRI, GMI, AQI, SGI, DEPI, SGAI are the values of the correspondingly named Beneish indicators (Beneish DSRI, Beneish GMI, etc.). ACCRUALS val is the value of the Beneish ACCRUALS feature, and LVGI val is the value of the Beneish LVGI feature.
A.7
Dataset Statistics
Descriptive statistics for the number of engineered features per firm-quarter report and reports per company (CIK) are presented in Table 5. The ”Number of Extended Features” refers to the count of non-zero values among all potentially derived financial features for a given report prior to final selection for the model, indicating the richness of available data per report. ”Important Tags” count refers to a predefined subset of raw US-GAAP tags deemed critical.
Table 5: Descriptive Statistics for Feature Counts per Firm-Quarter Report
Feature
Count
Mean
Median
Min
Max
Std Dev
Base numerical numbers (n important tags)*** Aggregated Measures (n aggregates) Change-base Features (n diff features)* Ratio-based Measures (n ratios) Beneish Features (n benish features)
43 9 16 45 9
20.60 1.35 11.21 19.72 2.62
21.0 1.0 11.0 21.0 2.0
11.0 0.0 8.0 1.0 0.0
33.0 7.0 15.0 39.0 9.0
4.01 1.36 0.77 6.88 1.57
All Features (n features)
122
55.50
56
20
93
12.35
*
Refers to a broader set of calculated differential values tracked during preprocessing, from which the 8 change-based measures listed earlier are a specific subset. *** Count of non-zero values for a predefined subset of 43 important US-GAAP tags.
A.8
Dechow Model Features
The Dechow model [Dechow et al., 2011], as implemented in our study, incorporates the following features, based on the implementation from https://github.com/jdonadio/FSFraud: • DECHOW RSST ACCRUALS (RSST Accrual): This measures discretionary accruals based on the Reverse-SalomonTeoh (RSST) model. It is calculated as: RSST Accruals =
∆W C + ∆N CO + ∆F IN Average Total Assets
where: – ∆W C is the change in working capital, calculated as: ∆W C = [∆Current Assets − ∆Cash and Short-term Investments] − [∆Current Liabilities − ∆Short-term Debt] – ∆N CO is the change in net non-current operating assets, calculated as: ∆N CO = ∆[Total Assets−Current Assets−Investments]−∆[Total Liabilities−Current Liabilities−Long-term Debt] – ∆F IN is the change in financing activities, calculated as: ∆F IN = ∆Short-term Investments − ∆[Long-term Debt + Short-term Debt + Preferred Stock] • DECHOW CH REC (Change in Receivables): This is the change in accounts receivable scaled by total assets, calculated as: ∆Accounts Receivable Change in Receivables = Average Total Assets • DECHOW CH INV (Change in Inventory): This is the change in inventory scaled by total assets, calculated as: Change in Inventory =
∆Inventory Average Total Assets
• DECHOW SOFT ASSETS (Soft Assets): This is the ratio of intangible assets and goodwill to total assets, indicating the proportion of ’soft’ or less tangible assets: Soft Assets =
Intangible Assets + Goodwill Total Assets
DECHOW CH CASHSALES (Change in Cash Sales): This is the change in cash sales, where cash sales are calculated as sales minus changes in accounts receivable: Change in Cash Sales = Sales − ∆Accounts Receivable • DECHOW CH ROA (Change in Return on Assets): This is the change in Return on Assets, indicating the trend in a company’s profitability relative to its assets: Change in ROA =
Earningst−1 Earningst − Average Total Assetst Average Total Assetst−1
• DECHOW ISSUANCE (Issuance): This feature indicates whether the company issued new shares or new debt in the current period. It is a binary variable:1, if equity or debt issuance occurred, 0 otherwise
B
Appendix B. Details on MD&A Reports Dataset
This appendix provides further details on the Management Discussion and Analysis (MD&A) sections used in our study. We elaborate on the characteristics of the raw MD&A data, the summarization process employed, the resultant Synthetically Summarized MD&A (SMD&A) dataset, and the associated costs for this data processing step.
B.1
Raw MD&A Sections
The raw MD&A sections were extracted using the API https://sec-api.io/. We subscribed for monthly plan which costs $55(https://sec-api.io/pricing and which is enough to download all the available quarterly MD&A sections. Specifically, we query the endpoint https://sec-api.io/docs/sec-filings-item-extraction-api to get the item part1item2 of Form-10Q which refers exactly to the desired MD&A sections. These sections are typically lengthy and contain a mix of textual narratives, financial figures, and sometimes tables, often embedded within HTML structures. Tokens’ count Distribution of Raw MD&A The distribution of token counts for the raw MD&A sections is depicted in Figure 5. These sections exhibit considerable variability in length.
Figure 5: Distribution of token counts in raw quarterly MD&A sections. The x-axis represents the number of tokens, and the y-axis represents the frequency.
Table 6 presents the descriptive statistics for the token counts of the raw MD&A sections in our dataset. Table 6: Descriptive statistics for token counts of raw MD&A sections.
Statistic Count Mean Standard Deviation Minimum 25th Percentile 50th Percentile 75th Percentile Maximum
Value 10,159 14,021 12,580 18 7,049 11,370 17,180 270,346
Sample Raw MD&A Below is an excerpt from a sample raw MD&A section, illustrating its typical structure and content. Note the presence of HTML entities. Sample Raw MD&A Excerpt Item 2. Management’s Discussion and Analysis of Financial Condition and Results of Operations
CAUTIONARY STATEMENT RELATING TO THE SAFE HARBOR PROVISIONS OF THE PRIVATE SECURITIES LITIGATION REFORM ACT OF 1995 This Quarterly Report contains forward-looking statements as that term is defined in the federal securities laws. The events described in forward-looking statements contained in this Quarterly Report may not occur. Generally, these statements relate to our business plans or strategies, projected or anticipated benefits or other consequences of our plans or strategies, financing plans, projected or anticipated benefits from acquisitions that we may make, or projections involving anticipated revenues, earnings or other aspects of our operating results or financial position, and the outcome of any contingencies. Any such forward-looking statements are based on current expectations, estimates and projections of management. We intend for these forward-looking statements to be covered by the safe-harbor provisions for forward-looking statements. Words such as \may," \will," \expect," \believe," \anticipate," \project," \plan," \intend," \estimate," and \continue," and their opposites and similar expressions are intended to identify forward-looking statements. We caution you that these statements are not guarantees of future performance or events and are subject to a number of uncertainties, risks and other influences, many of which are beyond our control that may influence the accuracy of the statements and the projections upon which the statements are based. Factors that could cause actual results to differ materially from those set forth or implied by any forward-looking statement include, but are not limited to, our ability to remain competitive with competitors, risks associated with the generic product industry, dependence on a limited number of suppliers, risks associated with healthcare reform and reductions in reimbursement rates, difficulty in predicting revenue stream and gross profit, industry and market changes, the effect of fluctuations in operating results on the trading price of our common stock, inventory levels, reliance on outside manufacturers, risks of incurring uninsured environmental and other industry specific liabilities, governmental approvals and regulations, risks associated with hazardous materials, potential violations of government regulations, product liability claims, reliance on Chinese suppliers, potential changes to Chinese laws and regulations, potential changes to laws governing our relationships in India , fluctuations in foreign currency exchange rates, tax assessments, changes in tax rules, global economic risks, risk of unsuccessful acquisitions, effect of acquisitions on earnings, indemnification liabilities, terrorist activities, reliance on key executives, litigation risks, volatility of the market price of our common stock, changes to estimates, judgments and assumptions used in preparing financial statements, failure to maintain effective internal controls, compliance with changing regulations , as well as other risks and uncertainties discussed in our reports filed with the Securities and Exchange Commission, including, but not limited to, our Annual Report on Form 10-K for the fiscal year ended June 30, 2011 and other filings. Copies of these filings are available at www.sec.gov. Any one or more of these uncertainties, risks and other influences could materially affect our results of operations and whether forward-looking statements made by us ultimately prove to be accurate. Our actual results, performance and achievements could differ materially from those expressed or implied in these forward-looking statements. We undertake no obligation to publicly update or revise any forward-looking statements, whether from
new information, future events or otherwise. NOTE REGARDING DOLLAR AMOUNTS In this quarterly report, all dollar amounts are expressed in thousands, except for per-share amounts. The following Management’s Discussion and Analysis of Financial Condition and Results of Operations (MD&A) is intended to provide the readers of our financial statements with a narrative discussion about our business. The MD&A is provided as a supplement to and should be read in conjunction with our financial statements and the accompanying notes. Executive Summary We are reporting net sales of $212,024 for the six months ended December 31, 2011, which represents a 22.3% increase from the $173,343 reported in the comparable prior period. Gross profit for the six months ended December 31, 2011 was $39,163 and our gross margin was 18.5% as compared to gross profit of $26,410 and gross margin of 15.2% in the comparable prior period. Our selling, general and administrative costs (SG&A) for the six months ended December 31, 2011 increased $6,073 to $27,097 from the amount we reported in the prior period. Our net income increased to $7,621, or $0.29 per diluted share, compared to net income of $1,628, or $0.06 per diluted share in the prior period.
Our financial position as of December 31, 2011 remains strong, as we had cash and cash equivalents and short-term investments of $28,700, working capital of $115,838 and shareholders’ equity of $161,571. Our business is separated into three principal segments: Health Sciences, Specialty Chemicals and Agricultural Protection Products. The Health Sciences segment is our largest segment in terms of both sales and gross profits. Products that fall within this segment include pharmaceutical intermediates, APIs, finished dosage form generic drugs and nutraceutical products. We typically partner with both customers and suppliers years in advance of a drug coming off patent to provide the generic equivalent. We believe we have a pipeline of new APIs poised to reach commercial levels over the coming years as the patents on existing drugs expire, both in the United States and in Europe. In addition, we continue to explore opportunities to provide a second-source option for existing generic drugs with approved abbreviated new drug applications (ANDAs). The opportunities that we are looking for are to supply the APIs for the more mature generic drugs where pricing has stabilized following the dramatic decreases in price that these drugs experienced after coming off patent. As is the case in the generic industry, the entrance into the market of other generic competition generally has a negative impact on the pricing of the affected products. By leveraging our worldwide sourcing, quality assurance and regulatory capabilities, we believe we can be an alternative economical, second-source provider of existing APIs to generic drug companies. On December 31, 2010, we acquired certain assets of Rising Pharmaceuticals, Inc. (\Rising") . We believe that the acquisition of Rising will establish another platform for our growth in our Health Sciences business by the expansion of our finished dosage form product offerings from both foreign and domestic
facilities as well as complementing our core strength of sourcing active pharmaceutical ingredients. The addition of Rising provides Aceto with a presence as a developer and marketer of our own brand of generic pharmaceuticals, the Rising brand. According to an IMS Health press release on May 18, 2011, "global spending for medicines will reach nearly $1.1 trillion by 2015, reflecting a slowing compound annual rate of growth of 3 { 6 percent over the next five years. This compares with 6.2 percent annual growth over the past five years. Lower levels of spending growth for medicines in the U.S., the ongoing impact of patent expiries in developed markets, continuing strong demand in pharmerging markets and policy-driven changes in several countries are among the key factors that will influence future growth, according to IMS Institutes new study, The Global Use Of Medicines Outlook Through 2015". Aceto supplies the raw materials used in the production of nutritional and packaged dietary supplements, including vitamins, amino acids, iron compounds and biochemicals used in pharmaceutical and nutritional preparations. Aceto’s identification of a change in the attitudes of Europeans towards nutritional products led to the decision to globalize this business and create an operating company to focus on it, Aceto Health Ingredients GmbH, headquartered in Germany. This globally structured business has become the model for all of our business segments, providing international reach and perspective for our customers. The Specialty Chemicals segment is a supplier to the many different industries that require outstanding performance from chemical raw materials and additives. Specialty Chemicals include a variety of chemicals which make plastics, surface coatings, textiles, fuels and lubricants perform to their designed capabilities. Dye and pigment intermediates are used in the color-producing industries such as textiles, inks, paper, and coatings. Many of our raw materials are also used in high-tech products like high-end electronic parts (circuit boards and computer chips) and binders for specialized rocket fuels. We continue to respond to the changing needs of our customers in the color producing industry by taking our resources and knowledge downstream as a supplier of select organic pigments. In addition, Aceto is a leader in the supply of diazos and couplers to the paper, film and electronics industries. According to a December 15, 2011 Federal Reserve Statistical Release, in the third quarter of calendar year 2011, the index for consumer durables, which impacts the Specialty Chemicals segment, grew at an annual rate of 11.4%. Item 3. Quantitative and Qualitative Disclosures About Market Risk Market risk is the risk of loss arising from adverse changes in market rates and prices, such as interest rates, foreign currency exchange rates and commodity prices. Our primary exposure to market risk is interest rate risk and foreign currency exchange rate risk. We do not use derivative financial instruments for trading or speculative purposes. We do not use any derivative contracts to hedge foreign currency or interest rate exposure. We seek to minimize foreign currency exchange rate risk through management of our current assets and liabilities which are denominated in foreign currencies. The principal foreign currencies to which we are exposed are the Euro, Indian Rupee and Chinese Yuan. Interest Rate Risk. Our interest expense is sensitive to changes in the general level of interest rates, as substantially all of our borrowings are at variable rates. Our exposure to interest rate risk relates primarily to our Amended and Restated Credit Agreement, as amended (the “Credit Agreement”). As of December 31, 2011, we had 0outstandingunderourrevolvingcreditf acilityand5,000 outstanding under our term loan facility. Each of these facilities bears interest at a variable rate based on LIBOR or the Base Rate (as defined in the Credit Agreement). A 100 basis point
increase in interest rates would increase our interest expense by $50 annually. Foreign Currency Exchange Rate Risk. We are exposed to foreign currency exchange rate risk related to our purchases and sales denominated in foreign currencies. Our foreign currency exchange rate risk is inherent in the sales and expenses of our foreign subsidiaries, which are denominated in their respective local currencies. The financial statements of our foreign subsidiaries are translated into U.S. dollars at exchange rates in effect at the balance sheet date for assets and liabilities and average exchange rates during the period for revenues and expenses. As a result, changes in exchange rates may affect the reported value of our foreign assets, liabilities, revenues and expenses, and could result in foreign currency translation gains or losses in our consolidated statements of operations. A hypothetical 10and Chinese Yuan would not have a material effect on our results of operations.
B.2
MD&A Summarization Process
To make the extensive textual data from MD&A sections more manageable for LLM processing while retaining core financial insights, we employed a summarization strategy using Qwen3 -32B model. This model was chosen for its long-context capabilities and efficiency. System Prompt for Summarization The following system prompt was used to guide the Qwen3 -32B model in summarizing the raw MD&A sections. The ‘quarter info‘ placeholder was dynamically filled with the specific quarter and year of the report (e.g., ”Q4 2023”). System Prompt for MD&A Summarization
You are a highly skilled financial analyst with deep expertise in summarizing corporate disclosures.\\ You will be provided with the ’Management’s Discussion and Analysis’ (MD\&A) section of a financial report for {quarter\_info}.\\ Your task is to summarize it following the instructions below: Extract and present the **distinct, and factual insights**, along with subjective statements, management commentary, and qualitative explanations, organized into the following sections: --** 1. Strategic Priorities and Initiatives**\\ { Summarize key strategies, corporate objectives, growth plans, restructuring efforts, and major initiatives discussed by management.\\ { Capture significant strategic shifts, operational transformations, ambitious targets, or business model changes.\\ { Highlight subjective language, including optimistic tone, vague descriptions of progress, or assertions lacking clear supporting evidence. ** 2. Operational and Segment Performance**\\ { Summarize operational results and segment-level performance, including production metrics, KPIs, challenges, and improvements.\\ { Pay special attention to:\\ { Unexplained variances in performance.\\ { Misalignment between narrative explanations and operational metrics.\\ { Subjective, vague, or generic explanations (e.g., \seasonality," \market dynamics," \operational excellence") without adequate quantification.\\ { Unusual operational trends, sales fluctuations, production shifts, or inventory movements. ** 3. Financial Results and Key Trends**\\ { Capture **all financial metrics**, including revenue, profitability, margins, cost drivers, liquidity trends, capital structure, and debt along with financial ratios.\\ { If the metrics are presented in tables, rewrite them in the section \Important Figures and Tables" instead of here.\\ { Also Include commentary on:\\
{ Revenue recognition patterns or timing shifts.\\ { Significant margin changes or cost structure shifts.\\ { Increases in accounts receivable, inventory, or other working capital components relative to sales without clear justification.\\ { Use of non-recurring items, adjustments, or changes in estimates that materially impact results.\\ { Use of non-recurring items, adjustments, or changes in estimates that materially impact results.\\ { Subjective rationalizations for financial outcomes (e.g., references to \strong demand" or \improved efficiencies") that lack numeric validation. ** 4. Identified Risks and Uncertainties**\\ { Summarize disclosed risks, including operational, supply chain, regulatory, competitive, legal, and macroeconomic risks.\\ { Capture both concrete risks and:\\ { Subjective assessments of risk severity.\\ { Ambiguous or hedged language (e.g., \may," \could," \uncertain").\\ { Shifts in tone, emphasis, or presentation of risks compared to prior periods. ** 5. Forward-Looking Statements and Guidance**\\ { Capture management’s expectations, forecasts, assumptions, and outlook for future periods.\\ { Highlight:\\ { Changes in guidance or underlying assumptions.\\ { Optimistic tone, hedging, or caveats (e.g., \expects," \believes," \aims").\\ { Whether forward-looking statements are grounded in quantifiable drivers or rely mainly on qualitative assertions. ** 6. Significant Changes, Events, or Developments**\\ { Summarize material recent or upcoming events affecting the business, such as mergers, acquisitions, divestitures, leadership changes, legal proceedings, regulatory actions, or external shocks.\\ { Note how management frames these events|whether impacts are clearly quantified or described with vague or qualitative language. ** 7. Important Figures and Tables**\\ { Extract key figures, tables, or financial data that are critical to understanding the MD\&A.\\ { For each table:\\ Recreate the exact table content in clean markdown format preceded by the table title. ** 8. Management Explanations and Justifications**\\ { Capture how management explains or justifies operational and financial results, risks, or variances.\\ { Pay attention to:\\ { Vague, broad, or overly generic justifications.\\ { Repetitive use of boilerplate terms (e.g., \market conditions," \execution excellence") without specific detail.\\ { Narratives that shift accountability to external factors or uncontrollable circumstances without precise quantification. ** 9. Accounting Estimates, Judgments, and Policy Changes**\\ { Summarize any disclosures related to:\\ { Changes in accounting policies, methodologies, or estimates.\\ { Adjustments to key assumptions (e.g., impairments, allowances, revenue recognition).\\ { Areas where significant management judgment materially affects reported results.\\ { Note whether explanations are clear, detailed, vague, hedged, or superficial. ** 10. Capital Allocation and Liquidity Management**\\ { Summarize commentary on:\\ { Cash management strategies, liquidity preservation, and debt management.\\ { Capital expenditures, share repurchases, dividend policies, and
financing activities.\\ { Highlight any:\\ { Indications of liquidity stress.\\ { Mismatches between optimistic narratives and defensive liquidity actions (e.g., drawing on credit lines despite claimed strong financial performance). ** 11. Legal, Regulatory, and Compliance Matters**\\ { Summarize discussions related to:\\ { Ongoing or pending litigation.\\ { Regulatory investigations or changes.\\ { Compliance risks, including ESG-related disclosures that have material financial implications.\\ { Note whether these issues are presented transparently, minimized, or framed with ambiguous language. --**Formatting Instructions:**\\ { Use section headers exactly as written above ie with the numbers and titles.\\ { Present each point as a bullet ({) under the appropriate section.\\ { Include both objective data and subjective commentary.\\ { Explicitly note subjective explanations, optimistic framing, hedging, or vague descriptions wherever they appear.\\ { Be precise and factual but the summary should be detailed\\ { Avoid redundancy; each bullet must convey a distinct, meaningful insight. **Critical Constraint:**\\ { Base the summary **strictly on the content explicitly stated in the MD\&A.**\\ { Do not include any external knowledge, assumptions, interpretations, or analysis beyond the document provided.\\ { Only output the summary | do not include any commentary, explanations, or meta-text.
B.3
Summarized MD&A (SMD&A) Sections
The summarization process resulted in the SMD&A dataset, consisting of condensed versions of the original MD&A narratives. Tokens’ count Distribution of SMD&A Figure 6 illustrates the token count distribution for the SMD&A sections. As intended, these summaries are substantially shorter than the raw MD&A sections.
Figure 6: Distribution of token counts in Summarized MD&A (SMD&A) sections. The x-axis represents the number of tokens, and the y-axis represents the frequency.
Table 7 provides descriptive statistics for the token counts of the SMD&A sections.
Table 7: Descriptive statistics for token counts of SMD&A sections.
Statistic
Value
Count Mean Standard Deviation Minimum 25th Percentile 50th Percentile 75th Percentile Maximum
10,159 3,807 1,250 1 2,960 3,666 4,476 7,914
Sample SMD&A An excerpt from a sample SMD&A is provided below. This illustrates the more structured and condensed format achieved through the summarization process. Sample SSMD&A Excerpt **1. Strategic Priorities and Initiatives** - The acquisition of St. Jude Medical, Inc. (St. Jude Medical) was completed on January 4, 2017, to expand Abbott’s presence in the cardiovascular and neuromodulation markets. - Abbott is reshaping its business portfolio through divestitures, including the sale of its vision care business (AMO) to Johnson & Johnson for $4.325 billion in cash. - The company is pursuing the acquisition of Alere Inc. (Alere) to expand its global diagnostics presence. The purchase price was reduced from $56.00 to $51.00 per share in April 2017, with the acquisition expected to close by the end of Q3 2017, subject to regulatory approvals. - Abbott is implementing cost improvement initiatives across various functions and businesses, partially offsetting increased expenses from the St. Jude Medical acquisition. - The company is investing in research and development (R&D), with R&D expenses increasing significantly due to the integration of the St. Jude Medical business. - Abbott has a share repurchase program authorized by its board in 2014 for up to $3.0 billion, in addition to $512 million remaining from a prior program. - Abbott increased its quarterly dividend by approximately 2% in 2017 compared to 2016, indicating a focus on shareholder returns. **2. Operational and Segment Performance** - Net sales for the Cardiovascular and Neuromodulation Products segment increased by 198.5% in the first six months of 2017 due to the St. Jude Medical acquisition. Excluding the acquisition and foreign exchange, sales in this segment decreased by 1.5% as lower coronary stent sales and a favorable 2016 royalty agreement resolution were partially offset by higher Structural Heart and endovascular sales. - Sales in the Established Pharmaceutical Products segment increased by 5.5% in the first six months of 2017. Excluding foreign exchange, sales in Key Emerging Markets increased 8.2% in the first half of 2017, driven by growth in Russia, China, and Latin America, partially offset by the impact of a new GST system in India. - Nutritional Products sales decreased slightly by 0.3% in the first six months of 2017, with International Pediatric Nutritionals declining by 8.0%. Challenging conditions in the Chinese infant formula market continued to impact international performance. - U.S. Pediatric Nutritionals increased by 7.7%, driven by momentum from recently launched infant formula products and growth in the PediaSure toddler brand.
- International Adult Nutritionals increased by 3.3% compared to the first half of 2016, while U.S. Adult Nutritionals decreased by 4.5% due to competitive and market dynamics. - Diagnostic Products sales increased by 3.7% in the first six months of 2017, with a 5.1% increase excluding foreign exchange, driven by share gains in Core Laboratory and Point of Care markets in the U.S. and higher international sales. - The Other category decreased by 26.0% in the first six months of 2017, reflecting the sale of AMO partially offset by double-digit growth in Abbott’s Diabetes Care business. - The decrease in Other Emerging Markets by 6.1% in the first six months of 2017 is attributed to the unfavorable impact of Venezuelan operations. Excluding Venezuela and foreign exchange, sales in Other Emerging Markets increased 4.1%. **3. Financial Results and Key Trends** - For the three months ended June 30, 2017, total net sales were $6,637 million, a 24.4% increase from $5,333 million in the same period in 2016, with a 25.3% increase excluding foreign exchange. - For the six months ended June 30, 2017, total net sales were $12,972 million, a 27.0% increase from $10,218 million in the same period in 2016, with a 27.7% increase excluding foreign exchange. - U.S. sales increased by 42.5% in the second quarter of 2017 and by 47.0% in the first six months of 2017. - International sales increased by 16.3% in the second quarter of 2017 and by 17.9% in the first six months of 2017. - Gross profit margin decreased from 54.4% in the second quarter of 2016 to 46.3% in the second quarter of 2017, and from 53.8% to 45.0% for the first six months of 2017, primarily due to higher intangible amortization and inventory step-up amortization from the St. Jude Medical acquisition. - R&D expenses increased by $165 million in the second quarter of 2017 and by $333 million in the first six months of 2017, driven by the addition of St. Jude Medical. - Selling, general, and administrative (SG&A) expenses increased by 22.7% in the second quarter and 32.6% in the first six months of 2017, primarily due to the St. Jude Medical acquisition and integration costs, partially offset by cost improvement initiatives. - Interest expense (income), net increased by $100 million in the second quarter and $279 million in the first six months of 2017 compared to 2016, due to the $15.1 billion in debt issued in November 2016 to finance the St. Jude Medical acquisition. - Taxes on earnings from continuing operations in the first six months of 2017 included $430 million of tax expense related to the gain on the sale of the AMO business. - Earnings from discontinued operations, net of tax, were $46 million in the first six months of 2017, primarily reflecting net tax benefits from the resolution of tax positions related to AbbVie’s operations prior to the 2013 separation. **4. Identified Risks and Uncertainties** - Abbott operates in highly competitive and regulated markets, with ongoing debate over healthcare product availability, delivery, and payment methods, which could adversely affect its operations. - The company faces risks related to the integration of the St. Jude Medical acquisition, including potential challenges in combining operations and achieving expected synergies. - Foreign exchange fluctuations, particularly in Venezuela, pose a risk to financial reporting and cash flows. - Regulatory and legal risks are present, including the FDA warning letter related to the Sylmar, CA manufacturing facility acquired from St. Jude Medical. - The acquisition of Alere is subject to regulatory approvals and
potential antitrust concerns, with the FTC and European Commission reviewing the transaction. - Alere is divesting certain businesses in connection with the regulatory review, including the Triage MeterPro and B-type Natriuretic Peptide assay businesses to Quidel Corporation and the subsidiary Epocal Inc. to Siemens Diagnostics Holding II B.V. **5. Forward-Looking Statements and Guidance** - Abbott expects to complete the acquisition of Alere by the end of the third quarter of 2017, subject to customary closing conditions and regulatory approvals, which are now due by September 30, 2017. - The company expects to maintain an investment-grade debt rating. - Abbott expects to fund cash dividends, capital expenditures, and other investments with cash flow from operations, cash on hand, short-term investments, and borrowings. - Abbott expects to use the modified retrospective method to adopt the new revenue recognition standard (ASU 2014-09) and does not expect it to have a material impact on its consolidated financial statements. - The company expects to evaluate the impact of recently issued accounting standards, including ASU 2017-07, ASU 2016-16, and ASU 2016-02, on its consolidated financial statements. **6. Significant Changes, Events, or Developments** - Abbott completed the acquisition of St. Jude Medical on January 4, 2017, for $23.6 billion, including $13.6 billion in cash and $10 billion in Abbott common shares. - Abbott sold its AMO segment to Johnson & Johnson for $4.325 billion in cash, completed on February 27, 2017, and recognized a pre-tax gain of $1.151 billion. - Abbott sold 50 million ordinary shares of Mylan N.V. in the first six months of 2017, generating approximately $1.9 billion in proceeds, reducing its ownership interest from 14% to 3.7%. - Abbott received a warning letter from the FDA in April 2017 regarding its Sylmar, CA manufacturing facility, which is part of the St. Jude Medical acquisition. - Abbott entered into a $2.8 billion term loan agreement in July 2017 to fund the acquisition of Alere. - Abbott commenced a tender offer to purchase its outstanding shares of Aleres Series B Convertible Perpetual Preferred Stock at $402 per share, subject to conditions.
**7. Important Figures and Tables** **Net Sales to External Customers (in millions)** **Three Months Ended June 30:** | Segment | 2017 | 2016 | Total Change | Impact of Foreign Exchange | Total Change Excl. Foreign Exchan |---------|------|------|---------------|-----------------------------|-------------------------------| Established Pharmaceutical Products | $1,021 | $981 | 4.1% | 0.6% | 3.5% | | Nutritional Products | $1,731 | $1,740 | (0.6)% | (1.1)% | 0.5% | | Diagnostic Products | $1,273 | $1,226 | 3.8% | (1.6)% | 5.4% | | Cardiovascular and Neuromodulation Products | $2,260 | $758 | 198.2% | (1.1)% | 199.3% | | Other | $352 | $628 | (44.0)% | (1.3)% | (42.7)% | | **Net Sales** | **$6,637** | **$5,333** | **24.4%** | **(0.9)%** | **25.3%** | | **Total U.S.** | **$2,360** | **$1,655** | **42.5%** | **42.5%** | | | **Total International** | **$4,277** | **$3,678** | **16.3%** | **(1.3)%** | **17.6%** |
**Net Sales to External Customers (in millions)** **Six Months Ended June 30:** | Segment | 2017 | 2016 | Total Change | Impact of Foreign Exchange | Total Change Excl. Foreign Exchan |---------|------|------|---------------|-----------------------------|-------------------------------| Established Pharmaceutical Products | $1,971 | $1,868 | 5.5% | 1.0% | 4.5% | | Nutritional Products | $3,373 | $3,411 | (1.1)% | (0.8)% | (0.3)% | | Diagnostic Products | $2,431 | $2,344 | 3.7% | (1.4)% | 5.1% | | Cardiovascular and Neuromodulation Products | $4,363 | $1,467 | 197.4% | (1.1)% | 198.5% |
| Other | $834 | $1,128 | (26.0)% | (1.3)% | (24.7)% | | **Net Sales** | **$12,972** | **$10,218** | **27.0%** | **(0.7)%** | **27.7%** | | **Total U.S.** | **$4,684** | **$3,186** | **47.0%** | **47.0%** | | | **Total International** | **$8,288** | **$7,032** | **17.9%** | **(1.0)%** | **18.9%** | **Preliminary Allocation of Fair Value of St. Jude Medical Acquisition (in billions):** | Item | 2017 | |------|------| | Acquired intangible assets, non-deductible | $15.0 | | Goodwill, non-deductible | $15.1 | | Acquired net tangible assets | $3.4 | | Deferred income taxes recorded at acquisition | ($4.6) | | Net debt | ($5.3) | | **Total preliminary allocation of fair value** | **$23.6** | **Assets and Liabilities Held for Disposition (in millions) as of December 31, 2016:** | Item | 2016 | |------|------| | Trade receivables, net | $176 | | Total inventories | $82 | | Prepaid expenses and other current assets | $266 | | Current assets held for disposition | $524 | | Net property and equipment | $130 | | Intangible assets, net of amortization | $150 | | Goodwill | $1,966 | | Deferred income taxes and other assets | $503 | | Non-current assets held for disposition | $2,753 | | **Total assets held for disposition** | **$3,266** | | Trade accounts payable | $145 | | Salaries, wages, commissions and other accrued liabilities | $108 | | Current liabilities held for disposition | $253 | | Post-employment obligations, deferred income taxes and other long-term liabilities | $34 | | **Total liabilities held for disposition** | **$287** | **8. Management Explanations and Justifications** - Management attributes the significant increase in Cardiovascular and Neuromodulation Products sales to the St. Jude Medical acquisition, but notes that sales excluding the acquisition and foreign exchange impact decreased by 1.5% due to lower coronary stent sales and the favorable 2016 royalty agreement resolution. - The decrease in International Pediatric Nutritionals is explained as being due to \challenging conditions in the Chinese infant formula market," a qualitative explanation without numeric validation. - The increase in U.S. Pediatric Nutritionals is attributed to \continued momentum of several recently launched infant formula products" and \growth of the PediaSure toddler brand," with no specific quantitative details provided. - The increase in Diagnostic Products sales is attributed to \share gains in the Core Laboratory and Point of Care markets in the U.S. and higher sales to various international markets," a broad explanation without specific market share or sales figures. - The decrease in the Other category is attributed to the \sale of the AMO business," partially offset by \double-digit growth in Abbott’s Diabetes Care business," with no further quantification of the Diabetes Care growth. - The increase in R&D expenses is explained as being due to the \addition of the acquired St. Jude Medical business," a straightforward explanation without further detail. - The increase in SG&A expenses is attributed to the \addition of the acquired St. Jude Medical business as well as the incremental expenses to integrate St. Jude Medical with Abbott’s existing vascular business," partially offset by \cost improvement initiatives," which is a general term without specifics.
- The decrease in working capital is explained as being due to the \use of cash to fund the cash portion of the St. Jude Medical acquisition, repayments of debt, pension contributions and dividend payments," with a partial offset from the sale of Mylan shares and business dispositions. **9. Accounting Estimates, Judgments, and Policy Changes** - Abbott received a warning letter from the FDA regarding its Sylmar, CA manufacturing facility, and has prepared a plan for corrective actions, which is progressing. - The preliminary allocation of fair value for the St. Jude Medical acquisition is based on estimates and may be subject to material changes as the valuation is finalized. - Abbott has recognized a $70 million credit to intangible amortization expense in the second quarter of 2017 due to measurement period adjustments to the value of intangibles. - Abbott is evaluating the impact of several new accounting standards, including ASU 2017-07, ASU 2016-16, ASU 2016-02, and ASU 2016-01, which will become effective in 2018 and 2019. - Abbott is currently evaluating the impact of ASU 2014-09 (revenue recognition) and expects to use the modified retrospective method for adoption, but has not yet quantified the potential impact. **10. Capital Allocation and Liquidity Management** - Abbott reduced its cash and cash equivalents from $18.6 billion at December 31, 2016, to $9.7 billion at June 30, 2017, primarily due to the St. Jude Medical acquisition, debt repayments, pension contributions, and dividends. - Net cash from operating activities increased by $1.109 billion in the first six months of 2017 compared to 2016, due to the favorable impact of the St. Jude Medical acquisition and reduced pension contributions. - Abbott has $5.0 billion in unused lines of credit available, which expire in 2019. - Abbott entered into a $2.8 billion term loan agreement in July 2017 to fund the Alere acquisition. - Abbott is maintaining an investment-grade debt rating and has a long-term debt rating of BBB by Standard & Poor’s and Baa3 by Moody’s. - Abbott is utilizing cash flow from operations, cash on hand, short-term investments, and borrowings to fund dividends, capital expenditures, and other business investments. **11. Legal, Regulatory, and Compliance Matters** - Abbott received a warning letter from the FDA in April 2017 related to its Sylmar, CA manufacturing facility, which was acquired as part of the St. Jude Medical acquisition. - The FDA inspection findings have not yet resulted in a material impact on financial results, and Abbott is implementing a corrective action plan. - Abbott is subject to regulatory scrutiny in multiple jurisdictions, and tax authorities in various jurisdictions regularly review its income tax filings. - The company expects the recorded amount of gross unrecognized tax benefits to decrease by $200 million to $350 million, including cash adjustments, within the next twelve months as a result of concluding various domestic and international tax matters. - Abbott’s U.S. federal income tax returns are settled through 2013, and St. Jude Medical’s federal income tax returns are settled through 2013 except for one item.
C
Appendix C. Details on AAER Dataset and Preprocessing
This appendix details the acquisition, preprocessing, and filtering pipeline for the Accounting and Auditing Enforcement Releases (AAERs) used to generate fraud labels for our study.
C.1
Raw AAER Data Acquisition and Initial Characteristics
Raw AAER data was programmatically downloaded as JSON objects using the sec-api.io service, specifically querying their AAER Database API endpoint7 . Approximately 3,300 AAERs were initially collected. Each JSON object contains structured information about an enforcement release. An example structure of a downloaded raw AAER JSON object is shown below: Sample Raw AAER JSON Object { "id": "c9ac87509126bd0f1f62c89346cae52d", "dateTime": "2004-02-25T09:21:21-05:00", "aaerNo": "AAER-1964", "releaseNo": ["LR-18595"], "respondents": [{ "name": "FOO", "type": "individual" }], "respondentsText": "FOO", "urls": [{ "type": "primary", "url": "https://www.sec.gov/enforcement-litigation/litigation-releases/lr-18595" }], "summary": "The SEC filed a complaint against FOO...", "tags": ["disclosure fraud", "financial reporting fraud"], "entities": [{ "name": "FOO", "type": "individual", "role": "defendant" }, { "name": "Just for Feet, Inc.", "type": "company", "role": "entity involved in the fraud", "cik": "918111", "ticker": "FEET" }], "complaints": [ "Ruttenberg was instrumental in the acquisition of fraudulent confirmations..." ], "parallelActionsTakenBy": ["United States Department of Justice", "..."], "hasAgreedToSettlement": false, "hasAgreedToPayPenalty": false, "penaltyAmounts": [], "requestedRelief": ["permanent injunction", "disgorgement", "..."], "violatedSections": ["Section 17(a) of the Securities Act of 1933", "..."], "otherAgenciesInvolved": [{"name": "United States Department of Justice", ...}] }
Each AAER JSON object includes a ‘primary url‘ field, which typically links to a detailed document (often a PDF or HTML page) on the SEC website. Figure 7 shows an example of the first page of such a document.
C.2
AAER Preprocessing Pipeline
The downloaded AAERs underwent a multi-step preprocessing pipeline to extract relevant information and filter them for suitability in our fraud detection task. Parsing and Initial Data Extraction The pipeline began by iterating through all downloaded JSON files. For each AAER instance, key fields such as ‘aaerNo‘ (cleaned to a standard format, e.g., ”AAER-XXXX”), ‘dateTime‘ (date part extracted), ‘tags‘, ‘summary‘, ‘complaints‘, the 7
https://sec-api.io/docs/aaer-database-api
Figure 7: Example first page of a primary document linked from an AAER release.
‘primary url‘, and company-specific details (name, role, CIK) from the ‘entities‘ list were extracted. This information was structured into a tabular format for further processing. Extraction of Fiscal Quarter Violation Information A critical challenge is that raw AAER JSONs do not directly provide the specific fiscal years and quarters during which the violations occurred. This temporal information is essential for linking fraud events to quarterly financial reports. To address this, we implemented an automated extraction process: 1. Content Retrieval: For each AAER, the content of the document linked by its ‘primary url‘ was fetched. This involved using web automation tools (like Selenium with ChromeDriver) to download PDF documents or render HTML pages, followed by text extraction (using libraries like PyPDF2 for PDFs and BeautifulSoup for HTML). This process was parallelized to expedite the processing of numerous documents. 2. Fiscal Quarter Identification: The extracted text content from each AAER document was then processed by the Qwen3 32B model. This model was tasked with identifying the precise quarters (e.g., ”2018q1”, ”2019q2”) and companies associated with the financial violations described in the text. The process was guided by the system prompt detailed below, designed to ensure consistent and accurate extraction. API interactions were managed with rate limiting to ensure robust performance. 3. The promt was tailored to put emphasis on earning mistatements, as it is what our work considers as Fraud. System Prompt for Fiscal Quarter Extraction from AAER Content You are a specialized AI agent tasked with meticulously extracting detailed information about **earnings misstatements** (financial statement violations that directly impact the calculation of reported earnings, income, assets, or liabilities) from U.S. Securities and Exchange Commission (SEC) Accounting and Auditing Enforcement Releases (AAERs). Your goal is to deconstruct complex legal and financial text into structured, factual data. ### **Objective** Your primary objective is to identify every quarter in which a **true earning misstatement** occurred, specify the company responsible, detail the specific types of earning misstatements based on predefined categories, and describe the fraudulent scheme that led to these misstatements. You must adhere strictly to the formats and rules defined below. ### **Key Definitions: ‘LIST_MISTATEMENT_TYPE‘** You must categorize all identified **earnings misstatements** using **only** the types from this predefined list. An "earnings misstatement" directly alters the reported financial performance or position (e.g., net income, assets, liabilities, equity balances). Violations related *solely* to disclosure failures that do not alter the numerical financial statements (e.g., failure to disclose related party relationships without affecting specific account balances, or control issues) should *not* be categorized here unless they clearly result in a numerical misstatement of an account listed below. *
*
*
*
*
**Revenue**: Overstating or understating sales or income. This includes premature revenue recognition, fictitious sales, or improper income classification directly impacting the income statement. **Other Expense/Shareholder Equity Account**: Manipulating expenses not directly related to cost of goods sold (e.g., operating, selling, general \& administrative expenses, R\&D), or directly misstating equity accounts (like retained earnings, common stock, additional paid-in capital) through improper accounting entries that affect net income or equity balances. Do not include in this category disclore fraud or governance issues that do not impact financial accounts. **Assets Valuation** : Improperly recognition of assets or their values. This includes inflating asset values (e.g., property, plant \& equipment, intangible assets) or failing to recognize impairments, which directly impacts the balance sheet. **Capitalized Costs as Assets**: Improperly recording expenses as long-term assets (e.g., property, plant \& equipment, intangible assets) to inflate current period income by reducing expenses. **Accounts Receivable**: Overstating the money owed by customers. This
*
*
*
*
*
*
*
includes recording fictitious sales, failing to write off uncollectible receivables, or otherwise inflating the asset balance. **Inventory**: Overstating the value of goods for sale. This includes counting non-existent inventory, improper valuation methods, or misclassifying costs. **Cost of Goods Sold (COGS)**: Understating the direct costs of production. Often linked to inventory manipulation (e.g., overstating inventory leads to understated COGS), directly impacting gross profit and net income. **Reserve Account**: Manipulating funds set aside for future contingent liabilities (e.g., warranty, litigation reserves, bad debt reserves). This includes understating reserves to boost current income or overstating them to create "cookie jar" reserves for future manipulation, directly impacting expenses or liabilities. **Liabilities**: Understating company obligations. This includes concealing debt, failing to record accrued expenses (e.g., unbilled services, payroll), or misclassifying liabilities to improve financial ratios or conceal obligations. **Marketable Securities**: Misstating the value of short-term investments. This includes improper valuation (e.g., failing to mark to market when required) or failing to recognize impairment losses, directly impacting asset values and potentially income. **Allowance for Bad Debt**: Understating the estimated uncollectible accounts receivable to inflate net receivables and income. This is a specific type of reserve manipulation. **Payables**: Understating money owed to suppliers. This includes delaying invoice recording, concealing vendor liabilities, or manipulating cut-off dates, directly impacting liabilities and potentially expenses.
--### **Input Format** You will receive a dictionary containing: ‘"aaerNo"‘: The unique identifier of the AAER. * ‘"content"‘: The full text of the AAER. * ‘"entities"‘: A list of dictionaries, each representing an entity (company * or individual) involved in the AAER. --### **Output Format** You must generate a JSON list of dictionaries. Each dictionary represents a single fraudulent scheme by a specific company in a specific quarter. ‘‘‘json [ { "quarter": "YYYYqQ", "is_fiscal_quarter": true, "fraud_scheme_description": "A concise, factual description of the earning misstatement mechanics, focusing on how the numerical financial statements were altered. You must also justify the selection of the misstatements indicated in the ’misstatements’ field. The justification should be provided only when the ’misstatements’ field is not empty. The justification should be in the form of a list of sentences, each explaining why a specific misstatement type was selected for that quarter. For example: \n- Revenue because the company recorded fictitious sales transactions.\n- Accounts Receivable because the inflated sales led to an overstatement of amounts owed by customers.", "misstatements": ["Type1", "Type2", "Type3", ...], "misstating_company": { "name": "Company Name", "role": "respondent",
"cik": "0001234567" } } ]
The extracted fiscal quarters were then associated with their respective AAERs in our structured dataset. Filtering and Refinement The dataset of AAERs, now augmented with extracted fiscal quarters, underwent several filtering steps to refine its suitability for our fraud detection task: • Date Filtering: To align with the availability of our financial features (which start from 2009), only AAERs with identified violation years from 2009 onwards were considered for the primary dataset used in the experiments. This comprehensive filtering process significantly narrowed down the set of AAERs to those most pertinent for training and evaluating models for company-level financial statement fraud detection.
C.3
Final Fraud Label Statistics
After the complete preprocessing and filtering pipeline: • From the initial ∼3,300 AAERs, the filtering steps (for relevance based on tags, direct company culpability, CIK presence, and violations occurring from 2009 onwards) resulted in 249 unique AAERs. • These 249 AAERs were then merged with our financial features dataset. Due to factors such as non-overlapping CIKs between the AAER dataset and the companies present in our financial reports database, and further filtering based on the completeness threshold for financial features, the number of AAERs contributing to fraud labels in our final experimental dataset was reduced to 137. • These 137 AAERs collectively identified 511 firm-quarters as fraudulent instances. These instances formed the positive class in our fraud detection experiments. This rigorous process ensures that the fraud labels used in our study are well-defined, temporally accurate, and directly linkable to the financial and textual data of the companies involved.
D
Appendix D: Final Dataset Construction and Data Splitting Methodology
This appendix outlines the procedures for constructing the final dataset used in our experiments and details the different data splitting strategies employed to evaluate our models under various conditions.
D.1
Final Dataset Construction
The creation of our final experimental dataset involved several key steps: merging data from different sources, addressing class imbalance, and ensuring data consistency. Data Merging The foundational step was the integration of three primary data sources: • Financial Data: As described in Section 3.1, this includes 122 engineered financial indicators derived from quarterly reports. • Summarized MD&A (SMD&A): Textual data obtained from summarizing MD&A sections, as detailed in Appendix B. • AAER-derived Fraud Labels: Binary fraud labels (fraud/non-fraud) for specific firm-quarters, processed as described in Appendix C. These datasets were merged based on common identifiers: Central Index Key (CIK), fiscal year, and fiscal quarter. Company names were standardized (converted to lowercase) before merging to ensure consistency. A mapping between CIKs and company names was also created for reference. Any records that did not have corresponding entries across these essential dimensions or lacked MD&A data were excluded. Handling Class Imbalance: Stratified Downsampling of Non-Fraud Cases Financial fraud is a rare event, leading to highly imbalanced datasets. To create a tractable dataset for experimentation while preserving a significant level of imbalance, we downsampled the non-fraudulent cases to achieve a target fraud rate of approximately 5.03% in the final dataset (as discussed in Section 3.4). This was performed carefully to maintain the underlying characteristics of the non-fraud data. The stratified downsampling algorithm is as follows: 1. Segregation: The merged dataset was divided into two subsets: fraudulent firm-quarter samples and non-fraudulent firmquarter samples. 2. Target Calculation: • Let Nf raud be the total number of fraudulent samples. • The target number of non-fraudulent samples (Nnon f raud target ) to achieve the desired overall fraud percentage 100−Pf raud (Pf raud = 5.03%) was calculated as: Nnon f raud target = Nf raud × Pf raud . 3. Proportional Group Sampling (Initial Pass): • The non-fraudulent subset was grouped by ‘year‘ and ‘sicagg‘. The ‘sicagg‘ field represents an aggregated Standard Industrial Classification code, corresponding to high-level industry sectors derived from the first two digits of the SIC code, based on classifications from official sources (e.g., https://siccode.com/). This grouping aims to preserve temporal and industry distributions. • A global downsampling factor (Fdownsample ) was computed: Fdownsample = Nnon f raud target /Nnon f raud original , where Nnon f raud original is the total count of non-fraudulent samples before downsampling. • For each (‘year‘, ‘sicagg‘) group within the non-fraudulent data: – The target number of samples for this group (Ngroup target ) was calculated by multiplying the original size of this group by Fdownsample . – Ngroup target was adjusted to be at least 1 (if the group was non-empty and Ngroup target was positive) and no more than the actual number of samples available in that group. – Ngroup target samples were randomly selected from this specific group. • All samples selected from these groups were collected. 4. Refinement Pass (If Target Not Met): • If the total number of non-fraudulent samples collected in the initial pass was less than Nnon f raud target , a refinement step was performed. • Groups that still contained unselected non-fraudulent samples were identified. • Additional samples were iteratively drawn from these groups, prioritizing those with more remaining available samples, until Nnon f raud target was reached or no more unselected samples were available. This ensures the target size is met more closely while still favoring the original distribution.
5. Final Dataset Assembly: The original set of Nf raud fraudulent samples was combined with the Nnon f raud target downsampled non-fraudulent samples to form the final dataset used for all experiments. A fixed random seed was used during the sampling process to ensure reproducibility. This stratified downsampling ensures that while the dataset is made more balanced, the non-fraudulent samples still reflect the temporal and sectoral diversity of the original population. Final Dataset Composition After merging and downsampling, the final dataset used for our experiments consists of 10,159 firm-quarter observations. This includes: • Fraudulent Samples: 511 (approximately 5.03%) • Non-Fraudulent Samples: 9,648 (approximately 94.97%) • Unique Companies: 5,658 It is important to note that a single company (identified by its CIK) can contribute multiple firm-quarter observations to the dataset, and these can include both fraudulent and non-fraudulent instances over different time periods. The dataset spans from 2009 to 2021. The overall distribution of samples per aggregated industry sector (based on ‘sicagg‘) in this final dataset is presented in Table 8. Table 8: Overall industry sector distribution in the final experimental dataset (after downsampling). Numbers represent firm-quarter instances.
Industry Sector (Aggregated SIC - ‘sicagg‘) Agriculture, Forestry, And Fishing Construction Finance, Insurance, And Real Estate Manufacturing Mining Public Administration Retail Trade Services Transportation & Public Utilities Wholesale Trade
D.2
Total Samples 39 110 2320 3789 640 5 426 1770 786 274
Data Splitting Strategies for FSFD Tasks
To evaluate model performance under different assumptions and levels of difficulty, we employed three distinct data splitting strategies, corresponding to the tasks defined in Section 4. For tasks involving cross-validation, 5 folds were used, and a fixed random seed was employed for fold generation to ensure reproducibility. Classic FSFD: Random K-Fold Cross-Validation This strategy represents the traditional approach to evaluating FSFD models. • Methodology: The final dataset of 10,159 firm-quarter observations was split into 5 folds using random sampling. Stratification was applied based on the ‘is fraud‘ label to ensure that each fold maintained approximately the same 5.03% fraud ratio as the overall dataset. • Characteristics: In this setup, observations from the same company can appear in both the training and testing sets of a given fold (though not the same observation). This allows the model to potentially learn company-specific patterns. • Fold Statistics (Averages over 5 Folds): – Test Set Size: ∼2032 samples – Fraud Samples in Test: ∼102 (Fraud Ratio: ∼5.03%) – Industry distribution in test folds (e.g., ‘Manufacturing‘ ∼758 samples with ∼52 fraud; ‘Services‘ ∼354 samples with ∼22 fraud) typically reflected the overall dataset distribution due to random sampling. • Illustration: See Figure 8.
Full Dataset (Firm-Quarters)
Fold 1
Fold 2
Fold 3
Fold 4
Fold 5
Test Train (F2,F3,F4,F5) Firm-quarters randomly assigned. Same company may appear in train and test portions of a fold. Process repeated for each fold as test set. Figure 8: Conceptual diagram of Classic FSFD (Random K-Fold) splitting. Each fold serves as a test set once, with the remaining folds as training.
Company-Isolated FSFD (CI-FSFD): Company-Based K-Fold Cross-Validation This strategy imposes a stricter evaluation by ensuring that companies seen during training are not present in the test set. • Methodology: The dataset was split into 5 folds at the company (CIK) level. All firm-quarter observations belonging to a specific company (which may include a mix of fraud and non-fraud instances for that company) were assigned entirely to one fold. The assignment of companies to folds was performed aiming to balance the number of fraud reports and total reports per fold, and preserve the industry sector distribution within each fold’s test set as much as possible. • Characteristics: This split tests the model’s ability to generalize to entirely unseen companies, preventing it from relying on idiosyncratic patterns of companies present in the training data. • Fold Statistics (Averages over 5 Folds): – Test Set Size: ∼2032 samples – Unique Companies in Test: ∼1132 – Fraud Samples in Test: ∼102 (Fraud Ratio: ∼5.03%) – Industry distributions were actively balanced. For example, across test folds, ‘Manufacturing‘ had ∼758 samples (with ∼52 fraud), and ‘Services‘ had ∼354 samples (with ∼22 fraud). • Illustration: See Figure 9.
Full Dataset (Grouped by Company)
Companies A
Fold 1 (Test) (e.g., Co. Group A)
Companies B
Companies C
Companies D
Companies E
Train for Fold 1 (e.g., Co. Groups B, C, D, E) Groups of companies are assigned to folds. If a company’s data (all its firm-quarter instances) is in the test set of a fold, none of its data appears in the training set for that fold.
Figure 9: Conceptual diagram of Company-Isolated FSFD (CI-FSFD) splitting. Companies (and all their associated firm-quarter instances) are assigned to folds.
E
Appendix E: Model’s Details and Hyperparameters
This appendix provides detailed model configurations and hyperparameters for all models used in our experiments: Logistic Regression, MLP, Random Forest (LightGBM), XGBoost, RCMA-adapted, and LLM-based models. For all models, hyperparameters were optimized using Hyperopt, targeting the maximization of the F1-score on a dedicated validation set (typically 10% of the training data for that specific fold/split). The ‘decision threshold‘ reported for classification models is the optimal threshold found on the validation set. It is important to note that the same set of optimized hyperparameters was applied across all folds for a given model and task (e.g., Classic FSFD or CI-FSFD).
E.1
Logistic Regression (MLP-Classifier with no Hidden Layers)
Our Logistic Regression baseline is implemented as a specialized case of the MLP Classifier with no hidden layers. Table 9: Optimized Hyperparameters for Logistic Regression.
E.2
Parameter
Value
Features Type Learning Rate Batch Size Dropout Rate Epochs Patience Oversample Standardize Decision Threshold
Dechow 0.1 64 0.459 2 20 True True 0.5
Multi-Layer Perceptron (MLP)
The MLP Classifier uses a feed-forward neural network architecture. Table 10: Optimized Hyperparameters for MLP.
E.3
Parameter
Value
Classic FSFD Task Features Type Hidden Dims Learning Rate Batch Size Dropout Rate
Financial (122 features) [512] 0.1 128 0.413
CI-FSFD Task Features Type Hidden Dims Learning Rate Batch Size Dropout Rate
Financial (122 features) [512, 512] 0.001 32 0.145
Common Parameters Epochs Patience Oversample Standardize Decision Threshold
2 20 True False 0.5
Random Forest (LightGBM Implementation)
For the Random Forest baseline, we utilized the Random Forest mode of the LightGBM library [Ke et al., 2017].
E.4
XGBoost
The XGBoost classifier [Chen and Guestrin, 2016] was optimized with Hyperopt.
Table 11: Optimized Hyperparameters for Random Forest (LightGBM).
Parameter
Value
Classic FSFD Task Features Type Learning Rate Max Depth Num Estimators Num Leaves Standardize
Financial (122 features) 0.0799 7 200 5 True
CI-FSFD Task Features Type Learning Rate Max Depth Num Estimators Num Leaves Standardize
Financial (122 features) 0.0996 78 200 50 False
Common Parameter Decision Threshold
0.5
Table 12: Optimized Hyperparameters for XGBoost.
E.5
Parameter
Value
Classic FSFD Task Features Type Learning Rate Max Depth Num Estimators Num Leaves Standardize
Financial (122 features) and Dechow 0.05 0 50 50 True
CI-FSFD Task Features Type Learning Rate Max Depth Num Estimators Num Leaves Standardize
Financial (122 features) 0.090 88 200 50 True
Common Parameter Decision Threshold
0.5
RCMA-adapted Model
The RCMA-adapted model architecture and training parameters were based on the work of Wang et al. [2023], with modifications for our SBERT-based text processing.
E.6
LLM-based Models
This section details the configuration for our Large Language Models, including Llama-3.1 8B and Fino1 8B, in various finetuning and zero-shot settings. All LLMs use LoRA for fine-tuning. Financial-only Models (Llama-3.1 8B and Fino1 8B) These models are fine-tuned exclusively on financial text. SMD&A-only Models (Llama-3.1 8B, Fino1 8B, Fino1 14B) These models are fine-tuned exclusively on Summary Management Discussion & Analysis (SMD&A) text.
Table 13: Optimized Hyperparameters for RCMA-adapted Model.
Parameter
Value
SBERT Model Name SBERT Output Dim Trainable SBERT Layers Max SMD&A Length Financial Features Num Financial Groups Financial Embedding Dim Text Embedding Dim Dropout Rate Learning Rate Batch Size Validation Batch Size Pos Weight Beta (for FocalLoss) Focal Gamma (for FocalLoss) Gradient Accumulation Steps
‘jinaai/jina-embeddings-v2-small-en‘ 512 1 8192 tokens 122 7 512 512 0.05 1e-4 8 8 0.75 2.0 4
Classic FSFD Task MLP Hidden Dims Epochs Patience Consistency Loss Weight Oversample
[128] 10 7 0.2 False
CI-FSFD Task MLP Hidden Dims Epochs Patience Consistency Loss Weight Oversample
[512] 20 10 0 True
Financial+SMD&A Models (Llama-3.1 8B and Fino1 8B) These models are fine-tuned on both Financial (122 features) and Summary Management Discussion & Analysis (SMD&A) sections. Zero-shot Model (Fino1 8B SMD&A) For the zero-shot evaluation, the Fino1 8B model was used without any fine-tuning (i.e., ‘num layers to finetune‘ is 0).
Table 14: Hyperparameters for LLM-based Models (Financial-only).
Parameter
Value
Llama-3.1 8B (Financial) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Max Context Batch Size Gradient Accumulation Steps Learning Rate LoRA Target Modules
‘unsloth/Llama-3.1-8B-unsloth-bnb-4bit‘ 8 8 0.05 32 1500 tokens 4 2 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘
Fino1 8B (Financial) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Max Context Batch Size Gradient Accumulation Steps Learning Rate LoRA Target Modules
‘TheFinAI/Fino1-8B‘ 8 8 0.05 32 1500 tokens 4 2 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘
Common Parameters Epochs (Classic FSFD Task) Epochs (CI-FSFD Task) Max New Tokens Only Completion Undersample Run Eval on Start
10 20 1 True True False
Table 15: Hyperparameters for LLM-based Models (SMD&A-only).
Parameter
Value
Llama-3.1 8B (SMD&A) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Max Context Batch Size Gradient Accumulation Steps Learning Rate LoRA Target Modules (Classic FSFD Task) LoRA Target Modules (CI-FSFD Task)
‘unsloth/Llama-3.1-8B-unsloth-bnb-4bit‘ 8 8 0.05 32 8500 tokens 8 1 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘ ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘, ‘lm head‘
Fino1 8B (SMD&A) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Max Context Batch Size Gradient Accumulation Steps Learning Rate LoRA Target Modules (Classic FSFD Task) LoRA Target Modules (CI-FSFD Task)
‘TheFinAI/Fino1-8B‘ 8 8 0.05 32 8500 tokens 8 1 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘ ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘, ‘lm head‘
Fino1 14B (SMD&A) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Max Context Batch Size Gradient Accumulation Steps Learning Rate LoRA Target Modules
‘TheFinAI/Fin-o1-14B‘ 8 8 0.05 40 8500 tokens 4 2 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘
Common Parameters Epochs (Classic FSFD Task) Epochs (CI-FSFD Task) Max New Tokens Only Completion Undersample Run Eval on Start Use Full Summary
10 20 1 True True False True
Table 16: Hyperparameters for LLM-based Models (Financial+SMD&A).
Parameter
Value
Llama-3.1 8B (Financial+SMD&A) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Learning Rate LoRA Target Modules
‘unsloth/Llama-3.1-8B-unsloth-bnb-4bit‘ 8 8 0.05 32 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘
Fino1 8B (Financial+SMD&A) Model URL LoRA R LoRA Alpha LoRA Dropout Layers to Finetune Learning Rate LoRA Target Modules
‘TheFinAI/Fino1-8B‘ 8 8 0.05 32 1e-4 ‘q proj‘, ‘v proj‘, ‘up proj‘, ‘down proj‘, ‘gate proj‘
Common Parameters Max Context (Classic FSFD Task) Max Context (CI-FSFD Task) Batch Size (Classic FSFD Task) Batch Size (CI-FSFD Task) Gradient Accumulation Steps (Classic FSFD Task) Gradient Accumulation Steps (CI-FSFD Task) Epochs (Classic FSFD Task) Epochs (CI-FSFD Task) Max New Tokens Only Completion Undersample Run Eval on Start Use Full Summary
9500 tokens 9500 tokens 8 4 1 2 10 20 1 True True False True
Table 17: Hyperparameters for Zero-shot Fino1 8B SMD&A Model.
Parameter
Value
Model URL Max Context Max New Tokens Batch Size Zero-Shot
‘TheFinAI/Fino1-8B‘ 9500 tokens 1 1 True
F
Appendix F: LLM System Prompts
This appendix details the system prompts used for the Large Language Model (LLM) based classifiers, corresponding to the different input data configurations: Financials Only (FIN), Synthetically Summarized MD&A Only (SMD&A), and combined Financials + SMD&A. For each configuration, the model was provided with a specific user prompt outlining the task, the context (industry sector), and the relevant data. The LLM was then fine-tuned to generate a single token representing the classification: ”YES” (indicating fraud) or ”NO” (indicating non-fraud) immediately following the prompt. The prompts were designed to fit within the model’s context window, with specific token allowances made for the variable data portions (financial strings or MD&A content). The token counts provided below are approximate estimates for the fixed textual parts of each prompt, based on the Llama-3 tokenizer, and exclude the tokens from placeholder content like ‘industry title‘, ‘financials str‘, or ‘mda content‘.
F.1
Prompt for Financials Only (FIN) Input
When using only financial data, the LLM was presented with the following prompt structure. LLM Prompt: Financials Only You are a financial forensic analyst. The company operates in the {industry_title} sector. Below are key financial indicators derived from its income statement, balance sheet, and cash flow statement: {financials_str} Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is engaging in Financial Manipulation Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"?
F.2
Prompt for SMD&A Only (SMD&A) Input
When using only the Synthetically Summarized MD&A text, the following prompt structure was employed.. LLM Prompt: SMD&A Only The company operates in the {industry_title} sector. Below is the summary of the Management Discussion and Analysis (MDA) section of the quarterly report: {mda_content} Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Manipulation Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"?
F.3
Prompt for Combined Financials + SMD&A Input
For the combined input scenario, the LLM received a prompt integrating both financial metrics and the SMD&A content. LLM Prompt: Financials + SMD&A The company operates in the {industry_title} sector. Here are financial variables derived from the income statement, balance sheet, and cash flow statement of the company. {financials_str} Also below is the structured summary of the Management Discussion and Analysis (MDA) section of the quarterly report: {mda_content}
Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"?
G
Appendix G: Sample Prediction Details
This appendix provides an illustrative example of a prediction made by our Fino-1 8B model using the combined Financial + SMD&A input. It shows the complete prompt provided to the model (truncated for brevity in this display, but the full content was used for prediction), the model’s generated answer, the ground truth label, and the associated prediction probabilities. The following example corresponds to a firm-quarter instance from the CI-FSFD task, where the model correctly identified a fraudulent case. Sample Model Input Prompt (Financials + SMD&A - Truncated for Display) The company operates in the RUBBER & PLASTICS FOOTWEAR sector. Here are financial variables derived from the income statement, balance sheet, and cash flow statement of the company. - Total Assets: $22,921,000,000 - Cash and Short-term Investments: $3,695,000,000 - Property, Plant, and Equipment: $4,688,000,000 - Degree Of Financial Leverage: 873% - Invested Capital Ratio: 20% - Cash To Total Asset: 16% - Debt Service Coverage: 36% - Financial LeverageIndex: 5.44% - Times InterestEarnedRatio: 9,275% - Current Asset To Revenues: 164% - Current Liabilities To Revenues: 76% - Short TermDebt To Revenue: 0.23% - Intangible Asset ToRevenue: 4.55% - LongtermLeverage: 15% - Asset Quality Index: 0.96 - Leverage Index: 0.99 ... Also below is the structured summary of the Summary Management Discussion and Analysis (SMD&A) section of the quarterly report: # 1.
Strategic Priorities and Initiatives
{ NIKE’s goal is to deliver value to shareholders by building a profitable global portfolio of branded footwear, apparel, equipment, and accessories businesses. { The company’s strategy is to achieve long-term revenue growth by creating innovative, must-have products, building deep personal consumer connections with its brands, and delivering compelling consumer experiences through digital platforms and at retail. { In fiscal 2018, NIKE introduced the Consumer Direct Offense, a new company alignment designed to allow NIKE to better serve the consumer more personally, at scale. { Through the Consumer Direct Offense, NIKE is focusing on the Triple Double strategy, with the objective of doubling the impact of innovation and increasing its speed to market and direct connections with consumers. --# 2.
Operational and Segment Performance
{ For the third quarter of fiscal 2019, NIKE Brand delivered 8% revenue growth, with 12% growth on a currency-neutral basis, driven by higher revenues across all geographies, footwear and apparel, as well as growth in most key categories, led by Sportswear and the Jordan Brand. { Converse revenues decreased 4% on a reported basis and 2% on a currency-neutral basis, primarily due to declines in the U.S. and Europe, partially offset by revenue growth in Asia. { In North America, on a currency-neutral basis, revenues increased 7% for the third quarter and first nine months of fiscal 2019, driven by growth in several
key categories for the quarter and nearly all key categories for the year-to-date period, led by Sportswear. { NIKE Direct in North America increased 6% and 7% for the third quarter and first nine months, respectively, as higher digital commerce sales and the addition of new stores more than offset an 8% and 4% decline in comparable store sales, driven by NFS performance. { In EMEA, on a currency-neutral basis, revenues grew 12%, driven by balanced growth across all territories and led by Sportswear and the Jordan Brand. { NIKE Direct in EMEA increased 15% and 14% for the third quarter and first nine months, respectively, due to comparable store sales growth, higher digital commerce sales, and new store additions. { In Greater China, on a currency-neutral basis, revenues increased ... --# 3.
Financial Results and Key Trends
{ For the third quarter of fiscal 2019, revenues increased 7% to $9.6 billion, and net income was $1.1 billion with diluted earnings per share of $0.68, compared to a net loss of $921 million and diluted loss per share of $0.57 for the same period in fiscal 2018. { Income before income taxes increased 11%, driven by revenue growth and gross margin expansion, partially offset by higher selling and administrative expense. { Gross margin increased to 45.1% for the third quarter of fiscal 2019, compared to 43.8% in the same period in fiscal 2018. { For the first nine months of fiscal 2019, revenues increased 9% to $28.9 billion, and net income was $3.0 billion, compared to $3.1 billion in the prior year. { Gross margin for the nine months ended February 28, 2019, was 44.4%, compared to 43.5% in the prior year. { Selling and administrative expense increased to $3.09 billion for the third quarter of fiscal 2019, representing 32.2% of revenues, compared to $2.767 billion and 30.8% of revenues for the third quarter of fiscal 2018. { For the first nine months of fiscal 2019, selling and administrative expense increased to $9.296 billion, representing 32.1% of revenues, compared to $8.391 billion and 31.5% of revenues in the prior year. ... --# 4.
Identified Risks and Uncertainties
{ The company is exposed to foreign currency market volatility, partly due to global trade uncertainty and geopolitical dynamics. { Foreign currency exposures arise from transactions denominated in non-functional currencies and the translation of foreign currency-denominated results into U.S. Dollars. { Argentina has been identified as a hyper-inflationary market, and the functional currency of the Argentina subsidiary was changed to U.S. Dollars in the second quarter of fiscal 2019. { The translation of foreign currency-denominated profits and foreign exchange rate fluctuations had an unfavorable impact on income before income taxes for both the third quarter and first nine months of fiscal 2019. { The functional currency change in Argentina did not have a material impact on the Company’s results of operations or financial condition, and management does not anticipate a material impact in future periods based on current rates. { The Company may face challenges in accessing credit markets or increased interest costs due to future volatility. { Foreign currency hedge gains and losses may impact operating performance, depending on actual market rates versus standard rates. ... --# 5.
Forward-Looking Statements and Guidance
{ The company remains committed to its long-term financial goals, and continues to see opportunities to drive growth and profitability despite foreign currency volatility. { NIKE Direct is expected to continue accelerating growth, driven by digital commerce and store expansion. { Investments in data and analytics, digital commerce platforms, and a new enterprise resource planning tool are part of the end-to-end digital transformation strategy. { The company plans to continue share repurchases under its new $15 billion four-year program, with funding expected from operating cash flows, excess cash, and debt proceeds. { Management believes that existing cash, cash equivalents, short-term investments, and cash generated by operations, along with access to external funding, will be sufficient to meet capital needs in the foreseeable future. ... --# 6.
Significant Changes, Events, or Developments
{ The Company completed the $12 billion share repurchase program authorized in November 2015 during the first nine months of fiscal 2019, repurchasing 192.1 million shares. { A new four-year, $15 billion share repurchase program was authorized in June 2018, and 43.7 million shares were repurchased under this program during the first nine months of fiscal 2019. { The functional currency of the Argentina subsidiary was changed to U.S. Dollars in the second quarter of fiscal 2019, due to hyper-inflationary conditions. ... --# 7.
Important Figures and Tables
Revenues for the Three Months Ended February 28, 2019 and 2018 | Period | Revenues (in millions) | % Change | % Change |--------|------------------------|----------|----------| 2019 | $9,611 | | | 11% | 2018 | $8,984 | 7% | | Revenues for the Nine Months Ended February 28, 2019 and 2018 | Period | Revenues (in millions) | % Change | % Change |--------|------------------------|----------|----------| 2019 | $28,933 | | | 11% | 2018 | $26,608 | 8% | | Gross Profit and Gross Margin for the Three Months Ended February 28 | Period | Gross Profit (in millions) | % Change | Gross Margin | |--------|----------------------------|----------|--------------| | 2019 | $4,339 | | | 45.1% | | 2018 | $3,938 | 10% | 43.8% | ... --# 8.
Management Explanations and Justifications
{ The increase in NIKE Brand revenues is attributed to growth across all geographies, footwear and apparel, and key categories like Sportswear and the Jordan Brand. { The decline in Converse revenues is explained by revenue declines in the U.S. and Europe, partially offset by growth in Asia. ...
--# 9.
Accounting Estimates, Judgments, and Policy Changes
{ Revenue recognition is based on transfer of control to the customer, with variable consideration for sales returns, discounts, and miscellaneous claims estimated and recorded as a reduction to revenues. ... --# 10.
Capital Allocation and Liquidity Management
{ Share repurchase activity increased significantly in fiscal 2019, with $3.386 billion spent on 43.7 million shares. ... --# 11.
Legal, Regulatory, and Compliance Matters
{ The Company has no off-balance sheet arrangements that have or are reasonably likely to have a material effect on financial condition, results of operations, liquidity, or capital resources. ... Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"?
Model Prediction and Ground Truth Model’s Generated Answer: NO Ground Truth Label: NO Prediction Probability for ”YES”: 0.0073334336280823 (Note: The decision threshold optimized on the validation set for this fold was applied to this probability to yield the binary prediction.) Instance Identifiers: • CIK: 320187 • SIC (Aggregated): 3021 (RUBBER & PLASTICS FOOTWEAR) • Quarter: 2019q3
This example illustrates how the model processes the combined textual and numerical information to arrive at a classification decision. The relatively low probability for a correct ”YES” prediction in this specific case, despite being above the decision threshold for this fold, highlights the challenging nature of the CI-FSFD task.
H
Appendix H: Detailed Performance Analysis
This appendix provides a more granular look at the performance of our LLM-based fraud detection framework (Fino-1 8B with SMD&A input) across different subgroups: individual companies, industry sectors, and performance on unseen companies in the classic setting. We present key metrics such as True Positives (TP - Detected Fraud), False Negatives (FN - Undetected Fraud), False Positives (FP - Non-Fraud Incorrectly Flagged as Fraud), Total Actual Fraud instances, Recall (TP / (TP + FN)), and Precision (TP / (TP + FP)).
H.1
Fine-Grain Labels Performance Analysis
Table 18 presents a detailed breakdown of the model’s performance on detecting different types of financial statement misstatements. The model exhibits varying recall across different misstatement categories, indicating that certain types of fraud are more challenging to detect than others. Table 18: Fraud Detection Performance by Misstatement Type (CI-FSFD Task). Model: Fino-1 8B with SMD&A.
Misstatement Type mis Reserve Account mis Capitalized Costs as Assets mis Cost of Goods Sold (COGS) mis Accounts Receivable mis Allowance for Bad Debt mis Liabilities mis Revenue mis Payables mis Other Expense/Shareholder Equity Account mis Assets Valuation mis Inventory
H.2
Detected Fraud (TP)
Undetected Fraud (FN)
Total Actual Fraud
Recall
AUC
7 7 15 25 1 19 51 17 55 10 2
4 8 24 52 3 59 160 56 210 85 36
11 15 39 77 4 78 211 73 265 95 38
0.636 0.467 0.385 0.325 0.250 0.244 0.242 0.233 0.208 0.105 0.053
0.819 0.796 0.669 0.756 0.955 0.689 0.707 0.698 0.645 0.536 0.565
Company-Level Performance Analysis (CI-FSFD)
The Company-Isolated FSFD (CI-FSFD) task evaluates the model’s ability to generalize to companies not seen during training. Performance at the individual company level can vary significantly. Table 19 lists the top 30 companies where the model demonstrated the best performance (highest recall) in detecting fraudulent quarters under the CI-FSFD setting. Table 20 shows the 30 companies where the model struggled the most (lowest recall). The visual representations of True Positives and False Negatives for the top 10 from these company groups are in Figure 10 and Figure 11 respectively.
Table 19: Top 30 Companies by Recall in Detecting Fraudulent Quarters (CI-FSFD Task), with False Positives and Precision. Model: Fino-1 8B with SMD&A. Industry Sector (Company Name) 3m company Akorn, inc. Hertz Ixia Jda software group, inc. Mcdermott international inc. Ocz technology group, inc. Roadrunner transportation systems, inc. Surgalign holdings, inc. Swisher hygiene inc. Taronis technologies, inc. Cognizant technology solutions corporation Kbr, inc. L3 technologies, inc. Tech data corp. Uti worldwide inc. Nci, inc. Newell brands inc. Quadrant 4 system corp. The kraft heinz co. Halliburton company Sciclone pharmaceuticals, inc. Corporate resource services, inc. Future fintech group inc. Mimedx group, inc. Synchronoss technologies, inc. Granite construction inc. Apex global brands inc. Evoqua water technologies corp. Gt advanced technologies inc.
Detected Fraud (TP)
Undetected Fraud (FN)
False Positives (FP)
Total Actual Fraud
Recall
Precision
9 2 2 1 2 1 2 6 2 1 1 2 1 1 1 2 7 3 7 4 3 3 2 2 6 9 5 1 1 1
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1 3 5 4 1 1 1
1 0 0 0 0 0 0 0 0 0 0 2 2 2 1 1 1 1 0 1 0 0 0 0 0 3 0 0 0 0
9 2 2 1 2 1 2 6 2 1 1 2 1 1 1 2 7 3 8 5 4 4 3 3 9 14 9 2 2 2
1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 0.875 0.800 0.750 0.750 0.667 0.667 0.667 0.643 0.556 0.500 0.500 0.500
0.900 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 0.500 0.333 0.333 0.500 0.667 0.875 0.750 1.000 0.800 1.000 1.000 1.000 1.000 1.000 0.750 1.000 1.000 1.000 1.000
Table 20: Worst 30 Companies by Recall in Detecting Fraudulent Quarters (CI-FSFD Task). Model: Fino-1 8B with SMD&A. Industry Sector (Company Name) African gold acquisition corp. Amyris, inc. Andeavor llc Argo group international holdings, ltd. Assisted living concepts, inc. Axesstel, inc. Barrett business services, inc. Belden inc. Biomet, inc. Blue earth, inc. Brixmor property group inc. Cantaloupe, inc. Celadon group, inc. Celsius holdings, inc. China valves technology, inc. Chs inc. Citigroup inc. Comscore, inc. Cpi aerostructures, inc. Dxc technology company Elanco animal health inc. Fmc technologies, inc. Fte networks, inc. General electric company General motors company Gtt communications, inc. Healthcare services group, inc. Home loan servicing solutions, ltd. Homestreet, inc. Iconix brand group, inc.
Detected Fraud (TP)
Undetected Fraud (FN)
False Positives (FP)
Total Actual Fraud
Recall
Precision
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
1 1 1 15 2 1 10 1 4 2 4 5 3 2 1 15 1 5 2 2 4 5 2 7 7 2 4 5 7 3
0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 2 0 0 0 0
1 1 1 15 2 1 10 1 4 2 4 5 3 2 1 15 1 5 2 2 4 5 2 7 7 2 4 5 7 3
0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000
0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000
Figure 10: Plot: Top 10 companies by number of correctly detected fraudulent quarters (True Positives vs False Negatives) in the CI-FSFD task. Model: Fino-1 8B with SMD&A.
Figure 11: Plot: Top 10 companies by number of undetected fraudulent quarters (True Positives vs False Negatives) in the CI-FSFD task. Model: Fino-1 8B with SMD&A.
H.3
Sectorial Performance Analysis (CI-FSFD)
Performance also varies when aggregated by industry sector. Table 21 details the detection performance, including false positives and precision, across major industry sectors for the CI-FSFD task. Table 21: Fraud Detection Performance by Industry Sector (CI-FSFD Task). Model: Fino-1 8B with SMD&A. Industry Sector
Detected Fraud (TP)
Undetected Fraud (FN)
False Positives (FP)
Total Actual Fraud
Recall
Precision
Construction Retail trade Services Transportation & public utilities Manufacturing Mining Wholesale trade Finance, Insurance, & Real Estate
6 4 34 8 61 4 4 0
4 6 74 21 198 15 20 52
20 25 206 47 277 10 42 57
10 10 108 29 259 19 24 52
0.600 0.400 0.315 0.276 0.236 0.211 0.167 0.000
0.231 0.138 0.142 0.145 0.180 0.286 0.087 0.000
Overall (CI-FSFD Task Total)
121
390
684
511
0.237
0.150
Note: The ”Overall” row aggregates TP, FN, FP, and Actual Fraud across all test folds for the CI-FSFD task for the listed sectors and calculates overall Recall and Precision from these sums.
Figure 12 visually represents the True Positives and False Negatives by sector.
Figure 12: Plot: Detected Fraud (TP) vs. Undetected Fraud (FN) cases per industry sector in the CI-FSFD task. Model: Fino-1 8B with SMD&A.
H.4
Performance on Unseen Companies in Classic FSFD Setting
While the Classic FSFD setting involves random splitting, a small fraction of companies in the test set of each fold might still be entirely unseen during the training phase for that specific fold. Analyzing performance on these truly ”unseen” companies within the classic random split provides insight into the model’s baseline generalization even when not explicitly forced by a company-isolated split. Table 22 summarizes key average metrics for the Fino-1 8B (SMD&A) model on these unseen company instances within the Classic FSFD’s 5-fold cross-validation. The very low average number of unseen fraudulent instances (1.6 per fold) in the Classic FSFD setting makes it difficult to draw firm conclusions about the model’s ability to detect fraud in entirely new companies from these specific metrics (F1, Recall, Precision).
Table 22: Average Performance Metrics on Unseen Companies within Classic FSFD Test Folds. Model: Fino-1 8B with SMD&A (Averages over 5 Folds).
Metric
Mean Value ± Std. Dev.
Avg. Test Samples per Fold Avg. Unseen Samples in Test Fold Avg. Test CIKs per Fold Avg. Unseen CIKs in Test Fold Avg. Total Fraud Samples in Test Fold Avg. Unseen Fraud Samples in Test Fold Avg. Fraud Rate among Unseen Samples
2604.2 ± 0.4 757.6 ± 18.4 2163.2 ± 9.4 672.6 ± 14.1 130.8 ± 15.7 1.6 ± 1.2 0.0021 ± 0.0015
F1 Score (on Unseen Samples) Recall (on Unseen Fraud Samples) Precision (on Unseen Samples) AUC Score (on Unseen Samples) Accuracy (on Unseen Samples)
0.0 ± 0.0 0.0 ± 0.0 0.0 ± 0.0 0.646 ± 0.280 0.9924 ± 0.0022
The performance metrics such as F1 score, Recall, and Precision for unseen fraud cases are not statistically significant due to the very low average number of unseen fraud samples (1.6 per fold) in this random splitting setting. This underscores the importance of dedicated CI-FSFD for robustly evaluating generalization.
I
Appendix I: Explainability
This appendix details the methodology employed for generating explanations of our LLM’s predictions, specifically focusing on the textual components of the input. Understanding which parts of the financial text contribute most to a fraud prediction is crucial for interpretability and trustworthiness in high-stakes domains like financial anomaly detection.
I.1
LRP-based Explanation
Our explainability approach is built upon the Layer-wise Relevance Propagation (LRP) [Achtibat et al., 2024] framework. LRP is a technique used to decompose the prediction of a deep neural network into contributions of its input features. It assigns a ”relevance score” to each input component (e.g., a token) indicating its importance to the final output. For a classification task, LRP propagates the prediction score backward through the network, layer by layer, until it reaches the input features. The core idea is to conserve the total relevance during propagation, ensuring that the sum of relevances at one layer equals the sum of relevances at the preceding layer. This property allows for a clear attribution of the final prediction to individual input elements. In our implementation, we leverage the LXT library, which provides an efficient and specialized LRP implementation for Transformer-based models. After the model makes a prediction (i.e., outputs logits for ”Fraud” or ”Not Fraud”), we backpropagate the relevance from the logit corresponding to the predicted class (or, more specifically, the ’Fraud’ logit, regardless of the prediction, to understand drivers of potential fraud) back to the input embeddings. The LRP rules applied ensure that the relevance scores accurately reflect the contribution of each token in the input sequence to that specific logit. The process for generating LRP explanations for each test sample is as follows: • The trained LLM is set to evaluation mode, and all its parameters are frozen ) for the input embeddings, which are set to requires grad=True to compute gradients for LRP. • For each test sample, the input prompt (containing the textual and numerical financial data) is tokenized and fed into the model. • The model performs a forward pass to obtain the logits for the FRAUD LABEL ID and NOT FRAUD LABEL ID tokens at the last position of the output sequence. • The gradient of the FRAUD LABEL ID logit (representing the unnormalized score for the ”Fraud” class) with respect to the input embeddings is computed. • The LRP relevance score for each input token is then calculated as the element-wise product of the input embeddings and their corresponding gradients, summed across the embedding dimensions. This yields a single relevance score for each token. • These raw relevance scores are then normalized by their absolute maximum to scale them between -1 and 1, facilitating easier interpretation.
I.2
Token-Level vs. Sentence-Level Relevance
Given the nature of our input data, which consists of long financial reports (averaging around 3800 tokens per context), providing token-level relevance scores directly to a human analyst can be overwhelming and impractical for actionable insights. A raw sequence of 3800 token relevance scores does not immediately highlight the key information at a glance. Therefore, focusing on individual token relevance is not the most pertinent approach for interpretability in this context. Instead, we emphasize sentence-level relevance. Financial analysts typically review reports section by section, and understanding which sentences or clauses are most indicative of fraud is significantly more valuable than knowing the exact contribution of every single token. Sentence-level aggregation allows for a higher-level summary of the model’s reasoning, making the explanations more digestible and actionable.
I.3
Sentence-Level Relevance Aggregation
To derive sentence-level relevance from the token-level scores, we implemented a simple yet effective aggregation method: • Summing of Relevance Scores: For every sentence, we sum the absolute relevance scores of all tokens belonging to that sentence. Using the absolute sum helps identify sentences that strongly contribute, either positively or negatively, to the fraud prediction. • Normalization and Ranking: The summed relevance scores for sentences are then normalized and ranked. Sentences with higher absolute summed relevance are considered more impactful on the model’s prediction. This aggregation provides a concise summary of the most relevant sentences within a lengthy financial disclosure, allowing users to quickly pinpoint suspicious statements or critical pieces of information that drove the model’s classification decision.
I.4
Highlighted Examples
For visual interpretability, we generate PDF heatmaps that highlight the most relevant portions of the input. These heatmaps use color intensity to represent the magnitude of a token’s (or aggregated word’s) relevance score, with different hues indicating positive or negative contributions to the fraud prediction.
Figure 13: Example of a highlighted financial report section indicating token-level relevance for a fraud prediction. Red indicates higher positive relevance towards a ”Fraud” prediction, while blue indicates negative relevance.