ConceptioArchivearXiv CS
arXiv CSopen access

Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach Besim Shala, Peter Mandl, Andreas Humpe, and Martin Häusl

arXiv:2607.14957v1 [cs.AI] 16 Jul 2026

University of Applied Sciences Munich, Germany {besim.shala, peter.mandl, andreas.humpe, martin.haeusl}@hm.edu

Abstract—Online firestorms are rapid collective escalations of highly negative user-generated content and may cause substantial reputational and economic damage. Existing detectors usually work with volume signals, sentiment scores, or predefined linguistic features. Such signals are useful, but they capture contextual meaning shifts in evolving discussion threads only indirectly. This paper proposes an LLM-based detection system with two operating modes. The first mode classifies complete Reddit threads retrospectively by combining local chunk-level assessments into a thread-level judgment. The second mode processes threads sequentially and issues early warnings when a sliding window exceeds calibrated thresholds. In this mode, the language model estimates three firestorm indicators: negativity share, escalation level, and contributor count. On a balanced Reddit dataset, the global mode achieves strong classification performance, while the early warning mode reaches high recall and detects escalating threads after only a small number of comments and distinct contributors. The results indicate that LLMs can be used not only for static judgment tasks, but also as repeated estimators in context-aware monitoring of social media discourse. Index Terms—Online firestorms, social media monitoring, early warning systems, large language models, LLM-as-a-Judge, crisis detection, Reddit, computational social science.

I. I NTRODUCTION Online firestorms, often referred to as “shitstorms” in German language discourse, are sudden waves of highly negative user generated content directed at organizations, brands, or individuals across social media platforms [1], [2]. They may escalate within hours and can erode both reputation and public trust [3]. Their consequences are not limited to perception. Evidence from CSR related firestorms indicates a moderate positive correlation between social media sentiment and stock market performance [4], which suggests that negative sentiment during such events may also be associated with measurable downside risk. Since firestorms usually rise and subside quickly, companies need to react before negative sentiment gains further momentum [5]. Timely detection is therefore a central precondition for crisis response, especially where short term brand damage can translate into longer lasting reputational harm [3]. Existing detection approaches still rely heavily on aggregated volume signals or predefined linguistic features computed at the post level [6]. This is problematic because escalation is not only a property of isolated posts. It emerges through relations between comments, changes in tone, and the accumulation of mutually reinforcing contributions. Response oriented frameworks have considered sequential dynamics at the firm

level [7], but early detection itself is still commonly tied to static feature thresholds rather than to the discourse sequence that precedes escalation. Sentiment analysis is also fragile in this setting. Sarcasm, irony, and implicit outrage can degrade classification accuracy, with documented case studies reporting only 74% accuracy and positive sentiment misclassification rates above 90% in highly sarcastic discourse [8]. It remains unclear how meaning shifts within coherent discussion threads can be converted into actionable warning signals. This paper addresses that problem with a large language model based detector that supports retrospective thread classification and sequential early warning. Both modes use the model as a context sensitive evaluator rather than as a keyword matcher or volume counter. The resulting research question is: How can large language models be leveraged in a sequential, context aware architecture to enable reliable, early detection of online firestorms in social media discussion threads? The proposed system makes two contributions: • It translates large language model based evaluation into a deterministic and calibrated decision rule for firestorm monitoring. By using contextual model assessments rather than only keyword, sentiment, or volume signals, it supports warning decisions that can be integrated into real time monitoring workflows. • It extends the LLM as a Judge paradigm [9] from static evaluation tasks to dynamic sequential monitoring. The model repeatedly estimates escalation signals over evolving discussion threads, linking firestorm theory to operational indicators such as negativity share, escalation level, and contributor count. This provides a compact basis for evaluating LLM supported firestorm detection under both retrospective and early warning conditions. II. R ELATED W ORK This section reviews three bodies of work that frame the present study: conceptual definitions of online firestorms, prior detection and measurement approaches, and early indicators that can be used for sequential warning. A. Defining and Characterizing Online Firestorms Online firestorms are characterized as sudden discharges of large quantities of negative messages, typically manifesting as negative word-of-mouth and complaint communication directed

at a person, company, or group [1]. They represent crowd- dimensions: width (interaction volume), height (intensity of based outrage, potentially devastating storms of emotional negative sentiment), and length (duration of public discourse) and aggressive indignation in social media [2]. Events of [16]. this kind are further marked by high message volume and C. Early Indicators of Firestorm Escalation a distinctly negative, indignant opinion climate [1], [10]. The literature also points to several signals that occur early in Participation is less the result of reflective evaluation than firestorm escalation. For system design, three recurring patterns of the cursory, reactive state characteristic of scrolling through are especially useful because they are theoretically grounded online communities, with the perception of being part of a and can be estimated from thread text. They are not treated as collective actor serving as the strongest driver of engagement a causal sequence, but as complementary warning signals for [11]. Importantly, firestorms are primarily characterized by the sequential mode developed in this paper. active attack strategies such as public complaining, negative The first is the degree of negativity in the discourse. word-of-mouth, and aggressive commentary rather than by A sharp increase in negative comments within short time silent boycott or patronage reduction [12], although heightened spans is considered a defining characteristic of firestorm wrongness judgments toward firms can simultaneously foster onset [1]. The emerging opinion climate is strongly negative both vindictive complaining and patronage reduction [13]. and indignant, distinguishing firestorm threads from ordinary For the purposes of this study, the term online firestorm refers to a discussion thread in which collective outrage, negative critical discussions through the density and intensity of hostile condensation, and escalating tonality are jointly observable, contributions rather than their mere presence [10]. The second indicator concerns linguistic and emotional following the criteria established in the foundational literature escalation. Moral-emotional language significantly amplifies [1], [2]. This operational definition guides both dataset labeling the diffusion of content in social networks, thereby accelerating and system design. collective outrage dynamics [17]. Early shifts in language use, specifically a decline in self-referential phrasing and a rise in B. Detection and Monitoring Approaches Prior work has translated firestorm indicators into several negatively charged, emotionally loaded expressions, can signal types of detection and monitoring systems. An automated an impending firestorm before volume-level escalation becomes real-time firestorm detector combining monitoring, sentiment visible [18]. The third indicator is the number of actively contributing analysis, and a statistical detection stage inspired by epidemicommenters. Firestorms are characterized not only by negative ological surveillance has been proposed, capturing activity tonality but by the participation of many distinct users within in short time intervals and comparing it against expected a brief period [10], consistent with the piling-on dynamic baselines to trigger an alarm when deviations are statistically that distinguishes collective outrage from isolated negative significant [14]. Firestorm potential in brand communities commentary [2]. has been investigated through text mining, operationalizing The three indicators considered here are negativity share, detection via high-arousal and low-arousal emotion indicators, escalation level, and contributor count. Together, they inform linguistic style matching, and tie strength between community the design of the sequential early warning mode. Rather than members, yielding concrete monitoring guidelines [7]. The claiming to identify a definitive or exhaustive set of firestorm early detection of company-related firestorms on Twitter has precursors, the goal is pragmatic: to select a small number of been examined by systematically evaluating feature variants that theoretically motivated, estimable signals that together enable change significantly within the first hours, outlining an hourly reliable early detection in practice. monitoring logic in which incoming posts are continuously assessed for anomalous deviations [6]. Lexical change in D. Research Gap and Positioning negative electronic word-of-mouth has been operationalized This literature leaves a specific gap at the level of sequenas an early indicator, demonstrating that measurable linguistic tial thread analysis. Existing approaches either depend on shifts precede network-level escalation signals such as retweets aggregated volume signals and predefined linguistic features, or mentions [15]. A framework distinguishing between trigger which only approximate contextual meaning shifts, or they features observable at onset and evolving firestorm features has are validated on narrow platform specific datasets with limited further been proposed, providing both early warning indicators transferability. To the best of our knowledge, context sensitive and metrics for tracking ongoing escalation [3]. LLM evaluation has not yet been examined as part of a sliding Beyond detection, researchers have also sought to measure window architecture for continuous thread level early warning. and categorize the severity of firestorms once they occur. The present study therefore treats firestorm detection as a Large-scale media analytics have been used to establish a sequence problem. Discussion threads are processed as evolving firestorm scale, showing that automated analysis can efficiently discourse structures rather than as collections of independent process large volumes of social media data but typically reaches comments. accuracies of only between 70% and 75% [8]. A standardIII. M ETHODOLOGY ized Social Media Firestorm Scale has since been proposed, drawing an analogy to the Saffir-Simpson hurricane scale The methodology follows the unit of analysis used by the and operationalizing firestorm severity along three measurable detector. It starts from complete Reddit threads, assigns one

TABLE I S UBREDDIT D ISTRIBUTION BY T HEMATIC C ATEGORY

Category Politics Consumer Entertainment News Science Meta

FS NFS 40 3 15 0 14 0 12 7 0 85 9 5

Gaming Total

10 100

0 100

Example Subreddits conservative, socialism anticonsumption, fuckamazon popculture, fauxmoi worldnews, nottheonion askscience, askacademia subredditdrama, changemyview chess, codwarzone

TABLE II D ESCRIPTIVE S TATISTICS BY C LASS

Metric FS (n=100) Avg. comments 176.7 (±129.3) Med. comments 136.0 Avg. contributors 123.8 (±92.7) Med. contributors 97.0 Avg. length (ch) 146.2 (±61.8) Med. length (ch) 131.4

NFS (n=100) 96.6 (±73.3) 80.5 66.4 (±54.0) 57.5 328.1 (±117.6) 314.2

FS = firestorm; NFS = non-firestorm; Avg. = arithmetic mean (M ); Med. = median; SD = standard deviation. Values in parentheses report standard deviations.

FS = firestorm; NFS = non-firestorm threads.

global label per thread, and then uses the same corpus to evaluate both retrospective recognition and early warning. A. Reddit Dataset

by consistently neutral discourse norms and fact-oriented discussion, providing a stable reference baseline. Firestorm threads, by contrast, were sampled across a broader spectrum of topical contexts to ensure that the detector is evaluated against diverse escalation patterns rather than communityspecific artifacts. Table II summarizes the structural properties of the dataset. Firestorm threads exhibit substantially higher participation volume, with nearly twice as many comments and unique contributors on average, while individual comments are markedly shorter on average (M = 146.2, SD = 61.8 vs. M = 328.1, SD = 117.6 characters). Here, M denotes the arithmetic mean and SD denotes the standard deviation. Each thread was labeled with a single, global binary annotation, distinguishing firestorm from non-firestorm cases at the thread level. The labeling followed the definitional criteria of collective outrage, negative condensation, and escalating tonality [1], [2]. An explicit annotation of the exact escalation onset was deliberately omitted, as such a point is empirically ambiguous and highly dependent on interpretive choices. This labeling strategy is consistent with principles of reliable content coding, which require that categorical decisions rest on clearly formulated criteria applied consistently [24].

The detector requires labeled discussion threads rather than isolated posts. Firestorm and non-firestorm cases therefore need to be represented in comparable numbers, and each thread must retain its sequential structure, including comment order and distinct contributors. The corpus also has to cover different topics so that performance is not driven by artifacts of a single community. The dataset consists of 200 publicly available Reddit discussion threads, equally divided into 100 firestorm and 100 nonfirestorm cases. Reddit was chosen as the data source because the platform organizes discussions as coherent comment threads within thematically bounded communities, i.e., subreddits, enabling the observation of collective escalation dynamics within clearly defined topical contexts [19]. As a text-centered platform, Reddit is particularly suited for the early detection of firestorms, as collective outrage dynamics can emerge through rapid and synchronous escalation of linguistic content [20]. Empirical studies further confirm that Reddit user data exhibits quality comparable to that of traditional research samples in B. System Architecture terms of reliability and validity [21]. The detector has two modes with the same thread level Data collection followed a passive approach, relying exclu- reference label but different decision points. The Global sively on already existing, publicly accessible content without Recognition Mode uses the complete thread and provides a intervening in ongoing discussions [22]. Thread selection retrospective performance baseline. The Early Warning Mode employed purposive sampling [23], guided by established defini- processes a thread while it grows and asks when a warning tions of online firestorms [1], [2]. Subreddits were exploratively can be issued. Numerical parameters were selected through screened across a spectrum of politically and thematically iterative empirical testing; they are practical optima for this diverse communities to ensure that clearly escalated threads, setup and are not claimed to be globally optimal. The system borderline cases, and uneventful discussions were represented. prompts were refined during development and are reported in Non-firestorm threads were drawn from predominantly neutral the Appendix. subreddits characterized by fact-oriented discourse and clear discussion norms (e.g., r/askscience, r/changemyview), serving C. Global Recognition Mode as reference cases to assess the detector’s specificity. The global mode treats the Reddit thread as a complete Table I provides an overview of the subreddit distribution document and returns one binary thread level classification. across both classes, grouped by thematic category. Because discussion threads frequently exceed the practical conThe asymmetric distribution of subreddits across classes text limits of current Transformer models, whose self-attention reflects the deliberate sampling strategy. Non-firestorm threads mechanism scales quadratically in compute and memory with were drawn from a small number of communities characterized sequence length [25], the input is segmented into chunks up to

Fig. 1. Global Recognition Mode pipeline with local LLM-based chunk assessment (Stage 1) and global thread-level classification (Stage 2).

Fig. 2. Two-stage Early Warning Mode with threshold calibration (Stage 1) and sequential warning detection via sliding windows (Stage 2).

12,000 characters each. This segmentation proceeds comment by comment in chronological order, preserving local discourse coherence within each chunk while ensuring full coverage of the thread. Research on chunk-based representations confirms that segmentation can preserve essential semantic content while reducing input size [26]. The overall architecture is illustrated in Fig. 1. Each chunk is then independently assessed by a large language model (LLM) operating under a fixed system prompt that encodes the firestorm definition and assessment criteria [1], [2] (Stage 1 in Fig. 1). The prompt-based assessment follows the paradigm of prompt-based learning, in which inputs are transformed into a prompt template with defined slots and the model output is mapped to the desired labels via a fixed answer space [27]. For each chunk, the LLM produces two outputs: a local binary judgment (firestorm / no-firestorm) and a concise summary capturing the conflict-relevant content of the segment. In the second stage (Fig. 1, right panel), all chunk summaries and their associated local labels are concatenated in chronological order into a compressed meta-description of the entire thread. This condensed representation is then submitted to the LLM in a final call that produces the global threadlevel classification together with a brief justification. This two-stage architecture follows the principle of hierarchical context merging, in which local segment representations are progressively integrated into a global representation [28]. Song et al. [29] demonstrate with HOMER that such hierarchical merging of chunk-level information improves long-context task performance compared to baselines, supporting the practical viability of this approach.

of increasing size are formed in fixed steps of 10 comments, up to a maximum of 100 comments. This approach draws on the overlapping chunking principle established in the Retrieval-Augmented Generation literature, where consecutive text segments share content at their boundaries to preserve contextual continuity across the analysis [30]. For each window, the LLM estimates three early indicators in the role of a standardized judge (LLM-as-a-Judge), following the paradigm in which LLMs produce structured numerical scores within a prompted evaluation context [9]. The three indicators are grounded in the communication science literature on firestorm precursors (see Section II): • Negativity share (neg_share): the estimated proportion of negative comments in the current window. • Escalation level (esc_level): the degree of reinforcement, piling-on, and intensification. • Contributor count (num_contrib): the number of distinct users whose comments are classified as conflictamplifying within the window. 1) Threshold Calibration via Grid Search: Based on the calibration data, a grid search [31] is performed over candidate threshold combinations for all three indicators plus a positive status signal (emerging or full firestorm). The grid search scanned the following candidate values:

D. Sequential Early Warning Mode The early warning mode adds a temporal perspective. It does not wait for the thread to be complete, but evaluates the growing discussion and triggers a warning once the calibrated conditions are met. The architecture is divided into two stages: a calibration pipeline and a runtime pipeline, as illustrated in Fig. 2. In the calibration pipeline (Stage 1 of Fig. 2), a balanced subsample of 80 threads (40 firestorms, 40 non-firestorms) was used to empirically derive a fixed decision rule. For each thread in this subsample, overlapping sliding windows

neg share ∈ {0.30, 0.40, 0.50, 0.60, 0.70}, esc level ∈ {0.20, 0.30, 0.40, 0.50, 0.60},

(1)

min contrib ∈ {3, 5, 8, 12, 15}. The grid therefore contains 125 candidate threshold combinations. The selected combination maximizes the following scoring function: score = 2 RecFS − FPR − λc̄det ,

(2)

where RecFS is the recall on the firestorm class, FPR is the false-positive rate, c̄det is the mean number of comments until detection, and λ is a mild penalty weight. This explicit prioritization of recall reflects the asymmetric cost structure of firestorm detection, where missing an escalation carries greater risk than triggering a false alarm [7]. The choice of thresholds on a separate calibration set follows the standard two-stage approach of first producing continuous scores and

Algorithm 1 Early Warning Runtime Pipeline Require: Thread T = (ci )N i=1 , thresholds θ = (τn , τe , τk ) Require: step size s = 5, maximum horizon H = 100 Ensure: warning flag warn and detection point d 1: F ← {emerging, full} ▷ firestorm statuses 2: warn ← false; d ← ⊥; pos ← s 3: while pos ≤ min(N, H) and ¬warn do 4: W ← (c1 , . . . , cpos ) ▷ current prefix window 5: (n, e, k, q) ← LLMJudge(W ) ▷ estimate indicators 6: negOK ← (n ≥ τn ); escOK ← (e ≥ τe ) 7: cntOK ← (k ≥ τk ); stateOK ← (q ∈ F ) 8: if negOK ∧ escOK ∧ cntOK ∧ stateOK then 9: warn ← true; d ← pos ▷ issue warning 10: end if 11: pos ← pos + s ▷ advance window 12: end while 13: return warn, d Note. n denotes negative share, e escalation level, k contributor count, and q thread status. The value d = ⊥ indicates that no warning was triggered.

TABLE III API C ALL C OMPLEXITY PER T HREAD Mode Global Global Global Early Warning Early Warning

Stage Chunk assessment Thread classification Total Calibration

Calls per thread ⌈L/C⌉ 1 ⌈L/C⌉ + 1 ⌊min(N, 100)/scal ⌋

Runtime

1 to ⌊min(N, 100)/srun ⌋

L = thread length in characters; C = chunk size (12,000 characters); N = number of comments; scal = calibration step size (10); srun = runtime step size (5). Runtime calls terminate early upon first threshold breach.

Table III summarizes the number of LLM API calls required per thread in each operating mode. In the global mode, computational cost scales linearly then selecting a threshold that maximizes a target metric on with thread length, as each chunk requires one independent validation data [32]. assessment call plus one final aggregation call. In the early Once the thresholds are fixed, the runtime pipeline (Stage 2 warning mode, the number of calls depends on both the thread of Fig. 2) applies them unchanged to previously unseen threads. length and whether the decision rule triggers early. In the best Each thread is processed in sliding windows of 5 comments case, a firestorm is detected after a single window evaluation; per step, up to a maximum of 100 comments. At each window in the worst case, the full sequence of windows up to the position, the LLM estimates the three indicators, and the fixed 100-comment horizon is processed without triggering. decision rule is evaluated. A warning is triggered as soon as all threshold conditions are jointly met; otherwise, the analysis B. Evaluation Metrics continues to the next window until either a warning is issued The evaluation is anchored at the thread level for both modes, or the thread ends. Algorithm 1 formalizes this procedure. with the global binary label (firestorm vs. non-firestorm) serving as ground truth. For the global mode, the system’s final threadIV. E XPERIMENTS level decision is compared against the ground truth label via the The experiments test the system in the two settings for standard confusion matrix (TP, FP, TN, FN). Precision, Recall, which it is designed. The first setting asks whether complete and F1-score are computed per class. Accuracy is reported as a Reddit threads can be classified reliably after the fact. The summary measure. Because the test set is balanced (100 threads second asks whether a warning can be issued early during per class), accuracy is not inflated by a dominant majority class, sequential processing. Evaluation therefore covers classification which makes it a meaningful aggregate in this setting [35]. Since the global mode does not involve a calibration step, all performance, API call complexity, and detection timeliness. 200 threads are used for evaluation. A. Experimental Setup For the early warning mode, evaluation is performed exAll experiments use OpenAI’s GPT-4o mini as the underlying clusively on the 120 threads (60 firestorm, 60 non-firestorm) language model. The model was chosen for pragmatic reasons: that were not used during threshold calibration, ensuring a the detector requires many repeated model calls, and GPT-4o strict separation between parameter tuning and performance mini was designed for cost efficient applications that chain or assessment. In addition to standard classification metrics, two parallelize calls at scale [33]. Models of this class can perform temporal measures quantify detection timeliness: complex assessment tasks from instructions alone and do not • Mean comments until detection: the mean number of require task specific fine tuning for the prompt based evaluation comments seen before the first warning is triggered across design used here [34]. all correctly detected firestorm threads. Model behavior is controlled through dedicated system • Mean contributing users until detection: the mean number prompts that define the assessment criteria and enforce of distinct contributing users at the point of first detection. structured JSON output via schema constraints, ensuring The false-positive rate (FPR) is reported separately as a that responses can be parsed automatically without manual dedicated measure of alarm burden. Together with Recall on intervention. In the global mode, the system prompt encodes the firestorm class, the FPR enables a pragmatic assessment the firestorm definition and specifies the output format for of the trade-off between early sensitivity and alarm reliability. both classification stages. In the early warning mode, a V. R ESULTS separate system prompt instructs the model to estimate the three numerical indicators together with a categorical status The results are reported separately for global thread classifiassessment. cation and early warning. Section VI then discusses how these

TABLE IV G LOBAL M ODE : C LASSIFICATION M ETRICS AND C ONFUSION M ATRIX

TABLE VI E ARLY WARNING M ODE : C ONFUSION M ATRIX

(a) Confusion Matrix Pred. FS 88 (TP) 5 (FP)

True FS True NFS

(b) Per-Class Metrics Class Prec. Firestorm 0.95 NFS 0.89

Rec. 0.88 0.95

Pred. NFS 12 (FN) 95 (TN)

F1 0.91 0.92

Supp. 100 100

FS = firestorm; NFS = non-firestorm. Overall accuracy: 0.915. TABLE V E ARLY WARNING M ODE : P ER -C LASS C LASSIFICATION M ETRICS Class Firestorm No-Firestorm

Prec. 0.82 0.98

Rec. 0.98 0.78

F1 0.89 0.87

Support 60 60

findings should be interpreted. A. Global Classification Table IV summarizes the classification performance and confusion matrix for the global recognition mode, evaluated on all 200 threads. Table IV reports both the confusion matrix and the per-class metrics. The global mode reaches an overall accuracy of 0.915. For the firestorm class, Precision is 0.95 and Recall is 0.88, which means that the system rarely assigns the firestorm label incorrectly but misses 12 of the 100 firestorm threads. The resulting F1-score is 0.91. For non-firestorm threads, Recall is higher at 0.95, while Precision is 0.89. The confusion matrix contains 88 true positives, 95 true negatives, 5 false positives, and 12 false negatives. This pattern shows that the global mode is somewhat conservative when assigning the firestorm label. B. Early Warning Tables V to VII present the classification metrics, confusion matrix, and temporal detection results for the early warning mode. The calibrated thresholds are reported in the text rather than in a separate table: min_neg = 0.4, min_esc = 0.2, and min_contrib = 3. The calibrated thresholds set the minimum negative share to 0.4, the minimum escalation level to 0.2, and the minimum contributor count to 3. Table V shows that Recall for the firestorm class reaches 0.98, meaning 59 out of 60 firestorm threads are correctly detected. Precision for the firestorm class is 0.82, reflecting the deliberate sensitivity bias introduced by the scoring function. The non-firestorm class exhibits the inverse pattern: Precision of 0.98 at Recall of 0.78. Table VI shows 59 true positives and 47 true negatives, with 13 false positives and only 1 false negative. The resulting FPR = 13/60 ≈ 0.217, meaning approximately 22% of non-firestorm threads trigger a warning. Table VII reports that on average the system triggers its first warning after 8.56 comments, involving a mean of 4.02 distinct contributing users at the point of detection. To

True FS True NFS

Pred. FS 59 (TP) 13 (FP)

Pred. NFS 1 (FN) 47 (TN)

FS = firestorm; NFS = non-firestorm. TABLE VII E ARLY WARNING M ODE : T EMPORAL D ETECTION M ETRICS Metric Mean comments until detection Mean comment ratio (det./thread avg.) Mean contributing users until detection

Value 8.56 4.8% 4.02

Comment ratio = 8.56 / 176.7 (mean firestorm thread length).

contextualize this figure, the mean firestorm thread in the test set contains 176.7 comments; detection at 8.56 comments thus occurs after approximately 4.8% of a thread’s eventual volume has materialized, indicating that the early warning mode responds well before the escalation pattern has fully developed. VI. D ISCUSSION AND L IMITATIONS This section discusses what the empirical results imply for firestorm detection, how they compare with prior approaches, and where the current validation remains limited. A. Discussion The results are encouraging in both operating modes, but they require a careful reading because the task combines classification, temporal detection, and operational risk. In the global mode, the hierarchical LLM-based classifier clearly exceeds the accuracy range between 70% and 75% reported for traditional automated sentiment analysis in firestorm monitoring [8]. The improvement is plausibly linked to two architectural choices. Chunk-level assessment preserves local discourse context that would be lost in comment-by-comment classification, a limitation also noted in contextual hate speech detection [36]. The subsequent merging step then combines these local signals into a thread-level judgment, in line with evidence that hierarchical merging can improve long-context performance [29]. This design makes it possible to represent piling-on behavior and tonal shifts across comment sequences, which dictionary-based and feature-engineered approaches [6], [7] capture only indirectly. The early warning mode addresses the harder operational question of timeliness. Its firestorm Recall of 0.98 shows that almost all escalating threads are flagged. The cost is a higher false-positive rate, but this trade-off is expected for early warning systems. Even under idealized conditions, early warning indicators involve a pronounced balance between sensitivity and false alarms [37]. In the firestorm setting, this sensitivity bias is reasonable when the cost of reviewing a

benign thread is lower than the potential reputational and [38]. The extent to which the observed detection performance financial cost of a missed escalation [7]. generalizes to platforms with different interaction structures 1) Structural Filtering and Detection Timeliness: The con- such as Twitter/X or Facebook remains an open question. tributor count threshold (min_contrib = 3) operationalizes Single-annotator bias. The ground-truth labels were asthe multi-actor criterion from firestorm theory [10]. A warning signed by a single annotator following the definitional criteria is not triggered by negativity alone; it also requires evidence of collective outrage, negative condensation, and escalating of collective dynamics. This adds a structural filter that pure tonality [1], [2]. While this is consistent with principles of sentiment assessment lacks. reliable content coding [24], the absence of independent second The temporal detection metrics further support the practical annotation means that systematic labeling biases cannot be fully value of the approach. Detection after an average of 8.56 ruled out. comments and 4.02 distinct contributing users indicates that Absence of escalation onset annotation. The deliberate the system responds to initial escalation signals well before omission of an exact escalation start point means that the broad participation has materialized. This aligns with the early warning mode cannot be validated against a precise theoretical observation that online firestorms are triggered by a temporal ground truth. The reported temporal metrics provide sudden, targeted collective brawl characterized by an immediate a pragmatic operationalization of detection timeliness but are shift towards aggressive language against a single entity [18]. approximations rather than precise measurements relative to Identifying the threat before broad participation is observable the true escalation moment. allows companies to consciously deploy the most effective Model dependence. The system relies on a specific LLM immediate mitigation strategy [5]. (GPT-4o mini). The stability of results across model versions, The comparison with prior work is therefore mainly concep- alternative providers, and prompt variations has not been tual. Volume-based detectors identify deviations from activity systematically evaluated. Since prompt-based assessments are baselines [14], while the present approach evaluates the content sensitive to instruction formulation [39], the exact threshold of the emerging discourse. Feature engineered approaches configuration may not transfer directly to other model variants. depend on predefined linguistic indicators [7]; the LLM-based However, the pipeline architecture is model-agnostic: the evaluation can also use implicit meaning, irony, and local calibration procedure can be re-applied to any model that context when judging whether a thread is escalating. produces structured numerical outputs. Reliability of LLM indicator estimates. The early warning B. Theoretical and Practical Contributions mode relies on a general-purpose LLM to estimate three numerAt the conceptual level, the study extends the LLM-as-a- ical indicators (neg_share, esc_level, num_contrib) Judge paradigm from static evaluation to sequential monitoring. from unstructured text. While structured JSON output conThe model is not asked to score a finished artifact once. Instead, straints and low-temperature sampling reduce variance, the it repeatedly estimates indicators over an evolving discussion accuracy of these estimates has not been validated against an and feeds these estimates into a deterministic decision rule. independent ground truth at the comment level. In particular, This connects qualitative firestorm theory with reproducible LLMs are susceptible to hallucination and instruction-following operational signals, namely negativity share, escalation level, inconsistencies [39], which may cause systematic over- or unand contributor count. derestimation of individual indicators. The observed calibration For organizations, the two modes map to different monitoring results provide only indirect evidence of sufficient estimability. tasks. The global recognition mode can be used retrospectively They show that calibrated thresholds perform well on heldto screen thread archives and identify past escalation events. out data, but they do not prove that each individual indicator The early warning mode is closer to an operational workflow: a is estimated accurately. A direct comparison between LLMwarning can surface a thread for human review, support a short estimated indicator values and human-annotated ground truth situation summary from the most recent window, and notify would strengthen confidence in the approach. roles such as communications, moderation, or legal teams. Text-only focus. The proposed approach relies exclusively Since thresholds are calibrated explicitly, the same procedure on textual content and does not incorporate author-level or can be repeated for organization specific data, community network-level signals such as social capital, follower reach, norms, or different risk tolerances. or tie strength between community members. Prior research suggests that such structural features can significantly influence C. Limitations the diffusion and escalation of firestorm content [40]. Several limitations are relevant when interpreting the results. They concern the data source, the annotation procedure, VII. C ONCLUSION temporal validation, model choice, and the range of signals available to the detector. This study evaluated an LLM-based firestorm detector on a Platform dependence. The dataset consists exclusively balanced dataset of 200 Reddit threads. The main finding is of Reddit discussion threads. Reddit-specific characteristics, that thread structure matters. When comments are treated as an including community norms, moderation practices, and vis- evolving discourse rather than as isolated posts, the system can ibility mechanisms, shape the observable discourse patterns capture piling-on dynamics, tonal shifts, and implicit escalation

patterns that are difficult to represent with predefined features or volume baselines alone. The global recognition mode achieves an F1-score of 0.91 and an accuracy of 0.915, which is above the accuracy range between 70% and 75% reported for conventional automated approaches [8]. The early warning mode detects 98% of firestorms after an average of 8.56 comments. This indicates that theoretically grounded early indicators can be estimated by a general-purpose LLM well enough to support calibrated and deterministic early warning. The findings therefore extend the LLM-as-a-Judge paradigm to a sequential monitoring context and provide a reproducible operationalization of qualitative firestorm theory. Future work should first test generalizability across further platforms and language contexts [38]. A systematic comparison across multiple LLMs would also clarify how sensitive the results are to model choice and prompt formulation [39]. The temporal evaluation could be strengthened with more finegrained escalation onset annotations. The decision rule could further be extended with uncertainty estimates, adaptive window sizes, or case based threshold calibration. Finally, the system should be studied in operational incident workflows, where warnings are routed to communications or moderation teams. In such settings, context sensitive LLM evaluation combined with explicit calibration offers a practical path toward earlier detection of collective escalation. ACKNOWLEDGMENT This work was supported in part by the Bavarian State Ministry of Economic Affairs, Regional Development and Energy under the BayVFP funding line Digitalization, program area Information and Communication Technology, through the project “NetiChecker” under grant number DIK-25070017//DIK0787/01. D ISCLAIMER Large Language Models were used as supporting tools for drafting selected text passages and for creating and revising figures. The authors retain full and sole responsibility for the content, cited sources, argumentation, and final version of this article. R EFERENCES [1] J. Pfeffer, T. Zorbach, and K. M. Carley, “Understanding online firestorms: Negative word-of-mouth dynamics in social media networks,” Journal of Marketing Communications, vol. 20, no. 1–2, pp. 117–128, 2014. [2] K. Rost, L. Stahel, and B. S. Frey, “Digital social norm enforcement: Online firestorms in social media,” PLOS ONE, vol. 11, no. 6, p. e0155923, 2016. [3] N. Hansen, A.-K. Kupfer, and T. Hennig-Thurau, “Brand crises in the digital age: The short- and long-term effects of social media firestorms on consumers and brands,” International Journal of Research in Marketing, vol. 35, no. 4, pp. 557–574, 2018. [4] F. Dias, P. Rita, N. António, and C. Vong, “Understanding the impact of online firestorms on financial performance in corporate social responsibility campaigns,” Marketing Intelligence & Planning, vol. 43, no. 6, pp. 1199–1219, 2025. [5] J. G. Qu, J. Yi, W. J. Zhang, and C. Y. Yang, “Silence is golden? Mitigating different types of online firestorms of Fortune 100 corporations on Twitter,” Public Relations Review, vol. 49, no. 5, p. 102391, 2023.

[6] K. Koch, A. Dippel, and M. Schumann, “Does my social media burn? – identify features for the early detection of company-related online firestorms on Twitter,” Online Social Networks and Media, vol. 25, p. 100151, 2021. [7] D. Herhausen, S. Ludwig, D. Grewal, J. Wulf, and M. Schoegel, “Detecting, preventing, and mitigating online firestorms in brand communities,” Journal of Marketing, vol. 83, no. 3, pp. 1–21, 2019. [8] K. Nuortimo, E. Karvonen, and J. Härkönen, “Establishing social media firestorm scale via large dataset media analytics,” Journal of Marketing Analytics, vol. 8, no. 4, pp. 224–233, 2020. [9] J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y. Wang, W. Gao, L. Ni, and J. Guo, “A survey on LLM-as-a-Judge,” arXiv preprint arXiv:2411.15594, 2024. [10] M. Johnen, M. Jungblut, and M. Ziegele, “The digital outcry: What incites participation behavior in an online firestorm?” New Media & Society, vol. 20, no. 9, pp. 3140–3160, 2018. [11] M. Gruber, C. Mayer, and S. A. Einwiller, “What drives people to participate in online firestorms?” Online Information Review, vol. 44, no. 3, pp. 563–581, 2020. [12] E. Delgado-Ballester, I. López-López, and A. Bernal-Palazón, “How harmful are online firestorms for brands? An approach to the phenomenon from the participant level,” Spanish Journal of Marketing – ESIC, vol. 24, no. 1, pp. 133–151, 2019. [13] T. K. H. Chan, Z. W. Y. Lee, D. Skoumpopoulou, and F. Situmeang, “Judging the wrongness of firms in social media firestorms: The heuristic and systematic information processing perspective,” Journal of the Association for Information Systems, vol. 25, no. 2, pp. 463–500, 2024. [14] B. Drasch, J. Huber, S. Panz, and F. Probst, “Detecting online firestorms in social media,” in Proceedings of the 23rd European Conference on Information Systems (ECIS), 2015. [15] W. Strathern, R. Ghawi, M. Schönfeld, and J. Pfeffer, “Identifying lexical change in negative word-of-mouth on social media,” Social Network Analysis and Mining, vol. 12, no. 1, p. 59, 2022. [16] K. Nuortimo, J. Harkonen, K. Breznik, and R. Hannes, “Developing a social media firestorm scale: From conceptualization to AI-assisted validation,” Journal of Marketing Analytics, 2025. [17] W. J. Brady, J. A. Wills, J. T. Jost, J. A. Tucker, and J. J. Van Bavel, “Emotion shapes the diffusion of moralized content in social networks,” Proceedings of the National Academy of Sciences, vol. 114, no. 28, pp. 7313–7318, 2017. [18] W. Strathern, M. Schoenfeld, R. Ghawi, and J. Pfeffer, “Against the others! Detecting moral outrage in social media networks,” in 2020 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 2020, pp. 322–326. [19] J. G. Qu, C. Y. Yang, A. A. Chen, and S. Kim, “Collective empowerment and connective outcry: What legitimize netizens to engage in negative word-of-mouth of online firestorms?” Public Relations Review, vol. 50, no. 2, p. 102438, 2024. [20] T. K. H. Chan, Z. W. Y. Lee, M. Pan, and K. Sun, “Understanding the drivers and outcomes of ideologically charged social media firestorms: The sociotechnical and social learning perspectives,” Journal of Management Information Systems, vol. 42, no. 3, pp. 737–766, 2025. [21] M. R. Jamnik and D. J. Lane, “The use of Reddit as an inexpensive source for high-quality data,” University of Massachusetts Amherst, Tech. Rep., 2017. [22] T. Rocha-Silva, C. Nogueira, and L. Rodrigues, “Passive data collection on Reddit: A practical approach,” Research Ethics, vol. 20, no. 3, pp. 453–470, 2024. [23] S. Campbell, M. Greenwood, S. Prior, T. Shearer, K. Walkem, S. Young, D. Bywaters, and K. Walker, “Purposive sampling: Complex or simple? Research case examples,” Journal of Research in Nursing, vol. 25, no. 8, pp. 652–661, 2020. [24] R. Artstein and M. Poesio, “Inter-coder agreement for computational linguistics,” Computational Linguistics, vol. 34, no. 4, pp. 555–596, 2008. [25] G. Qin, Y. Feng, and B. Van Durme, “The NLP task effectiveness of long-range transformers,” arXiv preprint arXiv:2202.07856, 2022. [26] Y. Li, S. C. Han, Y. Dai, and F. Cao, “ChuLo: Chunk-level key information representation for long document understanding,” arXiv preprint arXiv:2410.11119, 2024. [27] P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” arXiv preprint arXiv:2107.13586, 2021. [28] L. Ou and M. Lapata, “Context-aware hierarchical merging for long document summarization,” arXiv preprint arXiv:2502.00977, 2025.

[29] W. Song, S. Oh, S. Mo, J. Kim, S. Yun, J.-W. Ha, and J. Shin, “Hierarchical context merging: Better long context understanding for pre-trained LLMs,” arXiv preprint arXiv:2404.10308, 2024. [30] Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang, “Retrieval-augmented generation for large language models: A survey,” arXiv preprint arXiv:2312.10997, 2024. [31] J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” in Journal of Machine Learning Research, vol. 13, 2012, pp. 281–305. [32] Y. Nan, K. M. Chai, W. S. Lee, and H. L. Chieu, “Optimizing F-measure: A tale of two approaches,” arXiv preprint arXiv:1206.4625, 2012. [33] OpenAI, “GPT-4o mini: Advancing cost-efficient intelligence,” https: //openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence, 2024. [34] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman et al., “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023. [35] D. Chicco and G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, no. 1, p. 6, 2020. [36] J. M. Pérez, F. M. Luque, D. Zayat, M. Kondratzky, A. Moro, P. S. Serrati, J. Zajac, P. Miguel, N. Debandi, A. Gravano, and V. Cotik, “Assessing the impact of contextual information in hate speech detection,” IEEE Access, vol. 11, pp. 30 575–30 590, 2023. [37] C. Boettiger and A. Hastings, “Quantifying limits to detection of early warning for critical transitions,” Journal of The Royal Society Interface, vol. 9, no. 75, pp. 2527–2539, 2012. [38] N. Proferes, N. Jones, S. Gilbert, C. Fiesler, and M. Zimmer, “Studying Reddit: A systematic overview of disciplines, approaches, methods, and ethics,” Social Media + Society, vol. 7, no. 2, p. 20563051211019004, 2021. [39] Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba, “Large language models are human-level prompt engineers,” arXiv preprint arXiv:2211.01910, 2022. [40] H. Jöntgen, “Fueling the firestorm: Effects of social capital on users’ persuasiveness during online firestorms,” in Proceedings of the 28th European Conference on Information Systems (ECIS), Jun. 2020. [Online]. Available: https://aisel.aisnet.org/ecis2020 rp/179

Record · ID 373445 · SHA-256 e06c973db9672713
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.