CitePrism: Human-in-the-Loop AI for Citation Auditing and Editorial Integrity Gowrika Mahesh, Budanur Madappa Darshan Gowda, Kavana Gopladevarahalli Papegowda, Prajwal Basavaraj, Binh Vu, Swati Chandna, and Mehrdad Jalali* Applied Artificial Intelligence and Data Analytics, Department of Information SRH University Heidelberg, Heidelberg, Germany
arXiv:2605.16000v1 [cs.SI] 15 May 2026
*
Corresponding author: Mehrdad Jalali, [email protected]
Abstract—Editors and reviewers are expected to ensure that manuscripts cite relevant, accurate, current, and ethically appropriate literature, yet manuscript-level citation auditing remains largely manual, fragmented, and difficult to scale. Citation context—not reference counts alone—determines whether a citation substantiates a claim, while incomplete or inconsistent bibliographic sources can distort automated judgments. Bibliometric and classifier-based tools advance field-level or datasetlevel analysis, but rarely provide an integrated editorial workflow combining context extraction, metadata verification, self-citation review prompts, and human oversight. Large language models (LLMs) can interpret local citation neighborhoods, yet they require fact-checking, governance, and hybrid verification. We present CitePrism, a feasibility-stage hybrid decision-support prototype for editorial citation auditing that combines LLMassisted contextual reasoning, embedding-based semantic similarity, metadata verification, and integrity-oriented flags within a mandatory human-in-the-loop analyst workflow. CitePrism is not an autonomous misconduct detector, an automated accept/reject system, or a substitute for editorial judgment. In a singlemanuscript pilot validation (n = 104 references; pavement engineering), agreement with human binary relevance labels reached Cohen’s κ = 0.429 (moderate agreement). At operating threshold τ = 17, the prototype exhibited conservative screening in that case study: all human-labeled irrelevant citations were flagged, with additional false positives requiring analyst review. These findings indicate feasibility for editorial triage assistance only; they should not be interpreted as general editorial-performance validation. Multi-manuscript, multi-annotator studies remain required. Index Terms—citation auditing, editorial decision support, hybrid AI, large language models, research integrity, metadata verification, human-in-the-loop systems, publication ethics, scholarly communication
I. I NTRODUCTION Citations are infrastructure for scholarly trust. They anchor claims in prior evidence, make intellectual lineage visible, and enable readers and reviewers to scrutinize how authors situate new contributions within existing knowledge [1], [2]. When citation practice is weak—through irrelevant or superficial citations, distorted use of prior work, omission of foundational literature, excessive self-citation, or bibliographic inaccuracy—the interpretability and credibility of a manuscript may be compromised even when its core claims appear sound. Editorial citation checking is therefore important but undersupported. Editors and reviewers are expected to assess
whether each reference is relevant to its in-text claim, whether metadata are accurate, whether self-citation is proportionate, and whether key prior work has been acknowledged [3], [4]. This task is cognitively demanding: reference lists are long, reviewer time is limited, and citation problems are often subtle rather than obvious [2], [5]. Modern journal workflows face scale, reviewer fatigue, and integrity risks simultaneously. Submission volumes have increased across many fields [6], [7], intensifying triage burdens [8]. Citation padding, coercive citation, cartels, and metadata error remain concerns [9]–[12]. Editorial teams need accountable decision support that prioritizes attention without substituting for human judgment [13], [14]. Existing approaches remain fragmented. Bibliometric and network analyses map co-citation structure, collaboration, and thematic concentration at field level [15]–[17] but do not, by themselves, audit whether a specific in-text citation supports a local claim. Citation-intent and influence-classification research improves contextual and importance labeling [18]–[20]. Metadata infrastructures enable verification [21], yet source incompleteness and cross-database discrepancies can distort downstream analysis [12]. LLMs create new opportunities for contextual screening [22] but require human oversight, factchecking, and governance safeguards [14], [23]. A gap remains for a transparent, editorial-facing, manuscript-level workflow integrating citation-context extraction, hybrid relevance scoring, metadata verification, self-citation review prompts, missing-literature suggestions, and mandatory human oversight. CitePrism addresses this gap as a hybrid, human-supervised decision-support prototype— not an automated misconduct detector or autonomous editorial system. This article makes five contributions: 1) A manuscript-level hybrid citation-auditing framework for editorial decision support. 2) A transparent relevance-scoring model combining LLMassisted contextual reasoning and embedding-based semantic similarity. 3) A metadata-verification and integrity-flagging layer for DOI absence, metadata mismatch, missing abstracts, retraction signals, and self-citation review prompts.
4) A human-in-the-loop editorial analyst interface supporting threshold-based triage and citation-level rationale inspection. 5) Preliminary empirical validation on a 104-reference manuscript showing moderate agreement with human labels (κ = 0.429) and high sensitivity for irrelevantcitation detection at the selected operating point (τ = 17), with conservative false-positive behavior. II. E DITORIAL C ITATION AUDITING : P ROBLEM C ONTEXT AND R ESEARCH G AP Journal editorial workflows currently lack scalable, transparent, manuscript-level tools for citation auditing. Editors and reviewers cannot feasibly inspect every reference in depth, yet citation quality materially affects peer assessment and research integrity [3], [24]. Citation counts are insufficient. Bibliometric indicators and network maps summarize field structure, collaboration, and concentration [17] but do not establish whether a particular citation supports the claim in which it appears [2]. Citation context matters. Anderson and Lemken argue that rigorous review requires examining how cited works are used in context—whether citations are peripheral, substantive, supportive, critical, empirical, or potentially distorted [2]. Editorial auditing must therefore operate at claim level (intext neighborhoods), not only at bibliography level. Metadata quality matters. Citation analysis depends on complete and consistent bibliographic records. Rong et al. show substantial discrepancies between Web of Science and Crossref, including effects of merging sources on reference coverage, missing key nodes, and disciplinary variation in data quality [12]. Manuscript-level auditing must treat metadata verification as a prerequisite, not an afterthought. Self-citation detection is non-trivial. Self-citation is a normal part of scholarly communication but can be misused to inflate impact [4], [25]. Reliable detection requires authorname disambiguation; fuzzy matching alone is vulnerable to homonyms and variant spellings [25]. LLM-only approaches are risky. LLMs can articulate contextual fit but may hallucinate, encode bias, and produce unverifiable rationales without independent checks [14], [23]. Kasneci et al. emphasize continuous human oversight, critical evaluation, and governance for responsible LLM use [14]. Interpretable hybrid designs are needed. Citation influence and relevance classification benefit from transparent rationales and human-in-the-loop verification [20]. A practical editorial system should combine LLM-assisted reasoning, embedding similarity, metadata cross-checks, self-citation review prompts, and analyst-facing explainability. CitePrism is positioned to address this gap. Its objective is editorial triage assistance: helping analysts prioritize citations for closer inspection without automating accusations, rejections, or misconduct findings. III. R ELATED W ORK Table I summarizes how CitePrism relates to representative prior streams integrated in this review.
A. Citation quality, citation context, and scholarly trust Citation quality is increasingly treated as a researchintegrity issue [1], [9], [10]. Beyond ethics, how a work is cited matters. Anderson and Lemken propose citation context analysis (CCA) as a rigorous literature-review method that examines in-text citation passages to determine whether cited works are used peripherally or substantively, supported or critiqued, and whether knowledge claims have been empirically examined or distorted [2]. CCA goes beyond raw citation counts and bibliography inspection: it requires systematic attention to local usage at claim level. This perspective aligns with editorial citation auditing. A reference may appear in a bibliography yet fail to support the sentence in which it is invoked; conversely, a low-cited paper may be central to a specific argument. CitePrism operationalizes this insight through citation-context extraction (±1 sentences), hybrid relevance scoring, and a rationale viewer that lets analysts inspect model explanations against context and metadata—supporting human-supervised judgment rather than automated verdicts. B. Bibliometric data quality and metadata reliability Computational citation analysis depends on high-quality bibliographic and citation data. Rong et al. compare Web of Science (WoS) and Crossref in a large-scale case study, showing that the sources differ in coverage of high-impact literature, reference completeness, and network structure [12]. Merging datasets can improve citation-network completeness in some disciplines but may also introduce low-quality links and missing key nodes, with heterogeneous effects across fields. These findings directly motivate CitePrism’s metadataenrichment and metadata-mismatch detection layer. Parsed reference strings are cross-checked against OpenAlex and fallback APIs; discrepancies in title, year, or DOI are surfaced as analyst-review flags rather than silently overwritten [12], [21]. Manuscript-level auditing must remain aware that source-data incompleteness—including missing abstracts and incomplete Crossref records—can weaken both embedding and LLMbased relevance signals. C. Citation influence, citation intent, and interpretable citation classification Citation-intent corpora (e.g., SciCite, ACL-ARC) and frame-based models classify how authors cite prior work (background, method, comparison, etc.) [18], [19], [26]. Recent transformer-based approaches target citation influence or importance. Iqbal et al. present a SciBERT-based framework for citation influence classification with sparse rationale extraction, arguing that black-box classifiers are unsuitable for high-stakes evaluation and that human-readable rationales support verification [20]. CitePrism shares the emphasis on interpretability and human-in-the-loop review but differs in scope. Iqbal et al. focus on classifying citation importance within benchmark datasets; CitePrism targets manuscript-level editorial auditing,
combining hybrid relevance scoring (LLM + embeddings), metadata verification, self-citation flags, missing-citation suggestions, and configurable triage thresholds for editorial workflows. Citation influence and citation relevance to a local claim are related but not identical tasks—a limitation we discuss in Section VIII.
core contribution is manuscript-level triage: connecting citation context, metadata integrity, and hybrid relevance to editorial decision support rather than discipline-wide mapping.
D. Self-citation, author-name ambiguity, and integrity screening
CitePrism was designed around six principles: • Decision support, not automation. Outputs prioritize citations for human review; they do not issue misconduct findings or accept/reject recommendations. • Transparency. Per-citation scores, bands, rationales, and flags are inspectable; processing stages are logged. • Hybrid evidence. LLM-assisted reasoning is combined with embedding similarity and metadata verification. • Integrity-aware signaling. Self-citation, retraction, and metadata anomalies are surfaced as review prompts. • Configurable triage. Analysts adjust operating threshold τ for binary flagging separately from interpretive relevance bands. • Confidentiality awareness. Unpublished manuscripts may be processed via external APIs only under governed editorial policies; local deployment is recommended for sensitive workflows [14], [29].
Self-citation is common in cumulative research programmes but can be misused to inflate metrics [3], [4]. Ratio-only self-citation measures are insufficient because they ignore context and author identity. Mondal et al. propose self-citation detection combined with author-name disambiguation to assess researcher credibility on citations, noting that homonyms and spelling variants create false positives and false negatives when matching is naive [25]. Hybrid and ensemble strategies can improve reliability over single heuristics. CitePrism flags QUESTIONABLE SELF-CITE when author overlap coincides with low relevance scores, using fuzzy name matching (thefuzz) as a prototype approach. These outputs are review prompts, not misconduct determinations. Mondal et al.’s emphasis on disambiguation informs our limitation statement: robust editorial self-citation screening will require stronger identity resolution (e.g., ORCID-aware matching) in future versions. E. LLM-assisted scholarly workflows and responsible AI LLMs are being applied to literature synthesis, analytics, and review support [22], [27]. Kasneci et al. survey opportunities and challenges of LLMs for education, highlighting benefits for engagement and personalization alongside risks of bias, brittleness, and misuse [14]. They argue for human oversight, critical thinking, fact-checking, and clear governance— including transparency, privacy, and sustainable deployment considerations. Although Kasneci et al. do not address citation auditing directly, their governance framing applies to editorial LLM use. CitePrism treats LLM outputs as hypotheses: scores and rationales require cross-checking against embeddings and external metadata, and analysts retain accountability [14], [23], [28]. This hybrid posture contrasts with LLM-only screening pipelines. F. Bibliometric and network-based auditing perspectives Bibliometric mapping analyses co-citation, collaboration networks, thematic clusters, and concentration patterns to characterize research fields. Hassanein et al. illustrate this tradition in a bibliometric study of blockchain research in accounting and auditing, reporting collaboration structures, cocitation trends, thematic mapping, homophily, and Matthew Effect patterns [17]. Such work is valuable for field-level intelligence but does not substitute for per-citation editorial review. CitePrism optionally includes network-style diagnostics (e.g., venue concentration) as supplementary signals, yet its
IV. C ITE P RISM F RAMEWORK A. Design principles
B. System architecture Figure 1 situates CitePrism within a stylized editorial workflow from submission through decision, highlighting optional pre-review citation screening before formal peer review. Figure 2 shows the modular system architecture. C. Manuscript parsing and citation-context extraction Inputs: PDF manuscript; optional analyst configuration (reprocessing flags, threshold τ ). Outputs: Structured manuscript metadata; reference records; in-text citation contexts (target sentence plus ±1 neighboring sentences); processing status per stage. Following Anderson and Lemken’s emphasis on citation context rather than bibliography lists alone [2], CitePrism extracts in-text citation neighborhoods for each reference. PDF text is extracted with pypdf, with pdfminer.six as fallback; structured parsing uses Google Gemini 2.5 Flash. For robustness on long manuscripts, parsing follows a two-mode policy: single-call processing below 60,000 characters and segmented processing (up to 50,000 characters per segment) for longer inputs. Parsed artifacts are stored in SQLite for traceability and selective reprocessing. D. Metadata enrichment and verification In line with Rong et al.’s demonstration that WoS and Crossref differ in coverage and completeness [12], CitePrism treats metadata as a first-class audit object. Parsed references are enriched primarily via OpenAlex [21], with Crossref among the fallbacks. When abstracts are unavailable, a fourtier strategy applies: Semantic Scholar, Crossref, arXiv, and controlled publisher-page retrieval. A metadata consistency
TABLE I P OSITIONING C ITE P RISM RELATIVE TO PRIOR CITATION - ANALYSIS AND EDITORIAL - SUPPORT APPROACHES .
Stream / work
Main focus
Strength
Limitation for editorial auditing
CitePrism extension
CCA [2]
In-text citation usage and rigorous reviews
Claim-level interpretive depth
Context extraction + rationale inspection
Data quality [12]
WoS vs. Crossref completeness
Influence classification [20]
SciBERT + sparse rationales
Quantifies source discrepancies Interpretable importance labels
Selfcitation [25]
Disambiguation + credibility
Addresses name ambiguity
Responsible LLM use [14]
Governance and oversight
Bibliometric mapping [17]
Field-level networks/themes
Ethics framing for LLM deployment Macro structure and trends
Not an automated editorial workflow No manuscript UI or triage Benchmark classification, not editorial pipeline Not integrated with relevance/metadata Not citationspecific No per-claim manuscript audit
Manuscript-level screening support
CitePrism
Editorial citation-auditing support
Preliminary singlemanuscript validation
Combines rows above in one analyst-facing prototype
Editorial triage
Manuscript submission
Integrated hybrid workflow
CitePrism citation screening (optional)
Metadata verification + mismatch flags Hybrid scoring + editorial triage at τ
Fuzzy self-cite flags + relevance context
Hybrid verification + human accountability
Peer review
Editorial decision
Human analyst reviews flagged citations, rationales, and integrity prompts. No automated accusation or rejection. Fig. 1. Editorial problem context and positioning of CitePrism as optional, human-supervised citation screening before or alongside peer review.
check compares parsed and retrieved records using title similarity and year-tolerance rules. Suspected discrepancies, missing DOIs, missing abstracts, and retraction signals are retained as explicit analyst-review flags—reflecting that source-data incompleteness can distort automated relevance judgments. E. Hybrid relevance scoring CitePrism computes two complementary signals per reference: • Embedding score (RSembed ): Cosine similarity between manuscript and reference abstracts encoded with all-MiniLM-L6-v2 [30], rescaled to [0, 100].
LLM score (RSllm ): Structured batch judgments from Llama-3.1-8B-Instruct using manuscript abstract, citation neighborhood, and reference abstract, returning a numeric score, intent label, evidence snippet, and rationale. The fused relevance score is: •
RSfinal = 0.6 × RSllm + 0.4 × RSembed .
(1)
The 0.6/0.4 weighting is a prototype design choice reflecting greater reliance on contextual LLM reasoning while retaining an independent semantic anchor; it was not optimized in the present study and should be calibrated in future
Fig. 2. CitePrism system architecture: five processing stages, SQLite persistence (documents, processing logs, API cache), and an editorial analyst interface. Detailed interface screenshots are provided in Appendix A.
multi-manuscript validation. This design parallels interpretable citation-influence classification [20] but targets editorial relevance to local claims rather than dataset-level importance labels; the two tasks should not be conflated (Section VIII). Interpretive relevance bands (on RSfinal ): Relevant (≥ 70), Borderline (40 ≤ RSfinal < 70), Irrelevant (< 40). These bands support qualitative interpretation. Operating threshold τ : A separate, analyst-adjustable cutoff for binary triage (Flagged vs. Clean) in the review interface and evaluation module. τ does not redefine the three bands; it controls screening sensitivity for workflow prioritization. Figure 3 summarizes the hybrid scoring and risk-detection workflow.
F. Integrity-oriented risk flags and self-citation analysis Risk detection attaches review-oriented flags without altering RSfinal : RETRACTED, METADATA MISMATCH, MISSING DOI, QUESTIONABLE SELF-CITE, and missing-abstract warnings. Self-citation analysis combines fuzzy author matching (thefuzz) with overlap checks at author, team, and venue levels, informed by Mondal et al.’s observation that credible self-citation screening requires disambiguation beyond string overlap [4], [25]. Low-relevance self-citations are flagged for analyst review in line with COPE-informed interpretation [3]; flags are review prompts, not misconduct determinations.
G. Missing-citation suggestion The system proposes up to three candidate references not present in the bibliography, with short rationales, using manuscript title, abstract, and current reference list. Suggestions are generative hypotheses for expert verification and must not be treated as authoritative replacements for literature search. H. Human-in-the-loop analyst workflow The Streamlit-based interface supports upload, staged processing, side-by-side PDF inspection, threshold adjustment, citation tables with scores and flags, rationale viewing, diagnostics, and export. Following Iqbal et al.’s argument for interpretable rationales and human verification [20], and Kasneci et al.’s insistence on oversight and critical evaluation of LLM outputs [14], analysts are expected to inspect flagged items, compare rationales with citation context and metadata, override model judgments, and document decisions outside the tool. Interface screenshots, runtime logs, and extended diagnostics are provided in Appendix A (Figures A1–A15). V. E XPERIMENTAL S ETUP A. Case-study setting Validation was conducted as a pilot feasibility study limited to one manuscript. The test manuscript (Paper 1) is a machinelearning study on resilient modulus prediction in pavement engineering with 104 references. Human annotators assigned binary labels (1 = relevant; 0 = not relevant) using manuscript
Fig. 3. Hybrid relevance scoring and citation-risk detection workflow from manuscript ingestion through fused scoring, band assignment, integrity flags, and analyst triage at operating threshold τ .
context, producing gold_labels_paper1.csv. This design supports prototype feasibility assessment only; it does not establish general editorial performance, cross-domain robustness, or operational readiness for unsupervised deployment.
was not evaluated as an author-facing tool or as an automated integrity sanctioning mechanism.
B. Evaluation protocol The evaluation module aligns gold labels with scored references by reference identifier and computes Cohen’s κ [31], accuracy, precision, recall, and F1 at a selected operating threshold τ [32]. Cohen’s κ is emphasized because it adjusts for chance agreement. Three-band relevance categories remain defined on RSfinal independently of τ .
In this single-manuscript pilot, the prototype at τ = 17 achieved Cohen’s κ = 0.429, corresponding to moderate agreement on the Landis–Koch scale [32]. Accuracy was 0.721; macro-averaged F1 was 0.690; weighted F1 was 0.749 (Table II). Figure 4 shows the publication-style confusion matrix and metrics summary for this operating point. These values characterize one feasibility run and must not be read as proof of general screening accuracy across journals or domains. Confusion-matrix interpretation. All 21 human-labeled irrelevant citations were flagged (Flagged-class recall = 1.000). Twenty-nine human-labeled clean citations were also flagged (false positives). No human-labeled irrelevant citation was missed at this operating point. The system therefore exhibited conservative screening: prioritizing sensitivity to humanidentified irrelevant citations at the cost of additional analyst workload. These results should be interpreted as pilot-stage evidence of screening feasibility in one domain and annotation setting,
C. Human-in-the-loop review assumptions Binary metrics reflect analyst-triage behavior at τ , not editorial accept/reject decisions. The operating point τ = 17 was selected in the prototype evaluation interface as the point reporting κ = 0.429 for Paper 1; threshold calibration across disciplines and policies remains future work. D. Governance and deployment assumptions The case study assumes an editorial analyst (editorial staff member, integrity officer, or designated reviewer) with authority to inspect outputs and disregard false positives. The system
VI. R ESULTS
Paper 1 at = 17 ( = 0.429, accuracy=0.721)
21
0
Clean (1)
29
54
Human gold label
Flagged (0)
Flagged (0) Clean (1) CitePrism prediction Fig. 4. Pilot evaluation for Paper 1 at τ = 17: confusion matrix and summary metrics (κ = 0.429; n = 104).
TABLE II C LASSIFICATION METRICS FOR PAPER 1 AT OPERATING THRESHOLD τ = 17 (n = 104 REFERENCES ). Class
Precision
Recall
F1
Support
Flagged (0) Clean (1)
0.420 1.000
1.000 0.651
0.592 0.788
21 83
Accuracy Macro avg Weighted avg
0.710 0.883
0.721 0.825 0.690 0.721 0.749
104 104
Cohen’s κ = 0.429 (moderate agreement)
not as validation of general editorial performance, automated misconduct detection, or manuscript disposition.
VII. D ISCUSSION A. From citation counting to citation-context auditing Anderson and Lemken show that rigorous scholarship requires analyzing how cited works are used in context, not merely whether they appear in a reference list [2]. Iqbal et al. extend this logic computationally, demonstrating that citation influence classification benefits from interpretable rationales rather than opaque scores [20]. CitePrism translates these insights into an editorial workflow: extraction of local citation neighborhoods, hybrid relevance scoring, and analyst inspection of per-citation rationales. The single-manuscript pilot (κ = 0.429) suggests moderate alignment with human relevance judgments in one feasibility setting but does not resolve the inherent subjectivity of borderline citations or establish generalizable editorial performance.
B. Metadata quality as a prerequisite for reliable citation auditing Rong et al. demonstrate that WoS and Crossref can diverge substantially in coverage and network completeness, and that merged datasets introduce trade-offs between breadth and precision [12]. In Paper 1, 43 of 104 references lacked abstracts in the test run, weakening embedding and LLM signals. CitePrism’s metadata-mismatch flags are therefore not optional extras; they warn analysts when automated scores may rest on incomplete or inconsistent bibliographic ground truth. C. Self-citation flags as review prompts, not misconduct claims Mondal et al. highlight that self-citation detection for credibility assessment requires author-name disambiguation because homonyms and variants undermine naive matching [25]. CitePrism’s QUESTIONABLE SELF-CITE flag combines overlap heuristics with low relevance scores to prioritize review, in line with COPE-informed practice [3]. These signals must not be interpreted as accusations of citation manipulation; editors retain responsibility for contextual judgment, including legitimate programmatic self-citation. D. Why hybrid AI is preferable to LLM-only editorial screening Kasneci et al. caution that LLMs require continuous oversight, bias awareness, and fact-checking [14]. Iqbal et al. similarly argue against black-box classification in high-stakes settings [20]. CitePrism therefore combines LLM-assisted contextual scores with embedding similarity and metadata verification, exposing disagreements for analyst review. LLMonly editorial screening would risk hallucinated rationales and unverifiable triage decisions [23].
E. From bibliometric mapping to manuscript-level editorial decision support Bibliometric studies such as Hassanein et al. reveal field structure—co-citation, collaboration, thematic clusters, homophily, and Matthew Effect patterns [17]—but do not tell an editor whether a specific citation supports a claim. CitePrism complements macro mapping with micro-level triage: reference-quality checks, integrity prompts, and configurable operating thresholds for editorial workloads. F. Implications for journal editorial workflows In controlled pilot use—with mandatory human oversight— CitePrism could support: (i) editorial triage of manuscripts with unusually large or weak bibliographies; (ii) reviewer assistance through structured audit reports highlighting lowrelevance or metadata-anomalous citations; (iii) referencequality checks before acceptance; (iv) prioritization of questionable self-citation patterns for human review; (v) missingliterature suggestions as hypotheses for expert verification; and (vi) transparent audit logs for internal integrity processes. None of these uses should bypass human decision-making or author communication policies. G. Interpretation of conservative screening behavior At τ = 17, CitePrism favored false positives over false negatives: all 21 human-labeled irrelevant citations were flagged, with 29 false positives among clean labels (accuracy 0.721; weighted F1 0.749). For editorial triage, missing a problematic citation may be costlier than reviewing additional borderline cases. Threshold selection remains policy-dependent and fieldspecific. H. Future work Future validation should include: multi-manuscript and multi-domain studies; multiple independent annotators and inter-annotator agreement; stronger author-name disambiguation following Mondal et al. [25]; expanded metadata-source integration informed by Rong et al. [12]; field-specific calibration of τ and fusion weights; comparison against SciBERTstyle citation-influence baselines [20]; and local or private LLM deployment for confidential manuscripts [14]. VIII. L IMITATIONS Pilot scope. CitePrism is a feasibility prototype, not a validated production editorial system. Results derive from one 104-reference manuscript in pavement engineering; they must not be interpreted as general editorial-performance validation. Single manuscript and domain. Cross-disciplinary generalization is untested. Limited annotation setting. A single gold-label file was used; inter-annotator agreement and adjudication protocols were not evaluated. Interpretive citation context. Following Anderson and Lemken [2], relevance judgments are context-sensitive and may require domain expertise beyond automated scores.
Bibliographic source incompleteness. As Rong et al. show, Crossref and WoS differ in coverage and completeness [12]; missing abstracts (43/104 in Paper 1) and metadata errors can distort hybrid scoring. Self-citation and identity. CitePrism lacks robust authorname disambiguation; Mondal et al. demonstrate that credible self-citation analysis requires stronger identity resolution than fuzzy matching [25]. Task distinction. Citation influence/importance classification [20] is related to, but not identical with, editorial relevance to a local claim; metrics from one task may not transfer to the other. No SciBERT baseline comparison. The current validation does not compare CitePrism against SciBERT-style citationclassification baselines [20]. Prototype weighting and thresholding. The 0.6/0.4 fusion weights and τ = 17 were not systematically optimized. LLM fact-checking requirement. LLM rationales require analyst verification and must not be treated as authoritative [14], [23]. External APIs. Parsing and scoring rely on third-party services, raising privacy, reproducibility, and governance concerns [14], [29]. Generative suggestions. Missing-citation proposals require independent verification. Multi-manuscript validation needed. Larger studies with multiple annotators and manuscripts remain essential. IX. E THICAL , G OVERNANCE , AND D EPLOYMENT C ONSIDERATIONS Human oversight is mandatory [3], [14]. CitePrism must not automatically accuse authors of misconduct. Integrity flags— including self-citation prompts informed by Mondal et al.’s caution about identity ambiguity [25]—and low relevance scores are attention-prioritization signals only. Editors remain accountable; authors should not be penalized on automated scores alone [24]. LLM-generated explanations must be checked against citation context, metadata records, and external sources [12], [14], [20]. Kasneci et al. emphasize bias awareness, factchecking, transparency, privacy, and sustainable deployment; these principles apply directly to editorial use of LLMs on unpublished manuscripts [14]. Editorial offices need clear policies before pilot use: when screening occurs; which models and APIs are used; data minimization and retention; analyst training; and author communication pathways. Unpublished manuscripts require confidentiality safeguards; local or private deployment should be considered when third-party API processing is unacceptable [29]. CitePrism is intended for pilot-stage citation-auditing support in controlled editorial environments, not for unsupervised deployment or autonomous rejection workflows. X. C ONCLUSION Citation auditing is a core editorial-quality and researchintegrity problem, yet manuscript-level practice remains dif-
Manuscript + policy
CitePrism signals
Editorial analyst review
Journal policy and communication
Human accountability (no auto-sanctions) Fig. 5. Human-in-the-loop governance model for editorial deployment: automated signals inform analyst review; editors retain accountability and policy decisions.
ficult to scale with manual review alone. CitePrism contributes a transparent hybrid prototype for editorial decision support, combining LLM-assisted contextual reasoning, embedding similarity, metadata verification, integrity flags, and mandatory human-in-the-loop triage. It is not an autonomous misconduct detector or automated accept/reject system. A single-manuscript pilot (n = 104) showed moderate agreement with human labels (κ = 0.429) and conservative screening at τ = 17, capturing all human-labeled irrelevant citations in that run while generating additional false positives for analyst review. These findings support feasibility testing only and require multi-manuscript validation before stronger deployment claims. Future work requires multi-manuscript evaluation, multiple annotators, inter-annotator agreement analysis, disciplinespecific threshold calibration, ORCID-aware author disambiguation [25], comparison with SciBERT-style baselines [20], richer metadata governance informed by Rong et al. [12], and privacy-preserving deployment options [14]. Until such evidence is available, CitePrism should be understood as human-supervised decision support that augments—rather than replaces—editorial judgment. ACKNOWLEDGEMENTS The authors thank SRH University Heidelberg for academic support. AUTHOR C ONTRIBUTIONS (CR EDI T) Gowrika Mahesh, Budanur Madappa Darshan Gowda, Kavana Gopladevarahalli Papegowda, and Prajwal Basavaraj: Software, data curation, visualization, validation, and writing–original draft. Binh Vu: Methodological guidance, technical review, validation support, and writing–review and editing. Swati Chandna: Supervision, methodological guidance, and writing–review and editing. Mehrdad Jalali: Conceptualization, supervision, methodology, research-integrity framing, writing–review and editing, project administration, and corresponding-author oversight. C ONFLICT OF I NTEREST The authors declare no competing financial or non-financial interests relevant to this work.
DATA AND C ODE AVAILABILITY The source code for CitePrism is publicly available at: https://github.com/SRH-Heidelberg-University-ADSA/ CitePrism The repository includes the application source code, processing pipeline modules, Streamlit interface, SQLite schema, evaluation scripts, gold-label templates, README documentation, threshold guidance (Appendix C), hybrid-scoring pseudo-code (Appendix D), and a sample audit-report schema (Appendix E). Anonymized or synthetic demonstration materials are included where legally and ethically permissible. The copyrighted case-study manuscript (Paper 1) is not included. Interface screenshots and extended implementation evidence are provided in Appendix A of this document. AI AND T OOL -U SE S TATEMENT Large language model (LLM) systems were used only as components of the evaluated CitePrism prototype (document parsing and structured relevance reasoning). The manuscript text, analysis, interpretation, and final approval remain the responsibility of the human authors. All LLM-generated outputs discussed in this paper were subject to human review and are treated as decision-support signals rather than authoritative editorial judgments. R EFERENCES [1] M. Biagioli and A. Lippman, Gaming the Metrics: Misconduct and Manipulation in Academic Research. Cambridge, MA: MIT Press, 2020. [2] M. H. Anderson and R. K. Lemken, “Citation context analysis as a method for conducting rigorous and impactful literature reviews,” Organizational Research Methods, vol. 26, no. 1, pp. 77–106, 2021. [3] Committee on Publication Ethics, “COPE guidelines on authorship and contributorship,” 2023. [Online]. Available: https://publicationethics.org [4] M. Seeber, “Self-citations: Necessary evil or sign of limited scholarship?” Journal of the Medical Library Association, vol. 108, no. 4, pp. 658–661, 2020. [5] M. Kovanis, R. Porcher, P. Ravaud, and L. Trinquart, “The global burden of journal peer review and the literature on peer review,” Research Integrity and Peer Review, vol. 1, no. 1, p. 14, 2016. [6] L. Bornmann and R. Mutz, “Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references,” Journal of the Association for Information Science and Technology, vol. 66, no. 11, pp. 2215–2222, 2015. [7] M. Ware and M. Mabe, “The STM report: An overview of scientific and scholarly journal publishing,” International Association of Scientific, Technical and Medical Publishers, The Hague, Tech. Rep., 2015.
[8] F. Squazzoni, F. Grimaldo, and E. Mayson, “Publishing: Journals could share peer-review data,” Nature, vol. 546, no. 7657, pp. 209–211, 2017. [9] A. W. Wilhite and E. A. Fong, “Coercive citation in academic publishing,” Science, vol. 335, no. 6068, pp. 542–543, 2012. [10] I. J. Fister, I. Fister, and M. Perc, “Toward more reliable bibliometric statistics through detection of citation cartels,” Journal of the Association for Information Science and Technology, vol. 67, no. 3, pp. 625–634, 2016. [11] M. V. Simkin and V. P. Roychowdhury, “Read before you cite!” Complex Systems, vol. 14, no. 3, pp. 269–274, 2005. [12] G. Rong, Y. Chen, T. Koch, and K. Honda, “Assessing data quality in citation analysis: A case study of Web of Science and Crossref,” Journal of Informetrics, vol. 20, no. 1, p. 101775, 2026. [13] M. Breuning, J. Ishiyama, S. Fox, E. Gao, O. Ogunniyi, J. Padden, P. Reiter, and R. Whalen, “Why peer review failed during the COVID-19 pandemic,” PS: Political Science & Politics, vol. 54, no. 4, pp. 606–611, 2021. [14] E. Kasneci, K. Sessler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier, S. Krusche, G. Kutyniok, T. Michaeli, C. Nerdel, J. Pfeffer, O. Poquet, M. Schara, T. Seidel, M. Stadler, J. Weller, J. Kuhn, and G. Kasneci, “ChatGPT for good? on opportunities and challenges of large language models for education,” Learning and Individual Differences, vol. 103, p. 102274, 2023. [15] H. Small, “Co-citation in the scientific literature: A new measure of the relationship between two documents,” Journal of the American Society for Information Science, vol. 24, no. 4, pp. 265–269, 1973. [16] M. M. Kessler, “Bibliographic coupling between scientific papers,” American Documentation, vol. 14, no. 1, pp. 10–25, 1963. [17] A. Hassanein, K. B. Benameur, M. M. Mostafa, and W. Al-Shattarat, “Mapping the scientific research of blockchain technology in accounting and auditing: bibliometric analyses and a roadmap for future research,” Cogent Business & Management, vol. 12, no. 1, 2025. [18] S. Teufel, A. Siddharthan, and D. Tidhar, “Automatic classification of citation function,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2006, pp. 103–110. [19] A. Cohan, S. Feldman, I. Beltagy, D. Downey, and D. S. Weld, “Structural scaffolds for citation intent classification in scientific publications,” in Proceedings of NAACL-HLT, 2019, pp. 3586–3596. [20] A. Iqbal, M. I. U. Haq, A. Wahid, S. Muhammad, M. Yahya, and A. Shahid, “SciBERT-based interpretable framework for citation influence classification using sparse rationale extraction,” IEEE Access, 2026, early access. [21] J. Priem, H. Piwowar, and R. Orr, “OpenAlex: A fully-open index of the world’s scholarly works,” arXiv preprint arXiv:2205.01833, 2022. [22] Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al., “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, no. 3, pp. 1–45, 2024. [23] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023. [24] National Academies of Sciences, Engineering, and Medicine, Fostering Integrity in Research. Washington, DC: National Academies Press, 2017. [25] K. C. Mondal, S. Banerjee, A. Bhowmick, and R. Pal, “Identify Researchers’ Credibility on Citation Using Self-citation Detection by Author Name Disambiguation,” in Computational Intelligence in Communications and Business Analytics, ser. Communications in Computer and Information Science. Springer, 2026, pp. 424–438. [26] D. Jurgens, S. Kumar, R. Hoover, D. McFarland, and D. Jurafsky, “Measuring the evolution of a scientific field through citation frames,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 391–406, 2018. [27] J. Liu and N. H. Shah, “Can large language models replace human experts in systematic reviews?” Journal of Medical Internet Research, vol. 25, p. e50258, 2023. [28] Committee on Publication Ethics, “Artificial intelligence tools and resources,” 2023. [Online]. Available: https://publicationethics.org/resources/discussion-documents/ artificial-intelligence-tools-and-resources [29] D. B. Resnik and A. E. Shamoo, “Trust in science: A call for transparency and accountability,” Accountability in Research, vol. 27, no. 8, pp. 485–498, 2020.
Fig. A1. Application landing view with system status and active hybridanalysis configuration.
Fig. A2. Upload-and-process view with stage-level controls for standard or selective reprocessing.
Fig. A3. Live extraction and scoring logs for progress monitoring and troubleshooting.
[30] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proceedings of EMNLP-IJCNLP, 2019, pp. 3982–3992. [31] J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960. [32] J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,” Biometrics, vol. 33, no. 1, pp. 159–174, 1977.
A PPENDIX This appendix provides interface screenshots and runtime diagnostics referenced in the main text. Main-text Figures 1– 5 focus on editorial positioning, architecture, hybrid workflow, evaluation results, and governance.
Fig. A4. Pipeline status panel indicating completion states for parsing, enrichment, and scoring.
Fig. A5. Editorial analyst workspace with side-by-side PDF rendering and structured extraction review.
Fig. A8. Citation-level audit table with scores, binary flag status, self-citation markers, and missing-abstract indicators.
Fig. A9. Citation context and evidence viewer with rationale text and enriched metadata.
Fig. A6. Audit overview with headline indicators and adjustable operating threshold τ .
Fig. A10. Metadata consistency diagnostics highlighting parser-versusmetadata discrepancies.
Fig. A7. Summary panels: relevance distribution and analyst workload profile.
Three-band relevance is defined on RSfinal : Relevant (≥ 70), Borderline (40 ≤ RSfinal < 70), Irrelevant (< 40). These bands support interpretive reading of scores. Operating threshold τ controls binary Flagged vs. Clean triage in the analyst interface and evaluation module. Lower τ increases sensitivity (more flags); higher τ reduces analyst workload. In the Paper 1 pilot, τ = 17 yielded κ = 0.429 with full recall on human-labeled irrelevant citations in that run. Journals should calibrate τ against editorial policy (e.g., integrity screening vs. routine triage). for each reference r in manuscript.references: context = extract_citation_context(r)
Fig. A11. Citation-intent distribution derived from LLM outputs.
Fig. A15. Interactive citation-network visualization (manuscript–reference links and shared-author interconnections). TABLE A1 C ORE TECHNOLOGIES USED IN THE C ITE P RISM PROTOTYPE .
Fig. A12. Evaluation interface with gold-label upload and threshold control; all 104 references matched in the case study.
Component
Technology
PDF parsing LLM Scoring LLM Sentence embeddings Metadata primary source Metadata fallback Fuzzy matching Network visualization State management Report generation Frontend
Google Gemini 2.5 Flash Llama-3.1-8B-Instruct (HuggingFace) all-MiniLM-L6-v2 OpenAlex API Semantic Scholar, Crossref, arXiv, web retrieval thefuzz (Levenshtein ratio) PyVis, NetworkX SQLite (database/citeprism.db) fpdf2 (PDF), HTML Streamlit 1.53.0, Plotly
meta = enrich_metadata(r) # OpenAlex + fallbacks RS_embed = cosine(embed(manuscript.abstract), embed(meta.abstract)) RS_llm = llm_score(manuscript.abstract, context, meta.abstract) RS_final = 0.6 * RS_llm + 0.4 * RS_embed band = categorize(RS_final) # Relevant / Borderline / Irrelevant flags = integrity_checks(r, meta, RS_final) binary_flag = (RS_final < tau) { "manuscript_id": "hash", "reference_id": "ref_042", "RS_final": 28.5, "RS_llm": 22.0, "RS_embed": 38.2, "band": "Irrelevant", "flagged_at_tau": true, "tau": 17, "intent": "background", "rationale": "...", "flags": ["MISSING_ABSTRACT"], "self_cite": false
Fig. A13. Temporal currency analysis (69.2% of references from the most recent five years in Paper 1).
}
Fig. A14. Venue and author concentration summaries for diversity inspection.
TABLE A2 T HRESHOLD SENSITIVITY IN THE PAPER 1 PILOT (n = 104). ROW τ = 17a MATCHES THE PRIMARY RESULTS IN TABLE II. OTHER ROWS ARE RECOMPUTED FROM THE PUBLIC REPOSITORY SCORED EXPORT (P A P E R 1_ S C O R E D . J S O N ) AND G O L D _ L A B E L S _ P A P E R 1. C S V TO ILLUSTRATE SENSITIVITY; THEY ARE NOT INDEPENDENT VALIDATION RUNS .
τ
Acc.
P(Flag)
R(Flag)
F1(Flag)
Macro-F1
Wtd-F1
κ
#Flagged
10 15 17a 20 25
0.808 0.837 0.721 0.808 0.817
0.875 0.765 0.420 0.615 0.606
0.269 0.500 1.000 0.615 0.769
0.412 0.605 0.592 0.615 0.678
0.648 0.751 0.690 0.744 0.775
0.767 0.824 0.749 0.808 0.824
0.333 0.507 0.429 0.487 0.553
8 17 50 26 33