Overview of HIPE-2026: Person–Place Relation Extraction from Multilingual Historical Texts
arXiv:2606.25935v1 [cs.CL] 24 Jun 2026
Juri Opitz1[0000−0001−6892−4574] , Maud Ehrmann2[0000−0001−9900−2193] , Corina Raclé1[0009−0003−8842−8190] , Andrianos Michail1[0009−0004−1025−7851] , Matteo Romanello3,1[0000−0002−7406−6286] , and Simon Clematide1[0000−0003−1365−0662] University of Zurich, Switzerland École Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland 3 Swiss Federal Institute of Technology Zurich (ETHZ), Switzerland 1
2
Abstract. Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition of the HIPE evaluation series. Moving from named entity recognition and linking (HIPE-2020, HIPE-2022) to reasoning about relationships between entities, HIPE2026 targets two temporally grounded relation types: at, indicating that a person was present at a location at some point prior to a document’s publication date, and isAt, indicating presence contemporaneous with that date. This paper presents the results of the evaluation campaign, which confronted 17 participating teams with the challenges of historical language variation, OCR noise, and indirect contextual cues across three languages: French, German, and English. The datasets include historical newspaper text from the nineteenth and twentieth centuries, as well as a surprise-domain generalization set drawn from early modern French literary texts. A distinctive feature of HIPE-2026 is its three-fold evaluation framework, which assesses predictive accuracy, computational efficiency, and cross-domain generalization, reflecting the practical demands of large-scale historical document processing in the cultural heritage domain. Across more than 40 submitted runs, results reveal a wide range of strategies, from state-of-the-art large language models to lightweight task-specific classifiers, and highlight the trade-offs between accuracy, efficiency, and robustness inherent to historical relation extraction at corpus scale. System descriptions, datasets, and findings are presented and discussed, offering a detailed picture of the current state of temporally grounded relation extraction for historical documents. Keywords: Relation Extraction · Multilingual NLP · Historical Texts · Digital Humanities · Evaluation Campaign
1
Introduction
Historical documents are traces of the past, recording, among other things, what people did, where they went, and how they related to one another. The large-scale
2
J. Opitz et al.
digitization of historical newspapers and other cultural heritage collections has made these traces accessible at unprecedented scale, but typically yields texts that are noisy and marked by historical language variation, posing significant challenges for automatic processing. Robust information extraction from these sources – which span multiple languages, document types, and historical periods – is therefore a prerequisite for many downstream uses in digital history, computational humanities, and the social sciences [12,30,24,20,6,33,39]. Advancing and benchmarking such methods is the driving purpose of the HIPE evaluation series4 . HIPE-20265 is the third edition of the series [26]. While HIPE-2020 and HIPE-2022 focused on named entity recognition and linking in multilingual historical documents [10,11], HIPE-2026 moves the series toward deeper document understanding, advancing from the identification of entities to reasoning about the relationships between them. The central question is whether a text evidences a relation of physical presence between a person and a place, and whether that relation holds near the publication date of the document. In short: Who was where, and when? Such temporally grounded person–place relations underpin the reconstruction of individuals’ geographical and temporal trajectories, populating biographical knowledge graphs, supporting prosopographical research, and enabling spatial analysis of historical mobility. Advancing their extraction from historical documents is the core objective of HIPE-2026, in support of digital humanities scholarship. To this end, the task is deliberately designed to go beyond entity co-occurrence. On one side, apparent connections may be spurious: a person and a place mentioned in the same article may be unrelated, or linked only by affiliation rather than physical presence. On the other, genuine connections may be hidden: relevant evidence can be indirect (e.g. a person’s presence implied through an event whose location must be inferred from context), cross-sentential, or degraded by OCR errors. Beyond the quality of evidence, the task also requires temporal reasoning: systems must distinguish between a person’s presence at a place at some point in the past and their presence near the time of publication — a distinction that is often implicit in the text. HIPE-2026 therefore evaluates not only extraction of explicit locative statements, but also temporally grounded reasoning over text where evidence must be weighed rather than simply detected. Meeting this challenge reliably and at scale requires methods and models that are fit for cultural heritage processing: accurate enough to support scholarly use, scalable enough to process large historical collections, and versatile enough to generalize across domains and document types. For this reason, the campaign includes three complementary profiles: an accuracy profile on historical newspaper data, an efficiency profile that combines accuracy with computing resource metadata, and a generalization profile on a surprise French test set of literary and historiographical documents. The campaign attracted 17 participating teams, who submitted more than 40 runs across these profiles. 4 5
https://hipe-eval.github.io/ https://github.com/hipe-eval/hipe-2026
CLEF-HIPE-2026
3
This lab overview defines the task (Section 2), describes the data (Section 3), presents the evaluation framework and official metrics (Section 4), summarizes the participating systems (Section 5), reports the official results (Section 6), provides additional analyses (Section 7), and concludes with main lessons from the campaign (Section 8). All shared task artifacts are publicly available: datasets can be retrieved from https://github.com/hipe-eval/hipe-2026-data, the evaluation toolkit (including submissions and results) is available at https://github.com/ hipe-eval/hipe-2026-eval; Additionally, a paper with extended experiments will be released [27].
2
Task Definition
HIPE-2026 frames person–place relation extraction as a document-level classification task. Given a historical document and a set of candidate (person, location) pairs extracted from it, systems must determine what kind of relationship, if any, the text supports for each candidate pair. For each document, systems are given access to the full text, metadata including language and publication date, and up to 16 sampled candidate pairs. Each candidate pair consists of one person entity and one location entity drawn from the same document. Person and location entities are represented by clustered mentions and may optionally carry external Wikidata identifiers. Systems must assign labels for two relation types: at and isAt. – The at relation captures whether the document supports the claim that the person was at the place at some point before the publication date. It is a three-class task. The label TRUE indicates explicit textual evidence, PROBABLE indicates a plausible inference from contextual cues, and FALSE indicates absence of evidence or contradiction. – The isAt relation is a narrower temporal refinement of at. It captures whether the person is located at the place in the immediate temporal horizon of the document, operationalized in the guidelines as roughly up to one month before publication [25]. It is a binary task with labels TRUE and FALSE. Conceptually, isAt=TRUE presupposes a positive or at least plausible at relation. For example, if a text reports that a person is currently visiting Berlin, then both at and isAt should be positive. If the same text only mentions that the person visited Berlin the previous year, at may be positive while isAt should be FALSE. The official scorer nevertheless evaluates both relation types independently, so that inconsistent predictions can be measured rather than rejected. The PROBABLE label is central to the task design. It reflects an abductive view of interpretation, in which relevant meaning can be supported by the assumptions needed to make discourse coherent rather than by explicit surface statements alone [17]. In historical person–place relation extraction, such cases
4
J. Opitz et al.
Table 1. Overview of released training data and official test data. Label columns give counts for TRUE (T), PROBABLE (P), and FALSE (F). Split
Lang. Docs Pairs Pairs/doc
Years
at T
P
isAt F
T
F
Gold Train A
de en fr
34 35 35
466 307 478
13.71 1818–1948 135 8.77 1790–1960 125 13.66 1804–2018 181
Silver Train A
de en fr
64 23 293
893 206 4110
13.95 1798–1948 119 246 528 85 808 8.96 1790–1960 41 71 94 34 172 14.03 1798–2018 589 1010 2511 429 3681
Silver Dev A
de en fr
32 17 107
432 151 1498
13.50 1818–1948 41 147 244 29 403 8.88 1790–1920 29 54 68 18 133 14.00 1798–1988 179 367 952 127 1371
Gold Test A
de en fr
19 19 19
238 162 238
12.53 1808–1948 8.53 1800–1960 12.53 1804–1998
Gold Test B
fr
30
480
16.00 1542–1797 130
54 75 85
62 269 89 377 30 152 83 224 28 269 107 371
38 146 67 171 28 59 65 97 19 134 81 157 60 290
0 480
include event participation, institutional roles, travel reports, and geographically grounded narrative contexts.
3
Data
HIPE-2026 uses two evaluation domains. Domain A contains historical newspaper articles in German, English, and French, spanning the nineteenth and twentieth centuries, derived from the HIPE-2022 newspaper data, with manually created entity annotations and Wikidata links [11]. Domain B is a surprise test set consisting of French literature and historiographical works from the sixteenth to eighteenth centuries, included to assess out-of-domain generalization. Its entity annotations and Wikidata links come from the FreEMNER corpus (version 6) [14,13], which includes an extended and improved version of Presto core [3]. The at relationships were annotated on top of these for this shared task. Test A evaluates both at and isAt; Test B evaluates only at. The competition data (version 1.0) is archived on Zenodo6 and released on GitHub.7 Several preprocessing steps were applied to both domains. After conversion to the HIPE-2026 format8 , entity mentions without a Wikidata identifier, referred to as NIL mentions, were heuristically clustered at the document level if their character similarity exceeded 75%. Documents longer than 12,000 characters 6
https://zenodo.org/records/20615690 https://github.com/hipe-eval/HIPE-2026-data/releases/tag/v1.0 8 https://github.com/hipe-eval/HIPE-2026-data/blob/main/schemas/ hipe-2026-data.schema.json 7
CLEF-HIPE-2026
5
NOTICE. By virtue of a bil of sale issued by Tom Scarlett, Circuit Court Clerk for Putnam County, .ennessee, dated on the 6tU. day of January, 1920, I will expore to sale to tne highest bidder for cash on the 31st day of January, 1920, at 1:0; o’clock, P. M. at the Courthouse door in the Town of Cookeville, the following property, to-wit: A house and lot located in the 19th Civil District of Putnam County, Tennessee, containing one acre, more or less, bounded on the east by the lands of J. R. Watson; on the north by the T. C. R. K. Co.; on the west by the lands of Phy; on the south by the lands of J. R. Watson; levied on as the lands of D. Rittenberry, to satisfy a Judgment in favor of P. L. Ramsey ag-ih st him, with interest and all the cost.i of the case. This Jan. 6th, 1920. L. F. MILLER, Sheriff. Person
Location
Tom Scarlett Cookeville Tom Scarlett Putnam County J. R. Watson Tennessee D. Rittenberry Putnam County P. L. Ramsey Tennessee L. F. Miller Putnam County
at
at T/P/F
F 35.6/35.6/28.9 T 53.3/28.9/17.8 T 44.4/33.3/22.2 T 37.8/40.0/22.2 P 28.9/51.1/20.0 T 46.7/37.8/15.6
isAt isAt T/F F T T F T T
40.0/60.0 60.0/40.0 22.2/77.8 26.7/73.3 26.7/73.3 53.3/46.7
Fig. 1. Illustrative example from Test A: a newspaper excerpt published in the Putnam County Herald on 29 January 1920, together with its person–place relation annotations. Distribution columns report the percentages of submitted predictions over 45 predictions: T/P/F for at (TRUE/PROBABLE/FALSE) and T/F for isAt (TRUE/FALSE); bold values indicate the majority vote.
and statistical outliers in the number of locations or persons, defined as values more than three standard deviations above the mean, were then removed. From the remaining documents, person–location candidate pairs were selected from the Cartesian product, with a maximum of 16 pairs per document. Pairs were sampled with probability proportional to exp(−d/T ), where d is the text distance and T = 50 controls the degree of locality, favoring textually close pairs while still allowing occasional long-range ones. Detailed dataset statistics are provided in Table 1. The training and validation sets of Domain A were annotated by majority vote over three runs of GPT4.1. For each language, a subset of articles was subsequently human-annotated and released as Gold Train A data. A sample of 23 human-annotated articles was released early as illustrative examples for participants. The Domain A test set consists of 19 unreleased human-annotated articles per language. Since Domain B is used exclusively as test data, 30 documents were sampled and humanannotated for at, roughly matching the size of each Domain A language-specific test set. Human annotations were carried out by eight trained annotators, with at least two native annotators per language. Table 1 highlights three properties of the data. First, candidate relation extraction creates a substantial number of decisions per document: Test A contains 638 evaluated pairs across 57 documents, and Test B contains 480 evaluated pairs
6
J. Opitz et al.
[French Original] Si de ce vous esmerveillez : esmerveillez vous dadvantaige de la queue des beliers de Scythie : que pesoit plus de trente livres, & des moutons de Surie, esquelz fault (si Tenaud dict vray) affuster une charrette au cul, pour la porter tant elle est longue & pesante. Vous ne l’ avez pas telle vous aultres paillardes de plat pays. Et [la jument] fut amenee par mer en troys carracques & un brigantin jusques au port de Olone en Thalmondoys Lors que Grandgousier la veit, Voicy (dist il) bien le cas pour porter mon filz a Paris. Or ca de par dieu, tout yra bien. Il sera grand clerc on temps advenir. Si n’ estoient messieurs les bestes, nous vivrions comme clercs. Au lendemain apres boyre (comme entendez) prindrent chemin, Gargantua son precepteur Ponocrates & ses gens, ensemble eulx Eudemon le jeune paige. Et par ce que c’ estoit en temps serain & bien attrempé, son pere luy feist faire des botes fauves. Babin les nomme brodequins. Ainsi joyeusement passerent leur grand chemin : & tousjours grand chere : jusques au dessus de Orleans. [English Translation] If this astonishes you, be yet more astonished at the tails of the rams of Scythia, which weighed more than thirty pounds, and at the sheep of Syria, to which one must (if Tenaud speaks true) attach a cart to their hindquarters to carry it, so long and heavy is it. You have nothing of the sort, you hussies of the flatlands. [The mare] was brought by sea in three carracks and a brigantine as far as the port of Olonne in Talmondais. When Grandgousier saw it, he said: “Here is just the thing to carry my son to Paris. Come now, by God, all will be well. He will be a great scholar in time to come. Were it not for these gentlemen the beasts, we would live like scholars.” The next morning, after drinking (as you will understand), they took to the road: Gargantua, his tutor Ponocrates and his men, together with Eudemon, the young page. And since the weather was fair and mild, his father had him fitted with tawny boots — which Babin calls buskins. And so they passed merrily along their way, always in good cheer, as far as the outskirts of Orléans. Person
Location
Tenaud Scythie Grandgousier port de Olone Gargantua Paris Ponocrates Orleans Eudemon Orleans Babin Orleans
at
at T/P/F
F P F T T F
15.9/4.5/79.5 59.1/13.6/27.3 15.9/27.3/56.8 47.7/13.6/38.6 47.7/13.6/38.6 27.3/27.3/45.5
Fig. 2. Illustrative example from Test B: an excerpt from Rabelais’ Gargantua and Pantagruel (original French at the top and English translation below) and its person–place relation annotations. The distribution column reports the percentages of submitted at predictions over 44 predictions as T/P/F (TRUE/PROBABLE/FALSE); bold values indicate the majority vote.
CLEF-HIPE-2026
7
across 30 documents. Second, the class distributions are imbalanced: FALSE is the most frequent label, while PROBABLE is comparatively rare. Third, the surprise set is not merely another French test sample: it differs in both domain and period, and is therefore reported separately throughout the evaluation. To illustrate the annotation process, which requires fine-grained interpretation of contextual evidence, we provide an original English test set article in Figure 1. In the excerpt, person mentions are shown in bold and location mentions in italics. The annotations in Figure 1 illustrate several forms of evidence. The court clerk Tom Scarlett must be present in Putnam County shortly before the publication of the article in order to issue the bill of sale; therefore both at and isAt are set to TRUE. However, he may be present at any courthouse, not necessarily in Cookeville specifically, consequently both at and isAt are labeled as FALSE. Neighboring the estate for sale, the lands of J.R. Watson are indicated to lie within Tennessee, where he most likely also is present. It is important to note that at reflects the probability that an entity is located at a given place. Given a positive at relationship, isAt assesses whether the entity can be assumed to be physically present at that location within the time horizon of the article. The bill of sale is issued for D. Rittenberry’s lands, implying these lands lie within Putnam County; accordingly at is set to TRUE. Being involved in lawsuits, D. Rittenberry may not necessarily be present in Putnam County, thus isAt is marked as FALSE. The judgment was ruled in favor of P.L. Ramsey, suggesting that he is likely present in Tennessee. However, he might be a creditor from outside the state, therefore at is set to PROBABLE. If he indeed is located in Tennessee, he is most likely a fellow citizen of Putnam County, consequently isAt is marked as TRUE. Note the induced complexity due to long-range pronoun resolution: As indicated by the use of first-person pronouns and the signature of L. F Miller at the end of the article, L. F Miller is the sheriff of Putnam County and assumed present there, thus both at and isAt are set to TRUE. Notably, “I” does not refer to Tom Scarlett, despite his being textually closer. Since Domain B does differ significantly from the newspaper data of Domain A, we also provide an example text from Test B in Figure 2. Specifically, this is an excerpt from Gargantua and Pantagruel, a series of novels by François Rabelais written in the sixteenth century. The annotations of the surprise text may be explained as follows: There is no evidence Tenaud ever was in Scythia, only that he described Scythian rams, therefore at is marked as FALSE. It is unknown where Grandgousier is located. It is likely that he saw this mare at the port of Olonne, but the mare also could have been brought to his location, so we set at to PROBABLE to reflect this uncertainty. Gargantua is on his way to Paris but has not arrived, therefore at is FALSE for this relationship. Gargantua and his companions Ponocrates and Eudemon are however in the lands beyond Orléans, which indicates them having passed through Orléans. Thus, we mark their relationships with Orléans as TRUE. Babin is only mentioned in reference to what he calls a specific type of
8
J. Opitz et al.
boots, with no previous mention or any indication of being part of the travels, consequently we set at to FALSE.
4
Evaluation Framework
HIPE-2026 adopts three complementary evaluation profiles, accuracy, efficiency, and generalization – each targeting a different dimension of system performance. 4.1
Accuracy Profile
The official metric is macro-averaged recall, also known as balanced accuracy. For a label ℓ, recall is computed as Recall(ℓ) =
#correct predictions with gold label ℓ . #instances with gold label ℓ
The macro-averaged recall (MR) of a relation type is the arithmetic mean of the recalls of its labels. This choice gives equal weight to rare and frequent labels [23,31], which is important for less frequent labels such as at=PROBABLE and isAt=TRUE. For Test A, at is evaluated as a three-class problem and isAt as a two-class problem. The language-specific profile score is MacroRecallat + MacroRecallisAt . 2 The overall Test A ranking averages this score across German, English, and French for runs that submitted all three language files. Partial submissions are reported only in language-specific tables. For Test B, only at is evaluated, so the profile score is simply MacroRecallat . ScoreA =
4.2
Efficiency Profile
The efficiency profile ranks each run by combining its accuracy rank with resource ranks derived from organizer-normalized metadata: model parameter count and deployed model size. The official efficiency score is the mean of these three ranks: racc + rparams + rsize Reff = . 3 Lower efficiency scores are better. Runs with missing resource metadata are assigned the worst resource rank. 4.3
Generalization Profile
The generalization profile tests system accuracy on the surprise Test B, assessing the ability of systems to transfer beyond the historical newspaper domain to literary and historiographical texts of a different period and style. Only the at relation is evaluated, as a three-class problem, and the profile score is MacroRecallat .
CLEF-HIPE-2026
4.4
9
Scoring and Ranking
All profiles are scored using the evaluation toolkit available at https://github. com/hipe-eval/hipe-2026-eval. The toolkit validates submissions against the HIPE-2026 JSON schema and produces per-run diagnostic reports alongside the official scores. Systems are ranked within each profile independently; a run must cover all three languages of Test A to appear in the overall accuracy and efficiency rankings.
5
System Descriptions
Seventeen teams submitted at least one official run, alongside two organizerprovided baselines. The evaluation repository contains 185 prediction files in total. Most teams submitted complete Test A runs covering all three languages as well as Test B runs for the surprise set; one team submitted only an English Test A run. 5.1
Organizer Baselines
HIPE-Random-Baseline This baseline assigns labels by uniform random sampling from the label inventory of each relation type: TRUE, PROBABLE, and FALSE for at, and TRUE and FALSE for isAt. It does not use document text, metadata, entity information, or training data. HIPE-Ministral-Baseline This baseline is a minimal zero-shot system for person–place relation qualification, designed as a simple and reproducible reference point for the shared task. It uses Ministral-3-3B-Instruct-25129 , a 3-billionparameter instruction-following language model, executed locally in a GGUF setup with greedy decoding, i.e. temperature set to 0.0 [19]. The system formulates relation prediction as a prompted classification problem using a single instruction template. For each pre-identified person–location pair, the model is asked to assign labels for the two target relations, at and isAt, without taskspecific fine-tuning or few-shot examples. The system operates in a straightforward document-level loop over candidate pairs and produces structured JSON outputs. Its design intentionally avoids retrieval augmentation, external knowledge sources, and ensemble strategies, focusing instead on the behavior of a single generative model under deterministic decoding. The implementation is modular, allowing easy modification of the prompt, decoding settings, and inference routine. Overall, the baseline is intended as a transparent reference point that emphasizes simplicity and reproducibility, while providing a clean starting point for exploring more advanced prompting strategies, model adaptation techniques, or hybrid systems. 9
https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512
10
J. Opitz et al.
Table 2. Participating teams and coarse methodological characteristics of their official submissions. A check mark indicates that at least one submitted run from the team used the corresponding component. Supervised includes both classical supervised methods and task-specific adaptation of pretrained models, including fine-tuning, parameterefficient fine-tuning, and distillation. We use “–” to indicate missing information. Team
Country Runs LLM Superv. PLM enc. Feat./rules Ext. knowl.
Awakened BIU_NLP DS@GT FI-CODE FourBytes GippLab Hansel&Gretel INSA Lyon MaxFo-Ajie MILRIT ROSTI Rittik&Souvik Spinfo UMUTEAM VerbaNexAI I VerbaNexAI II whereami
RO IL US DE IN DE IN FR CN FR FR IN DE ES CO CO EG
3 3 3 3 1 3 3 3 3 3 3 1 3 3 2 3 2
HIPE-Random-Baseline CH HIPE-Ministral-Baseline CH
1 1
5.2
✓ ✓ ✓ – ✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓ ✓ – ✓ ✓ ✓
✓
✓
✓
✓ ✓ –
✓
✓ –
✓ ✓ ✓ ✓ ✓
✓
✓
✓ ✓ ✓
✓
✓ ✓ ✓ ✓
✓ ✓ ✓
–
✓ ✓ ✓ ✓ ✓
✓
Participating Systems
Awakened. The team explored both prompting-based and supervised approaches for multilingual person–place relation extraction [34]. Two submitted runs relied on Claude Sonnet 4 with carefully designed prompts and few-shot examples, while incorporating external knowledge from Wikidata, including biographical and geographical information associated with linked entities. Additional postprocessing rules enforced temporal consistency using birth and death dates. A third run employed a fine-tuned XLM-RoBERTa-large model augmented with knowledge-graph-derived and pattern-based features. The supervised model further incorporated a consistency-aware objective to discourage logically incompatible label assignments. The submitted runs therefore compared knowledgeenhanced prompting with feature-augmented multilingual encoder models. BIU_NLP. The team approached the task as a zero- and few-shot prompting problem without task-specific fine-tuning [18]. The team evaluated several open-weight instruction-following language models, including Gemma and Mistral variants. For each person–location pair, the models received a structured prompt requiring relation classification together with a short explanatory rationale. Experiments explored prompts written either in English or in the source document language, as well as different numbers of in-context examples. The submitted runs corresponded to different language model backbones and prompting configurations selected on development data.
CLEF-HIPE-2026
11
DS@GT. DS@GT proposed a lightweight relation extraction pipeline based on dependency graphs and engineered features [35]. Documents were first parsed with Stanza and transformed into document-level graphs enriched with crosssentence connections derived from entity linking, lexical similarity, and geographic containment information. From these structures, the system extracted a set of proximity-based, syntactic, and graph-derived features. Classification was then performed using either traditional machine learning ensembles or graph neural network architectures. Different combinations of graph-based and featurebased classifiers were selected for each language, enabling comparison between neural and non-neural prediction strategies while avoiding the use of pretrained language models during classification. FI-CODE. FI-CODE investigated three diverse approaches with a focus on computational efficiency and compact model size [16]. One system (run 2) employed a rule-based strategy that inferred relations from the textual distance between person and location mentions. Another system (run 1) adapted the ITER relation extraction architecture to a relation classification setting by providing entity annotations directly as input. The third system (run 3) explored promptbased inference with large language models, comparing alternative prompting styles, output formats, and model scales. Together, the submissions examined the trade-offs between heuristic methods, encoder-based models, and instructionfollowing language models. FourBytes. The team did not provide detailed information about their system. The reported efficiency metadata, including parameter count and disk size, is consistent with the XLM-RoBERTa base model [8], suggesting that the system may have used this model or a closely related architecture. GippLab. GippLab explored parameter-efficient fine-tuning of multilingual LLMs from the Qwen family [38,36]. Multiple model sizes were adapted using LoRA, with all linear layers receiving trainable adapters. Rather than training separate systems for individual languages, the team adopted a unified multilingual setup that jointly leveraged English, French, and German training data. The submitted runs corresponded to different Qwen3.5 model sizes selected from a broader set of experiments. This design emphasized scalable multilingual adaptation through lightweight fine-tuning. Hansel&Gretel. Hansel&Gretel investigated three complementary approaches spanning supervised fine-tuning, multi-agent reasoning, and prompted inference [32]. The first run fine-tuned a Qwen2.5 [37] instruction model using supervised generation of structured JSON outputs, complemented by Wikidata-based temporal consistency rules. The second run introduced a multi-agent framework in which specialized historian and geographer models produced independent assessments that were reconciled by an arbiter model. The third run relied on chain-of-thought prompting with a large reasoning model and dynamically retrieved Wikidata information. Across all runs, external biographical and geographic knowledge was incorporated to support temporal and spatial reasoning.
12
J. Opitz et al.
INSA Lyon. The INSA Lyon system combined handcrafted features, multilingual encoders, and large language models within a unified ensemble framework [22]. Candidate pairs were enriched with contextual and temporal information, including temporal expressions, tense-related cues, and lexical indicators. The resulting representations were processed by several model families, including feature-based classifiers, fine-tuned multilingual Transformer encoders, and zeroshot large language models. A dedicated decision layer aggregated the outputs through voting and stacking strategies. The submitted runs corresponded to the complete heterogeneous ensemble (run 1), a compact encoder-only baseline (run 2), and an ensemble configuration without large language model inference (run 3). MaxFo-Ajie. MaxFo-Ajie developed a prompt-based multilingual relation extraction pipeline using the GPT-5.5 model accessed through an OpenAI-compatible interface [5]. The system classified person–location pairs with structured prompts and language-matched few-shot examples drawn from the training data. The prompting strategy encouraged conservative predictions in cases of limited evidence and enforced a strict JSON output format. The submitted runs compared inline (run 1) and multi-turn document-level few-shot prompting (run 2), as well as pair-level inference (run 3). Additional components handled normalization, robust output parsing, and schema-compliant prediction generation. MILRIT. MILRIT submitted three systems representing different approaches to relation extraction [28]. Two runs were based on GLiREL [4], a general-purpose relation extraction framework. One variant fine-tuned GLiREL directly on the shared-task data (run 1), while another used GLiREL-generated relation scores as features for lightweight downstream classifiers (run 2). The third run employed a multilingual DeBERTa encoder augmented with entity markers, publicationdate information, and multi-view learning using large-language-model-generated input enrichments. The submissions therefore contrasted direct adaptation of a generalist relation extraction model with task-specific multilingual encoder architectures. ROSTI. ROSTI pursued a resource-efficient approach centered on logistic regression and sparse textual representations [7]. Sentences containing candidate entities were encoded using TF–IDF features derived from both sentence content and entity mentions. Separate classifiers were trained for the two target relations and for each language. Predictions were generated at the sentence-pair level and subsequently aggregated across all occurrences of a given person–location pair. The resulting system avoided external resources and large pretrained models while providing a compact baseline for multilingual historical relation extraction. Rittik&Souvik. Rittik&Souvik addressed the English-language setting using a transfer-learning approach based on BERT [29]. Person–location classification was formulated as a sequence-pair task through a question-answering-inspired
CLEF-HIPE-2026
13
input representation. To reduce overfitting in the low-resource setting, the team adopted a parameter-efficient fine-tuning strategy that updated only the upper layers of the encoder and the classification head. Class imbalance was addressed through cost-sensitive learning, with increased emphasis on minority labels during optimization. The system therefore combined lightweight adaptation with task-specific input reformulation. Spinfo. The team explored prompting strategies for open-weight large language models [9]. After evaluating several model families, the team selected a large reasoning model for all submitted runs. Person–location pairs were processed independently using structured prompts specifying task definitions, label inventories, and output constraints (run 1). Further experiments investigated the effect of temporal information extracted with HeidelTime and conditional prompting strategies that linked predictions across subtasks (run 2), and prompts translated into the document language (run 3).10 The submitted systems focused on understanding how prompting design influences multilingual historical relation extraction. UMUTeam. UMUTeam compared supervised multilingual classification and zero-shot large language model reasoning [15]. The supervised approach (run 1) used an XLM-RoBERTa cross-encoder with separate prediction heads for the two target relations. In parallel, two instruction-following language models were evaluated through structured prompting (run 2, run 3). All systems incorporated document metadata, publication dates, entity mentions, and local contextual evidence extracted around candidate pairs. The submitted runs therefore examined the relative strengths of fine-tuned multilingual encoders and prompt-based inference for historical relation extraction. VerbaNex AI I. VerbaNex AI I proposed a multilingual multitask architecture built on XLM-RoBERTa and enhanced with temporal reasoning components [2]. The system combined contextual representations with named entity recognition features, class balancing strategies, and hyperparameter optimization. During preprocessing, contextual evidence windows were extracted around candidate entities, while a lightweight Temporal Transformer introduced temporal positional information and semantic refinement mechanisms. The submitted approach was designed to capture both immediate and historical person–location associations in multilingual historical documents. VerbaNex AI II. VerbaNex AI II explored two complementary paradigms: an NLI-based encoder architecture and a self-distilled instruction-following language model [21]. The NLI system reformulated relation extraction as textual inference, pairing document content with hypotheses describing candidate relations. The submitted system relied on Nemotron-3-Nano-4B combined with 10
We were notified by the authors about potential issues in reproduction of run 3. For more information and extended experimentation we refer the reader to the Spinfo system description paper [9].
14
J. Opitz et al.
DSPy-based prompt optimization and self-distillation. Optimized prompting programs were used to generate high-quality training examples, which subsequently served for LoRA-based adaptation of the same model. The final system operated directly on multilingual OCR text and employed a unified inference pipeline across languages. whereami. The team developed a reasoning-oriented system based on multistage knowledge distillation [1]. The approach began with an exploration of prompting strategies across several large language models to identify a suitable teacher model. The teacher was then adapted through supervised fine-tuning and used to generate reasoning traces over the training corpus. These traces were subsequently distilled into a smaller Gemma-based student model through response-level supervision. To improve prediction consistency, the system incorporated rule-based constraints covering entity compatibility, relation directionality, contextual triggers, and locality constraints. 5.3
Discussion
The submitted systems illustrate the broad methodological diversity anticipated by the task design. Several teams adopted prompting-based approaches using large language models, either through proprietary systems (e.g., Awakened, MaxFo-Ajie) or open-weight models (e.g., BIU_NLP, Spinfo, UMUTeam, or FI-CODE). These approaches generally framed person–place relation extraction as an instruction-following or reasoning task, often relying on few-shot examples, structured output formats, and carefully engineered prompts. A second group focused on adapting open-weight generative models through supervised fine-tuning or parameter-efficient techniques such as LoRA, including GippLab, Hansel&Gretel, whereami, and VerbaNex AI II. These systems explored the extent to which task-specific supervision could be transferred into compact multilingual language models while retaining the flexibility of generative architectures. Another prominent line of work relied on encoder-based classification models, often built on multilingual Transformer encoders such as XLM-RoBERTa, DeBERTa, or NLI-oriented architectures. Examples include Awakened, INSA Lyon, MILRIT, Rittik&Souvik, UMUTeam, and VerbaNex AI I, which typically formulated the task as supervised pair classification and incorporated contextual, temporal, or entity-specific representations. In parallel, several teams investigated lightweight alternatives that emphasized efficiency and interpretability, including DS@GT’s graph-based feature extraction pipeline, ROSTI’s TF–IDF and logistic regression framework, and FI-CODE’s distance-based heuristics and compact relation classification models. Many submissions further augmented their core models with external knowledge or rule-based components. Wikidata-derived biographical and geographical information was frequently used to support temporal reasoning and consistency checking, while several systems incorporated post-processing rules, temporal constraints, entity-type validation, or retrieval-based enrichment. Finally, a number
CLEF-HIPE-2026
15
of teams explored ensemble and hybrid architectures that combined heterogeneous model families, most notably INSA Lyon, DS@GT, Hansel&Gretel, and MILRIT. Overall, the submissions demonstrate that the task can be addressed using a wide spectrum of techniques, ranging from highly efficient symbolic and statistical methods to large-scale reasoning models, reflecting the dual emphasis of the shared task on predictive quality and computational efficiency.
6
Main Results
6.1
Accuracy Profile on Test A
Table 3 reports the full official Test A ranking for complete submissions, including participant runs and organizer baselines. The official ranking is based on the mean of the three language-specific Test A scores. Spinfo achieved the highest Test A score with run 1 (0.748), followed by MaxFo-Ajie run 1 (0.700) and whereami run 1 (0.688). The leading run improves over the Ministral baseline by 0.166 absolute points. The language columns also show that no single team is uniformly dominant in every setting: Spinfo leads German and French, while MaxFo-Ajie leads English. The per-language winners in Table 4 suggest two additional patterns. First, high isAt macro recall was often crucial for the best German and French scores. Second, English was the only Test A language where the top run had similar at and isAt macro Recall, indicating a more balanced performance between the two relation types. 6.2
Generalization Profile on Test B
Table 5 reports the best runs on the surprise French literary test set. This profile evaluates only the at relation and is intentionally separate from Test A, because it changes both domain and time period. MaxFo-Ajie achieved the strongest generalization result: its three submitted runs occupy the first three ranks in the full ranking, with run 3 reaching 0.816 macro recall. This contrasts with Test A, where Spinfo was strongest overall. The difference indicates that in-domain newspaper performance and robustness to the literary surprise domain should be interpreted as related but distinct capabilities. 6.3
Efficiency Profile
Table 6 reports the full official efficiency profile, which combines each run’s accuracy rank, parameter-count rank, and model-size rank by their arithmetic mean. The efficiency profile changes the interpretation of the leaderboard. MILRIT run 3 ranks first under the official efficiency metric, despite not being the highest-accuracy run. FI-CODE and DS@GT_HIPE also rank highly because
16
J. Opitz et al.
Table 3. Full official Accuracy Profile ranking on Test A. Only runs submitted for all three Test A languages are included. The language columns report language-specific profile scores, i.e., the mean of at and isAt macro recall. Avg. is the official Test A score, obtained by averaging these three language-specific scores. Rank
T-Rank
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
1 2 3 4 5 6 7 8 9 10 11
12 13 14
15
16
Team Spinfo Spinfo MaxFo-Ajie Spinfo whereami whereami MaxFo-Ajie Awakened Awakened MaxFo-Ajie INSA Lyon gipplab Hansel&Gretel gipplab MILRIT UMUTEAM Ministral baseline VerbaNexAI II Hansel&Gretel BIU_NLP MILRIT Awakened Hansel&Gretel BIU_NLP VerbaNexAI II DS@GT_HIPE gipplab VerbaNexAI II VerbaNexAI I DS@GT_HIPE DS@GT_HIPE FI-CODE INSA Lyon INSA Lyon FI-CODE VerbaNexAI I ROSTI ROSTI UMUTEAM ROSTI BIU_NLP UMUTEAM FI-CODE MILRIT FourBytes Random baseline
Run
de
en
fr
Avg. ↑
1 3 1 2 1 2 2 3 1 3 1 2 3 1 3 2 1 3 2 2 1 2 1 3 2 1 3 1 2 2 3 2 3 2 3 1 3 2 3 1 1 1 1 2 1 1
0.7608 0.7710 0.6720 0.6860 0.7041 0.7072 0.6485 0.6746 0.6739 0.6335 0.6725 0.5941 0.5994 0.5957 0.5886 0.6305 0.5553 0.5697 0.5254 0.5050 0.5802 0.5417 0.5210 0.5451 0.5022 0.5160 0.5181 0.4808 0.5071 0.4937 0.5054 0.4678 0.4351 0.4435 0.4378 0.4488 0.5222 0.5192 0.4527 0.5196 0.4495 0.4167 0.4162 0.4365 0.3720 0.4058
0.7279 0.6808 0.7493 0.6427 0.7023 0.6992 0.7142 0.6415 0.5848 0.7003 0.5985 0.6174 0.6367 0.6143 0.5881 0.5384 0.5551 0.5812 0.5866 0.5762 0.4808 0.4989 0.5976 0.5101 0.5718 0.5103 0.5221 0.5607 0.4808 0.4401 0.4834 0.4624 0.5107 0.5128 0.4466 0.4757 0.4541 0.4431 0.4312 0.4335 0.4583 0.4702 0.4313 0.4353 0.4187 0.3733
0.7551 0.7349 0.6790 0.7383 0.6578 0.6435 0.6444 0.6854 0.7163 0.6295 0.6459 0.6697 0.6303 0.6322 0.6087 0.5880 0.6349 0.5877 0.6243 0.6529 0.6261 0.6076 0.5189 0.5617 0.4822 0.5162 0.4805 0.4597 0.4648 0.5171 0.4424 0.4902 0.4734 0.4561 0.5091 0.4640 0.3930 0.3896 0.4645 0.3849 0.4208 0.4357 0.4333 0.4073 0.4275 0.4355
0.7479 0.7289 0.7001 0.6890 0.6880 0.6833 0.6690 0.6671 0.6584 0.6544 0.6390 0.6271 0.6221 0.6141 0.5951 0.5856 0.5818 0.5795 0.5788 0.5781 0.5623 0.5494 0.5458 0.5390 0.5187 0.5142 0.5069 0.5004 0.4842 0.4836 0.4771 0.4734 0.4731 0.4708 0.4645 0.4628 0.4564 0.4507 0.4495 0.4460 0.4429 0.4408 0.4270 0.4264 0.4061 0.4049
CLEF-HIPE-2026
17
Table 4. Best language-specific Test A runs. Lang. Best run de en fr
Score at MR isAt MR
Spinfo run 3 0.771 MaxFo-Ajie run 1 0.749 Spinfo run 1 0.755
0.703 0.742 0.679
Runner-up
0.839 Spinfo run 1, 0.761 0.756 Spinfo run 1, 0.728 0.832 Spinfo run 2, 0.738
their resource ranks offset lower accuracy. The result of FI-CODE is particularly informative as a baseline assessment: its heuristic classifies pairs based on textual distance. While this strategy is extremely efficient, its moderate accuracy shows that distance biases are not strong enough to solve the task, especially because the data also contains many long-distance relations between persons and places. Conversely, the most accurate Test A system, Spinfo, ranks lower in the official efficiency profile because of its large reported parameter count and model size. This confirms the usefulness of reporting efficiency separately from accuracy for large-scale historical-text processing.
7
Additional Analysis
7.1
Baselines
On Test A, the Ministral baseline scored 0.582, clearly above the random baseline score of 0.405, but below the best participant run by a large margin. On Test B, the same baseline scored 0.506, again above random (0.363) but far behind the best surprise-domain run (0.816). These results make the baseline a reasonable reference point for zero-shot LLM prompting, while showing that task-specific system design and model size matters.
7.2
The Role of PROBABLE
The official ternary setup evaluates PROBABLE as a separate label for at. This makes the task more demanding than binary relation detection: systems must distinguish explicit evidence from plausible inference. Diagnostic files show that PROBABLE is often the hardest label. For example, on German Test A, Spinfo run 1 obtained high FALSE recall (0.952) but a lower PROBABLE recall (0.579). The Ministral baseline on the same file reached only 0.184 recall for PROBABLE. The generated binary analysis, where PROBABLE is mapped to TRUE for at, confirms that this distinction materially affects scores. In that setting, the top Test A score rises from 0.748 to 0.842, and the ordering of the best teams changes after the first position. We therefore treat binary at results as a useful sensitivity analysis, but not as a replacement for the official evidence-sensitive ternary task.
18
J. Opitz et al.
Table 5. Full official Generalization Profile ranking on Test B. The official score is at macro recall (MR) on the surprise-domain French test set. Accuracy is reported as an additional diagnostic and is not used for ranking. Rank T-Rank Team 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46
1 MaxFo-Ajie MaxFo-Ajie MaxFo-Ajie 2 Spinfo Spinfo 3 BIU_NLP Spinfo 4 whereami 5 gipplab 6 Awakened 7 Hansel&Gretel Awakened whereami Hansel&Gretel Hansel&Gretel gipplab BIU_NLP 8 VerbaNexAI II 9 UMUTEAM Awakened gipplab BIU_NLP 10 MILRIT VerbaNexAI II Ministral baseline 11 INSA Lyon MILRIT VerbaNexAI II INSA Lyon INSA Lyon 12 DS@GT_HIPE 13 ROSTI ROSTI 14 FI-CODE MILRIT 15 VerbaNexAI I DS@GT_HIPE ROSTI Random baseline DS@GT_HIPE UMUTEAM FI-CODE FI-CODE 16 FourBytes VerbaNexAI I UMUTEAM
Run MR ↑
Acc.
3 0.8163 0.8729 1 0.7945 0.8625 2 0.7712 0.8542 3 0.6984 0.8104 1 0.6910 0.7729 1 0.6837 0.8354 2 0.6674 0.7458 2 0.6665 0.8333 2 0.6647 0.7458 3 0.6613 0.7292 2 0.6349 0.6438 1 0.6338 0.7250 1 0.6325 0.8167 3 0.6187 0.7708 1 0.6107 0.6896 1 0.6085 0.7500 3 0.6000 0.6042 3 0.5724 0.7708 2 0.5723 0.5750 2 0.5509 0.5625 3 0.5382 0.6375 2 0.5265 0.4729 3 0.5152 0.5646 2 0.5076 0.6937 1 0.5062 0.5583 1 0.4705 0.6604 1 0.4679 0.6583 1 0.4419 0.6292 2 0.4231 0.6438 3 0.3986 0.6375 3 0.3919 0.5000 3 0.3840 0.5104 1 0.3773 0.5062 2 0.3755 0.3729 2 0.3742 0.5750 2 0.3726 0.4542 2 0.3721 0.5000 2 0.3660 0.4938 1 0.3628 0.3604 1 0.3626 0.3708 3 0.3620 0.3104 1 0.3580 0.4437 3 0.3546 0.4813 1 0.3445 0.2667 1 0.3346 0.5375 1 0.3333 0.6042
CLEF-HIPE-2026
19
Table 6. Full official Efficiency Profile ranking on Test A. The efficiency score is the mean of three component ranks: accuracy-profile score, parameter count, and model size. The final column reports the official Test A accuracy-profile score for reference. The table contains 43 runs rather than the 46 runs in the Accuracy Profile because the three MaxFo-Ajie runs opted out of the efficiency ranking. Tied overall ranks are shown only on their first row and ordered by decreasing accuracy-profile score. Ranking
Submission
Rank T-Rank Team Name 1 2 3 4 5 6
7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32
1 MILRIT 2 FI-CODE 3 DS@GT_HIPE DS@GT_HIPE DS@GT_HIPE MILRIT 4 ROSTI ROSTI ROSTI Random baseline 5 Awakened Ministral baseline 6 VerbaNexAI I 7 INSA Lyon 8 whereami whereami 9 VerbaNexAI II FI-CODE 10 UMUTEAM 11 gipplab VerbaNexAI I UMUTEAM VerbaNexAI II gipplab 12 FourBytes 13 Hansel&Gretel 14 Spinfo Spinfo gipplab INSA Lyon Spinfo Hansel&Gretel VerbaNexAI II MILRIT INSA Lyon FI-CODE Awakened Awakened Hansel&Gretel 15 BIU_NLP BIU_NLP UMUTEAM BIU_NLP
Efficiency Component ranks Run
Eff. ↓
3 2 1 2 3 1 3 2 1 1 2 1 2 2 1 2 3 1 1 2 1 2 1 3 1 1 1 3 1 3 2 2 2 2 1 3 3 1 3 2 3 3 1
9.7 10.3 10.7 12.0 12.3 13.7 13.7 13.7 13.7 15.0 15.3 15.7 15.7 16.0 16.7 17.0 17.0 17.0 17.3 17.7 18.0 18.3 18.7 19.7 19.7 20.0 20.3 20.7 20.7 20.7 21.0 21.7 21.7 22.0 22.7 23.0 23.7 24.0 24.3 24.7 24.7 26.0 31.0
Acc.
Acc. Param. Size Score ↑ 12 29 23 27 28 18 34 35 37 43 19 14 26 31 4 5 15 40 39 9 33 13 25 24 42 20 1 2 11 30 3 16 22 41 8 32 6 7 10 17 21 36 38
8 1 5 5 5 12 4 3 2 1 14 19 11 9 22 22 20 6 7 21 11 20 16 17 10 19 30 30 25 15 30 24 23 13 29 18 32 32 31 28 26 20 27
9 1 4 4 4 11 3 3 2 1 13 14 10 8 24 24 16 5 6 23 10 22 15 18 7 21 30 30 26 17 30 25 20 12 31 19 33 33 32 29 27 22 28
0.5951 0.4734 0.5142 0.4836 0.4771 0.5623 0.4564 0.4507 0.4460 0.4049 0.5494 0.5818 0.4842 0.4708 0.6880 0.6833 0.5795 0.4270 0.4408 0.6271 0.4628 0.5856 0.5004 0.5069 0.4061 0.5458 0.7479 0.7289 0.6141 0.4731 0.6890 0.5788 0.5187 0.4264 0.6390 0.4645 0.6671 0.6584 0.6221 0.5781 0.5390 0.4495 0.4429
20
7.3
J. Opitz et al.
System Behaviour on the Surprise Set
Figure 2 includes, alongside the gold annotations, the distribution of predictions submitted by all participating systems, offering a window into collective system behaviour on the surprise literary set. The Tenaud–Scythie pair (gold: FALSE) was correctly handled by most systems (79.5%), confirming that absent locative evidence for a named authority is generally well detected. More revealing are the error patterns. The Gargantua–Paris pair (gold: FALSE) attracted 27.3% PROBABLE predictions, suggesting that directional movement toward a destination is a systematic source of confusion. The Grandgousier–port de Olone pair (gold: PROBABLE) was assigned TRUE by a majority of systems (59.1%), confirming the tendency observed elsewhere to resolve locative uncertainty toward the more confident label. Finally, the Ponocrates and Eudemon–Orléans pairs (gold: TRUE) show that inferring physical passage through a location from narrative context is far from trivial, with 38.6
8
Conclusion
HIPE-2026 extends the HIPE evaluation series from named entity recognition and linking toward document-level relation understanding in historical text. The shared task focuses on temporally grounded person–place relation extraction, requiring systems to determine not only whether a person and a location are related, but also whether the evidence supports current presence near the publication date. This formulation moves beyond entity co-occurrence and challenges systems to perform contextual, temporal, and evidence-sensitive reasoning in noisy historical documents. The results show that the task is challenging but tractable under current modeling approaches. On the historical newspaper benchmark, the strongest systems substantially outperformed the organizer baselines, with the best run achieving a Test A score of 0.748 compared to 0.582 for the Ministral baseline and 0.405 for the random baseline. At the same time, the leaderboard remained highly competitive, with multiple teams achieving strong performance through markedly different methodological choices. The submissions ranged from prompted large language models and parameter-efficient fine-tuning to multilingual encoder architectures, graph-based methods, feature-engineered classifiers, and lightweight rule-based systems. Several observations emerge from the evaluation. First, language-specific performance patterns indicate that the two target relations pose different challenges. High isAt recall was particularly important for strong performance in German and French, whereas the best English systems achieved a more balanced trade-off between at and isAt. Second, the surprise-domain evaluation confirms that in-domain accuracy and cross-domain robustness are distinct capabilities. While Spinfo achieved the highest overall score on the historical newspaper benchmark, MaxFo-Ajie dominated the literary surprise set, demonstrating that successful transfer to substantially different genres and historical periods remains a separate challenge.
CLEF-HIPE-2026
21
The ternary formulation of at proved to be an important aspect of the task design. Distinguishing between TRUE and PROBABLE requires systems to separate explicit textual evidence from plausible contextual inference, making the task substantially more demanding than binary relation detection. The binary sensitivity analysis showed that collapsing PROBABLE into TRUE increases the best Test A score from 0.748 to 0.842 and changes the ordering of several leading systems. This confirms that modeling uncertainty and abductive inference is a central challenge of historical person–place relation extraction rather than a marginal edge case. The diversity of successful approaches is equally noteworthy. Many of the highest-ranked systems relied on large language models, either through prompting or supervised adaptation, suggesting that modern generative models are effective at integrating dispersed contextual and temporal evidence. However, competitive results were also obtained with encoder-based architectures and feature-driven approaches, indicating that strong performance does not depend on a single modeling paradigm. The task therefore provides a useful testbed for comparing reasoning-oriented and classification-oriented approaches under a common evaluation framework. The efficiency profile further highlights the importance of considering computational cost alongside predictive quality. Systems with the highest accuracy often relied on comparatively large models, whereas methods based on compact encoders, engineered features, or lightweight classifiers achieved substantially better efficiency rankings. In particular, the strong efficiency results of MILRIT, FI-CODE, and DS@GT demonstrate that practical large-scale deployment on cultural-heritage collections may favor different solutions than those that maximize accuracy alone. Reporting both dimensions therefore provides a more complete picture of system suitability for real-world historical-text processing. Overall, HIPE-2026 establishes a benchmark for evidence-sensitive and temporally grounded relation extraction in historical documents. The results show clear progress beyond entity recognition while also revealing persistent challenges involving temporal interpretation, implicit evidence, and domain transfer. Future work can investigate the remaining sources of error in temporal interpretation and plausible inference, extend the benchmark to additional relation types and historical genres, and explore approaches that improve cross-domain robustness without sacrificing computational efficiency. Acknowledgments. The CLEF-HIPE-2026 organizing team thanks the CLEF 2026 Conference and Evaluation Labs Committees for hosting the task and Simon Gabay for helping with the annotation of the surprise test data. This work is carried out within the framework of the Impresso – Media Monitoring of the Past project, funded by the Swiss National Science Foundation under grant No. CRSII5_213585 and by the Luxembourg National Research Fund under grant No. 17498891. Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article.
22
J. Opitz et al.
References 1. Aboelwafa, Y., Samir, A., Elmakky, N., Torki, M.: DistilledGemma: Balanced Efficiency-Accuracy for Person-Place Relation Extraction from Multilingual Historical Articles. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 2. Almanza González, D., Martinez-Santos, J.C., Puertas, E.: A Multitask Approach Based on RoBERTa, Temporal Transformer, and Person-Location Relation Extraction in Multilingual Historical Text. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 3. Blumenthal, P., Diwersy, S., Falaise, A., Lay, M.H., Sourvay, G., Vigier, D.: Presto, un corpus diachronique pour le français des XVIe-XXe siècles. In: atelier “ Les corpus annotés du français ”. Actes de TALN 2017, Orléans, France (Jun 2017), https://shs.hal.science/halshs-01585010 4. Boylan, J., Hokamp, C., Ghalandari, D.G.: GLiREL - generalist model for zeroshot relation extraction. In: Chiruzzo, L., Ritter, A., Wang, L. (eds.) Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). pp. 8230–8245. Association for Computational Linguistics, Albuquerque, New Mexico (Apr 2025). https://doi.org/10.18653/v1/2025. naacl-long.418, https://aclanthology.org/2025.naacl-long.418/ 5. Cao, H., Han, Z., Duan, X.: Few-Shot Prompting with Large Language Models for Multilingual Person–Place Relation Extraction from Historical Documents. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 6. Cardoso, S.D., Da Silveira, M., Pruski, C.: Construction and exploitation of an historical knowledge graph to deal with the evolution of ontologies. KnowledgeBased Systems 194, 105508 (2020) 7. Checchin, T., Guille, A., Guteherlé, N.: Some LLMs Are Smaller Than Others: A Lightweight Linear Model for Relation Extraction Applied to Historical Documents. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 8. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., Stoyanov, V.: Unsupervised cross-lingual representation learning at scale. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J. (eds.) Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 8440–8451. Association for Computational Linguistics, Online (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.747, https://aclanthology.org/2020.acl-main.747/ 9. Dick, A.K., Hermes, J., Reiter, N.: Scaffolding Large and Small Reasoning Models for Person-Place Relation Extraction. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 10. Ehrmann, M., Romanello, M., Flückiger, A., Clematide, S.: Overview of CLEF HIPE 2020: Named Entity Recognition and Linking on Historical Newspapers. In: Arampatzis, A., Kanoulas, E., Tsikrika, T., Vrochidis, S., Joho, H., Lioma, C.,
CLEF-HIPE-2026
23
Eickhoff, C., Névéol, A., Cappellato, L., Ferro, N. (eds.) Experimental IR Meets Multilinguality, Multimodality, and Interaction. pp. 288–310. Lecture Notes in Computer Science, Springer International Publishing, Cham (2020). https:// doi.org/10.1007/978-3-030-58219-7_21 11. Ehrmann, M., Romanello, M., Najem-Meyer, S., Doucet, A., Clematide, S.: Overview of HIPE-2022: Named Entity Recognition and Linking in Multilingual Historical Documents. In: Experimental IR Meets Multilinguality, Multimodality, and Interaction: 13th International Conference of the CLEF Association, CLEF 2022, Bologna, Italy, September 5–8, 2022, Proceedings. pp. 423–446. Springer-Verlag, Berlin, Heidelberg (Sep 2022), https://doi.org/10.1007/ 978-3-031-13643-6_26 12. Fokkens, A., Ter Braake, S., Ockeloen, N., Vossen, P., Legêne, S., Schreiber, G., et al.: BiographyNet: Methodological issues when NLP supports historical research. In: LREC. pp. 3728–3735 (2014) 13. Gabay, S., Clérice, T., Gille Levenson, M., Camps, J.B., Tanguy, J.B.: Freemcorpora/freemlpm: Freem lpm (lemma, pos- tags, morphology) corpus (2022). https://doi.org/10.5281/zenodo.6481300 14. Gabay, S., Ortiz Suarez, P., Bawden, R., Bartz, A., Gambette, P., Sagot, B.: Le projet FREEM : ressources, outils et enjeux pour l’étude du français d’ancien régime (the F RE EM project: Resources, tools and challenges for the study of ancien régime French). In: Estève, Y., Jiménez, T., Parcollet, T., Zanon Boito, M. (eds.) Actes de la 29e Conférence sur le Traitement Automatique des Langues Naturelles. Volume 1 : conférence principale. pp. 154–165. ATALA, Avignon, France (6 2022), https://aclanthology.org/2022.jeptalnrecital-taln.15/ 15. Gomez-Navalon, J., Bernal-Beltrán, T., Pan, R., García-Díaz, J.A., ValenciaGarcía, R.: UMUTeam at HIPE-CLEF 2026: Person-Place Relation Extraction from Multilingual Historical Documents. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 16. Griesbeck, L., Hennen, M., Babl, F., Geierhos, M.: FI-CODE@HIPE-2026: Efficient Person-Place Relation Classification from Distance Heuristics to Prompted Language Models. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 17. Hobbs, J.R., Stickel, M.E., Appelt, D.E., Martin, P.: Interpretation as abduction. Artificial intelligence 63(1-2), 69–142 (1993) 18. Keinan, R., Tsarfaty, R.: Where Were They? Person-Place Relation Extraction in Multilingual Historical Archives. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 19. Liu, A.H., Khandelwal, K., Subramanian, S., Jouault, V., Rastogi, A., et al.: Ministral 3 (2026), https://arxiv.org/abs/2601.08584 20. Lucchini, L., Sinatra, R., Emery, C., Panzarasa, P., Servedio, V.D.P., Riccaboni, M., Cattuto, C.: Following the footsteps of giants: Modeling the mobility of historically notable individuals. EPJ Data Science 8(7) (2019). https://doi.org/ 10.1140/epjds/s13688-019-0215-7 21. Morillo, A., Puertas, E., Martinez-Santos, J.C.: Person-Place Relation Extraction in Historical Texts: A Dual Approach with Fine-Tuned NLI Encoders and SLMs at HIPE-2026. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026)
24
J. Opitz et al.
22. Nguyen, S., Zheng, R., Nurbakova, D., Barrere, K.: INSA Lyon at HIPE 2026: A Modular Ensemble Pipeline for Person–Place Relation Classification in Historical Newspapers. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 23. Opitz, J.: A closer look at classification evaluation metrics and a critical reflection of common evaluation practice. Transactions of the Association for Computational Linguistics 12, 820–836 (Jun 2024), https://doi.org/10.1162/tacl_ a_00675 24. Opitz, J., Born, L., Nastase, V., Pultar, Y.: Automatic reconstruction of emperor itineraries from the regesta imperii. In: Proceedings of the 3rd International Conference on Digital Access to Textual Cultural Heritage. pp. 39–44 (2019) 25. Opitz, J., Ehrmann, M., Clematide, S., Corina, R., Boros, E., Michail, A., Romanello, M.: CLEF HIPE-2026 - Shared Task Participation Guidelines (May 2026). https://doi.org/10.5281/zenodo.20082076, https://doi.org/ 10.5281/zenodo.20082076 26. Opitz, J., Raclé, C., Boros, E., Michail, A., Romanello, M., Ehrmann, M., Clematide, S.: Clef hipe-2026: Evaluating accurate and efficient person–place relation extraction from multilingual historical texts. In: Campos, R., Jatowt, A., Lan, Y., Aliannejadi, M., Bauer, C., MacAvaney, S., Anand, A., Ren, Z., Verberne, S., Bai, N., Mansoury, M. (eds.) Advances in Information Retrieval. pp. 354–363. Springer Nature Switzerland, Cham (2026) 27. Opitz, J., Raclé, C., Michail, A., Romanello, M., Boros, E., Gabay, S., Ehrmann, M., Clematide, S.: Extended Overview of HIPE-2026: Evaluating Accurate and Efficient Person–Place Relation Extraction from Multilingual Historical Texts. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026). https://doi.org/10.5281/zenodo.20344461 28. Pham, N.H., Moreno, J.G., Doucet, A.: MILRIT at HIPE 2026: From Generalist Relation Extraction to Multi-View Learning for Historical Person-Place Relations. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 29. Ram, R., Soren, S.: Surgical Fine-Tuning and Cost-Sensitive Learning for LowResource Historical Relation Extraction: Jadavpur University at CLEF-HIPE 2026. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 30. Schich, M., Song, C., Ahn, Y.Y., Mirsky, A., Martino, M., Barabási, A.L., Helbing, D.: A network framework of cultural history. Science 345(6196), 558–562 (2014). https://doi.org/10.1126/science.1240064 31. Sebastiani, F.: An axiomatically derived measure for the evaluation of classification algorithms. In: Proceedings of the 2015 International Conference on the Theory of Information Retrieval. pp. 11–20 (2015) 32. Shekhawat, P.S., Gupta, R.: From Frozen Encoders to Multi-Agent LLM Pipelines: An Iterative Journey through Person–Place Relation Extraction in Multilingual Historical Texts. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026)
CLEF-HIPE-2026
25
33. Tamper, M., Kettunen, M., Mäkelä, E., Ruotsalo, T., Leskinen, P., Hyvönen, E.: BiographySampo: A linked open data service for prosopographical biography research. In: Proceedings of the Digital Humanities Conference (DH 2023). Graz, Austria (2023), https://seco.cs.aalto.fi/projects/biographysampo/ 34. Vasile, D.M., Apostol, E.S., Truâ, C.O.: Awakened at hipe-2026: Knowledgeaugmented prompting and logic-constrained fine-tuning for person–place relation extraction. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 35. Wesley, M.T.: Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 36. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al.: Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025) 37. Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al.: Qwen2.5 technical report. arXiv preprint arXiv:2412.15115 (2024) 38. Yazdani, S., Wahle, J.P., Gipp, B.: GippLab at HIPE 2026: Unlocking Person– Place Relation Extraction with the Qwen Model Family. In: Sánchez Salido, E., Barrón-Cedeño, A., García Seco de Herrera, A., MacAvaney, S., Struß, J.M. (eds.) CLEF 2026 Working Notes, CEUR Workshop Proceedings. CEUR-WS (2026) 39. Zhong, L., Wu, J., Li, Q., Peng, H., Wu, X.: A comprehensive survey on automatic knowledge graph construction. ACM Computing Surveys 56(4), 1–62 (2023)