arXiv:2609.15205v1 [cs.DB] 14 Sep 2026
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models Tong Li
Shuye Ding
Jiachuan Wang∗
Hong Kong University of Science and Technology Hong Kong SAR, China [email protected]
Hong Kong University of Science and Technology Hong Kong SAR, China [email protected]
Hong Kong University of Science and Technology Hong Kong SAR, China [email protected]
Yongqi Zhang
Shuangyin Li
Lei Chen
Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China [email protected]
South China Normal University Guangzhou, China [email protected]
Hong Kong University of Science and Technology Hong Kong SAR, China [email protected]
Bo Li Hong Kong University of Science and Technology Hong Kong SAR, China
Abstract Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong performance. Existing works have assumed the input texts to be table descriptions or specialized documents. However, these efforts have largely overlooked another prevalent category of texts, commonly found in news reports and social media: naturally occurring texts. Extracting tabular information from such texts poses two distinct challenges. First, high variability and the absence of explicit structural cues make fixed heuristic LLM prompts limited in precisely delineating extraction boundaries. Second, manually predefined schemas cannot capture open-ended, unseen attributes in naturally occurring text. In this paper, we propose a framework, TEAR, to address these challenges. It comprises two synergistic workflows: a Table Extraction Workflow that dynamically adapts instructions to overcome the limitation of heuristic instructions, and an Attribute Recommendation Workflow that discovers new attributes from texts to complement the heuristic schema. To our knowledge, TEAR is the first framework that supports automated text-driven attribute recommendation, enabling exploratory schema design for table extraction. To evaluate TEAR, we establish the benchmark for table extraction and attribute recommendation on naturally occurring texts, including two real-world datasets, manual annotations, ∗ Corresponding author.
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. Conference acronym ’XX, Woodstock, NY © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX
appropriate metrics, and baseline comparisons. Experiments show that TEAR achieves state-of-the-art performance on both tasks, and the recommended attributes effectively enhance extraction performance in exploratory scenarios.
CCS Concepts • Information systems → Information extraction; • Computing methodologies → Information extraction.
Keywords table extraction, attribute recommendation, text-to-table, large language models ACM Reference Format: Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li. 2026. TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX). ACM, New York, NY, USA, 22 pages. https: //doi.org/XXXXXXX.XXXXXXX
1
Introduction
Table extraction from texts, also known as text-to-table, is an emerging task focused on identifying semantic values from unstructured texts and organize them into tabular format [67], which unlocks the practical utility of texts for a wide range of downstream applications [26, 31, 43, 76], as well as simplifies data management [20, 71, 75, 77]. Previous works have developed into two lines, focusing on different categories of input texts. The first category is table descriptions, which are manually written or artificially generated according to well-defined tables, such as the NBA game summary and their original box scores in Figure 1(a). In this line of research, the tables exist natively, and description text can be obtained in bulk [4, 35, 44, 64]. Thus, with paired (text, table)s, researchers adopt a supervised paradigm, training generation models end-to-end to reconstruct the
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
Recover the box scores from its description. logical index
reconstruct
I need to archive the information according to the schema.
Losses Total points Wins Hawks 12 95 46 Magic 41 88 19 The Atlanta Hawks (46 - 12) beat the Orlando Magic (19 - 41) 95 - 88.
describes
In every document, there are (1) Plaintiff and Defendant; (2) Court Judges is about lending evidence, agreed lending amount, agreed repayment dates…; (3) Plaintiff can have multiple claims…(4) Note a valid agreement refers to an event in which the parties have entered into a written or oral agreement.
Civil Judgment of the XXX Court No.001 Plaintiff: XXX Defendant: XXX The case of XXX v. XXX regarding a private lending dispute was filed with this Court on March 29, 2016. XXX alleges: ...The plaintiff requests that XXX XXX argued: … The defendant requests that XXX
(a) Table descriptions.
(b) Specialized documents. What’s said about the accidents in these texts? Such as Victim name, Victim age, Victim status.
Attribute Recommendation
According to officials at the scene, a 15-year-old was injured, and the other, aged eighteen, was rescued from the vehicle by emergency crews and is said to be in a life-threatening condition. Both victims were transported to City General Hospital for treatment. Villanueva, currently at the Jersey Medical Center, told Dunn on Facebook Messenger that he came across eight undercover officers and then a man came out of nowhere, saw him first, and shot him.
Hospital name Information source Rescuer Encounter
instruction with specialized knowledge
Right! I would also like to extract for Hospital name. update schema and label sets
text-driven attributes (based on underlined texts) Table Extraction
Victim name Victim age Victim status \ 15-year-old injured \ eighteen life-threatening
Hospital name City General Hospital City General Hospital
Table Extraction
Victim name Victim age Villanueva \
Hospital name Jersey Medical Center
(c) Naturally occurring texts.
tables with heuristic schema
Victim status shot
updated tables with exploratory schema
Figure 1: Table extraction from different categories of input texts. original table from its description [32, 50, 67]. However, their input texts are generated under control and are dominated by the pre-existing tables, which are limited in reflecting the true difficulty of extracting from real-world texts. The second category is specialized documents, such as legal documents [5, 26] and biographies [5, 28]. These documents do not come with pre-aligned tables and involve more complicated contents, making annotating sufficient paired data for supervised learning expensive. Alternatively, researchers have shifted toward agentic extraction [7, 25, 26] powered by large language models (LLMs). This paradigm is known as in-context learning (ICL) [18], where the user depicts the task as instruction prompts that guide the LLMs to locate and extract relevant information without task-specific fine-tuning. The effectiveness of ICL stems not only from LLMs’ general semantic understanding, but also from how these documents are composed to facilitate information retrieval. Specifically, specialized documents usually follow established writing conventions to convey predefined knowledge, offering structural cues for human readability, which the LLM can also recognize and exploit. For example, in Figure 1(b), the prefix “Plaintiff:” or pattern “Party A v. Party B” are reliable cues that help readers and LLMs quickly identify important information.
absent from curated descriptions or specialized documents. Extracting tables from them, therefore, represents a pivotal advance of the field into realistic scenarios. Unlike previously studied input texts, n-texts are not organized around pre-existing tables or predefined knowledge. Instead, they unfold in a fluid, narrative style, exhibiting less literal consistency, which introduces new challenges for extraction. Challenge 1. Ineffective heuristic instruction. Resorting to LLM-based in-context learning for n-texts is appealing, given its semantic capability and data efficiency. Nevertheless, n-texts lack stable structural templates or cues in contrast to specialized documents. This high variability means that even carefully crafted instructions are often insufficient to delineate all possible extraction boundaries and convey nuanced task requirements [48, 55].
Our focus. Although effective for their assumed input texts, previous works have overlooked the extraction demands for naturally occurring texts [29, 34], the unscripted language produced for daily human communication, pervasive in news reports, social media discourse, and customer service interactions. Ubiquitous naturally occurring texts (n-texts) encapsulate the authentic, unfiltered information that flows through human interaction, a quality inherently
Challenge 2. Inadequate heuristic schema. A crucial step in extraction is understanding what information the texts contain, so as to determine the target schema, i.e., the attributes of interest. Existing methods typically work with heuristic schemas provided by human experts [25, 26], which may be satisfactory for their assumed inputs, where experts easily anticipate text contents. However, n-texts do not adhere to a stable informational template. Even
Example 1.1. Consider the attribute “Victim status” in Figure 1(c). Unlike the explicit plaintiff name in Figure 1(b), its scope is ambiguous: it may encompass only direct casualties (“injured”), or also extend to relocation (“transported”). Such ambiguity is not resolvable via LLM common sense but requires task specification. Exhaustive enumeration of analogous ambiguities in static instruction is not only impractical, but may instead dilute model attention to cause lost-in-the-middle [36].
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
if one were to navigate massive volumes of such texts and painstakingly summarize the observed contents into attributes, missing some attributes remains a risk. In other words, users confronting n-texts face an exploratory scenario: initially, they can only list some obvious attributes that immediately come to mind, yet still seek to uncover all the relevant attributes that will emerge from the texts ahead. Example 1.2. Consider the two texts in Figure 1(c). Though both are news reports about accidents, one details a rescue effort (“emergency crews”), and the other details a sudden, hostile encounter (“eight undercover officers”). Such contents diverge more sharply compared to those of the judgment in Figure 1(b), where each document follows a stable informational template to mention the plaintiff, defendant, case number, etc. Our proposals. In this paper, we propose a framework TEAR, short for Table Extraction with Attribute Recommendation, to address the above challenges for naturally occurring texts. To tackle the dilemma of heuristic instructions of Challenge 1, our idea is to dynamically select and insert demonstrative examples into the prompt for each text. These examples complement the abstract instructions with customized task specifications, enabling the LLM to perform extraction with clearer objectives. Also, retaining a small set of labeled data to supply such examples offers a practical trade-off between LLM usability and data efficiency [55]. Determining what constitutes an effective demonstration is non-trivial, because general text similarity [54] does not fully reflect the task utility of an example, such as its ability to clarify the extraction boundaries of a particular attribute. To this end, TEAR incorporates a proactive module that first performs a fuzzy prediction of what contents are likely to appear in the target table cells, and then retrieves examples that are informative for the specific extraction. To overcome the limitations of heuristic schemas discussed in Challenge 2 and support exploratory scenarios, TEAR introduces a novel attribute recommendation task, which automatically discovers new attributes in texts to assist in refining the prior schema, as in Figure 1(c). Although the capabilities of recent LLMs in processing open-ended information [1, 72] render this task possible, naively invoking LLMs and blindly accepting their outputs leaves the system vulnerable. Intuitively, an effective LLM-based attribute recommendation method should possess further abilities to (i) guide the LLM to propose attributes grounded in texts rather than ad-hoc fabrication; (ii) integrate the LLM’s differently articulated discoveries across multiple texts into a global candidate set; and (iii) prioritize the most salient candidates, shielding users from an undifferentiated, exhaustive list. Accordingly, we equip TEAR with components that operationalize these three intuitions. Finally, we establish an evaluation methodology tailored to this new setting. Prior work on table descriptions and specialized documents assumes the presence of logical indices to align and compare records [67], which are not applicable to n-texts because they offer no such convenience. We therefore introduce a new metric for extraction that assesses table semantics without relying on indices, along with a dedicated evaluation protocol for the novel attribute recommendation task. We accompany this evaluation with two n-text datasets, comprising 3,375 manually annotated ground-truth tables spanning two domains. Extensive experiments
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
on these datasets demonstrate that TEAR consistently outperforms state-of-the-art baselines on both extraction and recommendation tasks. To sum up, our main contributions are as follows: (1) We introduce an LLM-based framework TEAR for table extraction with attribute recommendation from naturally occurring texts. To the best of our knowledge, it is the first framework that could recommend new relevant attributes derived from texts under exploratory scenarios. (2) For table extraction, we propose a Proactive Demonstration Module that dynamically retrieves examples for adaptive instructions for each text. (3) For attribute recommendation, we propose a Discovery Mechanism that frames open-ended attribute discovery as table expansion to engage the LLM in proposing faithful new attributes, a Hybrid Integration Strategy that resolves duplicate LLM discoveries via diversity checking and contextualized analysis, and a Schema Coherence Score that ranks candidates by their holistic coherence with the initial schema. (4) We establish a new and fair evaluation methodology given the emergent evaluation difficulty from table extraction and attribute recommendation from naturally occurring texts. In the rest of this paper, we begin with a discussion of related work (Section 2) and the formal problem definition (Section 3). We then present TEAR, including an overview (Section 4.1), the table extraction workflow (Section 4.2), and the attribute recommendation workflow (Section 4.3). Next, we elaborate on the new evaluation methodology (Section 5). Lastly, we report the experiment results(Section 6) and conclusion (Section 7).
2 Related Works 2.1 Table Extraction from Texts Existing methods can be classified according to whether model parameters are updated. 1. Supervised paradigm. These methods treat table extraction as an end-to-end sequence generation task. They feed the text sequence into a language generation model [30, 56] and design different decoding strategies to generate rectangular tables as sequences, including row-by-row generation [67], parallel row generation [32], learnable cell ordering [50], and pointer-based decoding for the medical domain [76]. Researchers collect large amounts of paired (text, table) data and update the model parameters to minimize a loss function. Because these methods rely heavily on the training data distribution, they tend to perform well when the input texts are table descriptions that follow controlled patterns, but they struggle with n-texts, which exhibit higher variability and lack large-scale training data. 2. In-context learning paradigm. Recent works employ LLMs as general-purpose extractors via ICL [13, 38, 47]. Researchers convey the task to the LLM of the task through instructions rather than training it, thereby steering model behavior without any parameter updates. The research focus is on mimicking human-like behavior by decomposing the extraction task into subtasks (e.g., identifying entities, planning layout, and filling cells), and supplying each subtask with a dedicated instruction [1, 25, 26]. These heuristic instructions are coupled with specialized documents and tend to be effective because human
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
experts are themselves proficient at extracting knowledge from such documents. For instance, TKGT [26] translates human expertise in the legal domain into a knowledge graph and uses it to instruct the LLM. However, it is difficult to guarantee the success of this mode for n-texts, as the necessary template consistency and domain-specific heuristics are largely absent.
2.2
Open Information Extraction
Our attribute recommendation task is related to Open Information Extraction (OIE) [46, 77]. Among all subareas, the most relevant is novel slot detection [33, 68, 69], since a slot also manifests as a keyvalue pair. However, these works detect the existence of a new type, rather than revealing its semantic identity via canonical naming. Further, the slot types they defined are often distinguishable via surface value (e.g., a date vs. a name), hence are insufficient to represent distinct attributes that share similar values (e.g., “Victim name” or “Suspect name”). Other OIE tasks diverge farther from AR. For example, schema induction and event schema learning [6, 22, 37, 39, 72] discover unseen n-ary tuples representing predicates, entities, and relations, instead of extending known tuples with new dimensions; ontology learning [3, 66] establishes terminology, taxonomies, and axioms that are formal and universal rather than specific to input texts.
2.3
Other Table Extraction Tasks
Our task aims to identify semantic values in unstructured texts and align them into tables. We acknowledge that several other tasks are also termed “table extraction”, yet their intended application scenarios differ fundamentally from ours. 1. Table extraction by format mining. These methods rely on mining formatting patterns rather than deep semantic understanding, and thus cannot extract attribute values from completely unstructured, free-form texts. For example, web record extraction [9, 11, 58, 78] processes list or detail pages; tables repairing processes CSV files [10, 23] or PDF texts [61] with explicit delimiters; another recent work [2] considers text in heterogeneous data lakes where attributes appear in the fixed form of "<name>: <value>". 2. Table extraction by information integration. These works focus on reasoning or calculating for the attributes that are not directly stated in the texts. For instance, some count event occurrences [16]; others derive attributes such as lifespan or zodiac sign from extracted dates [5, 17]. Crucially, these approaches do not address the difficulty of extracting atomic attribute values from raw texts, especially highly variable n-texts. 3. On-demand table extraction. They extract tables in response to individual user queries rather than an overall schema, which maintains a corpus to locate answers [7] or retrieve relevant passages [8]. One advantage of our approach is that the pre-extracted tables enable a broad range of exact SQL-style operations, not only ad-hoc queries.
3
Problem Definition
We introduce the related notations. The naturally occurring text 𝑊 is a vanilla word sequence. A span [19] is a subsequence of its words, with P (𝑊 ) denoting the span space. Definition 3.1 (Value in Text). From the input text, a valid value for the extracted table is a list of non-overlapping spans. The value
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
space of the given 𝑊 , A (𝑊 ), is {⟨𝑎 1, 𝑎 2, . . . ⟩ | ∀𝑖, 𝑎𝑖 ∈ P (𝑊 ); ∀𝑖 < 𝑗, 𝑎𝑖 precedes 𝑎 𝑗 }.
(1)
Here, a list of spans is used for a value, instead of a single span, because the mentions can be non-consecutive in the naturally occurring text. Our basic settings keep the raw strings as values, allowing the LLM to tokenize and recognize them. For additional specialized usage, one could further convert raw strings into normalized formats by post-processing [60]. For example, converting "eighteen" to the numerical "18" for actual storing. Definition 3.2 (Schema). The extraction schema for naturally occurring text is a set of attributes. 𝑆 = {𝑠 1, 𝑠 2, ..., 𝑠 |𝑆 | }
(2)
For each attribute in the schema, its name is a string of regular naturalness [40], which contains complete words or acronyms in common usage, that are easy to understand by ordinary people and general LLMs. Meaningless or obscure symbols like "Column_1" or "REV" are not qualified. Definition 3.3 (Table). A table 𝐷 of 𝑛 records can be viewed as a list, 𝐷 = [𝐻, 𝑅1, ..., 𝑅𝑛 ], where 𝐻 = [ℎ 1, ℎ 2, ..., ℎ𝑚 ] is the header region that consists of 𝑚 unique attribute names. R = [𝑅1, ..., 𝑅𝑛 ] is the non-header region, where the 𝑖-th record 𝑅𝑖 = [𝐴𝑖1, 𝐴𝑖2, ..., 𝐴𝑖𝑚 ] is a list of 𝑚 values corresponding to 𝑚 attributes. Following previous works [16, 67], we define each table as a rectangular data structure with rows as records and columns as attributes. If the table follows the schema 𝑆, its header names 𝐻 ⊆ 𝑆. If the table is extracted from text 𝑊 , each value 𝐴𝑖 𝑗 ∈ A (𝑊 ), 𝑖 = [1..𝑛], 𝑗 = [1..𝑚]. We allow some values in the table to be an empty span list, i.e. |𝐴𝑖 𝑗 | = 0, indicating not mentioned. In a valid table 𝐷, we assume there is no entirely empty row or column to eliminate dummy structures. The table extraction (TE) task aims to output stable, task-specific predictions, not only subjectively reasonable ones, which is defined as follows. Definition 3.4 (Labeled Sample). Given a schema 𝑆, a labeled sample is a pair of text and table (𝑊 𝑙 , 𝐷 𝑙 ), where 𝐷 𝑙 is extracted from 𝑊 𝑙 and follows 𝑆. Definition 3.5 (Table Extraction). Given a schema 𝑆, labeled samples {(𝑊1𝑙 , 𝐷 𝑙1 ), (𝑊2𝑙 , 𝐷 𝑙2 ),...}, and unlabeled texts {𝑊1𝑢 ,𝑊2𝑢 ,...}, the method outputs the tables {𝐷 𝑢1 , 𝐷 𝑢2 ,...}. Let {𝐷ˆ 𝑢1 , 𝐷ˆ 𝑢2 , ...} be the ground truth, where 𝐷ˆ 𝑖𝑢 is extracted from 𝑊𝑖𝑢 and follows 𝑆. Given similarity metric 𝑔𝑡𝑒 of any two tables, Table Extraction aim to maximize 𝑔𝑡𝑒 (𝐷𝑖𝑢 , 𝐷ˆ 𝑖𝑢 ), ∀𝑖. The attribute recommendation (AR) task is defined as follows. Definition 3.6 (Attribute Recommendation). Given an inadequate schema 𝑆, labeled samples {(𝑊1𝑙 , 𝐷 𝑙1 ), (𝑊2𝑙 , 𝐷 𝑙2 ), ...}, the unlabeled texts {𝑊1𝑢 ,𝑊2𝑢 , ...}, and a recommendation budget 𝑘, the method outputs new attributes 𝑆 ′ that |𝑆 ′ | = 𝑘. Let 𝑆ˆ′ be the ground truth relevant set of attributes. Given a recall-based metric 𝑔𝑎𝑟 , Attribute Recommendation aim to maximize 𝑔𝑎𝑟 (𝑆 ′, 𝑆ˆ′ ). Note that these two tasks share the same input format, except for the recommendation budget 𝑘. Given the input, our TEAR framework accomplishes the TE task as other table extraction systems.
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Under exploratory scenarios with specified 𝑘, our framework additionally fulfills the AR task in response to the user requirements.
The header row sequence is:
4 The TEAR framework 4.1 Overview
The 𝑖-th data row sequence is:
Figure 2 provides an overview of our dual-workflow framework, which leverages the instruction-following ability of a backbone LLM to perform table extraction with attribute recommendation. (a) Table Extraction Workflow (TEW): Given an input text, the schema, and the labeled samples, we dynamically select demonstrative examples from the labeled samples and insert them into the prompt for adaptive instruction. (Step a1) We employ a surrogate model to make fuzzy predictions about the target table, generating a header preview and a value preview. (Step a2) For each labeled sample, we compute three utility scores by comparing its text to the input text, its table headers to the header preview, and its table values to the value preview. (Step a3) The most pertinent examples under each score are selected and prompted to the LLM together with the static instruction to obtain the output table. (b) Attribute Recommendation Workflow (ARW): Our ARW revisits the text to recommend attributes that are not yet present in the extracted table but are closely related to the existing schema. (Step b1) The LLM is instructed to expand the extracted table by adding new columns. (Step b2) We collect the newly expanded column headers from different tables, 𝑍 , to deduplicate and consolidate them into a global attribute candidate set 𝑍 ′ . (Step b3) For each candidate, we compute the semantic coherence of appending it to the heuristic schema to produce a ranked list.
The entire table sequence is:
4.2
𝑆𝑒𝑞(𝑅𝑖𝑙 ) = 𝑆𝑒𝑞(𝐴𝑙𝑖1 ) ⟨𝑠⟩ 𝑆𝑒𝑞(𝐴𝑙𝑖2 ) ⟨𝑠⟩...⟨𝑠⟩ 𝑆𝑒𝑞(𝐴𝑙𝑖𝑚 )
Table Extraction Workflow
Recalling Example 1.1, we posit that demonstrative examples should bridge the gap between the LLM’s general knowledge and task specification. Therefore, we propose a Proactive Demonstration Module that first proactively forecasts what contents are likely to appear in the target table, then uses these forecasts to query the labeled pool for examples that disambiguate their extraction. Specifically, it forecasts a header preview indicating the likely attribute names, and a value preview suggesting plausible textual spans in the non-header region. Since previews only highlight fuzzy cues to guide example selection, their generation is inherently easier and more noise-tolerant than producing a precise, complete table. We fine-tune a lightweight surrogate model for each dataset to generate them, which demands far fewer resources than end-to-end training. Proactive previews generation. For the input text 𝑊 𝑢 , we de-
4.2.1 fine its header preview 𝐻 𝑝 ⊆ 𝑆, and its value preview 𝑉 𝑝 ⊂ P (𝑊 𝑢 ). We use the labeled pool {(𝑊 𝑙 , 𝐷 𝑙 )} to construct the training data for the surrogate model 𝑀. Specifically, we serialize the table 𝐷 𝑙 into a sequence 𝑆𝑒𝑞(𝐷 𝑙 ), and optimize 𝑀 to generate this sequence autoregressively from 𝑊 𝑙 with the standard cross-entropy loss [62]. To serialize a table, we introduce three special tokens: ⟨𝑚⟩ separates spans in each value, ⟨𝑠⟩ separates cells in each row, and ⟨𝑛⟩ separates rows. Then, the sequence representation of a cell value 𝐴 = [𝑎 1, 𝑎 2, ..., 𝑎 |𝐴| ] is: 𝑆𝑒𝑞(𝐴) = 𝑎 1 ⟨𝑚⟩ 𝑎 2 ⟨𝑚⟩ ...⟨𝑚⟩ 𝑎 |𝐴|
𝑙 𝑆𝑒𝑞(𝐻 𝑙 ) = ℎ𝑙1 ⟨𝑠⟩ ℎ𝑙2 ⟨𝑠⟩...⟨𝑠⟩ ℎ𝑚
(3)
(4)
(5)
𝑆𝑒𝑞(𝐷 𝑙 ) = 𝑆𝑒𝑞(𝐻 𝑙 ) ⟨𝑛⟩ 𝑆𝑒𝑞(𝑅𝑙1 ) ⟨𝑛⟩ 𝑆𝑒𝑞(𝑅𝑙2 ) ⟨𝑛⟩ ...⟨𝑛⟩ 𝑆𝑒𝑞(𝑅𝑛𝑙 ) (6) During inference, given a new input text 𝑊 𝑢 , we obtain the predicted table sequence 𝑀 (𝑊 𝑢 ). Instead of reconstructing the full table, we probe it to obtain only 𝐻 𝑝 and 𝑉 𝑝 : for 𝐻 𝑝 , 𝑀 (𝑊 𝑢 ) is truncated at the first ⟨𝑛⟩ and split by ⟨𝑠⟩; for 𝑉 𝑝 , the sequence after the first ⟨𝑛⟩ is split by all special tokens. This yields coarse but informative previews that guide the subsequent example retrieval. 4.2.2 Extraction demonstration. Given the input text 𝑊 𝑢 and the generated previews 𝐻 𝑝 and 𝑉 𝑝 , we dynamically retrieve examples from the labeled pool based on three utility measures: header utility 𝑠𝑐ℎ , value utility 𝑠𝑐 𝑣 , and semantic utility 𝑠𝑐𝑠 . Header utility. Samples whose table headers are similar to 𝐻 𝑝 illustrate how to structure and populate attributes for the target table. We define the header utility of a labeled sample whose table 𝐷 𝑙 has headers 𝐻 𝑙 as 𝑠𝑐ℎ (𝐻 𝑝 , 𝐷 𝑙 ) = |𝐻 𝑝 ∩ 𝐻 𝑙 | + |𝐻 𝑙 \ 𝐻 𝑝 |/(|𝑆 | + 1).
(7)
The first term prioritizes samples that directly contain the anticipated headers; the second term, always less than 1, breaks ties by favoring samples with more additional headers, which provide broader contextual information. If 𝐻 𝑝 = ∅, meaning no header is previewed, Equation 7 reduces to |𝐻 𝑙 |/(|𝑆 | + 1), defaulting to retrieving samples with richer headers. Value utility. Samples whose table value spans are similar to 𝑉 𝑝 demonstrate how salient mentions or descriptive phrases are mapped into table cells. Let 𝑉 𝑙 = ∪𝑖 ∈ [1..𝑛],𝑗 ∈ [1..𝑚] 𝐴𝑙𝑖 𝑗 be the union of all spans in the labeled table 𝐷 𝑙 . We define the value utility as 𝑠𝑐 𝑣 (𝑉 𝑝 , 𝐷 𝑙 ) = F1(𝑉 𝑝 , 𝑉 𝑙 ; chrF𝛽),
(8)
where F1(·) is the standard F1-score (Equation 20) between two sets and chrF𝛽 [51, 52] is a widely used metric that measures pattern matching between two strings: chrF𝛽 = (1 + 𝛽 2 )
chrP · chrR , 𝛽 2 chrP + chrR
(9)
with chrP and chrR denoting character n-gram precision and recall, where we set 𝛽 by the previous default [53]. If 𝑉 𝑝 = ∅, meaning no value span is previewed, we will fall back to treating the original text 𝑊 𝑢 as a single long span in place of 𝑉 𝑝 in Equation 8. Semantic utility. Samples whose texts are overall semantically similar to 𝑊 𝑢 offer a comprehensive demonstration in style and content. We compute the semantic utility by the standard dense retrieval practice in RAG [21], where a deep embedding model [41], 𝐸𝑚𝑏 (·), is applied for normalized vectorization: 𝑠𝑐𝑠 (𝑊 𝑢 ,𝑊 𝑙 ) = 𝐸𝑚𝑏 (𝑊 𝑢 ) T · 𝐸𝑚𝑏 (𝑊 𝑙 ).
(10)
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
(b) Attribute Recommendation Workflow
(a) Table Extraction Workflow 𝑾𝒖
𝑺
text
schema
(𝑾𝒍 , 𝑫𝒍 ) labeled samples
(𝑾𝒍 , 𝑫𝒍 )
𝑾𝒖
𝑺
text
schema
labeled samples
extracted table Du
Proactive Demonstration Module Step a1: Proactive Previews Generation 𝑾 text
Step b1: Discovery Mechanism
[…],[…] 𝑽𝒑 𝑯𝒑 header preview value preview
Surrogate Model
𝒖
𝒑
𝑾
𝑯
[…],[…] 𝑽
𝑫𝒍
value utility scv
header utility sch
𝑺
𝑯𝒑
𝟏+𝜹 rank
𝑽𝒑
breadth-first greedy iteration JI: Your task is to judge
select top-𝑘 𝑡𝑒
𝑿𝒕𝒆
EI: Your task is to extract a table following the schema.
if the given elements are equivalent. low-diversity set
diveristy checking
attribute candidates
Step b3: Schema Coherence Score
Victim status, Victim name, Victim age,
LLM
LLM
contextualized analysis ∈𝑺 ∈ 𝒁′ ∈𝒁
𝑺
new attributes 𝑺′
Victim
0.64
𝑖=2 _job
0.24
𝑖 =3 ,
0.92
Sus
0.15
_vehicle 0.10
_type
0.05
Accident
0.21
0.01
_direction
0.02
𝑖=1
SP: Here is a schema ... the attributes are:
extracted table Du
pseudo-table 𝑫𝒖+
containing attribute proposals Z
Step b2: Hybrid Integration Strategy
𝑃∗ […],[…]
LLM
select top-𝑘 ar
Step a3: Extraction Prompting 𝑾𝒖
partition rank
(𝑾𝒍 , 𝑫𝒍 ) labeled samples
𝒑
𝑾𝒍 semantic utility scs
(𝑾𝒍 , 𝑫𝒍− , 𝑫𝒍+ )
compute utilities
Step a2: Extraction Demonstration 𝒖
𝑿𝒂𝒓 DI: Your task is to expand new columns.
𝑾𝒖
LLM
,
token probability
𝒔𝒄𝒔(“Victim vehicle”, 𝑺) = 0.64×0.10×0.92=0.059
Figure 2: Overview of TEAR. The LLM instructions are simplified for brevity, and the complete version is in the Appendix A.1. 4.2.3 Extraction prompting. For simplicity, we retrieve an equal number of examples under each utility type. The number 𝑘 𝑡𝑒 is a hyperparameter affected by the backbone LLM’s capability and can be tuned via validation performance. Let Xℎ , X𝑣 , X𝑠 be the top-𝑘 𝑡𝑒 samples indices according to 𝑠𝑐ℎ , 𝑠𝑐 𝑣 , 𝑠𝑐𝑠 , respectively. The Proactive Demonstration Module outputs demonstration set: 𝑋 𝑡𝑒 = {(𝑊𝑖𝑙 , 𝐷𝑖𝑙 )}𝑖 ∈ Xℎ ∪X𝑣 ∪X𝑠 .
(11)
We then use a natural language template with reasonable instructional phrases and formatting markers to wrap the text, schema, auxiliary previews, and retrieved demonstrations into a single prompt, the Extraction Instruction (EI). Figure 2 shows a conceptualized template, while a complete instantiation is in Appendix A.1. The final table extraction is then performed as 𝐷 𝑢 = 𝐿𝐿𝑀 (EI(𝑊 𝑢 , 𝑆, 𝐻 𝑝 , 𝑉 𝑝 , 𝑋 𝑡𝑒 )).
(12)
The above equation emphasizes what to populate the template, instead of the particular languages of EI.
4.3
Attribute Recommendation Workflow
We formulate the Attribute Recommendation (AR) task to support the emerging exploratory scenarios for n-texts, where the user provides only a heuristic schema as an initial direction, and the system is responsible for interacting with the texts to return a ranked list of new attributes. Importantly, AR does not produce a conclusive
attribute set, but rather exhibits a prioritized list serving as an automatic, intelligent summary of text-driven attributes. We have identified three essential abilities that an effective LLM-based AR method should possess (as in the Introduction), and our ARW instantiates corresponding components, as depicted in Figure 2(b). Briefly, it comprises three core components: a Discovery Mechanism that proposes new attributes from input texts via adaptive demonstration, a Hybrid Integration Strategy that efficiently resolves duplicates among generated proposals, and a Schema Coherence Score that ranks candidates by their relevance to the user’s initial schema. We detail each component in the following subsections.
4.3.1 Discovery Mechanism. To discover new attributes from an input text, we continue the idea of adaptive instruction, demonstrating the desired behavior through examples. Our Discovery Mechanism places attributes back into their naive context as column headers, and guides the LLM to expand the extracted table with new columns by drawing analogies to existing ones. By framing open-ended discovery as table expansion, we ground the LLM’s understanding of attributes in the structure of table columns, thereby preserving their original semantics more faithfully. The input to the mechanism is the text𝑊 𝑢 and the extracted table 𝑢 𝐷 , and the output is a pseudo-table 𝐷 𝑢+ which contains columns as new attribute proposals. A demonstrative example for this step
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
takes the form (text 𝑊 𝑙 , known table 𝐷 𝑙 − , new table 𝐷 𝑙+ ), which is constructed by splitting the columns of the sample table part 𝐷 𝑙 . Example 4.1. Suppose a labeled sample 𝐷 𝑙 has columns: “Victim name”, “Victim status”. To create a discovery example for an input text whose 𝐷 𝑢 has only a column “Victim name”, we split 𝐷 𝑙 into 𝐷 𝑙 − of column “Victim name”, and 𝐷 𝑙+ of “Victim status”. Note that although the attributes in the demonstrative “new" table belong to the existing schema 𝑆, we do not reveal 𝑆 to the LLM. This allows us to reuse the same labeled pool originally collected for 𝑆 to effectively demonstrate an open-ended discovery task. When retrieving the examples, we also update the previews with the header and value views derived from 𝐷 𝑢 . Let X̃ℎ and X̃𝑣 be the indices of the updated retrieved samples. The discovery demonstration set is 𝑋 𝑎𝑟 = {(𝑊𝑖𝑙 , 𝐷𝑖𝑙 − , 𝐷𝑖𝑙+ )}𝑖 ∈ X̃ℎ ∪ X̃𝑣 ∪X𝑠 .
(13)
Like the EI, we then construct the Discovery Instruction (DI), and the pseudo-table extraction is performed as: 𝐷 𝑢+ = 𝐿𝐿𝑀 (DI(𝑊 𝑢 , 𝐷 𝑢 , 𝑋 𝑎𝑟 )).
(14)
4.3.2 Hybrid Integration Strategy. There can be duplicates among individual discoveries. For instance, conceptually equivalent attributes may appear as "Hospital name" in one pseudo-table and "Medical center name" in another. LLMs are adept at detecting such duplicates through semantic reasoning over attribute names and their textual contexts. [14, 65]. We solicit a binary judgement {True, False} from the LLM over a small set of proposals at each time. This constrained format could effectively prevent the LLM from drifting into verbose or open-ended reasoning [36, 48], thereby preserving controllability. The input of the strategy are schema 𝑆 and proposals 𝑍 = ∪𝑖 𝐻𝑖𝑢+ \ 𝑆, and the output is a deduplicated subset 𝑍 ′ ⊆ 𝑍 . As 𝑍 grows, efficiency further draws our attention, as plainly invoking the LLM on all possible combinations becomes prohibitive. Therefore, we first introduce a lightweight diversity checking to identify a few plausible proposal subsets, and only submit those low-diversity subsets to the LLM for contextualized analysis. Diversity checking. For a group of attributes 𝑝 ⊆ 𝑍 ∪ 𝑆, we compute its diversity as the maximum Vendi Score [15, 42] over its subset: 𝑑𝑖𝑣 (𝑝) = max 𝑣𝑑𝑠 (𝑞), (15) 𝑞 ⊆𝑝
where 𝑣𝑑𝑠 (𝑞) ∈ [1, |𝑞|] is a continuous number that estimates the effective number of unique elements in 𝑞. The collection of lowdiversity subsets 𝑃 ∗ is defined as: 𝑃 ∗ = {𝑝 ⊆ 𝑍 ∪ 𝑆 ||𝑝 | > 1, 𝑑𝑖𝑣 (𝑝) < 1 + 𝛿 }, 1+𝛿 =
min
∀𝑠𝑖 ,𝑠 𝑗 ∈𝑆,𝑠𝑖 ≠𝑠 𝑗
𝑑𝑖𝑣 ({𝑠𝑖 , 𝑠 𝑗 }).
(16)
Here, 1 + 𝛿 is the minimum diversity between any two distinct schema attributes. If 𝑑𝑖𝑣 (𝑝) < 1 + 𝛿, 𝑝 approximates only a single attribute under the schema’s semantic granularity, and thus may accommodate duplicates. We could use a bottom-up search with pruning to find 𝑃 ∗ given that 𝑑𝑖𝑣 (·) is monotonic [42](see Appendix A.3).
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Contextualized analysis. For each proposal 𝑧 involved in some low-diversity subset, we obtain its context(𝑧) by prompting the LLM to produce a concise definition [72], based on its native texts and pseudo-tables (Detailed in Appendix A.4). Then, for each 𝑝 ∈ 𝑃 ∗ , a Judge Instruction (JI) is prompted to the LLM, 𝑟𝑠𝑝 = 𝐿𝐿𝑀 (JI(𝑆, {(𝑧, context(𝑧))}𝑧 ∈𝑝 )), 𝑟𝑠𝑝 ∈ {True, False}. (17) If the LLM responds that 𝑝 is duplicated, then we check whether 𝑝 contains a known attribute or not. If so, all other proposals are deemed naive repetitions of that known attribute and will be eliminated. If not, we choose the proposal in 𝑝 with the highest Schema Coherence Score (Section 4.3.3) and discard the rest. We could use earlier judgment to remove some proposals, so not every 𝑝 ∈ 𝑃 ∗ requires contextualized analysis. We hence design a breadth-first greedy iteration scheduling that dynamically updates 𝑃 ∗ to reduce LLM calls. As in Algorithm 1, each outer iteration selects a cover 𝑄 ⊂ 𝑃 ∗ of all remaining elements, and the inner loop sequentially submits 𝑄 to the LLM. This breadth-first design prunes confirmed duplicates early, shrinking 𝑃 ∗ and avoiding unnecessary exploration of nested sets that do not contain genuine duplicates. Complexity Analysis. We now briefly summarize the computational complexity of the Hybrid Integration Strategy, and refer to Appendix A.3 for detailed discussions and the Table 3 for the experiment report. In diversity checking, the number of candidate sets examined by a bottom-up search scales linearly with the output size |𝑃 ∗ |. In contextualized analysis, let |𝐵| be the number of elements at the start of an iteration, and 𝑁 ′ be the number of true duplicates. The number of LLM calls in that iteration is bounded by (ln |𝐵| + 1)(|𝐵| − 𝑁 ′ ). More duplicates yield a smaller bound and likely earlier True responses for more aggressive pruning. This property is beneficial: when duplicates are abundant, the algorithm eliminates large groups with few calls; when scarce, it degrades gracefully to exhaustive iteration. Even when the resource cannot afford the exhaustive iteration, early termination incurs little penalty since there are fewer duplicates in the first place. 4.3.3 Schema Coherence Score. It is necessary to prioritize attributes that meaningfully extend the user’s heuristic schema. For example, while “Reporter name” may be informative in a news article, it adds little value when the schema is designed to extract facts about criminal incidents rather than newsroom staffing. To quantify this notion of relevance, we introduce a Schema Coherence Score, which estimates how naturally an attribute candidate completes the description of the existing schema. Based on the language modeling principles [24, 49, 59], we compute this score as the conditional probability of the candidate’s name given the known schema. This formulation captures holistic coherence with the schema context. Formally, let 𝑧 be a candidate. We tokenize 𝑧 and append an end-of-name separator 𝑏 0 (a comma by default): 𝑡𝑘𝑛(𝑧) = 𝑇𝑜𝑘𝑒𝑛𝑖𝑧𝑒𝑟 (𝑧) + 𝑇𝑜𝑘𝑒𝑛𝑖𝑧𝑒𝑟 (𝑏 0 ).
(18)
We then prepend a Schema Prefix (SP) that enumerates the known attributes as the condition. As in Figure 2(Step b3), the score is 𝑠𝑐𝑠 (𝑧, 𝑆) = Π𝑖 𝑃𝑟𝑜𝑏 (𝑡𝑘𝑛𝑖 (𝑧)|SP(𝑆) + 𝑡𝑘𝑛 :𝑖 −1 (𝑧)),
(19)
where 𝑡𝑘𝑛𝑖 (𝑧) is the 𝑖-th token and 𝑡𝑘𝑛 :𝑖 −1 (𝑧) denote the tokens before it. Computing this score requires no decoding, only a single
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
Algorithm 1 breadth-first greedy iteration of 𝑃 ∗
predicted table 𝑫 true table 𝑫 ℎ3 ℎ1 ℎ2 ℎ 2 ℎ1 Victim gender Victim age Victim status Victim age Victim status eighteen life-threatening 𝑅1 15-year-old injured 𝑅1 male 𝑅2 \ 𝑅2 eighteen 15-year-old injured life-threatening 𝑅3 female eighteen badly injured (1) Header matching. (2) DF1(𝑅2 , 𝑅1 ) (4) structured-F1 (3) DF1(𝑅3 , 𝑅2 )
Require: attribute proposals 𝑍 , schema 𝑆, low-diversity set 𝑃 ∗ Ensure: attribute candidates 𝑍 ′ 1: 𝑍 ′ = ∅, 𝐵 ← 𝑍 ∪ 𝑆 2: while |𝐵| > 0 ∧ |𝑃 ∗ | > 0 do ⊲ Outer loop 3: 𝐵 ′ ← 𝐵, 𝑄 ← ∅ 4: while |𝐵 ′ | > 0 do ⊲ Find a greedy cover of elements. 5: 𝑞 = arg max𝑞 ∈𝑃 ∗ 𝑞 ∩ 𝐵 ′ 6: 𝑄 ← 𝑄 ∪ {𝑞}, 𝐵 ′ ← 𝐵 ′ \ 𝑞 7: end while 8: for 𝑞 ′ ∈ 𝑄 do ⊲ Inner loop 9: 𝑞 ← 𝑞′ ∩ 𝐵 10: if |𝑞| = 1 then 𝑍 ′ ← 𝑍 ′ ∪ 𝑞 11: else 12: Obtain 𝑟𝑠𝑝 for 𝑞 by Equation 17. 13: if 𝑟𝑠𝑝 then 14: if 𝑞 ∩ 𝑆 = ∅ then 15: 𝑍 ′ ← 𝑍 ′ ∪ {arg max𝑧 ∈𝑞 𝑠𝑐𝑠 (𝑧, 𝑆)} 16: end if ⊲ Exclude elements and sets. 17: 𝐵 ← 𝐵 \ 𝑞, 𝑃 ∗ ← {𝑝 ∈ 𝑃 ∗ |𝑝 ∩ 𝑞 = ∅} 18: else 𝑃 ∗ ← 𝑃 ∗ \ {𝑞} 19: end if 20: end if 21: end for 22: end while 23: 𝑍 ′ ← 𝑍 ′ \ 𝑆
forward pass through the LLM, making it substantially more efficient than generating language responses. Eventually, ARW ranks each 𝑧 ∈ 𝑍 ′ by 𝑠𝑐𝑠 (𝑧, 𝑆) and presents the top-𝑘 to the user.
5
Evaluation
Given that previous evaluation methodologies are unsuitable for our focus, we introduce new datasets and applicable metrics.
5.1
Datasets and Splits
The input texts of existing table extraction datasets (Appendix A.5) are table descriptions or specialized documents that cannot reflect the distribution of n-texts. To verify our focus, we present two real-world datasets, which together provide a total of 3,375 (text, table) pairs: • Incidents. The texts are gun-violence news reports, and tables capture victims, suspects, and accidents, with attributes such as “Victim name”, “Suspect name”, and “Accident address”. • Weather. The texts are weather forecasts for multiple countries, and tables include attributes such as “Weather frequency” and “Wind speed”. Both texts are news reports scraped by the CACAPO project [63], gathered without any extraction task in mind, and therefore exhibit realistic linguistic variability. The ground truth schema is derived from human responses to the 5W1H aspects (who, what, when, etc.) of each text, reflecting genuine text-driven needs instead of being arbitrarily carved. The original release contains some attribute annotations but lacks complete tables (values are not aligned to form multiple records). We hired university students to annotate complete tables while correcting errors (Appendix A.6).
pairwise 𝒇𝒗 𝐴21 𝐴22 𝐴23
pairwise 𝒇𝒌 ℎ1 ℎ2 ℎ3
𝐴መ11 \ 𝐴መ12 \
ℎ1 0.1 1 0.1 ℎ 2 0.2 0.1 1
1 0
0 1
pairwise 𝒇𝒗 𝐴31 𝐴32 𝐴33
𝐴መ 21 0 𝐴መ 22 0
1 0
0 0.1
pairwise DF1 𝑅1 𝑅2
𝑅1 𝑅2 𝑅3 0.1 0.8
1 .2
0.2 0.42
Figure 3: Evaluation example.
Besides the novel news datasets, we further adapt existing conversational texts, MultiWoz2.4 [70], to our focus: • Conversation. The texts are multi-topic, multi-turn human-human dialogues around services such as attractions, hotels, and restaurants. The dialogue states cover salient contents and can be converted to text-driven attributes, whose values may change as people change their requests or make clarifications. For table extraction, we sample 1000 texts and set ground truth labels as the final agreed-upon attribute values, while value drifts introduce intrinsic semantic noise. We use Conversation as a controlled stress test on verbosity and noise that does not diminish the contribution of the new datasets. The reason is that the dialogues have natural language patterns yet are oriented to a closed service ontology; thus, it is less variable than the new datasets, of which a statistical examination is in Appendix A.7. Table 1 summarizes the statistics, and the table annotations will be released. To reflect the application need under data scarcity, we randomly reserve 300 samples as the labeled set for all methods. For TE, we report performance using the full schema (Section 6.2). For AR, we further create three exploratory levels: we first identify low-frequency attributes that appear in <15% of samples, which non-experts would likely miss during heuristic schema design. We then randomly drop such attributes so that the heuristic schema lacks 20%, 35%, or 50% of all attributes. The absent attributes are completely hidden from the system, and the system is also unaware of the exploratory level (Section 6.3). We also evaluate the extracted table under these exploratory settings, comparing it against the table following the full schema (Section 6.4).
5.2
Metrics
5.2.1 Table extraction metrics. Let 𝐷 be the predicted table with headers 𝐻 and records R, and 𝐷ˆ be ground truth with 𝐻ˆ and R̂. We introduce table extraction metrics with the example in Figure 3. header-F1. Previous work use header-F1=F1(𝐻, 𝐻ˆ ; 𝑓 ) to evaluate whether the system correctly identifies the attributes mentioned in the text. Here, F1 is the standard F1-score between two sets, and 𝑓 is a given similarity function for comparing two elements (e.g.,
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Table 1: Benchmark statistics. #Token
#Record
|𝑆ˆ′ |
|𝑆 |
#Column
Datasets
#Text
Avg.
Max.
Avg.
Max.
Avg.
Max.
#Attr.
Level 1
Level 2
Level 3
Level 1
Level 2
Level 3
Incidents Weather Conversation
1,369 2,006 1,000
24.1 22.3 303
97 95 938
1.6 1.3 1
6 5 1
3.1 2.9 8.3
10 7 24
28 19 35
22 15 28
18 12 23
14 9 18
6 4 7
10 7 12
14 10 17
exact match, or soft string similarity). Formally, 1 ∑︁ max 𝑓 (𝑥, 𝑦), P(𝑋, 𝑌 ; 𝑓 ) = |𝑋 | 𝑥 ∈𝑋 𝑦 ∈𝑌 1 ∑︁ R(𝑋, 𝑌 ; 𝑓 ) = max 𝑓 (𝑥, 𝑦), |𝑌 | 𝑦 ∈𝑌 𝑥 ∈𝑋
(20)
F1(𝑋, 𝑌 ; 𝑓 ) = 2/(P(𝑋, 𝑌 ; 𝑓 ) −1 + R(𝑋, 𝑌 ; 𝑓 ) −1 ). Example 5.1. The header regions are 𝐻ˆ ={Victim age, Victim status}, 𝐻 ={Victim gender, Victim age, Victim status}. Given extract match 𝑓 , P(𝐻, 𝐻ˆ ; 𝑓 )=0.67, R(𝐻, 𝐻ˆ ; 𝑓 )=1, header-F1=F1(𝐻, 𝐻ˆ ; 𝑓 )=0.8. structured-F1. To evaluate whether the values are correctly extracted and aligned into a table, existing works largely follow the approach of Wu et al. [67], flattening a table into (header, index, value) triples. For instance, the table in Figure 1(a) uses team names as the logical index, yielding triples such as “(Losses, Hawks, 12)”. However, n-texts rarely provide a natural index column, and forcing one arbitrarily distorts the table’s native semantics. We therefore adopt a more fundamental view of a table: a set of records, where each record is a self-contained header-to-value dictionary. Under this view, comparing a predicted table to the ground truth reduces to matching two sets, for which standard F1 naturally solves. Crucially, this formulation respects the integrity of each record, and alignment is resolved through record matching, rather than an externally imposed index. We thus propose structured-F1: structured-F1=F1(R, R̂; DF1),
(21)
where DF1 (Dictionary F1) measures the similarity between two records. Specifically, each record 𝑅𝑖 is a dictionary mapping headers to values, with domain dom(𝑅𝑖 ) = 𝐻 and 𝑅𝑖 (ℎ 𝑗 ) = 𝐴𝑖 𝑗 . DF1 computes the similarity between 𝑅𝑖 and 𝑅ˆ 𝑗 in two steps: it first aligns headers using a similarity function 𝑓𝑘 , then aggregates the corresponding value similarities via 𝑓𝑣 . Formally, 1 ∑︁ ˆ , DP(𝑅𝑖 , 𝑅ˆ 𝑗 ; 𝑓𝑘 , 𝑓𝑣 ) = 𝑓𝑣 𝑅𝑖 (ℎ), 𝑅ˆ 𝑗 (arg max 𝑓𝑘 (ℎ, ℎ)) ˆ 𝐻ˆ |𝑅𝑖 | ℎ∈ ℎ∈𝐻 ∑︁ 1 ˆ 𝑅ˆ 𝑗 (ℎ) ˆ , DR(𝑅𝑖 , 𝑅ˆ 𝑗 ; 𝑓𝑘 , 𝑓𝑣 ) = 𝑓𝑣 𝑅𝑖 (arg max 𝑓𝑘 (ℎ, ℎ)), ℎ∈𝐻 |𝑅ˆ 𝑗 | ˆ ˆ ℎ∈ 𝐻
DF1(𝑅𝑖 , 𝑅ˆ 𝑗 ; 𝑓𝑘 , 𝑓𝑣 ) = 2/(DP(𝑅𝑖 , 𝑅ˆ 𝑗 ; 𝑓𝑘 , 𝑓𝑣 ) −1 + DR(𝑅𝑖 , 𝑅ˆ 𝑗 ; 𝑓𝑘 , 𝑓𝑣 ) −1 ). (22) Example 5.2. (1) We first compute the matched header to be invariant to the column order. (2) The matched position (bolded) is applied to aggregate the value similarity. For 𝑅2 and 𝑅ˆ1 , the empty value is ignored: DP=1, DR=1, and DF1(𝑅2, 𝑅ˆ1 )=1. (3) For 𝑅3 and 𝑅ˆ2 , DP=(0+1+0.1)/3=0.37, DR=(1+0.1)/2=0.55, so DF1(𝑅3, 𝑅ˆ2 )=0.42. (4) After computing for all 6 record pairs, P(R, R̂)=(0.8+1+0.42)/3=0.74, R(R, R̂)=(1+0.42)/2=0.71, structured-F1=0.73.
The complexity of structured-F1 is not higher than the previous evaluation (Appendix A.8). We use the same three string similarities as existing works: Extract Match (EM), chrF𝛽 (chrf) [52], and rescaled BERTScore (BS) [74]. For values with multiple spans indicating multiple mentions, the string similarity is first computed span-wise and aggregated via F1. 5.2.2 Attribute Recommendation Metrics. We evaluate AR quality using recall, as is standard in recommendation tasks. Let 𝑆 ′ = 𝑍 ′ [: 𝑘] denote the top-𝑘 attribute candidates. To aggregate performance across different 𝑘, we compute the recall curve R(𝑍 ′ [: 𝑘], 𝑆ˆ′, 𝑓 ) as a function of 𝑘, and report the Area Under the Curve (recall-AUC) as the overall metric. When comparing curves of different lengths, shorter curves are padded with their final value to ensure consistent normalization. A strong recommendation list reaches higher recall earlier, leading to a larger recall-AUC. For the similarity function 𝑓 , we use only chrf and BS because asking for an exact match under open-ended discovery is overly stringent for practical use.
6
Experiments
We conduct experiments to answer the following questions: Q1: How is the table extraction performance of TEAR ? Q2: How is the attribute recommendation performance of TEAR ? Q3: How do text-driven new attributes benefit table extraction? Q4: How efficient is TEAR? Q5: How are the ablation results of TEAR?
6.1
Setup
6.1.1 Baselines. For table extraction, we compare baselines from both the supervised and the ICL paradigms. Note that some other ICL approaches rely on knowledge about the specialized documents (e.g., external KGs [26] or type recognizing s [25]) to decompose extraction into subtasks and design dedicated instructions, making them difficult to apply to n-texts. 1. TRE [67] is one of the best supervised methods, which augments the sequence generation model with special table relation embeddings. We use its official implementation.1 . 2. MapMake [1] is a recent ICL method. We implement the one-shot version by selecting the labeled sample with the most headers and carefully writing the required reasoning steps. 3. RAG [41, 73]. We implement a strong baseline of dense RAG that could also dynamically choose examples to contrast with our Proactive Demonstration Module. As for the novel attribute recommendation task, there lacks existing methods, and we compare against three strong baselines that all prompt the LLM with retrieved examples to obtain attribute proposals, and then adopt different integration and ranking strategies. 1 https://github.com/shirley-wu/text_to_table
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
Table 2: Table extraction performance (%). The best results within each backbone are bolded, and the best in the same row are underlined. Dataset
Incidents
Weather
Conversation
Llama-3-8B-Instruct MapMake RAG TEAR
Qwen2.5-14B-Instruct MapMake RAG TEAR
Llama-3-70B-Instruct MapMake RAG TEAR
Metric
sim.
TRE
headerF1
EM chrf BS
86.7±0.8 91.2±0.7 91.9±0.7
47.4±0.8 54.8±0.9 55.3±0.9
81.5±0.1 87.9±0.1 88.5±0.1
87.8±0.3 92.5±0.3 93.1±0.3
68.1±0.3 74.9±0.2 75.8±0.2
83.3±0.1 88.2±0.0 88.9±0.0
89.1±0.4 92.8±0.2 93.4±0.2
68.1±0.2 74.3±0.3 75.0±0.1
86.8±0.0 91.6±0.0 92.3±0.0
90.3±0.1 93.9±0.1 94.5±0.1
structuredF1
EM chrf BS
64.5±0.1 74.9±0.5 78.8±0.4
25.7±0.4 34.1±0.3 42.5±0.2
61.2±0.2 73.6±0.1 76.8±0.1
69.1±0.2 79.4±0.1 82.9±0.1
40.9±0.4 54.1±0.2 60.5±0.4
62.0±0.0 76.8±0.0 74.9±0.0
70.6±0.1 82.0±0.2 82.4±0.2
44.8±0.3 56.5±0.1 61.9±0.2
69.8±0.0 80.0±0.0 81.5±0.0
73.3±0.0 82.4±0.1 84.6±0.1
headerF1
EM chrf BS
79.1±0.5 84.9±0.3 89.6±0.3
52.2±1.3 58.9±1.3 68.3±0.9
77.3±0.0 83.5±0.0 88.7±0.0
80.3±0.4 85.6±0.1 90.1±0.0
66.1±0.0 73.8±0.1 82.1±0.0
80.4±0.1 85.7±0.1 89.9±0.1
82.6±0.2 87.2±0.1 91.3±0.1
74.3±0.0 80.1±0.2 86.7±0.0
83.9±0.1 88.0±0.1 91.9±0.1
84.1±0.1 88.4±0.1 92.2±0.1
structuredF1
EM chrf BS
44.1±0.5 63.7±0.3 63.0±0.3
25.6±0.4 42.1±0.8 42.2±1.3
46.2±0.1 65.4±0.1 64.3±0.1
51.0±0.4 70.1±0.2 67.9±0.4
39.3±0.3 62.0±0.3 57.2±0.1
50.7±0.1 71.3±0.0 66.4±0.0
53.9±0.4 73.4±0.2 69.1±0.5
43.2±0.3 64.5±0.2 62.5±0.2
54.9±0.1 72.1±0.1 70.2±0.1
56.3±0.1 73.3±0.0 71.5±0.0
headerF1
EM chrf BS
90.0±0.5 94.7±0.3 94.3±0.3
52.7±2.0 62.2±1.9 61.5±1.9
84.6±0.1 91.9±0.1 91.3±0.1
88.7±0.2 94.5±0.1 93.9±0.0
76.3±0.1 85.3±0.2 84.4±0.1
87.4±0.0 93.4±0.0 92.8±0.0
90.7±0.1 95.3±0.0 94.9±0.1
78.9±0.0 86.3±0.1 85.9±0.1
86.8±0.0 93.4±0.0 92.9±0.0
90.1±0.1 95.1±0.1 94.6±0.1
structuredF1
EM chrf BS
80.9±0.6 85.4±0.5 87.0±0.3
44.4±1.9 49.3±1.9 53.6±2.1
76.9±0.1 81.2±0.1 83.8±0.2
82.4±0.1 86.1±0.1 88.4±0.1
66.2±0.1 71.5±0.1 75.8±0.2
81.1±0.0 85.1±0.0 87.6±0.0
85.1±0.2 88.3±0.2 90.3±0.1
69.4±0.0 74.8±0.1 77.8±0.0
80.8±0.0 84.6±0.0 86.7±0.0
84.6±0.1 87.8±0.1 89.7±0.2
4. Corpus Frequency Ranking (CFR) integrates and ranks attribute proposals according to their global frequency in the corpus. The intuition is that attributes mentioned more frequently across texts are more important. 5. Maximum Schema Similarity (MSS) ranks proposals by their maximum semantic similarity to the known schema attributes, 𝑠𝑖𝑚(𝑧, 𝑆) = max𝑠 ∈𝑆 𝐸𝑚𝑏 (context(𝑧)) · 𝐸𝑚𝑏 (context(𝑠)). The motivation is that newly discovered attributes should be semantically closer to those already known. 6. Direct LLM Reasoning (DIRECT) prompts an LLM to rank the proposals as the most suitable for extending the known schema, considering relevance and usefulness. 6.1.2 Implementation. For all compared methods, we employ two instruction-tuned open-source LLMs: Llama-3-8B-Instruct2 , Qwen2.514B-Instruct3 , and Llama-3-70B-Instruct4 . We choose medium-sized, open-source LLMs to reflect realistic deployment scenarios where large proprietary models may be too expensive or inaccessible. Notably, the challenges addressed in this work stem from the gap between the generalized capabilities acquired through pretraining and the specialized competencies required for the table extraction task, which cannot be overcome solely by switching to a larger or more capable LLM. The deep embedding model 𝐸𝑚𝑏 (·) is SFREmbedding-Mistral5 [41], a state-of-the-art open-source embedder. The surrogate model 𝑀 (·) is BART-Large6 [30]. For the Schema Prefix that ends with an enumeration of known attributes, we randomly sample 10 different permutations, use them to calculate 𝑠𝑐𝑠 (·), and 2 https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct 3 https://huggingface.co/Qwen/Qwen2.5-14B-Instruct 4 https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct
take the maximum score for each candidate. We will elaborate on the robustness of this implementation in Section 6.3. During validation, 100 samples are held out for early stopping and method tuning, and the remaining samples are used for training and retrieval. At inference, the validation set is also added for retrieval. The final number of shots is 𝑘 𝑡𝑒 =𝑘 𝑎𝑟 =5 for Incidents and Weather, 3 for Conversation; the final LLM decoding uses temperature 𝑇 𝑙𝑙𝑚 =5 and top-𝑝 𝑙𝑙𝑚 =0.5. All the experiments are conducted on a server equipped with Intel(R) Xeon(R) Gold 6240 CPU and one NVIDIA A800 (80GB Memory) for Llama-8B and Qwen, two NVIDIA A800s for Llama-70B. Each experiment is repeated 3 times with different random seeds, and we report the average and standard deviation. Our code is available at ... 7 .
6.2
Q1: Table Extraction Results
The TE performance is in Table 2 with the following key observations. (1) TEAR is consistently the best. This demonstrates that our method effectively selects more informative demonstrations by leveraging both textual and tabular characteristics. Among the baselines, RAG dynamically retrieves demonstrations, which is better than MapMake that uses fixed demonstrations, showing the critical role of adaptive demonstrations. (2) ICL paradigm methods exhibit a clear advantage over the supervised method, especially on structured-F1 that evaluates values. This agrees with the intuition that while patterns in header regions are relatively limited and are easier to learn via supervised training, the values in n-texts vary considerably with the input text, requiring a stronger generalization capability for correct extraction. Furthermore, larger LLMs are overall better, and this gap is more pronounced for MapMake and RAG. This indicates that providing well-selected examples with our
5 https://huggingface.co/Salesforce/SFR-Embedding-Mistral 6 https://huggingface.co/facebook/bart-large
7 The official repository is being prepared as part of the CR version.
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
TEARcan particularly boost the performance of relatively weaker LLMs. (3) The results also indicate that our new benchmarks are more challenging for the studied extraction systems than Conversation, as it is reported to have generally lower results, and the performance gap between different methods is larger. This aligns with our claim that the variability of n-texts is a more urgent challenge for existing systems, instead of the verbosity or value drifts characterized by Conversation.
6.3
Q2: Attribute Recommendation Results
The attribute recommendation performance is shown in Figure 4. (1) The results confirm the viability of using LLMs for AR. Particularly, on the Weather dataset, where the recall-AUC reaches an exceptionally high level in some cases, indicating that the LLMrecommended attributes semantically encompass nearly all groundtruth attributes. (2) Our method achieves the best overall performance. Under all comparisons, TEAR is the best on 48 out of 54 comparisons. (3) After our method, the baselines do not have an obvious second best, and the LLM inference, global frequency, and semantic similarity have their own advanced cases. (4) Trends across exploratory levels differ for different datasets. Recall that Levels 1 to 3 aim to discovering the remaining 20%, 35%, and 50% of attributes. On Weather, performance improves from Level 3 to Level 1, indicating that a more complete initial schema makes the task easier. In contrast, Level 1 is the hardest for Incidents. Examining the schema splits, we find that this is due to a few low-frequency attributes (e.g., “Number of rounds fired”) that are particularly difficult to discover precisely; consequently, the task becomes increasingly harder as the schema grows more complete.
6.4
Q3: Table Extraction with Text-Driven Attributes
We simulate the process of users expanding the schema with recommended attributes for table extraction, in order to demonstrate the effect of exploratory schema design on the overall information extraction system. Specifically, we take 𝑍 ′ [: 𝑘] ∩ 𝑆ˆ′ as the user-accepted text-driven attributes (the "checked" attributes in Figure 1(c)). We numerically estimate 𝑘 by the elbow point [57] of the 𝑠𝑐𝑠 (·), which reflects a realistic scenario where users only navigate top-ranked attributes instead of the complete list. Moreover, using set intersection to select attributes (i.e., only adopting a recommendation if its name exactly matches the ground truth) is a very conservative simulation. In practice, users would typically accept a recommendation as long as it is semantically close to their interests and appropriately formulated. Under each exploratory scenario, we first evaluate table extraction performance by running TEW. We then execute the ARW, update, and run TEW again to obtain the updated performance, denoted as TEAR*. Both results are compared against the ground-truth table following the complete schema. Results using Qwen are in Figure 5; those with Llama (Appendix A.9) exhibit similar trends. It shows that TEAR consistently outperforms the baselinem, and the interactive pipeline TEAR* achieves additional performance gains. Notably, on the Incidents dataset, while other methods degrade as the exploratory level increases, TEAR*
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Table 3: Hybrid Integration Strategy efficiency. Inci.
|𝑍 |
|𝑃 ∗ |
l.
t.(ms)
#Calls
True
False
1 2 3
207±3 203±4 218±13
182±15 130±1 157±20
4 4 4
17±10 13±8 22±11
81±7 79±6 91±6
29±6 23±2 23±2
52±4 56±7 68±6
Wea.
|𝑍 |
|𝑃 ∗ |
l.
t.(ms)
#Calls
True
False
1 2 3
279±9 238±4 212±21
10±3 380±124 315±122
3 4 4
12±0 37±3 28±5
9±2 198±51 163±40
4±1 30±3 33±5
5±1 167±53 130±36
Con.
|𝑍 |
|𝑃 ∗ |
l.
t.(ms)
#Calls
True
False
1 2 3
94±4 128±2 134±7
156±31 173±7 235±79
6 4 7
15±2 26±4 24±3
10±0 13±2 25±8
6±2 8±1 11±5
4±1 5± 1 14±6
maintains stable performance and even shows improvement in some settings, highlighting the benefits of text-driven attributes.
6.5
Q4: Efficiency
Figure 6 compares efficiency, where time is measured by running the LLM on a local research server, applying no acceleration, and the Input and Output tokens per text are counted only for LLM methods, For TE, the efficiency order is TRE>RAG>Ours>MapMake. For AR, results are averaged over three exploratory levels, and the efficiency order is CFR>MSS>DIRECT>Ours. The overhead of LLM methods primarily arises from the internals of the backbone, not the methods themselves. Compared to the high cost of fine-tuning or the extensive annotation effort required for supervised extraction models, our approach achieves strong performance with significantly lower demand for labels or computational resources. We further evaluate the efficiency of the Hybrid Integration Strategy in Table 3, where 1,2,3 are exploratory levels. Here, |𝑍 | is #attribute proposals, |𝑃 ∗ | is #low-diversity sets, l. is the size of the largest low-diversity set affecting pruning search iterations, t. is the pruning time. |𝑃 ∗ | is efficiently small, showing that many proposals are actually far from others and can be quickly excluded from duplicate analysis. And thanks to our breadth-first greedy scheduling (Algorithm 1), the number of LLM calls is even less.
6.6
Q5: Ablation Studies
6.6.1 Labeled Pool Size, Split and Intialization. We vary the size, split, and initialization of the labeled pool and report the structuredF1 with Llama3-8B in Figure 7 and Figure 8, while other LLMs and metrics have the same pattern. In Figure 7, we reserve a subset of 200 test samples and progressively enlarge the labeled pools. The upward trend in Incidents gradually saturates, while Weather exhibits continued improvement, indicating a higher demand for labels. In Figure 8, we use the same reserved test set and sample three different pools, which lead to similar results, demonstrating the robustness of TEAR against initialization. We further compare our default setting, where the same pool is used for learning 𝑀 and retrieval, with its disjoint setting, where the pool is split for learning and retrieval. The disjoint setting consistently underperforms our default setting, verifying that allocating a standalone retrievable set is unnecessary under data scarcity.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
Llama-3-8B-Instruct
Incidents
chrf
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
Weather
80 60 40
Conversation
80 60 40 80 60 40
Qwen2.5-14B-Instruct
BS
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
BS
80 60 40 80 60 40
chrf
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
80 60 40
BS
Llama-3-70B-Instruct
80 60 40 80 60 40
chrf
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
chrf
80 60 40
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 DIRECT CFR MSS Ours
structured-f1
header-f1
Figure 4: Attribute recommendation recall-AUC (%).
Incidents
EM
chrf
BS
EM
Weather chrf
BS
80 60 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 EM chrf BS 80 60 40
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
EM
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
MapMake-Qwen
TRE
RAG-Qwen
chrf
BS
Level 1 Level 2 Level 3 Level 1 Level 2 Level 3 Level 1 Level 2 Level 3
TEAR-Qwen
TEAR*-Qwen
Figure 5: Table extraction with text-driven attributes with Qwen backbone.
TE Runtime (103 s) & Latency (s) 6
8
64 4
4
4
2
TE Input (103) & Output (102)
6
6
4
0
8
2
22
AR Runtime (103 s) & Latency (s)
AR Input (103) & Output (102)
4
1.25 1.00 0.75 0.50 0.25 0.00
5 43 32 2 1 1 00
3 2 1
00 00 Incidents Weather Conversation Incidents Weather Conversation Incidents Weather Conversation Incidents Weather Conversation TRE (left) RAG (left) MapMake (left) TEW (left) DIRECT (left) CFR (left) MSS (left) ARW (left) TRE (right) RAG (right) MapMake (right) TEW (right) DIRECT (right) CFR (right) MSS (right) ARW (right) Figure 6: Efficiency comparison with Llama3-8B. Patterns for other LLMs are similar.
60 50
73.1
Incidents
77.7
80.0 78.5
68.7
66.6
62.1
200 300
500
82.2 80.0
82.1 80.7
69.8
69.6
700
EM
1000
Weather
80 70 68.0 69.7 65.4 67.0 60
72.1 68.3
72.1 69.1
50 48.0 51.5
53.1
54.3
500
700
40
chrf
200 300
BS
75.1 73.3 58.4
1000
Figure 7: Labeled pool size ablation with Llama3-8B.
Incidents 90 81.8 77.777.977.8 80.179.7 80 68.7 70 67.167.3 60 50 40 EM chrf BS pool 1 pool 2
structured-f1
70
82.0
structured-f1
80 77.0
structured-f1
structured-f1
90
Weather 90 80 69.870.669.9 65.866.265.1 70 60 51.5 50.348.0 50 40 EM chrf BS pool 3 Disjoint
Figure 8: Different labeled pools with Llama3-8B. 6.6.2 Proactive Demonstration Module. We ablate the demonstrations retrieved by each utility score, and the extraction performance
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
Incidents
Weather
82.1
structured-f1
structured-f1
Weather
80
79.5 75.8
81.2 78.1
82.1 78.7
82.2 78.7
70
65.6
67.9
68.8
68.4
103 per text structured-f1
Incidents 90
60
3
50 46.0 40 38.9 30 29.8 0 1
2 1 3
BS
5
7
EM
70
69.1
69.9
70.3
70.3
60
68.2
68.5
68.7
68.8
51.2
51.5
51.9
52.2 2
50 48.5 51.8 40
30 30.4 0 1
chrf
3 103 per text
structured-f1
78.7 79.0 80 74.8 74.2 74.6 71.3 70.8 67.7 70.3 66.4 68.7 70 64.6 68.8 63.6 65.0 63.0 63.5 60.3 60 58.4 48.8 51.9 50 44.0 45.0 40 EM chrf BS EM chrf BS Semantic Header Value Union Figure 9: Retrieve utility scores ablation with Llama3-8B.
1 3
5
Input Tokens
7
Figure 10: Retrieve shots ablation with Llama3-8B. Table 4: Influence of Hybrid Integration Strategy. Qwen
Δ chrf
Δ BS
Ratio (%)
Incidents Level 1 Incidents Level 2 Incidents Level 3 Weather Level 1 Weather Level 2 Weather Level 3 Conversation Level 1 Conversation Level 2 Conversation Level 3
0.009±0.004 0.007±0.001 0.005±0.006 0.001±0.001 0.002±0.015 -0.001±0.002 0.002±0.001 0.001±0.000 0.001±0.000
0.005±0.011 0.002±0.007 0.003±0.006 0.001±0.000 -0.007±0.009 0.005±0.003 0.002±0.001 0.001±0.000 0.001±0.000
16.4±1.9 12.9±0.5 12.4±1.3 1.3±0.1 13.9±0.6 17.1±0.1 8.4±3.3 8.3±2.7 12.0±5.2
is in Figure 9, showing that the header and value utilities derived by previews are more beneficial than the plain text semantic, and the union of three utilities is the best. We also ablate the number of retrieved examples, 𝑘 𝑡𝑒 , in Figure 10. Results show that only 1 shot could greatly boost the performance, and the performance reaches a high level after a few shots. The results of other backbones show similar patterns. 6.6.3 Hybrid Integration Strategy. We report the end-to-end improvement of Hybrid Integration Strategy and the Ratio of removed attribute proposals. Results with Qwen are in Table 4, and those of other models are similar. It shows that Qwen removes an average of 11% proposals while maintaining comparable recall-AUC, and that ratio for Llama3-8B is 15%, and Llama3-70B 7%. Moreover, we find that the duplicates predominantly occur among low-coherence proposals, which are at the tail of the recommendation list. This explains why integration brings only modest improvement on recallAUC (Δs). For each setting, we also sample 30 cases and expose the same context to a human, in order to evaluate the alignment between the LLM’s judgment on duplicates with human’s. Results are shown in Figure 12. Based on Cohen’s Kappa [27], Llama3-8B (𝜅=0.43-0.46) and Llama3-70 (𝜅=0.57-0.79) achieve moderate agreement, and Qwen (𝜅=0.73-0.87) achieves substantial agreement. It is hence a pragmatic design to appoint an LLM as the agent for attribute proposal deduplication.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
6.6.4 Semantic Coherence Score. To illustrate the quality of the highest-ranked attributes, namely when the budget 𝑘 is small, we further plot the recall curves for the top-k recommendations with Qwen in Figure 11, and those with Llama (Appendix A.9) show similar patterns. In contrast to the baselines, the results show that our Schema Coherence Score consistently pushes high-quality candidates toward the top of the list, substantially improving early recall, which is ideal for practical short-list applications. This difference is especially evident on the Weather dataset under Levels 1 and 2. We further study the effect of varying the schema attribute permutations succeeding the SP, shown as the shaded region (e.g., there will be 6 possible permutations for 3 attributes). This sensitivity arises because LLMs tend to attend more strongly to nearby context, the last few attributes, during next-token prediction. To mitigate this sensitivity, our full implementation (coherence-M) samples 10 random permutations and retains the maximum score for each candidate, as detailed in the implementation. This strategy successfully avoids underperforming orders and captures each candidate’s peak coherence across multiple contextualizations. 6.6.5 Manual Evaluation on Structured-F1. We conduct a human evaluation on 200 predictions across all benchmarks. Human evaluators are asked to compare the original prediction with a minimally perturbed version introducing a single semantic difference: either single-value corruption or value swap of two cells. The direction of structured-F1 change aligns closely with human preference (accuracy 90%, 95%, 92.5% for sim.=EM, chrf, BS.).
7
Conclusion
We introduce TEAR, a framework that leverages LLM in-context learning for table extraction with attribute recommendation. In the Table Extraction Workflow, we propose a Proactive Demonstration Module to customize demonstrative examples for each input, addressing the ineffectiveness of heuristic instructions by dynamically adapting to the high variability of n-texts. In the Attribute Recommendation Workflow, we design a Discovery Mechanism, a Hybrid Integration Strategy, and a Schema Coherence Score to openly discover, consolidate, and present a ranked list of text-driven new attributes to the user, overcoming the limitations of fixed heuristic schemas. This recommendation capability is the most distinctive feature of our method, tackling the exploratory schema design challenge of naturally occurring texts. We further establish the evaluation benchmarks on n-texts, including two new datasets, appropriate metrics, and a discussion of baselines. Experiments with three open-source LLMs show that TEAR achieves state-of-the-art performance on both tasks, and the recommended attributes effectively enhance table extraction in exploratory scenarios.
8
Discussion
1. AR feedback loop. We clarify that AR does not have an inherent convergence point that iterative refinement could reach because the open-ended candidate space is exhaustive if continuously prompted. As the first to establish the task, we focus on single-round quality, which provides known attributes to the system at the beginning. Multi-round interaction would require modeling the relationships among attributes provided across rounds and a more complex benchmark, which could be pursued in future work.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Recall (chrf)
Incidents 0.8 0.6 0.4 0.2
Level 1
Recall (BS)
0
10
0.8 0.6 0.4 0.2
20
Level 1
0
10
20
Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, and Bo Li
Incidents 0.8 0.6 0.4 0.2 30 0 0.8 0.6 0.4 0.2 30 0
Level 2
10
20
Level 2
10
CFR
20
Incidents 0.8 0.6 0.4 0.2 30 0 0.8 0.6 0.4 0.2 30 0
MSS
Level 3
10
20
Level 3
10
20
Weather 0.8 0.6 0.4 0.2 30 0 0.8 0.6 0.4 0.2 30 0
coherence-M
Level 1
10
20
Level 1
10
20
Weather 0.8 0.6 0.4 0.2 30 0 0.8 0.6 0.4 0.2 30 0
coherence
Level 2
10
20
Level 2
10
|S0|
20
Weather 0.8 0.6 0.4 0.2 30 0 0.8 0.6 0.4 0.2 30 0
Level 3
10
20
Level 3
10
20
30
30