Conceptio › Archive › arXiv CS
arXiv CSopen access

FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities Tian Du1,2,∗ , Tiantong Wu2 , Yafei Wang3 , Mengyu Liu1 , Xingyan Chen3 , Mu Wang3 1

arXiv:2609.17128v1 [cs.AI] 15 Sep 2026

Southwest University of Finance and Economics 2 Nanyang Technological University 3 Beijing University of Post and Telecommunication

Abstract Comprehensive structured data on inter-firm relationships is often scarce or inaccessible because many relationships are privately negotiated, selectively disclosed, and fragmented across proprietary databases. This scarcity hinders the discovery of collaboration opportunities, particularly for startups and small and medium-sized enterprises. Firm profiles are readily available, but collaboration potential cannot be inferred from business similarity alone, since similar firms may be competitors, whereas dissimilar firms may offer complementary products, technologies, channels, capabilities, or capital. We present FirmCORe (Inter-Firm Collaboration Opportunity Reasoning), a human-annotated benchmark for pairwise reasoning over weakly structured firm profiles, comprising 2,805 labeled firm pairs. Given two firm profiles, a model must determine whether the available evidence supports a collaboration opportunity and, for positive pairs, jointly predict its strength, primary collaboration type, and role direction. FirmCORe also provides parallel Chinese- and English-language evaluation sets containing identical instances and gold labels, enabling controlled analysis of input-language sensitivity. Experiments with representative locally deployed and hosted large language models (LLMs) show that the strongest model achieves a macro-F1 score of 74.51 for opportunity detection but only 61.57% exact match across all four output fields. Language effects vary across models, and high cross-language agreement can mask errors shared across languages. These results indicate that current LLMs are substantially more reliable at detecting broad collaboration opportunities than at identifying their specific types and role directions.

1

Introduction

Identifying potential inter-firm collaboration opportunities is important for industrial coordination and supply-chain resilience. However, comprehensive structured data on interfirm relationships remains scarce: many collaborations are established through private negotiations, disclosed only selectively, or recorded in proprietary databases. This data gap is particularly severe for startups and small and medium-sized enterprises, which often lack the resources to identify suitable suppliers, customers, and other potential collaborators. Existing transaction networks and structured supply–demand databases capture only a subset of realized relationships and ∗

Corresponding author.

provide even less information about collaboration opportunities that have not yet materialized. In contrast, firm profiles, including core business activities, products and services, and profile descriptions, are more widely available and therefore provide a practical source of evidence. Inferring collaboration opportunities from firm profiles is nevertheless challenging. Such profiles are often incomplete, weakly structured, and expressed at varying levels of granularity using heterogeneous or domain-specific terminology. More fundamentally, collaboration depends on capability complementarity rather than firm similarity (Hitt et al. 2000; Furlotti and Soda 2018). Similar firms may be competitors, whereas dissimilar firms may complement one another through products, technologies, distribution channels, operational capabilities, or capital (Mitsuhashi and Greve 2009; Mindruta et al. 2016; Greve et al. 2013). Reliable reasoning must therefore determine not only whether the available profile evidence supports a collaboration opportunity, but also its opportunity strength, primary collaboration type, and the role direction of the two firms. Existing research does not directly evaluate this form of reasoning. Corporate similarity methods primarily measure semantic or business relatedness (Davis and Aid 2022; AlMahri et al. 2026). Supply-chain link prediction typically relies on observed relational graphs (Cabrera et al. 2021; Tu et al. 2024). Partner recommendation is usually designed for candidate retrieval or ranking in specific application settings. None jointly evaluates whether two firm profiles provide sufficient evidence for a potential collaboration. Large language models (LLMs) offer a promising approach because they can integrate heterogeneous textual evidence and interpret domain-specific terminology (AlMahri et al. 2026; Zheng and Brintrup 2025; Sun et al. 2025). However, they may conflate relatedness with complementarity, infer plausible but weakly supported collaboration opportunities, or produce inconsistent predictions across opportunity existence, opportunity strength, collaboration type, and role direction. Consequently, the ability of LLMs to perform multidimensional inter-firm collaboration reasoning from firm profiles remains largely unknown. We investigate three research questions: (RQ1) Can LLMs distinguish evidence-supported collaboration opportunities from superficial business relatedness? (RQ2) Can they jointly recover opportunity strength, collaboration type, and

1 Benchmark Construction

▶ 2 Structured Pairwise Task

Heterogeneous Original Firm Records Firm

Main business

Objective: Profile-grounded Capability Complementarity —not similarity alone

firm Frim ... product summary

Firm name, industry, core business, products, firm summary

Stratified candidate-pair mining V4-Flash

Strong-signal

Borderline

Same-industry Hard

Random Non-edge

Industry Core business Products Services Summary

Industry Core business Products Services Summary

B→A A↔B

Reason over products, capabilities, resources, and business roles

Structured Prediction Yes · No

Strength

Weak · Strong ·

Type

Supply · Distribution · Technology · R&D · Capital ·

Human-annotated Dataset

Direction

A→B · B→A · A ↔ B ·

2 Annotators + Third-annotator Tie-breaking → Gold Labels

Confidence

Low · High

Pairwise Annotation

Qwen3.7-Max

GPT5.6Thinking

(Firm A, Firm B)

Parallel ZH–EN Dataset same firm-pairs · structure · gold labels

17 LLMs

Aligned prompts

Deterministic T=0

ZH / EN

Firm B

A→B

Opportunity

AI Quality Screening

Unified Zero-shot Protocol Local + API

Firm A

Unified Firm Profiles

▶ 3 LLM-based Evaluation

⊥ ⊥

⊥

Constraint-aware JSON + brief reason

Multi-dimensional Evaluation Opportunity detection

Macro-F1 · P/R/F1 · MCC

Structured relation prediction

Strength · Type · Direction

Joint consistency

Four-field Exact Match

Reliability / validity

Confidence-conditioned Errors Valid JSON · Field Constraints

Cross-lingual consistency

Δ Score · Agreement · Paired Flips Per-type Changes

Core Benchmark Probes Similarity ≠ Complementarity

Cooperation Mechanism

Directional Roles

Language Sensitivity

Error analysis across models, instances, categories, and languages

Figure 1: Overview of FirmCORe. The benchmark is constructed from unified firm profiles through stratified candidate mining, quality screening, blind human annotation, and tuple-level adjudication. Models then perform zero-shot structured prediction on parallel Chinese and English profile pairs under a unified evaluation protocol.

role direction? (RQ3) Would the structured predictions remain stable across firm profiles in different languages? Additional analyses of model output validity, consistency across different prediction fields, and the model’s self-reported confidence were also included. To address these RQs, we introduce FirmCORe (InterFirm Collaboration Opportunity Reasoning), a humanannotated benchmark comprising 2,805 gold-labeled firm pairs for structured reasoning about inter-firm collaboration opportunities. The main contributions of this paper are as follows: • We formulate inter-firm collaboration opportunity detection as a profile-based, capability-complementarity reasoning task that jointly predicts opportunity existence, opportunity strength, collaboration type, and role direction. • We introduce FirmCORe, a human-annotated benchmark containing 2,805 gold-labeled firm pairs derived from realworld firm profiles, together with parallel Chinese and English evaluation sets. • We systematically evaluate representative LLMs under a unified structured prediction setting and show that detecting broad collaboration opportunities is substantially easier than correctly identifying collaboration types, role directions, and complete four-field relation structures.

2

Related Work

Firm Similarity and Representation. Prior work uses business descriptions, industry attributes, and inter-firm graphs to support tasks such as similarity estimation, competitor retrieval, and industry classification (Cao et al. 2024; Yang et al. 2025). These methods primarily capture semantic or

structural relatedness between firms. However, similarity in products, business activities, or industry affiliation does not necessarily indicate complementary capabilities or an evidence-supported collaboration opportunity.

Link Prediction and Supplier Recommendation. Supply-chain link prediction and supplier recommendation leverage relational graphs, knowledge bases, firm attributes, and procurement signals to predict supplier–customer links or rank candidate firms for collaboration (Kosasih and Brintrup 2025; Li et al. 2025). The former relies primarily on observed inter-firm network structure, whereas the latter typically operates within specific platforms or application contexts. In contrast, FirmCORe uses only two weakly structured firm profiles, without relying on inter-firm networks or platform interactions, to distinguish capability complementarity from superficial relatedness.

LLM-Based Benchmarks for Business Reasoning. Recent benchmarks evaluate LLMs on economic reasoning, corporate tasks, and supply-chain knowledge or workflow execution (Quan and Liu 2024; Guan et al. 2026), primarily through individual questions, documents, simulated scenarios, or predefined operational workflows (Krumdick et al. 2024; Matlin et al. 2025). FirmCORe instead targets profilegrounded structured reasoning about inter-firm collaboration opportunities. Given two weakly structured firm profiles, it assesses whether LLMs can detect an evidence-supported collaboration opportunity and, for positive pairs, predict opportunity strength, collaboration type, and role direction.

3 3.1

The FirmCORe Benchmark

Task Definition

Each benchmark instance consists of a pair of firm profiles presented in a fixed randomized order as Firm A and Firm B. Each firm profile includes the firm name, industry, core business, products and services, and a brief profile summary. Each firm pair appears only once, and the reversed ordering is not included separately. For each pair (i, j), FirmCORe provides the following structured label: yij = (oij , gij , tij , rij ),

(1)

where oij indicates whether a collaboration opportunity exists, gij denotes its strength, tij denotes the primary collaboration type, and rij denotes the role direction. Opportunity existence is binary: oij ∈ {0, 1}, where oij = 1 indicates that a collaboration opportunity exists, and oij = 0 otherwise. Conditional on oij = 1, the remaining labels satisfy: gij ∈ {Weak, Strong}, tij ∈ {Supply, Distribution, Technology, R&D, Capital}, and rij ∈ {A2B : A ← B, B2A : B ← A, Bidirectional : A ↔ B}. gij = Strong denotes a direct and well-supported collaboration interface, whereas gij = Weak denotes a plausible but indirect or less clearly supported opportunity. The collaboration type represents the most direct and best-supported form of collaboration. Role direction is defined relative to the displayed A/B order and captures the primary flow of resources or capabilities. FirmCORe evaluates whether the supplied profiles support a potential collaboration, rather than whether the firms have collaborated in practice. Models must rely solely on the provided profiles. They also generate a brief rationale and a Low/High confidence label, which are analyzed separately from four-field exact match.

3.2

FirmCORe Framework

Figure 1 presents an overview of the FirmCORe framework, which is structured into three stages and five key components. A. Firm Profile Construction We consolidate records from heterogeneous data sources into unified firm profiles, each containing the firm name, industry, core business, products and services, and a textual summary. Records are matched based on identifiers and names, duplicate values are merged at the field level, and missing fields are omitted without imputation. We then remove records lacking valid firm names, deduplicate the resulting firm profiles, and exclude profiles that are too sparse for meaningful pairwise assessment. Data provenance, temporal and geographic coverage, privacy filtering, and release conditions are documented in Appendix C. B. Stratified Candidate Pair Mining To prevent the benchmark from being dominated by trivial negative samples, we mine candidate firm pairs from four strata: strongsignal pairs, borderline-signal pairs, same-industry hard pairs that exhibit business relatedness but lack explicit supply– demand complementarity, and randomly sampled non-edge pairs. Upstream relationship scores are computed from firm profiles and true available supply–demand information and

are used solely for candidate selection, never as ground-truth labels. We retain up to 3,000 firm pairs per stratum and remove duplicate pairs as well as pairs that differ only in ordering. We use DeepSeek-V4-Flash to screen for potential label leakage, privacy risks, and data-quality issues. Up to 750 pairs per stratum are selected for human annotation, while stratum membership, associated scores, and source identifiers are hidden from annotators. Detailed mining criteria and thresholds are provided in Appendix A. C. Human Annotation Each firm pair was independently annotated by two annotators under source-blind conditions (Bender and Friedman 2018). Annotators relied solely on the provided firm profiles and were not allowed to consult external sources. They labeled opportunity existence and, for positive pairs, opportunity strength, collaboration type, and role direction. A positive label required an evidencesupported collaboration interface involving products, services, technologies, channels, capabilities, resources, or capital. Industry similarity or lexical overlap alone was insufficient. For negative pairs, the remaining three fields were set to None (⊥). D. Gold Label Aggregation Labels were aggregated at the level of complete four-field tuples rather than by field-wise voting, which could produce internally inconsistent combinations that no annotator had selected. If the two initial annotators disagreed on the complete four-field tuple, a third annotator independently labeled the instance as a tie breaker. The gold tuple was accepted only when the third annotation matched one of the two initial tuples; cases with three distinct tuples or insufficient evidence remained unresolved and were excluded. Third-annotator adjudication was required for 30.55% of the samples. This process yielded 2,805 goldstandard instances, while the remaining 195 cases were excluded because the disagreements could not be resolved with sufficient confidence. The two initial annotators achieved an average field-level agreement of 84.26% and a full-tuple agreement of 69.45%. See Table 1 for detailed statistics. E. Structured Model Evaluation Let ℓ ∈ {zh, en} denote the input language. Each evaluation instance is expressed as  ∗  ∗ ∗ ∗ dℓij = xℓi , xℓj , yij , yij = o∗ij , gij , t∗ij , rij , (2) where xℓi and xℓj are the firm profiles presented as Firm A ∗ and B, respectively. The gold structured label yij contains opportunity existence, opportunity strength, collaboration type, and role direction. Given a language-aligned task instruction pℓ , a model fθ produces a raw textual response  ℓ zij = fθ pℓ , xℓi , xℓj . (3) A deterministic parser Π maps the raw response to   ℓ  ℓ ℓ ℓ ℓ Π zij = ŷij , ĉℓij , êℓij , ŷij = ôℓij , ĝij , t̂ℓij , r̂ij

(4)

ℓ where ŷij is the predicted four-field label, ĉℓij is the selfreported confidence, and êℓij is a brief rationale. A structurally valid prediction must satisfy

ôℓij = No

=⇒

ℓ ℓ ĝij = t̂ℓij = r̂ij = None.

(5)

Table 1: Statistics of the FirmCORe benchmark. Statistic

Value

Firm profiles Firms / industries Avg. profile length (characters)

3,230 / 1,114 179.09

Annotation outcomes Gold / unresolved pairs Positive / negative pairs Weak / strong opportunities

2,805 / 195 968 / 1,837 35 / 933

Collaboration types Supply / distribution Technology / R&D Capital

641 / 114 86 / 112 15

Role directions A2B / B2A / bidirectional

157 / 700 / 111

Annotation quality Field / tuple agreement Third-annotator involvement

84.26% / 69.45% 30.55%

Parallel ZH–EN sets Chinese / English instances 2,805 / 2,805 Pairs flagged for review 144 Note. Distribution denotes Marketing and Distribution; Technology denotes Technology Transfer; R&D denotes R&D and Co-development.

Let Iℓ denote the set of all evaluation instances and let = (i, j) ∈ I ℓ o∗ij = Yes denote the gold-positive subset. Opportunity detection is evaluated on all instances. Opportunity strength, collaboration type, and role direction ℓ . are evaluated on I+ Opportunity strength, collaboration type, and role direction are evaluated end-to-end on all gold-positive instances ℓ . If a model predicts No, or produces an invalid or missI+ ing relation attribute, the corresponding attribute prediction is counted as incorrect. Oracle-detection evaluation instead conditions on the subset of gold-positive instances that the model correctly detects as opportunities. The former measures complete pipeline performance, whereas the latter isolates relation-attribute reasoning after successful detection. For full structured prediction, we use four-field exact match: X   1 ℓ ∗ EMℓ = ℓ I ŷij = yij , (6) |I | ℓ ℓ I+

(i,j)∈I

where I[·] denotes the indicator function. For the parallel Chinese and English evaluation sets, cross-lingual consistency is measured by paired exact agreement:  1 X  zh en . (7) Agreezh,en = I ŷij = ŷij |I| (i,j)∈I

3.3

Chinese–English Parallel Evaluation Set

To assess model sensitivity to input language under controlled conditions, we constructed an English parallel evaluation set (Han et al. 2026). It contains the same 2,805 firm pairs, profile fields, and gold labels as the original Chinese

benchmark. Let P denote the fixed set of annotated ordered firm pairs, and let xzh k denote the Chinese profile of firm k. For each language ℓ ∈ {zh, en}, we define   zh (8) Dℓ = dℓij (i, j) ∈ P , xen k = T xk , where T (·) denotes the field-level translation function. Dzh and Den contain identical ordered firm pairs and gold labels, differing only in the language of the firm profiles. More details are shown in Appendix B. The English profiles were translated field by field with assistance from GPT5.6-Thinking and then selectively reviewed by human annotators. The translation process was designed to preserve the semantic scope, level of detail, and uncertainty of the Chinese originals. For recurring companies, consistent translations were used for firm names, industries, core business descriptions, products and services, and profile summaries. We flagged 94 distinct firm profiles appearing in 144 firm-pair instances for mandatory review because they involved potential ambiguities in terminology, firm names, abbreviations, business roles, or directional semantics. Annotators compared each flagged English field against its Chinese source and corrected any errors. This paired design enables instance-level comparison of model predictions while holding sample composition and gold labels constant. The complete Chinese and English prompts, output schema, and label mappings are provided in Appendix D.

4 4.1

Experiments

Experimental Setups

Models We evaluate 17 LLMs, ranging from 1.5Bparameter models to large mixture-of-experts systems. The benchmark covers both locally deployed open-weight models, including DeepSeek-7B, Llama3.1-8B, Llama3.2-3B, Qwen3-2B/4B/8B/Coder, Gemma3-12B, and Gemma4-26B, and hosted models, including GLM-5.1/5.2, Kimi-K2.7, Qwen3-VL-8B, Qwen3.6-27B, Qwen3.6-35B, Qwen3.6Max, and Qwen3.7-Plus. Implementation Each model received profiles of Firm A and B and a standardized prompt, and was asked to determine whether a collaboration opportunity existed, its strength and type, the directional roles, and the model’s self-reported confidence. All models used the same prompt template, output parser, label-normalization pipeline, and evaluation scripts. All models were evaluated in a zero-shot setting, with no examples or task-specific fine-tuning. Provider-specific reasoning modes were disabled. Each sample was evaluated once with a temperature of 0 and a maximum output length of 2,048 tokens. Hosted models were accessed through OpenAIcompatible chat completion APIs with a 180-second timeout. Transient connection failures, rate-limit errors, and server errors were retried up to four times. Local models were evaluated using Hugging Face Transformers. Evaluation Metrics We follow the structured evaluation protocol defined in Section 3.2. Opportunity detection is primarily measured by Macro-F1. Relation attributes are evaluated under both end-to-end and oracle-detection settings and complete predictions are assessed using four-field exact

5

Results and Analysis

Following the RQs, we first evaluate opportunity detection, then examine structured relation recovery and input-language sensitivity. We conclude with an analysis of output validity and self-reported confidence. Additional baseline comparisons, complete diagnostic results, and challenging failure cases are provided in Appendix E.

5.1

RQ1: Distinguishing Collaboration Opportunities from Relatedness

Table 2 reports performance on the collaboration opportunity detection. LLMs tend to conflate business relatedness with genuine collaboration potential. Qwen3.6-27B performs best, achieving a Macro-F1 of 74.51 with a 95% bootstrap confidence interval of 72.89–76.07. Its positive-class precision and recall are 59.99% and 84.71%, respectively. Several other models attain near-saturated recall but substantially lower precision. For example, Qwen3.6-35B and GLM-5.1 reach recalls of 97.62% and 98.86%, yet their precision is only 44.16% and 38.90%. This pattern indicates systematic overprediction once a model detects a seemingly plausible industrial or commercial connection. Table 2: Opportunity detection performance on Chinese. Model Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

51.50 51.42 50.82 37.64 36.50 34.48 32.22 25.77

64.32

66.42

58.77

61.96

58.85

58.59

58.36

61.04

56.25

61.65

59.76

61.37

44.43

59.03

30.47

52.89

41.77

48.70

44.27

48.80

42.81

51.04

36.29

49.45

41.83

47.76

36.74

52.21

36.07

50.92

14.46

40.68

37.58

42.99

19.38

44.85

44.75

48.11

40.41

45.49

40.76

43.11

37.08

44.33

41.68

49.91

29.92

49.54

36.60

36.90

11.58

41.01

35.07

37.14

9.38

39.76

35.67

29.15

13.74

39.65

34.80

28.02

9.38

39.31

34.43

12.34

8.09

39.31

R

F1

38.10 42.49 41.69 36.69 36.87 36.34 35.93 34.65

26.96 93.80 97.93 94.73 98.97 97.93 99.90 99.69

31.58 58.49 58.48 52.90 53.73 53.01 52.86 51.43

4.30 28.55 31.25 12.96 17.62 14.20 14.56 1.89

Qwen3.6-27B 74.51 59.99 84.71 70.24 52.25 54.40 89.46 67.66 47.88 Qwen3.6-Max 70.23 Qwen3.7-Plus 69.64 53.62 91.84 67.71 48.33 Qwen3-Coder 56.48 43.75 89.98 58.87 30.39 Qwen3.6-35B 56.04 44.16 97.62 60.81 36.37 GLM-5.2 53.78 42.97 97.21 59.59 33.57 Kimi-K2.7 52.87 42.55 97.42 59.23 32.87 Qwen3-VL-8B 44.46 39.32 96.80 55.92 23.49 GLM-5.1 43.22 38.90 98.86 55.83 24.67 Scores are percentages. N = 2,805 (968 positive, 1,837 negative). Incomplete prediction files are excluded, 95% CIs are provided in the machine-readable results.

Figure 2 further shows that models generally perform worse on same-industry hard pairs. For Qwen3.6-27B, Macro-F1 drops from 66.42 on randomly sampled non-edge pairs to 58.77 on same-industry hard pairs. This suggests that industry membership and semantic similarity often serve as decision shortcuts. Detailed comparisons with traditional and lightweight baselines, together with bootstrap uncertainty estimates, are provided in Appendix E. The best zero-shot LLM

80

60

40

20

0

Figure 2: Opportunity-detection Macro-F1 by candidategeneration stratum on the Chinese benchmark. Stratum prevalence differs substantially, so cross-stratum values are descriptive rather than controlled estimates of intrinsic difficulty. Table 3: End-to-end structured relation prediction on the Chinese benchmark. F1 denotes Macro-F1.

MCC

P

100

borderline random hard strong-signal Candidate-Generation Stratum

Model

Positive

Macro-F1

Qwen3.6-27B Qwen3.6-Max Qwen3.7-Plus Qwen3-Coder Qwen3.6-35B GLM-5.2 Kimi-K2.7 Qwen3-VL-8B GLM-5.1 Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

Macro-F1 (%)

match. We also report paired Chinese–English agreement, output validity, and confidence-based reliability metrics.

Strength

Type

Direction

Acc.

F1

Acc.

F1

Acc.

F1

Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

20.66 69.94 66.53 78.51 77.48 86.88 68.08 55.48

19.59 46.42 46.00 46.41 45.93 50.73 45.27 38.85

15.08 67.56 70.66 32.54 60.02 68.70 51.65 46.69

13.05 47.42 53.34 31.75 42.77 57.14 44.44 14.55

7.64 69.11 73.45 38.95 67.56 59.71 63.84 31.71

13.34 52.10 66.52 36.31 52.25 54.89 61.02 22.25

Qwen3.6-27B Qwen3.6-Max Qwen3.7-Plus Qwen3-Coder Qwen3.6-35B GLM-5.2 Kimi-K2.7 Qwen3-VL-8B GLM-5.1

64.67 76.34 72.42 86.98 75.62 75.21 88.64 60.85 58.16

44.62 49.87 46.74 46.70 48.04 50.00 53.08 41.83 42.32

60.12 66.32 68.18 64.67 70.87 67.56 74.28 59.61 65.50

51.05 51.97 58.10 46.65 54.53 54.25 60.23 40.07 56.28

63.95 68.39 72.73 68.08 68.90 67.87 77.38 68.08 71.80

56.11 58.61 67.04 59.54 63.38 59.87 67.71 54.61 63.72

outperforms all evaluated baselines, although the competitive supervised result indicates that task-specific decision patterns can also be learned effectively from labeled examples.

5.2

RQ2: Recovering Structured Collaboration Relations

Opportunity detection does not imply accurate recovery of the underlying relation structure. Table 3 reports complete per-model end-to-end results. Kimi-K2.7 obtains the highest Macro-F1 scores for collaboration type and

E2E Oracle

0 50 100 0 50 100 0 50 100 Strength Macro-F1 (%) Type Macro-F1 (%) Direction Macro-F1 (%)

Figure 3: End-to-end and oracle-detection Macro-F1 for opportunity strength, collaboration type, and role direction on the Chinese benchmark. End-to-end evaluation retains all gold-positive instances and counts a missed opportunity as an attribute error; oracle evaluation conditions on correctly detected gold-positive instances. 70

(57.6, 62.2)

(61.6, 60.1) (52.0, 59.2)

60

(40.2, 51.6)

(44.2, 51.1) (55.9, 53.9)

English EM (%)

50

(44.9, 44.2)

(37.2, 44.1) (41.5, 44.1)

40 (22.2, 28.0)

30

(26.2, 31.3)

(35.5, 30.5)

(20.7, 25.9)

20

(25.3, 21.8) (14.7, 14.5)

10 0

Llama3.2-3B DeepSeek-7B Llama3.1-8B Qwen3-4B Qwen3-8B Gemma3-12B Gemma4-26B Qwen3-2B Qwen3-VL-8B GLM-5.1

(13.0, 11.9) (4.9, 4.7)

0

10

20

30 40 Chinese EM (%)

50

60

70

GLM-5.2 Kimi-K2.7 Qwen3.6-35B Qwen3-Coder Qwen3.7-Plus Qwen3.6-27B Qwen3.6-Max

Figure 4: Chinese versus English four-field exact match across evaluated models. The dashed diagonal indicates equal performance in both languages.

role direction, reaching 60.23 and 67.71, despite achieving only 52.87 Macro-F1 for opportunity detection. In contrast, Qwen3.6-27B leads opportunity detection with 74.51 MacroF1 but reaches only 51.05 and 56.11 for collaboration type and role direction. Joint prediction remains substantially more difficult. Kimi-K2.7 achieves the highest positive four-field exact match of 61.98%, whereas Qwen3-2B obtains high overall exact match mainly by rejecting negative pairs and reaches only 4.65% exact match on gold-positive instances. Detailed per-type results in Appendix E show that Technology Transfer is particularly difficult and is frequently con fused with Supply and Production. Overall, current LLMs are considerably more reliable at identifying broad collaboration potential than at reconstructing its complete strength, mechanism, and role direction.

Where do structured-relation errors arise? Figure 3 separates errors caused by missed opportunities from errors in assigning relation attributes. The two scores are close for most high-recall systems. Kimi-K2.7, for example, detects 97.42% of gold-positive pairs and has a mean oracle–endto-end gap of only 1.11 points across the three attributes. Gemma4-26B, Qwen3.6-35B, GLM-5.2, and GLM-5.1 likewise have mean gaps below one point. Their remaining errors therefore arise mainly after an opportunity has already been detected. In contrast, Qwen3-2B covers only 26.96% of gold positives and its mean gap reaches 19.21 points, making missed detection the dominant bottleneck for that model. Qwen3.6-27B also shows a non-negligible 5.52-point gap despite leading opportunity detection overall. Complete joint and oracle-detection results are reported in Appendix E.2. Appendix F further decomposes positive-pair errors and analyzes systematic collaboration-type confusion and roledirection bias. Both correct

Qwen3.6-Max Qwen3.6-27B Qwen3.7-Plus Qwen3-Coder Qwen3.6-35B Kimi-K2.7 GLM-5.2 GLM-5.1 Qwen3-VL-8B Qwen3-2B Gemma4-26B Gemma3-12B Qwen3-8B Qwen3-4B Llama3.1-8B DeepSeek-7B Llama3.2-3B

Both incorrect

EN only

ZH only

Local API

0 50 100 0 Paired exact agreement (%)

20 40 60 80 Paired correctness outcome (%)

100

Figure 5: Paired Chinese–English prediction agreement and correctness outcomes. Paired agreement measures tuple identity and does not imply correctness. Qwen3.6-Max Qwen3.6-27B Qwen3.7-Plus Qwen3-Coder Qwen3.6-35B Kimi-K2.7 GLM-5.2 GLM-5.1 Qwen3-VL-8B Qwen3-2B Gemma4-26B Gemma3-12B Qwen3-8B Qwen3-4B Llama3.1-8B DeepSeek-7B Llama3.2-3B

-1.24

-1.69

-5.68

-12.44

8.36

-6.25

-0.63

-9.27

50

-1.79

-4.32

0.10

-6.46

1.65

-5.88

-2.68

-9.22

40

-9.34

30

-1.49

20

25.94

10

-5.70

0

-0.95 -3.17 -1.67

-2.31 6.35 2.60

-1.48

-1.53

1.38

3.78

0.47

-4.40

0.77

-2.14

0.03 2.27 2.25 5.99 0.05 1.02

-0.32 -0.96

1.65

-4.33

0.00

3.18

-9.71

-2.56

-4.00

1.48

0.00

1.90

-3.28 -2.25

0.00

-0.09

0.00

0.23

-0.43 -0.59 -1.19 -2.90

25.36

16.67

-2.21

30.00

9.90

28.53

5.80

6.58

1.90

-17.04

11.33

-19.89

50.60

12.50

-0.48

29.33

Su Di Te R& str ppl D ibu chnol y ogy tio n

Ca pi

-15.15

3.14

6.34

0.17

20.27

-2.09

-8.73

-13.19

5.44

-1.30

30.42

-29.62

5.23 5.69

39.18

-2.25 3.21

-4.18 3.27

21.78

-6.24 -0.55 -3.83

14.08

-2.51

-0.59

4.03

1.36

0.00

-0.68

-5.22

-3.83 -4.94 -2.40

0.00

-17.95

5.95

-2.90

-12.05

-22.40

8.36

-2.30

-7.01

-9.42

-10.29

-0.32

0.00

-3.50

-5.56

3.52

t al

A2 B

-9.51

-4.61

-1.35 -4.66

B2 A

-3.94

7.63

EN − ZH F1 (points)

Llama3.2-3B Gemma3-12B Qwen3-4B Llama3.1-8B DeepSeek-7B Gemma4-26B Qwen3-8B Qwen3-2B GLM-5.1 Qwen3-VL-8B Kimi-K2.7 GLM-5.2 Qwen3.6-35B Qwen3-Coder Qwen3.7-Plus Qwen3.6-Max Qwen3.6-27B

−10

-9.72

−20

0.00

Bid i

rec t

ion al

Figure 6: English-minus-Chinese class-level F1 differences for collaboration types and role directions. Positive values indicate higher performance under English input.

5.3

RQ3: Sensitivity to Input Language

Input-language sensitivity. Figure 4 shows that English input does not yield a consistent advantage. For example, four-field exact match increases by 11.34 points for Qwen3.635B but decreases by 4.96 points for Qwen3-8B, while Qwen3.6-27B remains relatively stable (−1.43 points). Complete model-level results are reported in Appendix Table 8. Figure 5 further shows that similar aggregate scores can conceal substantial instance-level changes. Moreover, high cross-language agreement does not necessarily indicate robust reasoning: Qwen3.6-Max produces identical predictions in 79.57% of paired instances, but is correct in both languages on only 54.62%, indicating that some errors are repeated across languages. Category-level variation. Figure 6 shows that language effects also vary across collaboration types and role directions. English improves some categories for particular models while degrading the same categories for others, and no relation category exhibits a consistent language advantage across all systems. Input-language sensitivity is therefore model-, instance-, and category-dependent rather than a uniform performance shift.

5.4

Reliability Analysis

Valid outputs are not necessarily reliable. Most hosted models produce syntactically valid JSON, but their crossfield consistency varies substantially. Qwen3-Coder achieves the lowest violation rate of 0.21%, whereas several models exceed 10%. Self-reported confidence is also only partially informative: although Qwen3.6-27B, Qwen3.7-Plus, and Qwen3.6-Max exhibit large High–Low exact-match gaps, their high-confidence subsets still contain 31.86–39.64% errors. Complete output-validity and confidence-conditioned results are reported in Appendix E.5.

6

Conclusion

We introduce FirmCORe, a human-annotated benchmark for evaluating structured reasoning over inter-firm collaboration opportunities from weakly structured firm profiles. The task requires models to distinguish capability complementarity from superficial relatedness and jointly predict the existence, strength, type, and role direction of each opportunity. Experiments on 2,805 parallel Chinese–English samples show that current LLMs can identify broad collaboration potential but remain far less capable of reconstructing the full relational structure. They often overpredict opportunities based on mere relatedness, conflate distinct collaboration mechanisms, and misidentify the direction of resource flows. Language effects vary across models and instances, while high cross-lingual consistency may reflect the same errors being repeated across languages rather than genuinely robust reasoning. FirmCORe therefore provides a controlled testbed for diagnosing and improving profile-grounded, multidimensional inter-firm reasoning in future models.

7

Limitations

FirmCORe evaluates potential collaboration opportunities inferred from provided firm profiles, rather than realized

partnerships or their commercial viability. Its pairwise setup does not address large-scale partner retrieval or ranking. Our experiments are further limited to zero-shot prompting and single deterministic runs, leaving fine-tuning, prompt sensitivity, tool use, and multi-run variability unexplored.

References AlMahri, S.; Xu, L.; and Brintrup, A. 2026. Enhancing supply chain visibility with knowledge graphs and large language models. International Journal of Production Research, 64(6). Bender, E. M.; and Friedman, B. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6: 587–604. Cabrera, S. C.; Pishchulov, G.; Sampaio, P.; Mehandjiev, N.; Liu, Z.; and Kununka, S. 2021. An approach and decision support tool for forming Industry 4.0 supply chain collaborations. Computers in Industry, 125: 1–16. Cao, L.; von Ehrenheim, V.; Granroth-Wilding, M.; Anselmo Stahl, R.; McCornack, A.; Catovic, A.; and Cavalcanti Rocha, D. D. 2024. Companykg: A large-scale heterogeneous graph for company similarity quantification. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4816–4827. Davis, C.; and Aid, G. 2022. Machine learning-assisted industrial symbiosis: Testing the ability of word vectors to estimate similarity for material substitutions. Journal of Industrial Ecology, 26(1): 27–43. Furlotti, M.; and Soda, G. 2018. Fit for the task: Complementarity, asymmetry, and partner selection in alliances. Organization Science, 29(5): 837–854. Greve, H. R.; Mitsuhashi, H.; and Baum, J. A. 2013. Greener pastures: Outside options and strategic alliance withdrawal. Organization Science, 24(1): 79–98. Guan, S.; Liu, Y.; and Cao, L. 2026. SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management. In Findings of the Association for Computational Linguistics: ACL 2026, 7526–7550. Han, W.; Zhang, Y.; Chen, Z.; Pechenizkiy, M.; Fang, M.; Zheng, Y.; et al. 2026. Mubench: Assessment of multilingual capabilities of large language models across 61 languages. In Findings of the Association for Computational Linguistics: ACL 2026, 16163–16192. Hitt, M. A.; Dacin, M. T.; Levitas, E.; Arregle, J.-L.; and Borza, A. 2000. Partner selection in emerging and developed market contexts: Resource-based and organizational learning perspectives. Academy of Management journal, 43(3): 449– 467. Kosasih, E. E.; and Brintrup, A. 2025. Towards trustworthy AI for link prediction in supply chain knowledge graph: a neurosymbolic reasoning approach. International Journal of Production Research, 63(6): 2268–2290. Krumdick, M.; Koncel-Kedziorski, R.; Lai, V. D.; Reddy, V.; Lovering, C.; and Tanner, C. 2024. BizBench: A quantitative

reasoning benchmark for business and finance. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8309–8332. Li, Y.; Ko, H.; and Ameri, F. 2025. Integrating graph retrieval-augmented generation with large language models for supplier discovery. Journal of Computing and Information Science in Engineering, 25(2): 021010. Matlin, G.; Okamoto, M.; Pardawala, H.; Yang, Y.; and Chava, S. 2025. Finance language model evaluation (flame). In Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2 ), 880–926. Mindruta, D.; Moeen, M.; and Agarwal, R. 2016. A twosided matching approach for partner selection and assessing complementarities in partners’ attributes in inter-firm alliances. Strategic Management Journal, 37(1): 206–231. Mitsuhashi, H.; and Greve, H. R. 2009. A matching theory of alliance formation and organizational success: Complementarity and compatibility. Academy of management journal, 52(5): 975–995. Quan, Y.; and Liu, Z. 2024. Econlogicqa: A questionanswering benchmark for evaluating large language models in economic sequential reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024, 2273–2282. Sun, Q.; Zheng, J.; Jin, B.; Chen, L.; and Peng, Y. 2025. InterCorpRel-LLM: Enhancing Financial Relational Understanding with Graph-Language Models. arXiv preprint arXiv:2510.09735. Tu, Y.; Li, W.; Song, X.; Gong, K.; Liu, L.; Qin, Y.; Liu, S.; and Liu, M. 2024. Using graph neural network to conduct supplier recommendation based on large-scale supply chain. International Journal of Production Research, 62(24): 8595– 8608. Yang, B.; Zhang, B.; Cutsforth, K.; Yu, S.; and Yu, X. 2025. Emerging industry classification based on BERT model. Information Systems, 128: 102484. Zheng, G.; and Brintrup, A. 2025. Enhancing supply chain visibility with generative AI: an exploratory case study on relationship prediction in knowledge graphs. International Journal of Production Research, 1–23.

A

Candidate-Mining Details

For firm ei , let Si and Di denote its supply and demand label sets, and let Ai = Si ∪ Di . We define   |U ∩ V | , |U ∪ V | > 0, (9) J(U, V ) = |U ∪ V |  0, otherwise, and compute J(Di , Sj ) + J(Dj , Si ) . 2 (10) These signals are used only for candidate mining and are not treated as gold labels. Sim(i, j) = J(Ai , Aj ), Comp(i, j) =

Stratum Strong-signal Borderline Same-industry Hard Random Non-edge

Criterion Upstream score ≥ 90 Upstream score in [60, 75) Same industry, Sim(i, j) ≥ 0.03, and Comp(i, j) ≤ 0.12 No edge in the upstream candidate graph

Table 4: Candidate-mining strata. Same-industry hard pairs are not presumed to be negative. The upstream candidate graph and scores are generated by DeepSeek-V4-Flash from firm profiles and published collaboration-opportunity information. We retain at most 3,000 pairs per stratum, remove cross-stratum and reversed duplicates, and randomly assign pairs satisfying multiple criteria to one stratum, as shown in Table 4. For scored edges, the Firm A/Firm B order follows the fixed randomized order. Qwen3.7-Max is subsequently used only to screen for potential label leakage, privacy risks, and data-quality issues. At most 750 pairs per stratum are retained for human annotation. Candidate strata, upstream scores, and source-record identifiers are hidden from annotators.

B

Chinese–English Parallel Translation

The original firm profiles and annotations in FirmCORe were collected and constructed in Chinese. To support crosslingual evaluation, we translated the full dataset into English using a controlled translation and quality-assurance protocol. The objective was not to produce a literal word-by-word translation, but to preserve the business meaning, entity identity, relational direction, and evidential boundaries of each original record. The resulting English version contains the same 2,805 annotated firm pairs as the Chinese version and preserves all gold labels without modification. Profile-level translation. Each firm profile was decomposed into five fields: entity name, industry, core business, products or services, and firm summary. Translation was performed at the profile level rather than independently at each occurrence. After removing repeated profiles, 2,256 unique profiles were identified and assigned a single English representation. The same translation was then reused whenever the profile appeared in multiple firm pairs. This procedure prevents a Chinese entity or business description from receiving inconsistent translations across different examples.

Business descriptions were translated into natural professional English while preserving the scope of the source text. In particular, we avoided introducing capabilities, products, qualifications, or business relationships that were not explicitly supported by the Chinese profile. Long Chinese nominal expressions were reorganized when necessary to improve readability, but their semantic content was not expanded. Enumeration boundaries were also retained so that separate products, services, and technical capabilities were not incorrectly merged. Entity-name normalization and verification. Entity names required special treatment because the dataset contains listed companies, private firms, public institutions, universities, government departments, branches, stores, research centers, projects, brands, and other non-standard organizational labels. For internationally recognized firms and institutions, we preferentially used the English name published on the entity’s official website or in official corporate materials. Examples include Alibaba Group, Tencent, Huawei, Chinese Academy of Sciences, and Tsinghua University. When no verifiable official English name was available, we produced a standardized English rendering based on the Chinese name and its organizational form. Such renderings were not treated as official legal names. For local firms whose names contained distinctive Chinese brand terms, transliteration was retained where a literal translation would obscure the entity identity. Descriptive translation was used for generic institutional names, such as research centers, administrative offices, branches, and service platforms. We recorded the provenance and confidence of entityname translations using three review levels. Verified names were supported by official English-language sources. Review recommended was assigned when the translation was linguistically reliable but the official English name could not be independently confirmed, or when the record referred to a branch, department, project, or non-corporate organization. Review required was assigned when the Chinese source name was abbreviated, malformed, generic, combined multiple entities, or otherwise insufficient for unique entity resolution. Domain terminology. A controlled terminology glossary was maintained for recurring concepts in corporate registration, manufacturing, supply chains, technology transfer, and public-sector administration. Terms were translated according to their functional meaning rather than by surface lexical correspondence. The five collaboration labels were also translated using fixed expressions throughout the dataset: Supply and Production Collaboration, Marketing and Distribution Collaboration, R&D and Co-development Collaboration, Licensing and Technology Transfer Collaboration, and Capital and Equity Collaboration. Direction labels were preserved explicitly. Thus, A2B indicates that Firm A provides products, services, technology, capabilities, or other resources to Firm B, whereas B2A indicates the reverse direction. No direction label was inferred or altered during translation. Preservation of source uncertainty. The Chinese profiles occasionally contained incomplete names, conflicting

industry descriptions, historical company names, or inconsistencies between the listed industry and business summary. Translation was not used to silently repair these source-level issues. Instead, the English version preserves the available information and, when necessary, explicitly signals that the source profile is inconsistent or that the legal entity cannot be uniquely identified. This design prevents translation from artificially improving the informational quality of one language version and ensures that Chinese–English comparisons remain valid. Quality assurance. The translated data underwent several automatic and manual checks. We verified that every original pair had a corresponding English record, that repeated Chinese profiles received identical English translations, and that all collaboration labels, scores, and role directions were unchanged. We additionally checked for missing English fields, residual Chinese characters in translated fields, inconsistent English names for the same Chinese entity, and malformed profile structures. High-risk entity names and source inconsistencies were retained in dedicated review fields. These procedures resulted in a traceable bilingual dataset in which each English profile can be directly aligned with its Chinese source, translation status, and review note.

C

Data Statement

Data source and acquisition date. The source data underlying FirmCORe was obtained from the Xunfu platform through an operational data export generated in June 2026. The export contains core organization and business records, supplementary profile information, upstream candidate matching relations, and AI generated labels. We use these records to construct unified organization profiles, intermediate supply and demand tables, candidate pairs, source blind annotation sets, and the final gold standard benchmark. Due to commercial access restrictions and licensing constraints, we release only the final gold labeled benchmark data. Data sources by field. The raw platform fields used in FirmCORe include organization names, selected industry and organization type attributes, and available business or contact related metadata. Derived fields include profile length, label overlap, business similarity, supply and demand complementarity, and shared label sets. These fields are computed using deterministic scripts. Other intermediate fields, including text summaries, AI generated labels, supply and demand labels, candidate relation scores, and explanatory fields, originate from upstream AI or platform systems. FirmCORe is therefore not an unmodified collection of corporate registration records. Entity matching, deduplication, and filtering. Organization profiles are constructed primarily by aligning names across platform tables, supplemented by available organization identifiers. Duplicate field values are merged, records with missing or invalid organization names are removed, and duplicate profiles are deduplicated. Profiles that contain insufficient business information for meaningful pairwise assessment are also excluded. During candidate pair construc-

tion, we remove self pairs, duplicate unordered pairs, reversed duplicates, pairs containing invalid profiles, and pairs containing fields that could reveal candidate generation signals. The current entity matching process relies mainly on organization names and deterministic rules rather than a comprehensive entity resolution system. Candidate pair construction and representativeness. FirmCORe is not uniformly sampled from all possible organization pairs. Candidate pairs are deliberately mined from four strata: strong signal upstream pairs, borderline signal upstream pairs, same industry hard pairs, and random non-edge pairs. We first construct a larger candidate pool for each stratum, then apply AI based quality screening and source blind human annotation. Up to 750 pairs are selected from each stratum, producing an annotation pool of 3,000 pairs. After complete tuple aggregation and disagreement resolution, 2,805 instances receive definitive gold labels. This stratified design covers decision regions ranging from obvious to difficult cases and prevents the benchmark from being dominated by trivial random negatives. Privacy and sensitive information. The original platform exports contain fields that are unsuitable for direct public release, including phone numbers, email addresses, authentication tokens, session identifiers, national identification numbers, instant messaging accounts, and other potentially sensitive metadata. FirmCORe does not release the original exports. Human annotation and model evaluation are based on minimized organization level business profiles that exclude direct contact details, authentication information, and internal platform fields. We also remove fields such as reason, action_plan, and raw_json, since they may contain sensitive content or directly reveal upstream system judgments. These measures reduce unnecessary exposure but do not replace independent legal, privacy, and data governance review. Any public release should therefore be limited to the minimum information required to reproduce the benchmark task. Label semantics and intended use. A positive label in FirmCORe indicates that, under the annotation guidelines, the two provided organization profiles contain sufficient evidence to support a plausible collaboration opportunity. Positive instances are further labeled with opportunity strength, primary collaboration type, and role direction. These labels do not indicate that the organizations have collaborated in practice. A negative label indicates that the available profiles do not provide sufficient evidence for a collaboration opportunity under the annotation guidelines. Since relevant capabilities or constraints may be absent from the profiles, a negative label should not be interpreted as evidence that collaboration is impossible in practice. FirmCORe is intended to evaluate profile based relation reasoning and structured relation prediction. Known biases and limitations. FirmCORe inherits several potential biases. First, the source data primarily reflects Chinese commercial and institutional contexts. The English evaluation set is translated from the same Chinese profiles rather than independently collected from naturally occurring

English organization profiles. Second, the gold labels contain fewer positive than negative instances. Positive instances are also concentrated in strong opportunities, the Supply and Production category, and one displayed role direction. These distributions may affect aggregate accuracy and class specific performance. We therefore report Macro-F1 and per class metrics in addition to accuracy. Finally, organization profiles reflect business information available at a particular point in time. They may become outdated as products, capabilities, and organizational roles change.

D

Prompts

We use language-aligned prompts for the Chinese and English evaluation sets. The two prompts share the same task definition, label space, consistency constraints, and JSON output schema. The Chinese prompt requires a brief rationale in Chinese, while the English prompt requires a brief rationale in English. We provide the complete English prompt below. English Prompt You are an annotator specializing in inter-firm collaboration opportunities. Given the profiles of two industry entities, Company A and Company B, determine whether the provided information supports a potential collaboration opportunity between them. Base your judgment strictly on the provided profiles. Do not introduce external knowledge or assume facts that are not stated or reasonably supported by the text. Output JSON only. Do not output Markdown, additional explanations, or any content outside the JSON object. The annotation fields are defined as follows. 1. has_opportunity • Yes: The profiles provide sufficient evidence for a plausible collaboration opportunity between A and B. • No: The profiles do not provide sufficient evidence for a clear collaboration or capability-complementarity relationship. 2. opportunity_score • 1: Weak opportunity. Cooperation is plausible, but the relationship is indirect, the evidence is limited, or the collaboration chain is relatively long. • 2: Strong opportunity. One party’s products, services, technologies, channels, capabilities, resources, or capital can relatively directly support the other party. 3. cooperation_type Select exactly one of the following: • Supply and Production Collaboration • Marketing and Distribution Collaboration • Licensing and Technology Transfer Collaboration • R&D and Co-development Collaboration • Capital and Equity Collaboration 4. role_direction • A2B: Under the selected cooperation type, Company A primarily provides Company B with the relevant products,

services, technologies, capabilities, resources, channels, or capital. • B2A: Under the selected cooperation type, Company B primarily provides Company A with the relevant products, services, technologies, capabilities, resources, channels, or capital. • Bidirectional: The opportunity primarily involves mutual exchange, joint development, or two-way value creation. 5. confidence • 1: Low confidence. • 2: High confidence. 6. reason Provide one concise English sentence explaining the evidence for the judgment. Consistency requirements • If has_opportunity = No, then opportunity_score, cooperation_type, and role_direction must all be None. • If has_opportunity = Yes, then opportunity_score must be 1 or 2, and neither cooperation_type nor role_direction may be None. • A strong opportunity must be supported by a relatively direct product, service, technology, channel, capability, resource, or capital relationship. • Business similarity, shared industry membership, or lexical overlap alone is insufficient evidence for a positive label. Company A {object_a_profile} Company B {object_b_profile} Return exactly one JSON object using the following schema: { "has_opportunity": "Yes / No", "opportunity_score": "1 / 2", "cooperation_type": "Supply and Production Collaboration / Marketing and Distribution Collaboration / Licensing and Technology Transfer Collaboration / R&D and Co-development Collaboration / Capital and Equity Collaboration", "role_direction": "A2B / B2A / Bidirectional", "confidence": "1 / 2", "reason": "One concise English sentence" }

E E.1

Additional Results

Baselines and Detection Uncertainty

Table 5 compares the best evaluated zero-shot LLM, Qwen3.6-27B, with trivial priors, lexical and embeddingbased similarity methods, rule-based signals, the upstream scoring system, and a lightweight supervised classifier. The trivial baselines expose the effect of class imbalance: AlwaysYes reaches 100% positive recall but only 25.66 Macro-F1, whereas Always-No and the majority-tuple baseline obtain

Table 5: Comparison of opportunity-detection baselines on the Chinese benchmark. Baseline

Macro-F1

Positive P

Positive R

Positive F1

MCC

Always-No Majority-Prior Always-Yes

39.57 39.57 25.66

0.00 0.00 34.51

0.00 0.00 100.00

0.00 0.00 51.31

0.00 0.00 0.00

LSA-Cosine

43.67

33.55

64.46

44.13

-2.84

Com.-Threshold Industry-Rule

46.43 33.29

67.29 9.11

7.44 7.33

13.40 8.13

13.73 -33.12

Upstream-Score

67.29

64.78

46.18

53.92

36.37

TF-IDF-Cosine J.S.

47.71 47.18

37.17 34.00

74.38 54.75

49.57 41.95

8.36 -1.21

L.S.C

71.69

59.89

70.04

64.57

43.92

Best zero-shot LLM (Qwen3.6-27B)

74.51

59.99

84.71

70.24

52.25

39.57 Macro-F1 without identifying any positive instance. Unsupervised similarity and rule-based methods remain below 48 Macro-F1. Industry similarity performs particularly poorly, confirming that shared industry membership alone is not a reliable proxy for collaboration potential. Supply– demand complementarity attains relatively high precision (67.29%) but very low recall (7.44%), indicating that direct label overlap is overly conservative and misses implicit capability complementarity. The upstream score provides a stronger reference at 67.29 Macro-F1, and the lightweight supervised classifier reaches 71.69. Qwen3.6-27B achieves the best overall result, with 74.51 Macro-F1, 52.25 MCC, and a comparatively balanced positive precision and recall of 59.99% and 84.71%. These results show that opportunity detection benefits from semantic reasoning beyond lexical similarity and manually defined matching signals, while the competitive supervised result indicates that task-specific decision patterns can also be learned effectively from labeled examples. Figure 7 complements the point estimates with instancebootstrap uncertainty. Qwen3.6-27B remains the strongest opportunity detector at 74.51 Macro-F1 (95% CI: 72.89– 76.07), followed by Qwen3.6-Max and Qwen3.7-Plus. These intervals quantify instance-level uncertainty, but they do not capture prompt sensitivity or run-to-run variation because each model was evaluated once under deterministic decoding.

E.2

Joint and Oracle Structured Prediction

Joint prediction. Table 6 evaluates whether opportunity existence, strength, collaboration type, and role direction are recovered simultaneously. Kimi-K2.7 achieves the highest positive four-field exact match of 61.98%, followed by Qwen3-Coder at 56.10%. In contrast, Qwen3-2B obtains an overall exact match of 51.98% but only 4.65% exact match on gold-positive instances because it detects only 26.96% of them. Overall exact match can therefore be dominated by correct rejection of negative pairs and should be interpreted together with positive exact match and positive coverage.

Qwen3.6-27B Qwen3.6-Max Qwen3.7-Plus Qwen3-Coder Qwen3.6-35B GLM-5.2 Kimi-K2.7 Qwen3-VL-8B GLM-5.1 Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

Local Hosted

0

20

40 60 Opportunity Macro-F1 (%)

80

100

Figure 7: Opportunity-detection Macro-F1 with 95% instance-bootstrap confidence intervals on the Chinese benchmark. Detection-conditioned analysis. Table 7 reports relationattribute performance conditional on correctly detecting a gold-positive pair. Coverage must be read together with the conditional scores: a high oracle score at low coverage characterizes only the subset already detected and is not a directly comparable end-to-end ranking. For most high-recall systems, the oracle–end-to-end gap is small, indicating that their remaining errors arise mainly when assigning strength, collaboration type, or role direction after an opportunity has already been detected. Missed detection is instead the dominant bottleneck for Qwen3-2B. Detailed type-confusion and direction-bias analyses are provided in Appendix F.

E.3

Complete Cross-Lingual Results

Table 8 provides the complete model-level statistics underlying the cross-lingual analysis. English input produces neither a universal gain nor a universal decline. Models with similar Chinese and English aggregate exact-match scores may nev-

Model

Overall

Pos. EM

95% CI

Cov.

Cond. Attr. EM

MCF

Qwen3-8B Gemma4-26B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B Deepseek-7B Qwen3-2B

35.47 34.47 22.25 20.71 14.69 4.92 12.98 51.98

47.62 47.11 44.11 43.49 30.89 14.15 13.12 4.65

[44.42, 50.72] [44.01, 50.31] [41.01, 47.21] [40.39, 46.59] [28.10, 33.78] [11.98, 16.43] [11.05, 15.29] [3.41, 5.99]

93.80 97.93 98.97 97.93 99.90 99.69 94.73 26.96

50.77 48.10 44.57 44.41 30.92 14.20 13.85 17.24

3.00 3.09 3.04 3.13 2.83 2.34 2.45 0.70

Kimi-K2.7 Qwen3-Coder Qwen3.6-Max Qwen3.7-Plus Qwen3.6-35B GLM-5.2 Qwen3.6-27B GLM-5.1 Qwen3-VL-8B

41.50 44.92 57.61 55.90 40.21 37.22 61.57 25.35 26.17

61.98 56.10 52.27 51.65 50.31 47.11 45.14 38.95 37.81

[58.78, 64.98] [53.00, 59.30] [49.17, 55.37] [48.45, 54.86] [47.11, 53.51] [43.90, 50.31] [42.05, 48.24] [35.95, 41.94] [34.81, 40.91]

97.42 89.98 89.46 91.84 97.62 97.21 84.71 98.86 96.80

63.63 62.34 58.43 56.24 51.53 48.46 53.29 39.39 39.06

3.38 3.10 3.01 3.05 3.13 3.08 2.73 2.94 2.85

Table 6: Joint structured prediction performance. Positive four-field exact match requires opportunity existence, strength, collaboration type, and role direction to be simultaneously correct on every gold-positive instance. Detection-conditioned threeattribute exact match is computed only for gold-positive instances predicted as opportunities and is reported together with positive coverage. Note. All values except MCF are percentages. Overall denotes four-field exact match over all instances. Pos. EM denotes four-field exact match on gold-positive instances. Cov. denotes positive coverage, i.e., the proportion of gold-positive instances predicted as opportunities. Cond. Attr. EM denotes detection-conditioned three-attribute exact match. MCF denotes mean correct fields.

ertheless show substantial asymmetry between the ZH-only and EN-only subsets. Moreover, paired agreement is consistently higher than the proportion of instances answered correctly in both languages, demonstrating that part of the apparent cross-language stability is attributable to the same incorrect tuple being repeated across the two input conditions.

E.4

Robustness to A/B Firm-Order Permutation

To test whether predictions depend on presentation order, we construct a direction-balanced variant of the Chinese test set. The variant retains all 2,805 firm pairs and preserves opportunity existence, opportunity strength, and collaboration type. We reverse the displayed order only for gold-positive pairs with a unidirectional relation. The original test set contains 157 A2B and 700 B2A instances. Using a fixed random seed, we swap 427 pairs, including 78 original A2B instances and 349 original B2A instances, yielding an approximately balanced distribution of 428 A2B and 429 B2A pairs. Under ideal equivariance, a swapped pair should preserve opportunity existence, strength, and type while reversing only A2B and B2A. Predictions for unswapped pairs should remain unchanged. Table 9 reports results for the nine models with predictions on both the original and balanced variants. Averaged across models, exact equivariance is 76.47% and direction equivariance is 85.90%. Thus, 23.53% of predictions fail to preserve the expected structured equivalence, and 18.91% change not only direction but also opportunity existence, strength, or collaboration type. Order sensitivity is therefore not confined to the directional label. A/B permutation does not systematically overturn previ-

Original correct -> Balanced wrong

Original wrong -> Balanced correct

+0.001

Qwen3.6-Max Qwen-VL-8B

+0.012

Qwen3.6-35B

+0.010 +0.003

Qwen3.7-Plus

+0.030

Qwen3.6-27B

+0.023

GLM-5.1

+0.044

Qwen3-Coder +0.019

Kimi-K2.7

+0.051

GLM-5.2 −0.06

−0.04

−0.02

0.00 0.02 Rate over all samples

0.04

0.06

0.08

Figure 8: Changes in four-field exact-match correctness after A/B firm-order permutation. Leftward and rightward bars denote Correct→Wrong and Wrong→Correct transitions, respectively; annotations report net correction rates.

ously correct predictions. As shown in Figure 8, an average of 2.46% of instances change from correct to incorrect, whereas 4.61% change from incorrect to correct, for a net improvement of 2.15 points. The balanced variant should therefore be interpreted as a stress test for order sensitivity rather than as evidence against the main conclusions drawn from the original benchmark.

E.5

Output Validity and Confidence Reliability

Table 10 jointly reports raw JSON validity, normalized parsing success, missing-field and invalid-label rates, cross-

Table 7: Oracle-detection conditional performance. Model Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

Coverage

Strength Macro-F1

Type Macro-F1

Direction Macro-F1

Oracle–E2E Gap

26.96 93.80 97.93 94.73 98.97 97.93 99.90 99.69

46.40 47.99 46.50 47.72 46.17 51.27 45.30 38.92

29.37 49.89 54.46 32.32 43.70 57.70 44.45 14.56

27.85 54.87 67.23 37.24 52.51 55.54 61.04 22.29

19.21 2.27 0.78 0.94 0.48 0.58 0.02 0.04

Qwen3.6-27B 84.71 48.68 57.23 62.43 5.52 Qwen3.6-Max 89.46 52.80 57.17 63.43 4.32 Qwen3.7-Plus 91.84 48.82 60.94 70.30 2.73 Qwen3-Coder 89.98 49.15 49.86 62.94 3.02 Qwen3.6-35B 97.62 48.62 55.63 64.35 0.88 GLM-5.2 97.21 50.71 55.61 60.76 0.99 Kimi-K2.7 97.42 53.81 61.71 68.85 1.11 Qwen3-VL-8B 96.80 42.56 41.15 55.56 0.93 GLM-5.1 98.86 42.59 57.15 64.11 0.51 Oracle includes only gold-positive instances correctly detected as Yes. Coverage is opportunity recall on gold-positive instances; the gap is the mean oracle minus E2E Macro-F1 across three attributes.

field constraint violations, and confidence-conditioned exact match. Most hosted models produce syntactically valid outputs, but comparable parsing success does not imply comparable structural consistency. Qwen3-Coder records the lowest violation rate at 0.21%, followed by Qwen32B (0.46%), Qwen3.7-Plus (1.03%), Qwen3.6-27B (1.43%), and Qwen3.6-Max (1.89%). Common errors include retaining non-empty relation attributes after predicting No and omitting required fields after predicting Yes. Self-reported confidence provides useful risk stratification for some models but is not a calibrated probability. Qwen3.627B, Qwen3.7-Plus, and Qwen3.6-Max show High–Low exact-match gaps of 67.05, 59.88, and 56.57 points, respectively, yet their High-confidence subsets still contain 31.86–39.64% errors. Interpretation also depends strongly on coverage: some models assign nearly all instances to the High-confidence subset, making the Low-confidence estimate unstable. Confidence should therefore be treated as a model-specific diagnostic and reported together with Highconfidence coverage.

F

Error Analyses

Evaluation alignment. All analyses in this section use the instance-level prediction table aligned to the 2,805 gold pair identifiers. Duplicate outputs, API failures, missing predictions, invalid labels, and cross-field violations are retained and handled explicitly. Cross-lingual statistics are computed only on firm pairs for which both Chinese and English predictions are available for the same model.

F.1

Error Decomposition

Figure 9 decomposes predictions on the 968 gold-positive pairs into complete tuple correctness, opportunity false negatives, and attribute errors conditional on a positive opportunity prediction. Kimi-K2.7 obtains the highest positive four-

DeepSeek-1.5B Qwen3-2B DeepSeek-7B Llama3.2-3B Gemma3-12B Qwen3-VL-8B GLM-5.1 Qwen3-4B Llama3.1-8B Qwen3.6-27B Gemma4-26B GLM-5.2 Qwen3-8B Qwen3.6-35B Qwen3.7-Plus Qwen3.6-Max Qwen3-Coder Kimi-K2.7

Complete success Opportunity false negative Attribute error

0

20

40

60

80

100

Figure 9: Prediction outcomes on the 968 gold-positive Chinese instances. Each prediction belongs to one of three mutually exclusive categories: (i) complete tuple correct; (ii) opportunity false negative, where the gold label is Yes but the model predicts No; or (iii) attribute error after a positive opportunity prediction, where the model predicts Yes but at least one of strength, collaboration type, or role direction is incorrect, missing, or invalid.

field exact match, recovering the complete tuple for 600 of 968 pairs (62.0%). It predicts the opportunity-existence label as No for 25 gold-positive pairs (2.6%). For another 343 pairs (35.4%), it correctly predicts Yes but returns at least one incorrect, missing, or invalid relation attribute. Qwen3.6-27B is the strongest opportunity detector in the main evaluation, yet its positive four-field exact match is 437/968 (45.1%): 148 gold-positive pairs (15.3%) are predicted as No, while 383 (39.6%) receive a positive opportunity prediction with at least one erroneous relation attribute. By contrast, Qwen3-2B predicts No for 707 gold-positive pairs (73.0%) and recov-

Table 8: Paired Chinese–English evaluation. Paired Agreement measures tuple identity rather than correctness. Model

ZH EM

EN EM

∆EM

Paired Agreement

Both Correct

ZH Only

EN Only

Qwen3-2B Gemma4-26B Gemma3-12B Qwen3-8B Qwen3-4B Llama3.1-8B DeepSeek-7B Llama3.2-3B

51.98 44.15 14.69 35.47 20.71 22.25 12.98 4.92

59.22 51.11 14.47 30.52 25.92 27.95 11.94 4.67

+7.24 +6.95 -0.21 -4.96 +5.20 +5.70 -1.03 -0.25

75.08 68.60 50.09 39.47 34.08 31.27 12.80 7.52

49.73 39.83 11.27 21.78 12.66 14.40 3.64 1.18

2.25 4.32 3.42 13.69 8.06 7.84 9.34 3.74

9.48 11.28 3.21 8.73 13.26 13.55 8.31 3.49

Qwen3.6-Max 57.61 62.21 +4.60 79.57 54.62 2.99 7.59 Qwen3.6-27B 61.57 60.14 -1.43 77.33 55.83 5.74 4.31 Qwen3.7-Plus 55.90 53.87 -2.03 75.90 49.66 6.24 4.21 Qwen3-Coder 44.92 44.21 -0.71 66.38 37.29 7.63 6.92 Qwen3.6-35B 40.21 51.55 +11.34 61.96 36.93 3.28 14.62 Kimi-K2.7 41.50 44.10 +2.60 61.68 35.58 5.92 8.52 GLM-5.2 37.22 44.10 +6.88 61.60 33.16 4.06 10.94 GLM-5.1 25.35 21.82 -3.53 55.12 18.18 7.17 3.64 Qwen3-VL-8B 26.17 31.34 +5.17 49.41 22.03 4.14 9.30 Scores are percentages; ∆EM = EN − ZH. Paired Agreement measures predicted-tuple identity across the Chinese and English versions and does not imply correctness. Complete pairs have N = 2,805, although usable pair counts may vary by model. Correctness denotes four-field exact match.

ers the complete tuple for only 45 (4.6%). Opportunity false negatives therefore dominate the errors of low-recall models, whereas high-recall systems are limited primarily by collaboration-type and role-direction errors after predicting that an opportunity exists.

F.2

Systematic Failure Modes

Relatedness mistaken for complementarity. Unsupported opportunity predictions are the most frequent error at the model–instance level. Among the 1,836 gold-negative pairs misclassified by at least one Chinese model, 1,773 (96.6%) are misclassified by three or more models. Moreover, 18,551 of the 23,621 false-positive decisions (78.5%) are assigned High confidence. At the pair level, both same-industry hard pairs (644/644) and random non-edges (626/627) attract at least one false positive. The recurrence across models indicates that broad industry proximity, shared terminology, or market adjacency often triggers a plausible collaboration narrative even when the supplied profiles do not identify a sufficiently specific product, technology, channel, capability, or capital flow. Opportunity false negatives on implicit complementarity. Across the 18,392 gold-positive model–instance decisions, models predict No in 1,903 cases (10.3%), producing opportunity false negatives. At the pair level, 804 of 968 positive pairs receive at least one false-negative prediction, and 209 receive false-negative predictions from at least three models. These errors are not confined to low-signal candidates: at least one model predicts No for 299/358 borderline pairs (83.5%) and 347/447 strong-signal pairs (77.6%). The pattern suggests that models do not consistently integrate evidence distributed across products, services, trading capabilities, channels, and business roles. When complementarity is expressed indirectly or with limited lexical overlap, the model

may assign the negative opportunity label even though the supplied profiles support a positive relation. Collaboration-type confusion. Type-related errors have three distinct forms. Across all gold-positive model–instance decisions, models predict No in 1,903 cases. In another 1,458 cases, they predict Yes but return an invalid or missing type, and in 4,495 cases, they predict Yes and return a valid type that differs from the gold label. Figure 10 isolates the third category. Technology Transfer is particularly unstable: among valid type predictions for Technology instances, only 6.5% are correct, while 48.9% are assigned Supply and Production and 23.4% R&D and Co-development. Distribution and R&D are also frequently mapped to Supply and Production (27.6% and 25.6%, respectively). These errors suggest that models often recognize a broad basis for collaboration but fail to distinguish the resource being exchanged: a delivered product or service, an existing technology, or a jointly developed capability. Role-direction bias and reversal. The gold direction distribution is highly imbalanced: 157 pairs are A2B, 700 B2A, and 111 Bidirectional. An always-B2A classifier therefore reaches 72.31% accuracy without performing pair-specific direction reasoning. Figure 11 shows that model biases are heterogeneous. Qwen3-8B assigns 79.9% of its valid direction predictions to B2A, compared with 72.3% in the gold data. In contrast, DeepSeek-7B predicts Bidirectional for 33.4% of valid outputs, nearly three times the gold share of 11.5%, while Llama3.2-3B assigns 69.1% to A2B and never predicts Bidirectional. Kimi-K2.7 attains the strongest direction Macro-F1 (67.71) and 77.38% accuracy, only 5.07 points above the majority-direction baseline.

Table 9: Prediction stability under A/B firm-order permutation. All rates are percentages. The balanced set contains the same 2,805 firm pairs as the original Chinese test set, with 427 gold-positive unidirectional pairs shown in reversed order. (a) Prediction equivariance Exact Eq. ↑

Direction Eq. ↑

Overturned ↓

Severe ↓

Swapped Pos. Overturned ↓

Qwen3.6-Max Qwen3-VL-8B Qwen3.6-35B Qwen3.7-Plus Qwen3.6-27B GLM-5.1 Qwen3-Coder Kimi-K2.7 GLM-5.2

91.37 88.63 86.35 81.21 81.89 51.98 80.25 62.82 63.74

94.44 93.05 89.77 91.27 89.55 73.23 84.99 77.79 79.04

8.63 11.37 13.65 18.79 18.11 48.02 19.75 37.18 36.26

6.74 8.09 8.48 16.93 16.33 38.86 13.80 31.12 29.84

26.00 39.81 34.19 28.10 36.07 41.92 22.25 25.53 38.17

Average

76.47 85.90 23.53 18.91 (b) Changes in four-field exact-match correctness

Model

Model Qwen3.6-Max Qwen3-VL-8B Qwen3.6-35B Qwen3.7-Plus Qwen3.6-27B GLM-5.1 Qwen3-Coder Kimi-K2.7 GLM-5.2

32.45

Correct→Wrong ↓

Wrong→Correct ↑

Net Correction ↑

Both Correct ↑

Both Wrong ↓

1.96 1.11 1.78 2.71 1.64 3.10 2.00 5.49 2.35

2.03 2.35 2.82 3.03 4.67 5.38 6.38 7.38 7.42

0.07 1.25 1.03 0.32 3.03 2.28 4.39 1.89 5.06

55.69 25.06 38.43 53.19 59.93 22.25 42.92 36.01 34.87

40.32 71.48 56.97 41.07 33.76 69.27 48.70 51.12 55.37

Average 2.46 4.61 2.15 40.93 52.01 Note. Exact equivariance requires the complete structured prediction to remain unchanged for unswapped pairs and to undergo only the expected A2B/B2A reversal for swapped pairs. Direction equivariance evaluates only the direction field. Severe denotes changes involving opportunity existence, opportunity strength, or collaboration type rather than direction alone. Correctness-transition rates and the Both Correct and Both Wrong columns are computed over all 2,805 pairs. Swapped Pos. Overturned is computed only over the 427 reversed gold-positive unidirectional pairs. Net Correction equals Wrong→Correct minus Correct→Wrong.

F.3

Cross-Lingual and Confidence-Related Failures

High cross-lingual agreement can reflect shared errors rather than robust reasoning. As shown in Table 12, the same incorrect tuple is produced in both languages for 21.5% of Qwen3.6-27B pairs, 25.0% of Qwen3.6-Max pairs, and 26.2% of Qwen3.7-Plus pairs. For Gemma3-12B, identical wrong tuples account for 38.8% of pairs, while a further 43.3% receive different incorrect tuples. These outcomes distinguish two failure modes: language-invariant reasoning shortcuts and language-sensitive decision changes. Neither pattern, by itself, identifies translation quality as the cause. Self-reported confidence is similarly informative only as a model-specific diagnostic. Qwen3.6-27B labels 90.20% of Chinese predictions as High confidence, yet 31.86% of that subset is incorrect under four-field exact match. The corresponding high-confidence error rates are 34.60% for Qwen3.6-Max and 39.64% for Qwen3.7-Plus. Thus, the binary confidence label separates risk for some models but remains far from a calibrated probability of correctness.

F.4

Representative Failure Cases

The cases in Table 13 illustrate three recurring boundaries. First, profile-level relatedness can support a plausible narrative without supplying the specific interface required by the gold label. Second, implicit opportunities may require combining a broad capability with a downstream business role, making them difficult to recover from lexical overlap alone. Third, semantically adjacent mechanisms remain difficult to separate even when the model correctly predicts the opportunity-existence label as Yes. In HSAI_00848, the representative prediction changes both the primary collaboration mechanism and the resource-flow direction. RNEG_01776 further shows that High confidence does not prevent a narrower type error: the direction is correct, but Supply and Production is replaced by Marketing and Distribution. These cases show that identifying broad collaboration opportunities is easier than reconstructing the complete relational structure.

Table 10: Output validity and self-reported confidence reliability on the Chinese benchmark. Output validity Model Qwen3-2B Qwen3-8B Gemma4-26B DeepSeek-7B Llama3.1-8B Qwen3-4B Gemma3-12B Llama3.2-3B

Strict JSON ↑

Parse ↑

100.00 0.00 99.43 0.00 100.00 0.00 0.00 99.32

100.00 96.54 100.00 99.07 100.00 99.04 100.00 99.32

Confidence reliability

Missing ↓ Invalid ↓ Violation ↓ High Cov. EM@High EM@Low 0.46 24.21 2.64 10.20 5.13 16.93 8.81 12.73

0.00 3.46 0.00 0.93 0.00 0.96 0.00 0.68

0.46 20.75 2.64 9.27 5.13 15.97 8.81 12.05

42.14 69.73 90.34 77.86 92.62 95.58 90.45 85.42

36.55 33.13 38.12 5.82 16.44 21.67 11.79 5.72

63.22 46.14 0.37 31.03 95.12 0.00 42.16 0.26

Gap

High Err. ↓

−26.67 −13.01 37.75 −25.22 −78.69 21.67 −30.38 5.46

63.45 66.87 61.88 94.18 83.56 78.33 88.21 94.28

Qwen3.6-27B 100.00 100.00 1.43 0.00 1.43 90.20 68.14 1.09 67.05 31.86 Qwen3.6-Max 99.89 99.89 1.99 0.11 1.89 86.24 65.40 8.83 56.57 34.60 Qwen3.7-Plus 100.00 100.00 1.03 0.00 1.03 92.55 60.36 0.48 59.88 39.64 Qwen3-Coder 100.00 100.00 0.21 0.00 0.21 98.86 44.32 96.88 −52.55 55.68 Qwen3.6-35B 99.96 99.96 13.12 0.04 13.08 97.54 40.94 11.76 29.17 59.06 GLM-5.2 100.00 100.00 8.13 0.00 8.13 60.78 56.72 7.00 49.72 43.28 Kimi-K2.7 99.96 99.96 2.71 0.04 2.67 83.71 49.23 1.75 47.48 50.77 Qwen3-VL-8B 98.97 98.97 16.19 1.03 15.15 98.93 26.41 100.00 −73.59 73.59 GLM-5.1 100.00 100.00 15.29 0.00 15.29 28.27 51.45 15.06 36.39 48.55 Correctness is four-field exact match. Gap is EM@High minus EM@Low. Strict JSON is measured on untouched raw output; Parse is measured after deterministic normalization. Missing, Invalid, and Violation are error rates over all predictions. “–" indicates zero coverage or unavailable conditioning. Confidence is a two-level self-report, not a calibrated probability.

Supply

7705

1211

167

640

7

Distribution

451

1034

13

53

16

70

gold_type

50 Technology

588

243

83

275

7

R&D

376

91

33

950

1

40 30

Row-normalized share (%)

60

20 Capital

55

Supply

24

5

Distribution Technology pred_type

10

R&D

115

10

Capital

Figure 10: Collaboration-type confusion aggregated over 17 models on gold-positive Chinese instances for which the model predicts Yes and returns a valid collaboration type. Cell annotations report aggregated prediction counts, while cell colors encode row-normalized percentages. Goldpositive instances predicted as No, together with predictions containing invalid or missing type outputs, are excluded from the matrix.

Qwen3.7-Plus Qwen3.6-Max Qwen3.6-35B Qwen3.6-27B Qwen3-VL-8B Qwen3-Coder Qwen3-8B Qwen3-4B Qwen3-2B Llama3.2-3B Llama3.1-8B Kimi-K2.7 Gemma4-26B Gemma3-12B GLM-5.2 GLM-5.1 DeepSeek-7B DeepSeek-1.5B Gold distribution

A2B B2A Bidirectional

0

20

40

60

80

100

Figure 11: Distribution of valid role-direction predictions on gold-positive Chinese instances for which the model predicts Yes. Models are ordered by their predicted B2A share, and the gold distribution is included as a reference. Goldpositive instances predicted as No and positive predictions with invalid or missing directions are excluded from the normalization.

Table 11: Aggregate failure modes on the Chinese benchmark. Counts are model–instance decisions aggregated over 17 models. Denominators include all eligible decisions for the corresponding task. Failure mode

Operational criterion

Frequency

Main pattern

Unsupported opportunity Gold opportunity is No, but the model 23,621/34,903 (67.7%) prediction predicts Yes. Opportunity false negative

Gold opportunity is Yes, but the model 1,903/18,392 (10.3%) predicts No.

Valid but incorrect type

The gold label is Yes, the model predicts 4,495/18,392 (24.4%) Yes, and the returned valid type differs from the gold type. Valid but incorrect direc- The gold label is Yes, the model predicts 5,092/18,392 (27.7%) tion Yes, and the returned valid direction differs from the gold direction.

Broad industrial or lexical relatedness is frequently treated as sufficient evidence for a concrete collaboration interface. Cross-field evidence for complementarity is sometimes insufficiently integrated, leading the model to assign the negative opportunity label. Predictions collapse toward frequent or semantically adjacent mechanisms, especially Supply and Production. Errors reflect both majority-class attraction and model-specific overprediction of A2B or Bidirectional relations.

Table 12: Representative paired Chinese–English outcomes. Values are percentages of usable paired instances. “Identical wrong” denotes the same incorrect four-field tuple in both languages. Model Qwen3.6-27B Qwen3.6-Max Qwen3.7-Plus Kimi-K2.7 Gemma3-12B Llama3.2-3B

N

Both correct

Identical wrong

Different wrong

ZH only

EN only

2,805 2,805 2,805 2,805 2,805 2,805

55.8 54.6 49.7 35.6 11.3 1.2

21.5 25.0 26.2 26.1 38.8 6.3

12.6 9.8 13.7 23.9 43.3 85.2

5.7 3.0 6.2 5.9 3.4 3.7

4.3 7.6 4.2 8.5 3.2 3.5

Table 13: Representative diagnostic cases. Evidence summaries are derived only from the supplied firm profiles. Pair

Pattern

Key profile evidence

Gold / representative Diagnostic interpretation prediction

op- A: telecommunications services, No / Yes–Strong– Broad ICT overlap is treated broadband, IoT, cloud, and data cen- Supply–B2A as a specific supply interface, ters. B: ICT infrastructure, network although neither profile states equipment, enterprise solutions, cloud, a concrete demand relationand smart devices. ship. MLAI_01557 Opportunity false A: retail of furniture, home- Yes–Strong– The model must connect a negative improvement materials, appliances, Distribution–B2A / broad trading capability to bathroom products, and lighting. B: No a downstream retail channel despite limited direct product broad import–export trade in agricultural products, machinery, textiles, and overlap. other goods. Commercial-channel cues HSAI_00848 Type and direction A: media distribution, advertising, Yes–Strong– confusion brand promotion, and digital market- Technology–B2A obscure both the technical ing. B: digital-marketing platforms, ad- / Yes–Strong– mechanism and the annotated resource-flow direction. delivery systems, analytics, and related Distribution– technical services. Bidirectional RNEG_01776 High-confidence A: property sales, leasing, and consult- Yes–Strong–Supply– The direction is preserved, type error ing. B: real-estate development, invest- A2B / Yes–Strong– but service provision is conment, asset management, and property Distribution–A2B flated with channel distribumanagement. tion under High confidence.

HSAI_00165 Unsupported portunity

Record · ID 919450 · SHA-256 b899c06b7f70b91b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.