Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision– Language Models Jaeyoung Kim
MADI
Eunseok Kim
arXiv:2607.05268v1 [cs.CV] 6 Jul 2026
MADI
Dongsuk Jang
MADI
Abstract Whether a hyperbolic representation model uses its geometry cannot be read off its curvature √ parameter. What matters is the dimensionless operating point cρ at which its embeddings sit, and whether the radial and cone machinery is active there. We develop a battery of necessary-condition diagnostics and apply it to three published hyperbolic vision–language families—MERU, HyCoCLIP, and PHyCLIP—across released checkpoints and controlled interventions on a fixed current-GRIT snapshot. The audit identifies three failure modes. First, curvature is not an active resource in any audited configuration. The operating point √ stays near-Euclidean (H(u) ≈ 1; no evaluated converged checkpoint reaches cρ > 1). Releasing the curvature floor moves scalar curvature and norms but keeps the operating point near-Euclidean, with no substantial downstream degradation across seeds. Second, the cone and traversal machinery is measured inoperative. Entailment cones are inactive, saturated, or misaligned, and graded traversal fails under controlled readouts across all families. Directed radial depth is a bounded non-detection: external parent–child norm ordering shows no signal above shuffle-null controls at quantified sensitivity. The only surviving signal, a small directed residual on the models’ native box→caption relation, remains non-operative. Third, hierarchy-looking evaluations are underdetermined. Taxonomy correlations are carried by angular distance, and coarse-retrieval gains track box/compositional supervision, not curvature. A mechanistic account explains why: the entailment objective admits a lowcurvature, √ wide-cone shortcut. A parameter-free aperture identity (cones saturate exactly when cρ ≤ 2K) locates the edge at which every entailment-trained unclamped run—across families and seeds—settles. Entailment-off runs show no arrest there and are still contracting at every probed horizon. The shortcut is thus the dominant accelerator of collapse rather than its sole cause: contrastive/alignment training alone also fails to hold curvature high. We conclude that these formulations, as released, do not instantiate the radial/cone mechanism their geometry motivates. We distill the battery into a five-number geometry report that future hierarchy claims can adopt.
1
Introduction
Vision-language models are increasingly expected to reason not only about visual similarity, but also about abstraction: an image of a “golden retriever” is also an image of a “dog”, an “animal”, and an “entity”. This We used Claude (Anthropic) as a general-purpose assistive tool for code debugging, numerical cross-checking, organization of results, and manuscript polishing. All research questions, experimental design choices, statistical procedures, interpretations, and scientific claims are the authors’ responsibility.
1
has motivated a growing line of hyperbolic vision-language models, including MERU (Desai et al., 2023), HyCoCLIP (Pal et al., 2025), and PHyCLIP (Yoshikawa & Matsubara, 2025), which replace Euclidean CLIPstyle representations (Radford et al., 2021) with negatively curved spaces. The motivation is compelling: hyperbolic spaces have exponential volume growth and are known to embed tree-like structures with low distortion (Bridson & Haefliger, 1999; Sarkar, 2011; Nickel & Kiela, 2017; Ganea et al., 2018). In this view, general concepts should lie closer to the origin, specific concepts should move toward the boundary, and entailment cones should encode asymmetric abstraction relations. But do current hyperbolic vision-language models actually instantiate this hierarchy mechanism? This question is more subtle than asking whether hyperbolic VLMs perform well on standard downstream benchmarks. A model can improve retrieval or classification through better contrastive alignment, compositional supervision, calibration, or regularization, without using nonlocal hyperbolic geometry. Conversely, a model can exhibit taxonomy-like semantic similarity in its angular structure without encoding directed radial hierarchy. Existing evaluations often conflate these possibilities: leaf-level classification does not test abstraction, symmetric taxonomy-distance correlations do not distinguish angular similarity from radial depth, and zero entailment violation can arise from saturated cones rather than learned hierarchy. We therefore conduct a structural audit of public and from-scratch hyperbolic VLMs. Our study covers three representative model families—MERU, HyCoCLIP, and PHyCLIP—including seven released checkpoints and matched current-GRIT interventions. We separate two forms of evidence. First, we analyze released checkpoints as fixed artifacts, asking whether the public models exhibit the geometry and hierarchy mechanisms claimed for them. Second, because GRIT is a URL-derived corpus whose accessible subset changes over time, we train matched baselines and interventions on a fixed current-GRIT snapshot; all causal comparisons are made only within this current snapshot. Our findings are fourfold. First, curvature is not a separately identifiable or performance-critical degree of √ freedom: across families and floor settings the operating point stays near-Euclidean ( cρ ≈ 0.2–0.3), and unclamping the curvature floor changes the parameterization—c and the norms—but not the operational geometry, while downstream metrics show no substantial degradation (Section 4). Second, the cone and traversal machinery is measured inoperative—apertures saturated or misaligned, graded traversal failing under controlled readouts—while directed radial depth is a bounded non-detection: external parent–child ordering shows no signal above shuffle-null controls at quantified sensitivity (Section 5). Third, a gradientlevel mechanism explains why natural interventions fail to install hierarchy: the entailment objective admits a low-curvature shortcut that widens cone apertures and suppresses violations without learning order. On the probed trajectories the entailment gradient pushes curvature down at every collapse-phase step (three seeds, both gradient-probed families), outweighing the contrastive gradient and any added depth signal. Removing the entailment term does not restore curvature—contrastive/alignment training alone also fails to hold it up—so the shortcut is the dominant accelerator of collapse rather than its sole cause (Section 7). Fourth, the evaluations usually read as evidence of hierarchy are underdetermined: the stronger taxonomy-distance correlation of hyperbolic checkpoints is carried by angular distance, and raw directed norm-ordering scores are confounded by marginal prompt statistics unless shuffle-controlled, so prior positive evidence should not be read as active radial or cone-based hierarchy without direct geometry diagnostics. These results do not imply that hyperbolic geometry is useless for vision-language learning. They show that the current evidence for active hyperbolic hierarchy in published formulations is insufficient. Our contribution is not a new hyperbolic VLM. Instead, we provide a diagnostic framework and a mechanistic explanation for why published formulations fail to activate the hierarchy mechanism their geometry motivates. Our contributions are fivefold: 1. We show that MERU, HyCoCLIP, and PHyCLIP checkpoints operate in a near-Euclidean regime, √ and that releasing the curvature floor preserves or lowers the image-side dimensionless radius cρ. 2. We show that standard downstream metrics are decoupled from the entailment/order mechanism: curvature can collapse, cone constraints can saturate, and norm ordering can disappear while retrieval, compositionality, and zero-shot metrics remain comparable under matched comparison. 2
3. We give a mechanistic explanation for curvature collapse: the entailment objective admits a lowcurvature shortcut—reducing curvature widens cone apertures and suppresses violations without learning order. A per-loss gradient decomposition (in the two gradient-probed families) shows the entailment term drives curvature down more strongly than any added depth signal can counteract. An entailment-off ablation shows this shortcut is the dominant accelerator rather than a necessary cause, since contrastive/alignment training alone also fails to hold curvature up (Section 7). 4. We provide direct hierarchy diagnostics—radial parent-child ordering, cone activity, traversal robustness, taxonomy angular/radial decomposition, and shuffle-controlled directed tests—and, aside from one small, non-operative exception (Section 5.1), find no evidence that the published formulations instantiate the active nonlocal hyperbolic hierarchy. 5. We identify design requirements for hyperbolic VLMs: curvature-identifying supervision, controlled radial representation learning, entailment objectives that do not reward low-curvature wide-cone shortcuts, and evaluation protocols distinguishing angular organization from radial hierarchy. Future methods may activate these mechanisms through adaptive entailment, norm regularization, or alternative hierarchy operators. Whether they succeed will turn on whether nonlocal curvature, radial depth, active cones, and operational traversal actually emerge.
2
Related Work
2.1
Hyperbolic representations and vision-language models
Hyperbolic geometry is a natural model for hierarchical data: its exponential volume growth embeds tree-like structures with low distortion (Sarkar, 2011), and Poincaré or Lorentz spaces have been used to embed symbolic taxonomies and partial orders (Nickel & Kiela, 2017; Ganea et al., 2018), typically associating hierarchy with a radial structure—general concepts near the origin, specific concepts near the boundary—and encoding asymmetric order through entailment-cone aperture and containment (Ganea et al., 2018). Hierarchy is not unique to hyperbolic geometry, however: order embeddings model entailment as an asymmetric partial order without negative curvature (Vendrov et al., 2015), and high-dimensional Euclidean embeddings represent WordNet-like trees competitively (Bansal & Benton, 2021). This motivates the distinction central to our work: apparent semantic hierarchy in a representation does not by itself imply that the model encodes it through nonlocal hyperbolic geometry. Building on these foundations, recent vision-language models encode abstraction in image-text representations. MERU introduces hyperbolic contrastive learning and motivates the radial coordinate as an abstraction axis (Desai et al., 2023); HyCoCLIP adds box-level compositional supervision and intra-/inter-modal entailment losses (Pal et al., 2025); and PHyCLIP factorizes the representation into product hyperbolic components for taxonomic and compositional structure (Yoshikawa & Matsubara, 2025). These models differ in architecture and supervision but share one geometric hypothesis: curvature and entailment cones should induce directed semantic hierarchy. We audit this hypothesis directly—asking not whether they improve benchmarks, but whether the released checkpoints and matched interventions exhibit the advertised mechanism. 2.2
Task-level analyses of hyperbolic VLMs
Ibrahimi et al. (2024) analyze released hyperbolic CLIP/MERU-style checkpoints and find improvements over Euclidean CLIP on spatial awareness, ambiguity resolution, out-of-distribution detection, and taxonomydistance correlation. However, such task-level differences do not by themselves identify the operating geometric mechanism. In particular, symmetric taxonomy-distance correlations can be driven by angular semantic organization rather than radial hierarchy. We return to this in Section 6.1, where we decompose the taxonomy signal and find it carried almost entirely by angular distance rather than radial hierarchy. 3
2.3
Euclidean and hierarchy-aware alternatives
A parallel line of work studies hierarchy in vision-language models without relying on hyperbolic geometry. Order embeddings model image-caption entailment as a partial order (Vendrov et al., 2015), and hierarchy-aware CLIP variants introduce explicit label or taxonomy supervision in Euclidean representation spaces (Geng et al., 2023). EuCLIP (Chou & Alam, 2024) further observes that Euclidean CLIP variants can match or outperform hyperbolic alternatives and reports that the learned curvature consistently collapses to its minimum clamp value in hyperbolic VLMs. Our work is complementary but distinct: whereas EuCLIP observes the collapse, we identify the gradient-level mechanism behind it—a low-curvature, wide-cone shortcut in the entailment objective—and provide necessary-condition diagnostics that test whether curvature is a separately identifiable or performance-critical degree of freedom. 2.4
Angular objectives and concurrent stabilization methods
Other recent works modify or stabilize the hyperbolic objective itself. Ramasinghe et al. (2024) propose an angle-based objective (Accept the Modality Gap, ATMG) that preserves a cross-modal modality gap rather than forcing image and text embeddings close in hyperbolic distance, motivated by the concern that geodesic proximity alignment may disrupt latent hierarchical structure. Angular stabilization is not equivalent to radial hierarchy, however: a model can improve angular alignment without establishing nonlocal curvature, parent-child radial ordering, active cone hierarchy, or monotonic traversal. Norm control in hyperbolic space is a recognized training concern—Guo et al. (2022) clip large embedding norms to avoid vanishing gradients √ at large radii—complementary to our finding that the trained operating point instead sits at small cρ. Most directly related is the concurrent work ARGENT (Huynh et al., 2026), which independently identifies a related entailment-cone instability: because cone aperture is inversely coupled to the parent norm, a model can minimize parent norms until the aperture degenerates into a half-space, collapsing the intended hierarchy. ARGENT responds with a method—an adaptive entailment loss that removes the norm-to-cone coupling, paired with a norm regularizer—and reports gains over a HyCoCLIP baseline. Our work is complementary along two axes. First, the√two analyses reach the same aperture degeneracy by different routes: √ the halfaperture ω ∝ arcsin 2C/( κ ∥ỹ∥) (ARGENT’s curvature κ is our c) depends on the product κ ∥ỹ∥, and either factor can drive it to π/2; ARGENT acts on the parent norm, whereas we isolate the gradient-level mechanism that drives the curvature factor down (Section 7). Second, our contribution is a set of necessarycondition diagnostics rather than a stabilized model. Since ARGENT’s checkpoints and code were not available at audit time, we do not audit it directly: whether a stabilized entailment loss activates nonlocal hyperbolic geometry or merely reorganizes angular structure within a near-Euclidean regime remains open.
3
Audit Setup and Diagnostics
Our goal is not to propose a new hyperbolic vision-language model, but to audit whether existing hyperbolic VLM formulations instantiate the geometric mechanism that motivates them. We therefore define a set of necessary-condition diagnostics for active hyperbolic hierarchy. These diagnostics are not intended to be a complete benchmark for every possible notion of hierarchy. Rather, they test the specific radial and cone-based mechanism that motivates current hyperbolic VLMs: nonlocal curvature should be used, general concepts should lie closer to the origin than specific concepts, entailment cones should encode asymmetric order, and the representation should support operational movement along hierarchy. 3.1
Released checkpoints and current-GRIT interventions
We separate two sources of evidence throughout the paper. First, we audit released checkpoints as fixed artifacts. This includes public MERU, HyCoCLIP, and PHyCLIP checkpoints. These analyses ask whether the models released by prior work exhibit these geometry and hierarchy mechanisms under our diagnostics. They do not require us to reconstruct the original pretraining corpus. Second, we perform controlled interventions on models trained from scratch on a fixed snapshot of the Grounded Image–Text Pairs (GRIT) dataset (Peng et al., 2024). GRIT is distributed as URL-referenced web 4
data, and its accessible subset shrinks over time as source links expire and some samples fail to decode. Documented to contain 20.5M pairs with 35.9M box annotations, the snapshot we obtained (crawled 2026-04-13) comprises 2,051 shards holding 13.1M image–text pairs and 25.0M parent-box annotations (exact counts in Appendix C, Table 16). We train on all shards for ≈29 effective epochs (500K iterations × batch 768 = 384M sample exposures). Our claims rest on matched within-snapshot interventions rather than on reproducing a particular training scale: baseline and variant models are trained on the same current-GRIT snapshot with the same optimizer, batch size, training budget (collapse-probe horizons excepted; Appendix C), data order, and hyperparameters (Table 17) except for the modified variable. Released checkpoints are shown only as historical references; all causal comparisons are made within the same current-GRIT snapshot. 3.2
Geometry diagnostics
Curvature alone does not determine whether a representation uses hyperbolic geometry. In Poincaré or Lorentz models, the deviation from local Euclidean behavior depends on the dimensionless product √ (1) u = c ρ, √ where c is the curvature parameter and ρ = ∥x∥ denotes the relevant radial coordinate. We write cρ throughout; all √ symbols are collected in Appendix A. We report not only learned curvature, but also radius ρ, the product cρ, and the local distortion factor H(u) =
sinh(u) . u
(2)
H(u) quantifies the radial nonlinearity of the hyperbolic geodesic relative to its Euclidean approximation at the tangent space: H(u) ≈ 1 corresponds to a locally Euclidean regime, while H(u) growing substantially above one signals that the hyperbolic-Euclidean discrepancy becomes order one. We do not interpret H(u) as a direct estimator of exponential-volume √ growth. We treat it as√a per-sample proxy for local distortion. Because H is a monotone function of u = cρ alone, H(u) and cρ are two views of the same quantity, not independent measurements. We report both as the natural unit for different parts of the analysis (the √ dimensionless radius for the operating point, H(u) for the resulting geodesic deviation). We use cρ > 1 as a simple marker for whether embeddings reach radii at which this nonlinearity becomes non-negligible. At u = 1 the geodesic deviates from its Euclidean baseline by H(1) − 1 ≈ 17.5%, and a 10% deviation is already reached by u ≈ 0.76 (Figure 1). The threshold is therefore a conservative marker for substantial nonlinearity, not a discontinuous transition. Because √ our reported ρ is the Lorentz √ spatial norm, the corresponding √ dimensionless geodesic radius is asinh( cρ). In the observed regime ( cρ up to ≈ 0.3) this differs from cρ by under 1.5% (asinh(0.3) ≈ 0.296), so the near-Euclidean conclusion is unchanged under either convention. We do not report Gromov δ-hyperbolicity: δ measures global tree-likeness of the whole metric, whereas our √ question is the local operating regime, for which cρ is the direct per-sample diagnostic. These diagnostics distinguish curvature collapse from geometry use. A model may have small curvature and small radius, resulting in locally Euclidean behavior. It√may have small curvature but larger radius, √ preserving the same effective geometry through the product cρ. Or it may enter a regime where cρ is large and H(u) deviates substantially from one, in which case the radial nonlinearity of hyperbolic geometry is functionally engaged. Our claims concern this effective geometry, not curvature as an isolated scalar. 3.3
Hierarchy diagnostics
We evaluate three hierarchy notions: lexical taxonomy, visual-semantic granularity, and model-native entailment/order. Radial ordering. Hyperbolic hierarchy is commonly motivated by a radial interpretation: general concepts should lie closer to the origin and specific concepts farther away. For a directed pair (p, c), where p is a parent and c a child, radial consistency is ( 1, ρc > ρp , RadialOrder(p, c) = (3) 0, otherwise, 5
√ Figure 1: Local distortion factor H(u) = sinh(u)/u as a function of the dimensionless radius u = cρ. The factor stays near one in the locally-Euclidean regime (shaded; deviation < 10%, exited only beyond u ≈ 0.76) and rises away from one as u → 1, where the geodesic deviates √ from its Euclidean baseline by H(1) − 1 ≈ 17.5% (dashed line). All audited checkpoints operate at cρ ≈ 0.2–0.3 (dotted line at the representative 0.2; Section 4), well inside the near-Euclidean regime and far from where hyperbolic nonlinearity becomes non-negligible. where ρc , ρp are the child and parent radii and ties are counted as 0. We use this as a necessary-condition diagnostic for the radial-depth hypothesis, not as a universal hierarchy metric. Entailment cone activity. For models using entailment cones, low violation rate alone is not evidence of hierarchy. It may indicate that the model has learned valid order relations, but it may also arise from saturated apertures that make the constraint trivial. We therefore report aperture distributions, active violation rates, directed-pair containment, and, where possible, the gradient contribution of entailment terms (Section 7). When apertures saturate at π/2, we interpret zero violation as inactive or trivial entailment rather than as successful hierarchy learning. Traversal robustness. If a learned representation supports hierarchy, moving along the proposed radial or geodesic direction should produce monotonic semantic changes. We therefore evaluate traversal trajectories by retrieving nearest captions or labels along interpolation paths. We report monotonicity, perfect traversal rate, collapse rate, and diversity of terminal retrievals. Traversal is a model-native diagnostic: it asks whether the representation supports an operational path from general to specific concepts. Product-factor hierarchy. For PHyCLIP, our goal is not to evaluate factor disentanglement in general. We ask a narrower question: whether individual factors or the product geometry support directed radial ordering or monotonic hierarchical traversal. Factor-wise results are therefore interpreted as hierarchy diagnostics, not as a complete analysis of product-factor semantics. 3.4
Evaluation diagnostics
Several evaluations can appear hierarchy-sensitive while failing to test the hyperbolic mechanism itself. We therefore separate three sources of signal that hierarchy-looking metrics conflate—angular semantic organization, radial hierarchy, and pair-specific directed structure—and define one diagnostic for each. Angular/radial taxonomy decomposition. A common hierarchy-looking metric is the correlation between taxonomy distance and embedding distance over class labels. This metric is symmetric: it measures whether semantically related labels are close, but does not determine whether hierarchy is encoded radially. 6
To decompose the signal, we compare a cosine-only regression dtree (i, j) ∼ β0 + β1 dcos (i, j)
(4)
dtree (i, j) ∼ β0 + β1 dcos (i, j) + β2 |ρi − ρj | .
(5)
to a cosine-plus-norm regression
We report the incremental explanatory power 2 2 2 ∆Rnorm = Rcos +norm − Rcos .
(6)
If taxonomy correlation is radial, norm differences should add measurable explanatory power beyond cosine distance. Since pairwise distances are not independent, we use class-level bootstrap or permutation-based procedures rather than treating all class pairs as independent samples. Shuffle-controlled directed tests. Raw directed norm consistency can be confounded by marginal norm distributions or prompt-length effects. For example, if fine-class prompts have slightly larger norms than their CIFAR-defined superclass prompts, many parent-child inequalities may hold even for random pairings. We therefore compare real directed pairs against shuffle-null controls: ∆pair = Orderreal − Eshuffle [Ordershuffle ] .
(7)
A raw ordering score is not interpreted as hierarchy unless it exceeds the shuffle-null distribution. Where prompt length can affect norms, we use length-matched prompts or residualized norms as sensitivity checks. As a sanity check, Appendix F applies the same radial and shuffle-controlled diagnostics to a synthetic hyperbolic tree with planted radial and angular hierarchy; the diagnostics recover the planted structure and collapse under radius- or angle-shuffled controls. Downstream metrics as decoupling probes. Retrieval, zero-shot classification, compositional tests, and hierarchy-aware penalties are useful measures of representation quality. However, they do not by themselves establish active hyperbolic hierarchy. We therefore use downstream metrics primarily as decoupling probes: if standard metrics remain comparable or improve while curvature collapses, cones saturate, or radial ordering disappears, then those metrics do not require active hyperbolic hierarchy.
4
Curvature Is Not Separately Identified from Radial Scale
We first ask whether current hyperbolic VLMs operate in a nonlocal hyperbolic regime. The answer is negative across released checkpoints and controlled interventions. Curvature either binds to the published floor or decreases when the floor is released, while the effective geometry remains near-Euclidean. 4.1
Released checkpoints remain near-Euclidean
Across released MERU, HyCoCLIP, and PHyCLIP checkpoints, curvature is at the implementation floor or fixed low-curvature setting, consistent with the curvature collapse reported by Chou & Alam (2024). However, the more important observation is not the scalar value of c alone. The effective operating radius √ cρ remains small across modalities and model families, yielding local distortion factors close to one and zero samples in the nonlocal regime. That curvature and radial scale are interchangeable—the same geometry can be written with large c and small ρ or the reverse—is classical for constant-curvature spaces (Sala et al., 2018; Gu et al., 2019). What we √ add is the empirical finding that trained hyperbolic VLMs sit at a near-Euclidean operating point, so it is cρ, not the scalar c, that characterizes them. Table 1 summarizes the released-checkpoint geometry diagnostics. The full per-modality and per-factor table is provided in Appendix B. Across √ all released checkpoints, the median local distortion factor is near one, and no model enters the regime cρ > 1. This holds throughout the distribution, not only at the median: 7
Table 1: Released hyperbolic VLM checkpoints operate in a near-Euclidean regime. Representative aggregate statistics; full per-modality values in Table 8 (Appendix B). The median H(u) and range are per-sample (pooled), so the upper endpoints cannot be recomputed from the per-modality means of Table 8 (PHyCLIP’s range reflects its widest subspace). √ Model family Curvature floor Median H(u) Range of H(u) % cρ>1 MERU HyCoCLIP PHyCLIP
0.1 0.1 0.1 (per-subspace)
1.009 1.005 1.006
[1.005, 1.014] [1.002, 1.007] [1.003, 1.023]
0% 0% 0%
√ the 99th-percentile and maximum image-side cρ stay at or below 0.37 for every released checkpoint, and √ no sample reaches even the cρ ≈ 0.76 distortion threshold (Appendix B.2, Table 9). Thus, the released √ representations do not operate in the nonlocal regime ( cρ > 1) that motivates hyperbolic hierarchy. Together, these diagnostics limit the interpretation of curvature collapse. The issue is not merely that c is small or clamped: the learned embeddings also remain at small dimensionless radius, so the effective operating geometry is close to Euclidean. Notably, the from-scratch current-GRIT baselines converge to the √ same dimensionless radius as the released checkpoints ( cρ ≈ 0.20 for HyCoCLIP at both scales; Tables 1– 2), despite differing in initialization and data availability. The near-Euclidean operating point is thus a recurring empirical regime across these runs rather than an artifact of any single one. 4.2
Unclamping curvature changes parameterization, not geometry
To test whether the curvature floor is suppressing a useful geometric degree of freedom, we train matched current-GRIT baselines and curvature-unclamped variants. Here “baseline” refers to a model we train from scratch on the current-GRIT snapshot under the published settings. Baseline and clampOff differ only in the curvature clamp. In the clampOff variants, the lower curvature clamp is relaxed from 0.1 to 0.001, while all other training settings are kept fixed. Controlled interventions are run at ViT-S and ViT-B. ViT-L is audited only at the released-checkpoint level (Table 8). √ Table 2: Curvature unclamping changes c and radial norms but not the operating geometry: cρ stays in the near-Euclidean ≈ 0.2–0.3 band and H(u) near 1, so the nonlocal-regime fraction is 0% for all. ViT-B rows are mean ± s.d. over three seeds (0, 37, 42). ViT-S rows are single-seed (seed 0). Per-seed values are in Table 15; measured on 500 ImageNet (Deng et al., 2009; Russakovsky et al., 2015) and 500 COCO (Lin et al., 2014) val images (image-side). c
√
ρ baseline
clampOff
baseline
cρ clampOff
H(u)
Model
baseline
clampOff
baseline
clampOff
MERU-S MERU-B HyCoCLIP-S HyCoCLIP-B PHyCLIP-S PHyCLIP-B
0.1000 0.1000 0.1000 0.1000 0.1000 0.1000
0.0419 0.921 1.052 0.291 0.215 1.014 1.008 0.0289±.0007 0.939±.004 1.049±.003 0.297±.001 0.178±.002 1.015±.000 1.005±.000 0.0112 0.635 1.839 0.201 0.194 1.007 1.006 0.0093±.0003 0.638±.000 2.019±.030 0.202±.000 0.195±.000 1.007±.000 1.006±.000 0.0155 0.644 1.508 0.204 0.188 1.007 1.006 0.0144±.0001 0.650±.000 1.562±.008 0.206±.000 0.187±.000 1.007±.000 1.006±.000
Table 2 shows the central result. Releasing the curvature floor changes c and embedding norms, but does not move the models into a nonlocal hyperbolic regime. The effect on the dimensionless radius is familydependent. In HyCoCLIP, c decreases by roughly an order of magnitude and norms increase, preserving √ cρ to within about 3.5%. In MERU, the compensation is incomplete and unclamping moves the model to an even lower dimensionless radius. PHyCLIP falls between the two. In every case, however, the model stays well inside the near-Euclidean regime: across all families and sizes in these current-GRIT runs the local distortion factor remains at or below 1.015 and the nonlocal-regime fraction remains 0% (Figure 2). 8
√ Figure 2: Dimensionless radius cρ for current-GRIT baseline (filled circles) and clampOff (open diamonds) across the three families and both ViT sizes (image √ side; seed-0 values from Table 15). Arrows trace the baseline→clampOff shift. Every model sits at cρ ≈ 0.2–0.3—deep inside the locally-Euclidean region √ (shaded) and far from the threshold cρ = 1 (dashed)—so releasing the floor moves c and ρ but keeps the operating point near-Euclidean: HyCoCLIP barely moves (inset), MERU drops further, PHyCLIP between. Table 2 reports the corresponding seed means. 4.3
Downstream metrics remain comparable under curvature collapse
If curvature were a performance-critical resource, relaxing the floor should substantially alter downstream behavior. Instead, standard metrics remain comparable when the geometry moves to near-Euclidean. Table 3: Current-GRIT matched comparisons (clampOff minus baseline), averaged over three seeds (0, 37, 42) and reported as mean ± standard deviation. Retrieval deltas (COCO, Flickr; R@5) are percentage points. ZSC is zero-shot ImageNet top-1 accuracy (percentage points). SC is SugarCrepe overall accuracy in points (0–100). For the WordNet hierarchical metrics, TIE and LCA are distances (lower is better) while J, PH, RH are similarity/precision/recall scores on a 0–1 scale (higher is better). All hierarchical deltas are within a few thousandths to about a tenth (TIE is largest, at +0.084) and consistent in sign. Per-seed absolute values are in Tables 10, 11, and 12. Downstream (∆) Model
T2I COCO I2T COCO
MERU-B −0.04±0.48 HyCoCLIP-B +0.96±0.63 PHyCLIP-B +0.64±0.22
T2I Flk
I2T Flk
WordNet hierarchy (∆) ZSC IN
SC
TIE↓
LCA↓
J↑
PH↑
RH↑
−0.47±0.23 −0.05±0.17 +0.00±0.89 −0.46±0.24 +0.12±0.12 +0.049±0.039 +0.007±0.021 −0.0038±0.0025 −0.0016±0.0019 −0.0036±0.0024 +1.60±0.43 +0.72±0.40 +1.00±0.56 −0.35±0.29 +1.36±0.22 +0.084±0.027 +0.029±0.028 −0.0053±0.0016 −0.0033±0.0019 −0.0053±0.0012 +0.63±0.48 +0.43±0.35 +0.17±0.65 −0.29±0.61 +0.76±0.79 +0.049±0.030 +0.003±0.022 −0.0041±0.0016 −0.0019±0.0016 −0.0040±0.0009
Table 3 reports current-GRIT matched comparisons for MERU-B, HyCoCLIP-B, and PHyCLIP-B. The positive HyCoCLIP-B deltas are consistent across all three seeds, but we read these structural interventions as decoupling evidence—performance does not degrade when curvature collapses—not as a claim that clampOff is a better or significantly improved model. The decoupling is direct. For HyCoCLIP-B clampOff, curvature drops by an order of magnitude to c ≈ 0.009 and the representation becomes near-Euclidean, yet COCO and Flickr (Young et al., 2014) retrieval and SugarCrepe (Hsieh et al., 2023) move in the positive direction relative to baseline, while ImageNet zero-shot (−0.35±0.29) and VL-Checklist (Zhao et al., 2022) shift within the seed spread. (VL-Checklist is positive at seed 0 but seed-variable, with a three-seed mean ranging from −0.5 to −4.3 points across subtypes.) Four of the five WordNet (Miller, 1995) hierarchical metrics (TIE, J, PH, RH) shift by small but consistent-sign amounts (e.g. TIE +0.084±0.027); none shows substantial degradation across the three seeds. For MERU-B these metrics show only small shifts (≤ 0.5 points, some sign-consistent) across the three matched seeds. PHyCLIP-B likewise shows only small shifts under clampOff, with retrieval and SugarCrepe moving in the positive direction, the ImageNet metric within the seed spread, and the WordNet 9
hierarchy metrics again shifting by small, consistent-sign amounts (e.g. RH) of negligible magnitude. Across all three families, collapsing curvature toward the Euclidean regime does not cost substantial downstream performance in this matched setting: the WordNet hierarchy metrics carry a small, consistent-sign cost of negligible magnitude, while the core retrieval, zero-shot, and compositionality metrics are flat to positive. This conclusion is scoped√to the regime we can reach: both the baseline and the clampOff models already operate near-Euclidean ( cρ ≈ 0.2–0.3), so within that regime curvature is not separately performancecritical. We cannot probe a stable trained high-curvature regime in the audited final checkpoints: early √ training transients can enter cρ > 1, but every converged audited configuration contracts to the nearEuclidean band—and that contraction is itself one of our findings, not a gap in the comparison. Released checkpoints often obtain higher absolute downstream scores than current-GRIT baselines. We treat them as historical references only, since they may differ in data availability. Together, these results indicate that curvature √ is not identifiable separately from radial scale: across families, sizes, and floor settings, the operating point cρ stays in the same near-Euclidean range (≈ 0.2–0.3), whether √ the curvature change is absorbed by compensating √ norm growth (HyCoCLIP, which nearly preserves √cρ) or only partially compensated (MERU, where cρ falls further). In either case it is the combination cρ, not the scalar c, that governs the operating geometry (Tables 1–2). Curvature moves, but because the operating point either holds or drops rather than entering the nonlinear regime, the scalar c is not separately recoverable from downstream behavior. This identifies the first of two distinct failures. Curvature is not the active geometric resource in any audited configuration: no converged run remains in the nonlinear hyperbolic regime. This does not by itself imply that hierarchy is absent: radial depth ordering can in principle be expressed even in a near-Euclidean space (Vendrov et al., 2015; Bansal & Benton, 2021), so the two failures are logically separable—flatness does not entail the absence of order—even though, as Sections 7.3–7.4 argue, they may share related objective-level causes—the entailment shortcut and the absence of a curvatureidentifying contrastive/alignment signal. We therefore ask separately, in the next section, whether the hierarchy mechanism the geometry motivates appears anywhere in the representation—in radial ordering, cone activity, or traversal—regardless of curvature.
5
Direct Diagnostics Do Not Detect an Operative Hierarchy
We distinguish lexical hierarchy, model-native cone/order structure, and operational traversal. Across these diagnostics, no audited formulation shows evidence of an operative, graded directed radial or cone-based hierarchy. 5.1
External radial ordering is not detected; native structure is detectable but non-operative
Hyperbolic VLMs are motivated by the radial-depth hypothesis: general concepts should be closer to the origin and specific concepts farther away. We evaluate this hypothesis using directed parent–child pairs given by the CIFAR-100 (Krizhevsky et al., 2009) fine→coarse label hierarchy: each fine class is paired with its CIFAR-defined superclass, yielding 100 (parent=coarse, child=fine) pairs. WordNet is not used to construct these pairs. It enters only in Section 6.1 for the taxonomy-distance correlation. For each pair, we ask whether the child embedding has larger norm than the parent embedding, and we compare the resulting consistency against a shuffle-null distribution over randomly permuted pairings, following the criterion of Section 3.4. A raw norm-ordering rate can be driven by a marginal effect—all specific prompts being larger-normed than all general ones on average—which need not reflect any learned parent-child relation: it can arise from prompt length, level marginals, or a global specific-versus-general norm offset. The shuffle null removes exactly this marginal component, leaving the pair-specific question: do true parent-child pairings carry more radial order than exchangeable pairings with the same norm marginals? This pair-specific version matches the radial-hierarchy claim these models actually make—a child sits deeper than its own parent, not merely deeper than the average parent—so we take the pair-specific excess as the honest target throughout. 10
Figure 3: Radial parent–child consistency (ViT-B released) against each model’s 10,000-permutation shuffle null (±1σ, ±2σ bands; null mean as vertical tick; ±1.6σ detection threshold dashed). The three hyperbolic models lie within ±1σ of their nulls. The only crossing is the Euclidean CLIP-B baseline (z = +1.73)—a model that by construction encodes no radial hierarchy—so a crossing at this level reflects the near-unitnorm marginal artifact rather than radial depth, and it does not survive a multiple-comparison correction (Table 29). Raw consistencies span 22–61%, but each sits at its own shuffle null, not above it.
Across released checkpoints, no hyperbolic family’s parent-child norm consistency exceeds its shuffle null (Figure 3). MERU, HyCoCLIP, and PHyCLIP all fall within the non-detection band (|z| < 1.6) of their shuffle nulls. The only checkpoint to cross it is the Euclidean CLIP baseline (z = +1.73), which by construction encodes no radial hierarchy, so its crossing marks the noise floor rather than a radial signal. Nor does any fall significantly below its null (most negative: PHyCLIP, z = −0.67). By the shuffle-controlled criterion, none of the hyperbolic models uses the radial coordinate as a reliable semantic-depth axis. The same holds at the remaining released scales: every ViT-S checkpoint and both ViT-L hyperbolic checkpoints stay within the band, the largest excursion being PHyCLIP-L (z = −1.59). This non-detection is bounded, not a power artifact: a graded positive control matched to the diagnostic’s design (same 100 pairs and shuffle procedure) recovers a planted pair-specific radial ordering at 80% power once its shuffle-controlled excess reaches ∆pair ≈ 13 percentage points (Appendix F.1), and every audited model stays well below this sensitivity—the largest absolute shuffle-controlled gap among released checkpoints is about 5 pp (PHyCLIP-L), the other hyperbolic released models within about 2 pp, so even the Euclidean CLIP-B crossing is a sub-detection-magnitude artifact. Across the current-GRIT runs, individual seeds cross ±1.6 in either direction, but no model’s three-seed mean exceeds the threshold in the radial-depth direction, the positive single-seed excursions are seed-inconsistent at the per-seed false-positive rate (the one seed-consistent pattern, PHyCLIP-B baseline’s below-null z ≈ −2, points against radial hierarchy and does not survive correction), and—decisively—every such model fails the operative graded traversal (Section 5.3). Full results are in Table 29. This is not to say the models lack semantic structure. They may still encode angular semantic similarity. Sibling or related concepts can be close in angle. The point is narrower: on this external taxonomy the radial coordinate is not aligned with the parent-child abstraction direction, in any family, above chance (the native relation is treated separately just below, where a small detectable-but-non-operative residual does appear). A model-native relation test partially qualifies this result. The diagnostic above uses the CIFAR100 fine→coarse label hierarchy, an external taxonomy rather than the relation the box-supervised models were trained on. We therefore repeat the directed radial test on the native GRIT box→full caption structure (hereafter NG2). Each sample pairs a full-image caption (the more specific concept, predicted to have larger norm) with its box-level captions (the more general concept, predicted smaller), following HyCoCLIP’s compositional convention (Pal et al., 2025). We use the same shuffle-controlled estimand as the external test above, with a sample-level null that re-pairs each full caption with boxes from other samples, preserving the marginal norm distributions of fulls and boxes so that only pair-specific structure survives (n = 1000 11
pairs, R = 500 permutations). We control caption length (a length-matched subset with |∆tok| ≤ 3 and a length-residualized variant) and partial out cosine distance, as in Section 6.1. After these controls, angular distance remains the dominant predictor in every model, and a radial supplement 2 beyond cosine is at most small. On the released checkpoints it is negligible: ∆Rnorm ≤ 0.00040 across the released ViT-B checkpoints (PHyCLIP-B not distinguishable from zero, p = 0.41; the largest, MERUB, is 0.00040), consistent with the external-taxonomy null above. Under the clampOff checkpoints it is larger but seed-unstable—spanning 0–0.0023 for HyCoCLIP-B (Mantel-significant at two of three seeds, indistinguishable from zero at seed 42, p = 0.285) and 0.00016–0.00035 for PHyCLIP-B (significant at all three; Table 30)—whereas the directed one-bit signal is seed-robust (full-set z between +7.6 and +10.0), so what varies with seed is precisely the graded component an active radial axis would require. Even the largest supplement (0.0023, HyCoCLIP-B clampOff seed 37) is only 0.23% of total distance variance on the native relation. Three bounds scope this reading. The released-versus-clampOff difference is cross-sectional, not a matched intervention, so it cannot determine whether removing the floor reveals or rescales the supplement. The sample-level null preserves the full/box norm marginals, so a within-sample visual-specificity gap that lengthmatching does not capture could produce the norm ordering without any graded depth coordinate. And a significant directed z certifies only a single reliable bit—that a full caption tends to sit outside its own box (z ≈ +8.4–+9.7 for the box-trained models, +9.3 for MERU-B despite no box supervision, all far above Euclidean CLIP-B’s +3.0)—not the graded axis the mechanism claims: the cosine-controlled increment 2 ≈ 0 and strict traversal monotonicity is 0%. The negative result is thus refined rather than is ∆Rnorm overturned: no operative radial axis under either the external or the native test, and a small, statistically detectable native supplement that remains secondary to angular structure and below what an active radial hierarchy would require. We verify the detectable-versus-operative distinction directly on the NG2 relation under three retrieval readouts (Appendix E.5, Tables 31 and 32). Stepping outward along the radial coordinate with the angle held fixed returns an essentially constant caption—a methods fact rather than model evidence: this readout uses cosine retrieval, which is norm-invariant, so its ≈ 0 rank correlation cannot expose a radial axis in any model. A geodesic (ambient straight-line) readout does produce a positive correlation, but two controls show it is generic to straight-line interpolation rather than a hyperbolic-hierarchy signature. First, a mismatchedtarget variant stays well above the shuffle null. Second, the Euclidean CLIP baseline produces the largest geodesic correlation of all seven models (ρ = 0.71), exceeding every box-trained hyperbolic checkpoint— despite having no radial structure, since it fails the per-pair direction precheck at chance (distinct from the pooled z) and its cosine-controlled increment is negligible. The NG2 direction therefore carries a strong pairwise norm signal but no operative abstraction axis under any informative readout: positive-but-generic under the controlled geodesic readout, never strictly monotonic, and with a cosine-controlled radial increment of ≈ 0. Full per-model values are in Appendix E.5.
5.2
Entailment cones are inactive, saturated, or misaligned
Entailment cones are intended to encode asymmetric semantic order. However, the observed cone behavior is fragile. Under the default threshold, some intra-modal entailment terms are almost trivially satisfied and provide little effective gradient. The text-side cone is already saturated at π/2 at baseline in every family, so the text→image violation rate is low to begin with (0–0.4% for the current-GRIT HyCoCLIP baselines, 2.7–5.1% for MERU; 9–14% for PHyCLIP at baseline, where larger image norms place some embeddings just outside the saturated cone). Under clampOff the PHyCLIP rate collapses (≤ 0.2%) as embedding norms compress, while MERU’s stays in low single digits and HyCoCLIP’s at zero. Under clampOff the image-side aperture often saturates too—with MERU near the saturation edge as the exception discussed below—yet the image→text rate stays pinned near 100% because text embeddings still fall outside the (now π/2) cone. Either way the cone constraint is vacuous rather than satisfied, so a low or zero violation rate evidences cone trivialization, not learned semantic hierarchy. Conversely, when entailment thresholds are tightened, violation rates become high but the geometry and traversal diagnostics remain unchanged. 12
Table 4 summarizes representative cone behavior. In HyCoCLIP, the text-side cone is saturated already at baseline (text→image violation 0%) and stays saturated under clampOff, so the vacuous constraint is the baseline condition rather than a clampOff-induced drop. The image-side cone additionally saturates under clampOff (11% → 100% image-side), but the image→text violation stays at 100% as text embeddings continue to fall outside the cone. MERU shows that non-saturation fails just as completely: its image-side apertures remain geometrically non-saturated at baseline (0.73 rad), yet text embeddings fall outside the active image cone in virtually every sample (image→text violation rate near 100% on matched GRIT pairs). Under clampOff MERU sits right at the saturation edge—MERU-B saturates to π/2 at all three seeds, while MERU-S remains non-saturated (1.175 rad)—the knife-edge regime at the saturation edge (Section 7.3). The cone is thus geometrically active without the embeddings entering it—a saturated aperture and a nearempty one fail for opposite reasons but with the same outcome. Neither non-saturation nor zero violation is therefore sufficient evidence of operational hierarchy. We show in Section 7.3 that this is not incidental: the entailment objective itself admits a low-curvature solution that widens apertures and suppresses violations without learning semantic order. Low violation rates must therefore be interpreted jointly with aperture and gradient diagnostics. This pattern is family-universal. PHyCLIP shows the same image-cone signature as MERU—text embeddings never enter the active image cone (image→text violation ≈ 100% at both baseline and clampOff), regardless of saturation state. A second, distinct signal appears on the text-cone side: PHyCLIP imageside apertures saturate from 41–45% at baseline to ≈ 96% under clampOff, and in the same clampOff transition embedding norms compress so that the text–image angle distribution shrinks, driving the complementary text→image violation rate down per size (PHyCLIP-S 13.8% → 0.2%, PHyCLIP-B 9.5% → 0.05%; Appendix B.4). The image-side saturation and this text→image collapse are co-symptoms of the same curvature collapse rather than one causing the other—saturation acts on the image-cone (image→text) test while embedding compression acts on the text-cone (text→image) test—and both routes leave the cone unable to encode order, the aperture-saturation signature we trace to the low-curvature shortcut in Section 7.3, now visible across all three families (the per-loss gradient decomposition is probed directly for MERU and HyCoCLIP; for PHyCLIP we observe the aperture signature rather than decomposing its gradients).1 Across both saturation regimes the cone fails to encode learned order—the image-side constraint is non-informative (text never enters the image cone), the text-side rate reflects embedding compression rather than order. Full per-model diagnostics are in Appendix B.4.
Table 4: A saturated text-side cone makes the text→image constraint vacuous (low t→i violation) while the image→text rate stays near 100%. Neither evidences active order structure. Representative ViT-B models are shown, plus the MERU-S clampOff exception, with all apertures at seed 0. Cones are measured on the same 256-sample GRIT batch as Table 13. Model
Image aperture
Text aperture
Interpretation
MERU-B baseline
0.73 rad
π/2 saturated
MERU-B clampOff
π/2 saturated
π/2 saturated
MERU-S clampOff
1.175 rad
π/2 saturated
HyCoCLIP-B baseline
1.42 rad
π/2 saturated
HyCoCLIP-B clampOff
π/2 saturated
π/2 saturated
Image-side cones are geometrically non-saturated, but traversal still collapses. Saturated at all three seeds, settling just below the 2K edge (Section 7.3); no operational hierarchy. √ Operating point above the edge ( cρ = 0.217 > 2K, single seed); the non-saturated image cone is consistent with the edge criterion (Section 7.3). Cone constraints remain nontrivial but do not yield traversal hierarchy. Both saturate; t→i constraint vacuous (≈ 0), i→t at 100%—inactive either way.
1 Every PHyCLIP factor shows a nonzero image–text norm separation (Cohen’s d ≥ 0.2 in all 64 factors), so no factor is collapsed to a constant or unused. This establishes only that the factors carry some variance, not that the variance is hierarchy-relevant: a norm gap between modalities is a modality-separation effect, not a parent–child one. Consistent with this, PHyCLIP’s directed radial consistency remains at chance (42%, Section 5.1). The factors are not dead, but what they encode is not radial hierarchy.
13
5.3
Traversal collapses instead of moving along hierarchy
A stronger operational test is traversal. If radial or geodesic movement corresponds to abstraction, then traversing the learned space should produce monotonic semantic changes. Instead, traversal collapses across families, scales, and factors (summarized across interventions in Table 19). Across MERU and HyCoCLIP current-GRIT baselines, norm monotonicity is low, perfect monotonic traversal is never observed, and terminal retrievals collapse to a single or very small set of captions. For these external-pool walks, retrieval at each interpolation step uses the model’s own metric—Lorentz distance for the hyperbolic models, not the norm-invariant cosine (the native-relation readouts of Appendix E.5 use cosine)—and monotonicity denotes the per-step fraction of steps moving in the predicted specific→general direction (chance 50%). Steps on which the retrieved caption is unchanged—frequent under the hub collapse below—count against monotonicity, so the raw rate conflates wrong-direction moves with sticky no-moves. Perfect traversal is a trajectory-level statistic whose random baseline is far lower (≈2−K for K steps), so 0% perfect traversal is conservative against either reference. In current-GRIT interventions, the same pattern appears: curvature unclamping changes c and norms, but traversal remains collapsed. PHyCLIP requires a stronger check because its hierarchy may be distributed across product factors rather than visible only in the aggregate representation. We therefore evaluate each of the 64 factors separately. The result is uniformly negative: across PHyCLIP-B and PHyCLIP-L, and across ImageNet, COCO, and Flickr30k retrieval pools, no factor exceeds the 50% per-step reference and no factor achieves a perfect traversal (Table 21). The aggregate (full product-distance) representation fares no better, with monotonicity in the 8–19% range, like the per-factor mean (21.3–25.6%). Since sticky no-move steps depress these raw rates, we read them as corroborating rather than load-bearing. Thus neither the aggregate representation nor any individual factor supports monotonic hierarchical traversal, and the product factorization does not reveal hidden monotonic hierarchy at either level. Qualitative traversal examples can therefore mislead: early steps may retrieve paraphrases or visually related captions, but systematic traversal reveals a later collapse phase in which trajectories converge to a few hub captions. This is not COCO-specific—on HyCoCLIP’s own Flickr30k demo pool (Pal et al., 2025), 200 image→root traversals on the released HyCoCLIP-B reach only 13 distinct terminals, 79% funnelling into two degenerate hubs (Appendix D.3). Low monotonicity and terminal collapse are the per-step and terminal views of the same step-wise stickiness. Because that stickiness registers as ties, the evidence hierarchy is as follows. Load-bearing: the graded geodesic readout with its mismatched-target and Euclidean controls—on which the Euclidean CLIP baseline attains the largest correlation of all seven models (ρ = 0.71), so positive geodesic movement is generic to straight-line interpolation rather than hyperbolic hierarchy—together with the near-zero cosine-controlled radial increment (Tables 32, 30). Corroborating: the tie-confounded per-step monotonicity. Symptomatic and pool-dependent: terminal collapse (Appendix D.3). The per-step monotonicity is pool-independent (structured label and relation pools keep diverse terminals yet still fail it) and corroborates these. Under every readout, and independent of the retrieval pool, the trajectories do not realize an operational hierarchy from general to specific concepts.
5.4
Threshold changes affect regularization, not hierarchy
Finally, we test whether activating intra-modal entailment by changing the threshold η recovers hierarchy. Lowering η from 1.2 to 0.7 changes downstream metrics in a scale-dependent way, but leaves curvature at the floor, the local distortion factor near one, and traversal collapsed. We report the full intervention in Appendix C.3. Taken together, the radial, cone, and traversal diagnostics find no evidence of the radial/cone mechanism in the representation, and the natural threshold intervention does not install it. Yet prior work has reported hierarchy-like signals from these same models. The next section asks what those evaluations measure. 14
6
Hierarchy-Looking Evaluations Are Underdetermined
The previous section shows that direct hierarchy diagnostics fail. We now ask why prior evaluations can nevertheless suggest that hyperbolic VLMs are more hierarchical. We find that several hierarchy-looking metrics are underdetermined: they capture useful angular semantic organization or bulk norm effects, not directed radial hierarchy. 6.1
Taxonomy-distance correlation is angular rather than radial
Prior work reports that hyperbolic VLMs can show stronger correlations between taxonomy distance and embedding distance. We reproduce this qualitative positive signal on CIFAR-100. Hyperbolic checkpoints obtain higher WordNet path-distance correlations than Euclidean CLIP across all sizes. Multiple hyperbolic variants exceed r ≈ 0.45, with HyCoCLIP-B clampOff reaching the strongest correlation (r = 0.510 at seed 0) despite its curvature collapsing to c ≈ 0.009 (Table 28).
Table 5: CIFAR-100 taxonomy-distance correlation is carried by angular semantic distance. The taxonomy r is the Pearson correlation between embedding cosine distance and WordNet tree distance. The decomposition then asks whether radial norm differences add explanatory power beyond that angular organization. Norm differences add negligible explanatory power beyond cosine distance. Values use manually disambiguated WordNet synsets for 13 ambiguous CIFAR-100 classes (Table 26). Under the naive last-token heuristic, Pearson r decreases by 0.01–0.10 across models but the decomposition conclusion is unchanged (Table 27). HyCoCLIP-B baseline/clampOff rows are seed means over 0/37/42. The full model list is in Table 28. Model
Setting
Taxonomy r
2 Rcos
2 Rnorm-only
2 ∆Rnorm
pperm
CLIP-B MERU-B HyCoCLIP-B PHyCLIP-B PHyCLIP-L HyCoCLIP-B HyCoCLIP-B
released released released released released baseline clampOff
0.370 0.489 0.465 0.456 0.456 0.462 0.508
0.137 0.239 0.217 0.208 0.208 0.214 0.258
0.001 0.001 0.005 0.002 0.003 0.001 0.001
+0.003 +0.001 +0.001 +0.000 +0.001 +0.002 +0.001
0.25 0.34 0.50 0.74 0.50 0.44 0.44
However, a stronger correlation need not reflect active radial hierarchy. Table 5 decomposes the signal into 2 angular and radial components, reporting both the norm-only regression Rnorm-only and the incremental 2 2 2 ∆Rnorm = Rcos +norm − Rcos . Cosine distance explains the taxonomy correlation, while norm differences add no detectable explanatory power beyond cosine: the norm-only R2 is small in isolation (≤ 0.007) and the 2 incremental ∆Rnorm ≤ 0.003 across the decomposed models. Mantel permutation tests detect no significant norm contribution at α = 0.05 in any of the 44 model×mapping combinations we evaluate (minimum pperm = 0.125; full per-model r and Mantel p values are in Table 27). The decomposition is unchanged under either WordNet mapping—the manually disambiguated synsets used here, or the naive last-token 2 ≤ 0.003 throughout (Appendix E.1, Table 27). The heuristic: Pearson r shifts by 0.01–0.10 but ∆Rnorm decomposition runs over same-level leaf pairs, where radial depth differences are structurally small. It bounds the radial share of the leaf-level taxonomy signal, not depth coding in general. A sensitivity analysis bounds the null: the largest audited seed-mean increment (MERU-B baseline, 0.0029; per-seed values reach 0.0054, still 2.4× below threshold) sits at the Euclidean CLIP noise floor and more than four times below the ≈ 0.013 detectability threshold (Appendix F.1; Table 28), and the same analysis shows a norm signal that merely re-encodes the coarse partition is undetectable because cosine already captures it—norm–cosine redundancy, not an absence of norm structure per se. This reconciles prior positive evidence with our diagnostics. Hyperbolic VLMs can learn useful angular semantic structure, and this can improve symmetric taxonomy-distance correlation. But symmetric semantic distance is not the same as directed radial hierarchy. 15
6.2
Directed norm scores require shuffle controls
Raw directed norm-ordering scores can also mislead: a marginal norm offset—specific prompts being highernorm than general ones on average—inflates the apparent ordering even without any learned parent–child structure, which is why our radial diagnostic (Section 5.1) reads directed consistency only against a shuffle null rather than as a raw rate. Under that control, no audited model’s directed score survives above chance on the external CIFAR-100 fine→coarse hierarchy (Figure 3, Table 29). The only surviving signal is the small pair-specific residual on the models’ own native relation (Section 5.1). Published evidence of this form—e.g., the observation that box embeddings sit closer to the origin than full-image embeddings (Pal et al., 2025)—is exactly what the null control adjudicates: its pair-specific component is real (Table 30, z ≈ +9.7) but remains secondary to angular structure and non-operative as a hierarchy readout. The lesson generalizes beyond our own diagnostic: any directed-looking norm score must be reported against a null control. Without one, a model can appear hierarchy-like purely because of marginal prompt statistics rather than learned parent-child relations—a confound that, as Section 5.1 shows, accounts for essentially all the apparent directed signal on that external hierarchy in current hyperbolic VLMs. 6.3
Leaf-level benchmarks do not validate hyperbolic hierarchy
Many standard benchmarks used to validate VLMs are leaf-level discrimination tasks. ImageNet classification, for example, evaluates separation among leaf classes, not abstraction across levels. Hierarchy-aware penalties such as TIE or LCA incorporate taxonomy into scoring, but improved scores can be driven by angular semantic similarity among related leaves rather than by directed radial hierarchy. Following the decoupling logic of Section 3.4, these leaf-level and hierarchy-aware scores cannot by themselves establish active hyperbolic hierarchy: a model can improve on them while geometry remains near-Euclidean and our direct hierarchy diagnostics fail. 6.4
Multi-granularity retrieval separates supervision from geometry
We evaluate retrieval across WordNet abstraction depths using ancestor-defined query centroids—each formed by averaging the embeddings of an ancestor concept’s descendant leaf classes—against ImageNet validation images. This test is more hierarchy-sensitive than leaf-level classification: coarse queries should retrieve broad semantic groups, while fine queries should retrieve narrow leaf-level classes. The retrieval picture varies sharply with abstraction depth. At the coarsest levels, all models remain weak in absolute terms: for the root-like entity level (depth 1, only two qualifying queries), even the best AP is 0.128. At fine depths (d ≥ 11) the four families are nearly tied at matched scale (ViT-B: AP 0.57–0.60). At coarse depths (d ≤ 5), however, HyCoCLIP and PHyCLIP substantially outperform CLIP and MERU at every scale, reaching AP ≈ 0.33–0.34 at ViT-B compared to 0.18–0.19 for the non-box-supervised models. This coarse-depth advantage is not produced by active nonlocal hyperbolic geometry, on two independent grounds. First, hyperbolic geometry alone does not deliver it: the one hyperbolic-but-non-box model, MERU, stays close to Euclidean CLIP at coarse depths across ViT-S/B/L, while only the box/compositional HyCoCLIP and PHyCLIP improve substantially (Table 34; Appendix G). Second, the winning models sit at a near-Euclidean operating point and fail every activity diagnostic—external radial ordering, cone activity, and traversal—so we find no evidence that their hyperbolic mechanism contributes to the gain. The advantage thus reflects stronger angular semantic organization from compositional supervision, not radial depth. The models’ native box-caption radial residual is real but tiny and non-operative. We do not isolate box supervision from other pipeline differences. The gain tracks the box/compositionally supervised families. Across Sections 6.1–6.4, the same pattern recurs: each hierarchy-looking signal, once decomposed or controlled, reduces to angular semantic organization, marginal prompt statistics, or box/compositional supervision rather than active radial geometry. Sections 4 and 5 showed the radial/cone mechanism is not activated and goes undetected under our diagnostics. This section shows the positive evidence for it was underdetermined. What remains is the mechanistic question: given that curvature collapses and the cones fail to encode order, why do the natural interventions meant to install hierarchy fail to do so? Section 7 answers 16
this by tracing the gradients that move curvature and by ablating the entailment term to test whether the contrastive/alignment objective can stabilize a nonlocal operating point.
7
Why Simple Interventions Fail
The previous sections show that current models neither enter the nonlocal hyperbolic regime nor exhibit the radial/cone hierarchy under our diagnostics. We now ask the mechanistic question: why do the natural interventions meant to install hierarchy fail? We examine these routes in order of increasing directness. Two fail analytically: pairwise depth ranking has no direct curvature gradient (Section 7.1), and naive curvaturecoupled depth supervision opens a norm-growth escape route (Section 7.2). The central result (Section 7.3) is that the hierarchy-motivated entailment objective itself admits a low-curvature shortcut, so curvature collapse is better explained as a property of the loss than an optimization accident. A c-only diagnostic traces the dominant curvature-lowering pressure on the full-objective trajectory to the entailment term. An entailment-off ablation (Section 7.4) then shows the shortcut is not the only route to low curvature—the contrastive/alignment objective alone also fails to stabilize a nonlocal operating point—while activating the entailment term more strongly (Appendix C.3) likewise does not install hierarchy. 7.1
Pairwise depth ranking has no direct curvature gradient
A natural intervention is to add a pairwise depth-ranking loss encouraging deeper concepts to have larger radii. However, if this loss is defined only on embedding norms, it has no direct dependence on curvature: Ldepth = max(0, m − (ρc − ρp )).
(8)
Therefore ∂Ldepth /∂c = 0. In from-scratch training, this loss can change norms but cannot identify curvature. 7.2
Naive curvature coupling opens a norm-growth escape route
A second intervention is to couple depth supervision to the dimensionless radius, for example by replacing √ cρ. This creates a direct curvature path, but it also opens an uncontrolled norm-growth the score with √ path: cρ can be increased either by raising the curvature c or by growing the embedding norm ρ, so a model under such supervision can satisfy the depth objective by inflating norms rather than by adjusting curvature, leaving the curvature collapse unaddressed. This does not mean curvature-coupled objectives are impossible in principle. It shows that curvature-coupled hierarchy supervision must be paired with explicit radial-growth control. Without such control, the depth objective is satisfiable by norm growth alone. This is precisely the path our c-only diagnostic (Section 7.3) removes, by detaching the norm so that the depth gradient reaches curvature directly. 7.3
The entailment loss has a low-curvature shortcut
The first two interventions fail for fixable reasons—one lacks a curvature gradient, the other lacks norm control. The third reveals a deeper obstacle: even when a curvature gradient is present and norm growth is blocked, the entailment objective itself drives curvature down. This mechanism helps explain the negative results in Sections 4–5, and we isolate it directly in the probed checkpoints and batches. To separate curvature identification from radial norm growth, we run a c-only diagnostic: √ h = c stopgrad(ρ).
(9)
This preserves a direct gradient from the depth objective to curvature while blocking the depth loss from changing feature norms. The diagnostic is not a hierarchy-learning method. It tests whether a depth signal can identify curvature once the norm-growth path is removed. We stress that depth supervision is not part of the training objective of any audited model. The depth gradient is measured purely as a counterfactual. We report it only for HyCoCLIP, whose depth pairs come from box-level parent captions. MERU has no box-level supervision, so the same counterfactual is not comparable there. 17
√ The implementation was verified to be additive rather than ratio-based, so c does not cancel. Finite differences confirm a nonzero curvature gradient. Nevertheless, curvature still collapses to the floor. A per-loss gradient decomposition reveals why. In representative full-objective checkpoints and batches, the contrastive term pushes curvature upward while it carries signal, whereas the entailment term pushes curvature downward by a comparable or (usually) larger magnitude. We use this diagnostic as a mechanism probe, not a population-level estimate: absolute magnitudes vary across checkpoints and batches, but the entailment c-down sign is consistent throughout the collapse phase, and, on the full-objective trajectory, the contrastive c-up sign is reliable until the gradient vanishes to noise near the floor. The depth-loss gradient on curvature depends on whether active parent–child pairs already satisfy ρchild > ρparent , and is therefore batch-dependent. In our HyCoCLIP-S baseline c-only implementation check (Table 37, a deterministic eightdraw probe), the entailment term pushes curvature down in all eight draws while the depth term stays small and curvature-up, the entailment magnitude exceeding the depth by 89–4500×. Even with the norm-growth path removed, the depth signal is therefore far too weak to counteract the low-curvature shortcut. We measure the per-loss curvature-gradient decomposition directly on the ViT-B collapse-probe trajectories (Figure 4), across three seeds (42, 37, 23; 23 replaces the 0 used elsewhere, on a dedicated collapse-probe set) for HyCoCLIP-B and MERU-B. During the collapse phase—the first ∼7–8k steps, while c falls from 1.0 to the floor—the pattern is consistent across all six runs: the entailment gradient pushes curvature down at 100% of collapse-phase steps and exceeds the contrastive push in magnitude at ≥ 96.7% of steps, by roughly two to three orders of magnitude (median ∼150–850× across the runs; raw entailment 0.1–0.7 against a contrastive push of order 10−3 ). Once c reaches the floor both terms subside: the contrastive gradient decays to noise (∼10−4 –10−3 ) with no stable sign, while the entailment gradient stays weakly but consistently c-down—the residual pressure that, when the floor is relaxed (clampOff), drives c further down to its saturation-edge equilibrium. We do not report a post-floor entailment/contrastive ratio: with the contrastive term at noise it reflects a vanishing denominator rather than the mechanism. The depth counterfactual—the one loss term specific to HyCoCLIP—is overwhelmed at this scale too (Table 37): across the three seeds it sits ∼300× below the entailment gradient, and its sign tracks batch ordering (net c-up, but c-down on roughly a third of collapse-phase steps), so it cannot counteract the shortcut. The collapse-phase pattern holds for MERU-B as well, which has no box-level depth term, so the shortcut does not depend on depth supervision. Consistent with the self-limiting picture, c reaches the floor at step 7500 for HyCoCLIP-B and 7750 for MERU-B (representative seed 42, Figure 4), and the entailment c-down pressure has largely subsided—once curvature can fall no further, the gradient that drove it down has nowhere left to push. This explains why the full objective favors the low-curvature basin. Entailment cones have aperture 2K . ω(ρ) = arcsin min 1, √ cρ
(10)
As foreshadowed in Section 5.2, reducing c widens the aperture, which reduces violations in the training (text→image) direction—realized jointly with norm compression—without requiring semantic hierarchy to be learned. Thus the hierarchy-motivated entailment objective contains a low-curvature, wide-cone shortcut, and the gradient decomposition in Figure 4 confirms that during collapse the entailment term drives c downward by a magnitude that exceeds the contrastive gradient on this full-objective trajectory, with the c-down sign pattern preserved under objective weighting. This shortcut is self-limiting, which explains why, in the clampOff runs, curvature settles above the relaxed 0.001 clamp rather than running to it. The entailment c-down pressure is generated by violated pairs. Its magnitude is non-monotonic over training: it rises to a peak as curvature falls and violations are most active, then falls once apertures saturate and fewer pairs violate the cone constraint. During the collapse phase, on this full-objective trajectory, the measured contrastive gradient runs counter to this descent while it carries signal—with smaller or comparable magnitude early and much smaller magnitude later. Once c is pinned at the floor, the contrastive gradient is at noise level with no stable sign and should not be read as a persistent restoring force. The equilibrium is therefore consistent with a balance at the saturation edge: below it, apertures saturate and the entailment c-down pressure vanishes; above it, violations are active and entailment pushes c down. Whether this balance is a stable attractor—whether the c-down pressure just above the edge is steep enough to dominate the contrastive push—is a dynamical question we leave to future 18
Figure 4: Full-objective gradient attribution during curvature collapse, across families (ViT-B; representative seed 42, first 10k steps, where c moves). Top: learned curvature c falls from 1.0 to the clamp floor (0.1, marked); the shaded region is the collapse phase. Bottom: per-loss curvature gradient ∂(loss)/∂(log c) on a symlog axis, raw before objective weighting (positive = c-down). For both MERU-B and HyCoCLIP-B, the entailment term (red) pushes curvature down throughout the collapse phase (0.1–0.7), the contrastive term (blue) runs counter to it on this trajectory with smaller magnitude, and the c-only depth counterfactual (green, HyCoCLIP only) stays one to four orders of magnitude smaller. The objective weight λe = 0.2 scales the entailment curve down but preserves its c-down sign. Per-seed sign fractions are in Appendix H.2.
work. The analysis below establishes only where the saturation edge lies, and that curvature settles there rather than at the floor. √ The location of this boundary follows in closed form. The half-aperture is ω(ρ) = arcsin min{1, 2K/( c ρ)} with K the fixed constant in the aperture formula (K = 0.1 in all three audited implementations, including PHyCLIP’s per-factor cones; distinct from the√curvature floor that coincidentally also equals 0.1 at baseline), so a cone saturates (ω = π/2) exactly when c ρ ≤ 2K = 0.2. The clampOff equilibria land at this edge: with the floor relaxed to 0.001, HyCoCLIP-B settles √ at c ≈ 0.009 with ρ ≈ 2.06 (Table 14, measured on the same GRIT shards as the Table 4 cones), giving c ρ ≈ 0.195 < 0.2, i.e. apertures at saturation, matching the saturated cones √ of Section 5.2. The same holds across all audited clampOff cases: HyCoCLIP-B and HyCoCLIP-S (both cρ ≈ 0.195), MERU-B (0.178–0.182 across seeds), and PHyCLIP-S and PHyCLIP-B (0.188)2 all sit below 2K with median apertures observed at π/2 (image-side saturation 95.6% for PHyCLIPB, 95.5% for PHyCLIP-S, ≥ 99% for the others), as the aperture identity requires. MERU-S (0.217 > 0.2, single seed) sits above the edge and its image-side cone remains non-saturated √ (1.175 rad, Table√4), consistent with the identity. Because the half-aperture is a monotone function of cρ, that a run at cρ < 0.2 has 2 The MERU and HyCoCLIP √cρ values here are the GRIT measurements (Table 14), matching the GRIT shards of the
√ cones. PHyCLIP’s cρ = 0.188 is the per-factor GRIT measurement (0.1878/0.1881 for the two sizes). The ImageNet/COCO value (Table 2) is 0.187–0.188, and the two coincide to within 0.002, as do the ImageNet and GRIT √ values for HyCoCLIP and MERU, so the eval set does not affect the saturation classification, which depends only on whether cρ < 0.2.
19
saturated cones is an identity, not an independent prediction. The empirical, falsifiable content is where the equilibria land. That every audited clampOff run—across families with opposite baseline cone regimes, and including the non-saturated MERU-S case—settles at or near the 2K edge (from both sides) rather than at the relaxed 0.001 floor is strong evidence that the equilibrium is set by aperture saturation rather than by the clamp. Three further observations sharpen this. First, the entailment objective creates an equilibrium at the edge wherever the edge √ is reachable: already under the 0.1 floor at 500k steps, the HyCoCLIP-B and PHyCLIP-B baselines sit at cρ = 0.202 and 0.206. Second, MERU-B baseline (0.297) is the clamp-pinned case: the c = 0.1 floor halts descent before ρ reaches the edge, and releasing the floor lets it continue to ≈0.18—so the equilibrium is set by the edge when reachable and by the clamp when not. Third, √ without entailment no such arrest appears: all six λe =0 runs are still contracting at their probe horizons ( cρ = 0.33–0.35 at ∼11k steps, slopes −0.035 to −0.048 per 1k steps), and the extended seed-0 runs in both families pass through the edge at ≈21–22k and continue smoothly to 0.103–0.106 at 40k with no inflection or arrest at 0.2 (Section 7.4). The λe =0 horizons (11–40k) and the clampOff endpoints (500k) are not a matched-horizon comparison. The contrast is between arrest at the edge and the absence of any stationary√point, not between endpoint values. MERU-B clampOff is seed-consistent: across three seeds it settles at cρ = 0.176–0.180 < 0.2 with the image cone saturated, and all three seeds remain near-Euclidean (H(u) ≈ 1.005; per-seed geometry in Table 15). The informative fact is not that saturation tracks the identity—it must—but that every equilibrium lands within 0.024 of the edge itself. The collapse-phase gradients that drive c down are not in tension with this above-floor equilibrium: the saturation they produce is precisely what removes the entailment pressure as the equilibrium is reached. Tables 14 and 13 thus record the same fact in two coordinates: the curvature at which c stops places every run at or near the 2K edge—just below it for the saturating runs, just above for the non-saturating ones. The evidence for this shortcut is of three kinds, of√decreasing independence from training noise. First, it is analytic: the aperture relation ω(ρ) = arcsin 2K/( cρ) means that lowering curvature mechanically widens cones and suppresses violations, independent of any run. Second, it is an equilibrium-location match: across families with opposite baseline cone regimes—and seed-consistently in the MERU-B triplet—the clampOff equilibria settle at or near the 2K aperture-saturation edge rather than at the relaxed 0.001 floor, while the entailment-off runs show no arrest there, contracting through the edge with no stationary point in either family (Section 7.4), so the arrest at√the edge is specific to the entailment objective, not a per-run accident. (That saturation coincides with cρ ≤ 2K is an aperture-formula identity, not an independent measurement.) Third, it is an empirical gradient probe: the per-loss decomposition on the full-objective collapse-probe trajectories shows the entailment term pushing curvature down and the contrastive term up, consistently across checkpoints and for MERU-B (no depth term), replicating across all three seeds (Appendix H.2). The first two kinds of evidence carry the claim. The gradient signs corroborate the direction without being load-bearing for the conclusion. It is worth stating plainly how this mechanism connects to the released checkpoints, since the two are established by different means. The causal demonstration—intervening on the curvature floor and reading off the resulting aperture and gradient behavior—is exercised only on the from-scratch current-GRIT clampOff runs, because the released checkpoints are all pinned at the floor and cannot themselves be intervened on. What bridges the two is not a second causal experiment but two seed- and corpus-independent facts: the √ from-scratch baselines converge to the same dimensionless √ operating point as the released checkpoints ( cρ ≈ 0.2–0.3, Section 4.1), and the saturation criterion cρ ≤ 2K is parameter-free, so it applies to any checkpoint regardless of how it was trained. The released checkpoints occupy the same near-threshold, near-Euclidean operating regime, and the cone sides that the aperture identity implies are saturated are observed saturated (text-side cones uniformly at π/2; image-side cones sit at the edge, since released image √ cρ is at or just above 0.2). The clampOff runs let us watch the objective drive a model to that boundary. The mechanistic account of why curvature collapses is thus validated interventionally on the current-GRIT runs and connected to the released artifacts through the shared, parameter-free operating-point criterion rather than through a claim that the released checkpoints were themselves intervened on.
20
The depth term faces a difficulty in either direction. When the active-pair radial ordering is partially correct (ρchild > ρparent for the majority of active pairs), the depth gradient is c-up but small in magnitude (Table 37). When the radial ordering is wrong-signed or absent, the depth term flips to c-down and contributes nothing positive. In either regime, allowing radii to update through the depth term reopens the norm-growth path (Section 7.2). Successful objectives must therefore jointly learn radial ordering, identify curvature, and control radial growth—three requirements the audited objectives do not satisfy together. The shortcut identified here dominates curvature descent on the full objective, but the next section shows it is not the only route to low curvature: removing the entailment term entirely still fails to preserve a nonlocal operating point (Section 7.4)—entailment is the dominant accelerator, not a necessary cause. 7.4
Entailment-off training does not stabilize curvature
The gradient decomposition attributes the dominant curvature-lowering pressure to the entailment term on the full-objective trajectory. A natural question is whether that term is also necessary for collapse: if the cone loss is the cause, removing it should leave curvature high. We test this directly by setting the entailment weight to zero (λe =0) and retraining from scratch as otherwise-matched current-GRIT collapse probes (∼11k steps; seed 0 extended to 40k in both families; per-variant horizons in Appendix C), for both ViT-B families across three seeds (0/37/42). For MERU-B this leaves a contrastive-only objective; for HyCoCLIP-B it is an entailment-off objective retaining the box-level compositional and non-entailment alignment terms. Removing entailment does not stabilize curvature (Figure 5, Table 6). In every run curvature still collapses to the clamp floor and the operating point keeps contracting through the probe horizon (below). This is consistent with the full-objective gradient decomposition (Figure 4): the contrastive c-up gradient there carries signal only transiently and is not a persistent restoring force, so once it decays nothing holds curvature up. Entailment is not necessary for collapse in these settings. What it supplies is a substantial acceleration of it: with the cone loss active (λe =0.2) curvature reaches the floor roughly 2600–2800 steps earlier than without it (floor at ∼7.3–7.5k vs. ∼9.8–10.3k), consistent across families and seeds at matched clamp and protocol (not seed-paired—Table 6). Table 6: Curvature-collapse summary for the entailment-off ablation (λe =0) versus the full objective (λe =0.2), ViT-B, matched clamp floor (0.1); means over three seeds ({0, 37, 42} for λe =0 and {23, 37, 42} for λe =0.2; the comparison is same-clamp and same-protocol, not seed-paired). Curvature reaches the floor in every setting. Entailment only accelerates the descent. Floor step is the first iteration with c ≤ 0.105 (the stricter c ≤ 0.1001 cutoff of Table 36 shifts these by up to ∼250 steps). c@k is curvature at step k. Model
λe
floor step
c@3k
c@5k
c@7k
MERU-B MERU-B HyCoCLIP-B HyCoCLIP-B
0.2 0 0.2 0
7500 10233 7250 9833
0.570 0.761 0.575 0.766
0.237 0.423 0.230 0.422
0.115 0.237 0.107 0.230
The √ entailment-off runs also expose what the non-entailment objective does on its own. Early in training nonlocal regime (Figure 5, bottom; peak ≈ 1.6–1.7 at step ∼1.3–1.8k; √cρ transiently overshoots into the cρ > 1, both√families, all seeds)3 and then contracts as both c and ρ fall: at the ∼11k probe horizon all six runs sit at cρ = 0.33–0.35 with terminal slopes of −0.035 to −0.048 per 1k steps—still contracting, the endpoints transient rather than stationary. Extended to 40k (seed 0, both families)4 , the runs pass through the 2K=0.2 edge at ≈21–22k and decline smoothly to 0.106 (MERU-B) and 0.103 (HyCoCLIP-B) at 40k with no stationary point—the edge is a reference landmark only here, since without the cone term it has no dynamical role for these runs. Once c is pinned at the floor, this continued contraction is entirely ρ-side. The 3 Operating-point values are the image-side √cρ, computed from the logged mean radial norm (rho_img_mean) during training. 4 The HyCoCLIP-B extension resumes its seed-0 probe from the 11k checkpoint. The MERU-B extension is a fresh contiguous run. Both use a 40k cosine learning-rate horizon, whereas the ∼11k probes inherit the 500k base schedule. At matched steps the MERU-B extension reproduces the base-schedule seed-0 trajectory (logged to 38k) to within ∼1%, so the schedule choice does not affect the trajectory in this window.
21
Figure 5: Removing entailment does not stabilize curvature. Top: learned curvature c during training for the full objective (λe =0.2, solid) and the entailment-off ablation (λe =0, dashed), ViT-B, at matched clamp floor (0.1). In both families curvature collapses to the floor with or without entailment; entailment only reaches the floor ∼2600–2800 steps earlier (markers). Bottom: the image-side operating √ point cρ for the entailment-off runs (shown for λe =0 only; the √ λe =0.2 collapse-probe runs did not log the radial norm) transiently overshoots into the nonlocal regime ( cρ > 1, peak ≈1.6–1.7) and then contracts; beyond the ∼11k probe horizon (dotted vertical line) the thin dashed curves follow seed 0 alone, crossing the 2K=0.2 edge at ≈21–22k and still declining at 40k (0.103–0.106) with no arrest (Section 7.4). Curves are means over three seeds up to the ∼11k probe horizon (bands: min–max; seeds {0, 37, 42} for λe =0 and {23, 37, 42} for λe =0.2); same-clamp and same-protocol, not seed-paired.
objective thus reaches the nonlocal operating point its geometry is built for but does not hold it—unlike the entailment-on clampOff arrest at the edge (0.18–0.22 at 500k; Section 7.3): absent a curvature-identifying signal, contrastive alignment leaves curvature underidentified and the operating point contracts. EuCLIP’s sweep already includes entailment-free hyperbolic runs whose scalar curvature likewise ends at the clamp floor (Chou & Alam, 2024). The operating-point trajectory is what that endpoint conceals—the overshoot, the contraction without a stationary point, and the absence of any arrest at the 2K edge. Two caveats bound this picture. First, the curvature collapse is not a weight-decay artifact—the learned curvature is excluded from weight decay—and the λe =0 gradient probe shows a weak but predominantly downward contrastive curvature gradient through the collapse phase, so the descent is driven rather than a passive drift. The contrastive gradient’s sign is thus trajectory-dependent—c-up on the full objective (Figure 4), predominantly c-down once entailment is removed. Second, the accompanying radial-norm contraction is not cleanly attributable to √the objective: the shared recipe applies weight decay (0.2) to the encoder, which shrinks ρ—and hence cρ—after the warmup transient, so weight decay and the logit temperature remain unruled confounds for the norm side of the contraction. A weight-decay-reduced run 22
would isolate this. The ablation’s load-bearing conclusion—that the non-entailment objective does not hold curvature high or pin a nonlocal operating point—does not depend on that isolation. The two failures are therefore separable and asymmetric. Contrastive/alignment underidentification supplies only a weak, unstable pressure that fails to stabilize a nonlocal operating point. The entailment shortcut adds a much stronger aperture-driven descent that accelerates collapse. Entailment is the dominant accelerator on the full-objective trajectory, not the sole possible route to low curvature—which is why a future model must fix the contrastive/alignment objective (Section 7.5, item 2), not merely the entailment loss. 7.5
Design requirements for future hyperbolic VLMs
None of these failures shows that hyperbolic geometry is the wrong tool for vision-language learning. What they pin down is why current published formulations leave the radial/cone mechanism dormant, and what a future model would have to demonstrate to claim otherwise. A hyperbolic VLM claiming hierarchy should report at least the following diagnostics: 1. Nonlocal geometry activation: learned curvature, radial norms, tion factors, and the fraction of embeddings in the nonlocal regime.
√
cρ distributions, local distor-
2. Curvature identifiability: whether curvature gradients come from contrastive, entailment, or hierarchy losses, and whether curvature can change without being absorbed by radial norm scale. In particular, whether contrastive alignment stabilizes a nonlocal operating point rather than letting √ cρ contract toward the Euclidean basin (Section 7.4). 3. Active entailment: cone aperture distributions, active violation rates, containment of true directed pairs, evidence that zero violation is not due to saturated apertures, and whether the order objective can be satisfied without lowering c, shrinking parent radii, or saturating apertures. 4. Directed hierarchy: parent-child radial ordering evaluated against shuffle-null controls and prompt-length confounds. 5. Operational hierarchy: traversal monotonicity, collapse rate, and terminal retrieval diversity. 6. Evaluation decomposition: separation of angular semantic organization from radial hierarchy in taxonomy-distance or hierarchy-aware metrics. We distill this contract into a report format that authors and reviewers can adopt directly: The five-number geometry report. A hyperbolic model claiming radial/cone hierarchy should report, per modality (and per factor for product spaces): √ 1. Operating √ point: the cρ distribution √ (median, p95/p99, maximum, % above the 10%-distortion √ marker cρ > 0.76, and % above cρ > 1. Our main tables give the median and the cρ > 1 fraction, with the upper-tail quantiles in Appendix B.2). [here: Tables 1, 2] 2. Saturation state: cone half-aperture distribution and % saturated at √ π/2, with the operating cρ vs. 2K for aperture point’s position relative to the objective’s closed-form saturation edge ( √ arcsin 2K/( cρ) ). [Table 13] 3. Directed violations, both directions: active violation rates at η=1 for parent→child and child→parent cones, so trivially satisfied constraints are visible. [Table 13] 4. Shuffle-controlled radial excess: directed parent–child ordering minus its permuted-pairing null (∆pair , z), not the raw rate. [Table 29] 2 5. Radial increment beyond angle: ∆Rnorm over a cosine-only regression, with a Mantel permutation p. [Table 5]
Where training access permits, additionally attribute the curvature gradient by loss term (∂L/∂ log c sign and magnitude; Figure 4). Operational hierarchy (traversal monotonicity, collapse rate, and terminal retrieval diversity) is the operational complement to these five geometry numbers and is reported separately in Tables 19 and 31. Items are necessary, not sufficient. 23
These are necessary-condition tests, not a complete benchmark for all hierarchy representations: a model may improve angular organization, stabilize cross-modal alignment, or reduce entailment violations without using active radial or cone-based hierarchy. Jointly, however, the criteria are far more constraining than any single test: a model that activates nonlocal geometry, identifies curvature independently of radial scale, maintains active and contained cones, and passes the directed and operational tests has closed each escape route by which the audited models pass the diagnostics they do pass. Passing all of them would still not prove the mechanism active—hierarchy could be realized in some form none of these tests detects—but a model reporting all of them gives readers the information needed to judge: the contract specifies what a hierarchy claim should report. One might object that the audit is self-sealing: if our mechanism (Section 7.3) explains why no model leaves the near-Euclidean basin, then a negative hierarchy result is guaranteed by construction rather than measured. Two features of the design rule this out. First, the diagnostics are not blind to a positive: the graded positive controls (Appendix F.1 and the planted-tree control of Appendix F) show that the directed and taxonomy tests do fire when a pair-specific radial ordering is present, and collapse to chance only under the shuffle null—so a model that did instantiate the mechanism would register on them. Second, the nondetections are bounded, not open-ended: the MDE analysis quantifies the smallest shuffle-controlled signal 2 we would have caught (∆pair ≈ 13pp for the radial test, ∆Rnorm ≈ 0.013 for the taxonomy test), and the audited hyperbolic and box-trained models’ shuffle-controlled excesses fall well below that threshold (largest 2 ≤ 0.0029 vs. the 0.013 taxonomy ≈ 5pp vs. the ≈ 13pp radial MDE; largest taxonomy increment ∆Rnorm MDE), so the null readings are bounded absences, not failures to look. The negative is therefore a measured absence within a quantified sensitivity envelope, on diagnostics demonstrated to detect the signal when it exists—not an artifact of restricting attention to a regime where the signal cannot appear. That the models do not reach the nonlinear regime is itself one of our findings (Section 4), with a mechanistic account of why. It is a property of the audited objectives, not an assumption built into the tests.
8
Conclusion
We asked whether public hyperbolic vision–language models use the geometry they are built around, and found two distinct, separable failures. First, curvature is not an active geometric resource in any audited configuration: across MERU, HyCoCLIP, and √ PHyCLIP—released and trained from scratch—the dimensionless operating point stays near-Euclidean ( cρ ≈ 0.2–0.3, H(u) ≈ 1, no evaluated checkpoint in the nonlocal regime), and releasing the curvature floor moves scalar curvature and norms while the operating point stays in the near-Euclidean band. Second, the cone and traversal machinery is measured inoperative— apertures saturated or misaligned, graded traversal failing under controlled readouts—while directed radial depth is a bounded non-detection: external parent–child ordering shows no signal above shuffle-null controls at quantified sensitivity, leaving only a tiny, non-operative residual on the models’ own native relation. A gradient-level mechanism explains why the natural interventions do not repair this—the entailment √ objective itself admits a low-curvature, wide-cone shortcut, and under a parameter-free aperture relation ( cρ ≤ 2K) curvature settles at the aperture-saturation edge in every entailment-trained unclamped run. Entailment-off ablations further show that fixing the cone loss alone would be insufficient: contrastive/alignment training also fails to hold curvature up (the norm side of the contraction is confounded with weight decay), though the entailment shortcut is the dominant full-objective accelerator. These results do not show that hyperbolic geometry is the wrong tool for vision–language learning. They show that current published formulations leave its mechanism dormant, and that the evidence usually read as hierarchy is carried by angular structure or box/compositional supervision effects rather than radial depth. To make future claims checkable rather than assumed, we distill our diagnostics into a five-number geometry report (Section 7.5)—operating point, cone-saturation state, directed violations, shuffle-controlled radial excess, and radial increment beyond angle—paired with the operational-hierarchy traversal statistics. A model that reports these lets readers judge whether its geometry is active, not merely present. 24
Reproducibility Statement We release the full audit suite as supplementary material (Appendix I): measurement scripts, the raw perrun outputs behind every table, checkpoint SHA-256 hashes, current-GRIT training configurations, and the radius/curvature/aperture convention documents. The released MERU, HyCoCLIP, and PHyCLIP checkpoints are identified by their hashes; the from-scratch current-GRIT checkpoints are available on request and will be released publicly on publication. Every reported table can be re-derived from the persisted outputs alone—the weights and configurations are needed only to regenerate those outputs from scratch.
References Sameer Bansal and Adrian Benton. Comparing euclidean and hyperbolic embeddings on the wordnet nouns hypernymy graph. In Proceedings of the Second Workshop on Insights from Negative Results in NLP, pp. 49–53, 2021. Martin R Bridson and André Haefliger. Metric spaces of non-positive curvature, volume 319. Springer Science & Business Media, 1999. Jason Chuan-Chih Chou and Nahid Alam. Embedding geometries of contrastive language-image pre-training. In European Conference on Computer Vision, pp. 399–416. Springer, 2024. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009. Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shanmukha Ramakrishna Vedantam. Hyperbolic image-text representations. In International Conference on Machine Learning, pp. 7694– 7731. PMLR, 2023. Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. Hyperbolic entailment cones for learning hierarchical embeddings. In International conference on machine learning, pp. 1646–1655. PMLR, 2018. Shijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen, and Yongfeng Zhang. Hiclip: Contrastive language-image pretraining with hierarchy-aware attention. arXiv preprint arXiv:2303.02995, 2023. Albert Gu, Frederic Sala, Beliz Gunel, and Christopher Ré. Learning mixed-curvature representations in product spaces. In International Conference on Learning Representations (ICLR), 2019. Yunhui Guo, Xudong Wang, Yubei Chen, and Stella X. Yu. Clipped hyperbolic classifiers are superhyperbolic classifiers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality. Advances in neural information processing systems, 36:31096–31116, 2023. Chuong Huynh, Hossein Souri, Abhinav Kumar, Vitali Petsiuk, Deen Dayal Mohan, and Suren Kumar. Argent: Adaptive hierarchical image-text representations. arXiv preprint arXiv:2603.23311, 2026. Sarah Ibrahimi, Mina Ghadimi Atigh, Nanne Van Noord, Pascal Mettes, and Marcel Worring. Intriguing properties of hyperbolic embeddings in vision-language models. Transactions on Machine Learning Research, 2024. Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014. George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995. 25
Maximilian Nickel and Douwe Kiela. Poincaré embeddings for learning hierarchical representations. Advances in neural information processing systems, 30, 2017. Avik Pal, Max Van Spengler, Guido D’Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. Compositional entailment learning for hyperbolic vision-language models. In International Conference on Learning Representations, volume 2025, pp. 87371–87399, 2025. Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, and Furu Wei. Grounding multimodal large language models to the world. In International Conference on Learning Representations, volume 2024, pp. 51575–51598, 2024. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PmLR, 2021. Sameera Ramasinghe, Violetta Shevchenko, Gil Avraham, and Ajanthan Thalaiyasingam. Accept the modality gap: An exploration in the hyperbolic space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 27263–27272, 2024. Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. Frederic Sala, Christopher De Sa, Albert Gu, and Christopher Ré. Representation tradeoffs for hyperbolic embeddings. In International Conference on Machine Learning (ICML), 2018. Rik Sarkar. Low distortion delaunay embedding of trees in hyperbolic plane. In International symposium on graph drawing, pp. 355–366. Springer, 2011. Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. Order-embeddings of images and language. arXiv preprint arXiv:1511.06361, 2015. Daiki Yoshikawa and Takashi Matsubara. PHyCLIP: ℓ1 -product of hyperbolic factors unifies hierarchy and compositionality in vision-language representation learning. arXiv preprint arXiv:2510.08919, 2025. Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the association for computational linguistics, 2:67–78, 2014. Tiancheng Zhao, Tianqi Zhang, Mingwei Zhu, Haozhan Shen, Kyusong Lee, Xiaopeng Lu, and Jianwei Yin. Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations. arXiv preprint arXiv:2207.00221, 2022.
26
A
Notation
Table 7 collects the symbols used throughout the paper.
Symbol c ρ = ∥x∥ √ u √= c ρ cρ > 1 H(u) = sinh(u)/u K ω(ρ) 2 Rcos 2 ∆Rnorm z ∆pair pperm η NG2
Table 7: Notation used throughout the paper. Meaning Learned curvature parameter. Radial coordinate: the Lorentz spatial norm (distance from the origin). Dimensionless operating point. Marker for the nonlocal (operative-hyperbolic) regime. Local distortion factor; H ≈ 1 is near-Euclidean and grows in the nonlocal regime. Fixed √ constant in the cone-aperture formula (K = 0.1); a cone saturates when cρ ≤ 2K. Entailment-cone half-aperture. Taxonomy-distance variance explained by cosine (angular) distance. Incremental R2 of embedding norm beyond cosine distance. Shuffle-null standardized score (directed radial test). Shuffle-controlled pair-specific excess (radial ordering). Mantel / shuffle permutation-test p-value. Entailment activation threshold. The model-native GRIT box→full-caption relation (its own training relation).
27
B
Full Geometry Diagnostics
B.1
Released-checkpoint per-modality local distortion
Table 8: Full released-checkpoint geometry diagnostics. All released checkpoints remain in the near√ Euclidean regime, with 0% of samples satisfying cρ > 1. Statistics use 500 ImageNet val images and 500 COCO val captions. PHyCLIP values are per-subspace, aggregated over 64 factors. √ Model Modality c ρ cρ H(u) MERU-S MERU-S MERU-B MERU-B MERU-L MERU-L HyCoCLIP-S HyCoCLIP-S HyCoCLIP-S HyCoCLIP-S HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-L PHyCLIP-L PHyCLIP-L PHyCLIP-L
image text image text image text full image box image full text box text full image box image full text box text full image box image full text box text full image box image full text box text
0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100 0.100† 0.100† 0.100† 0.100† 0.100† 0.100† 0.100† 0.100†
0.816 0.547 0.834 0.578 0.882 0.599 0.635 0.630 0.389 0.319 0.636 0.630 0.385 0.324 0.681 0.647 0.442 0.365 0.690 0.652 0.442 0.370
0.258 0.173 0.264 0.183 0.279 0.190 0.201 0.199 0.123 0.101 0.201 0.199 0.122 0.102 0.215 0.205 0.140 0.115 0.218 0.206 0.140 0.117
1.0111 1.0050 1.0116 1.0056 1.0130 1.0060 1.0067 1.0066 1.0025 1.0017 1.0068 1.0066 1.0025 1.0018 1.0078 1.0071 1.0033 1.0023 1.0080 1.0072 1.0033 1.0024
† PHyCLIP uses per-subspace curvatures. The value shown is the mean across 64 subspaces (see §B.6).
28
B.2
Operating-point upper tail:
√
cρ quantiles √ √ Tables 1 and 2 report median/mean cρ and the nonlocal-regime fraction ( cρ > 1). To rule out rare highradius samples a median could √ hide, Table 9 adds the upper tail—95th/99th percentiles and maximum—of the per-sample image-side cρ for every√ audited checkpoint. Across every audited checkpoint (over 500 ImageNet val √ images), the largest single cρ is 0.367 (released √ PHyCLIP-B). Local distortion H(u) reaches 10% only at cρ ≈ 0.76, and the nonlocal regime only at cρ = 1 (Section 3, Figure 1). The observed maximum is below half the former and just over a third of the latter. PHyCLIP’s wider tail reflects its widest subspace. √ Table 9: Upper tail (percentiles and maximum) of the per-sample image-side cρ, on the same 500-image ImageNet is not a model seed). In every √ val subset as Tables √ 8 and 2 (sampling seed 42 selects the images; it √ row % cρ > 0.76 and % cρ > 1 are 0 (omitted), and medians match the cρ of those tables to within 0.002 (per-sample medians vs. the means reported there). Released rows are the as-published checkpoints (Seed ‘—’). Baseline/clampOff are the from-scratch current-GRIT runs. Baseline and ViT-S rows use the representative seed 0. ViT-B clampOff is reported per training√seed (0/37/42). Consistent with Table 2, MERU-B clampOff is seed-consistent just below the 2K edge ( cρ = 0.177–0.180 across seeds). Even its maximum, 0.189, stays far below 0.76. PHyCLIP quantiles are over its 64 per-factor values per sample. Model
Setting
Seed
median
p95
p99
max
MERU-S MERU-B MERU-L HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
released released released released released released released
— — — — — — —
0.258 0.264 0.279 0.200 0.201 0.213 0.216
0.264 0.270 0.287 0.203 0.204 0.257 0.262
0.267 0.273 0.290 0.204 0.205 0.279 0.287
0.269 0.278 0.292 0.205 0.206 0.367 0.351
MERU-S MERU-B HyCoCLIP-S HyCoCLIP-B PHyCLIP-S PHyCLIP-B
baseline baseline baseline baseline baseline baseline
0 0 0 0 0 0
0.291 0.296 0.201 0.201 0.202 0.204
0.299 0.305 0.203 0.203 0.239 0.242
0.302 0.308 0.204 0.204 0.257 0.263
0.306 0.310 0.207 0.205 0.317 0.314
MERU-S HyCoCLIP-S PHyCLIP-S MERU-B MERU-B MERU-B HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B
clampOff clampOff clampOff clampOff clampOff clampOff clampOff clampOff clampOff clampOff clampOff clampOff
0 0 0 0 37 42 0 37 42 0 37 42
0.216 0.194 0.187 0.180 0.177 0.179 0.195 0.194 0.195 0.187 0.187 0.187
0.221 0.197 0.199 0.185 0.181 0.184 0.197 0.197 0.197 0.200 0.199 0.200
0.224 0.197 0.205 0.187 0.183 0.186 0.198 0.197 0.197 0.206 0.206 0.206
0.224 0.198 0.222 0.189 0.184 0.187 0.199 0.198 0.198 0.225 0.225 0.229
29
B.3
Full per-seed downstream and evaluation results
This section is reproducibility material: the complete per-seed absolute metrics underlying Table 3 (full precision, seeds 0/37/42, baseline and clampOff), from which each reported delta can be recomputed directly. Tables 10, 11, and 12 cover retrieval and WordNet hierarchical metrics, the 16-task zero-shot suite, and the per-subtype compositional results, respectively. Across all metrics and models, curvature unclamping produces only small, seed-comparable shifts with no substantial degradation. Table 10: Reproducibility material: per-seed absolute zero-shot retrieval (Recall@5/@10, points) and WordNet hierarchical-classification metrics, current-GRIT baseline and clampOff (seeds 0, 37, 42, full precision), backing the Table 3 deltas. TIE/LCA are distances (lower better). J/PH/RH are on 0–1 (higher better). Text→Image COCO
Image→Text
Flickr
COCO
Hierarchical Classification
Flickr
WordNet
Model / setting
seed R@5 R@10 R@5 R@10 R@5 R@10 R@5 R@10 TIE↓ LCA↓
MERU-B baseline
0 37 42 0 37 42
55.28 55.75 55.09 55.50 55.15 55.34
66.40 67.27 66.49 66.99 66.77 66.56
81.38 82.08 81.66 81.48 81.84 81.64
88.28 88.66 88.72 88.22 88.60 88.80
68.52 69.84 68.58 68.28 69.14 68.12
78.82 79.76 78.34 79.12 79.00 78.40
89.10 90.10 90.40 88.80 91.10 89.70
94.60 94.90 94.80 94.30 95.40 94.00
3.872 3.903 3.864 3.886 3.994 3.907
2.330 2.331 2.307 2.313 2.352 2.325
0.7707 0.8426 0.8440 0.7685 0.8413 0.8417 0.7710 0.8438 0.8424 0.7687 0.8428 0.8410 0.7618 0.8378 0.8355 0.7682 0.8422 0.8409
0 37 42 HyCoCLIP-B clampOff 0 37 42
57.38 56.31 57.33 57.64 57.44 58.82
68.39 67.28 68.40 69.07 68.69 69.53
82.94 82.60 83.50 84.12 83.16 83.92
89.70 89.54 89.96 90.28 89.84 90.74
69.36 69.34 70.44 71.40 70.92 71.62
79.50 79.74 80.38 80.68 80.98 81.00
91.40 90.50 91.80 92.30 92.10 92.30
96.30 95.10 95.80 95.40 96.70 95.30
3.299 3.395 3.282 3.353 3.498 3.378
2.086 2.125 2.077 2.083 2.173 2.118
0.8054 0.8684 0.8670 0.8002 0.8642 0.8640 0.8070 0.8692 0.8687 0.8017 0.8673 0.8627 0.7947 0.8599 0.8590 0.8002 0.8647 0.8620
PHyCLIP-B baseline
56.84 57.01 56.79 57.34 57.54 57.68
68.18 67.95 68.13 68.41 68.76 68.53
82.92 82.60 82.96 83.64 82.64 83.50
89.78 89.28 89.46 89.98 89.34 90.08
69.94 70.56 69.64 70.98 70.66 70.40
79.96 79.56 79.76 80.50 80.16 80.34
91.70 90.80 90.60 91.20 91.00 91.40
95.40 95.10 94.80 95.60 95.40 95.70
3.346 3.342 3.284 3.375 3.377 3.367
2.111 2.124 2.070 2.106 2.110 2.097
0.8034 0.8665 0.8661 0.8042 0.8659 0.8677 0.8067 0.8697 0.8671 0.8009 0.8655 0.8629 0.8002 0.8650 0.8627 0.8010 0.8660 0.8632
MERU-B clampOff
HyCoCLIP-B baseline
PHyCLIP-B clampOff
0 37 42 0 37 42
30
J↑
PH↑
RH↑
Table 11: Per-seed absolute zero-shot classification accuracy (top-1, %) across 16 datasets. Current-GRIT baseline and clampOff, seeds 0, 37, 42. Model / setting
seed
IN
C10 C100 SUN
Cal
STL Food CUB Cars Airc Pets Flwr DTD Euro RES C211
MERU-B baseline
0 37 42 0 37 42
37.32 75.31 45.11 49.45 72.24 92.80 48.89 8.44 37.06 73.96 42.79 50.30 72.37 93.50 49.64 10.73 37.39 75.28 44.00 50.05 73.25 92.55 47.90 9.44 37.12 70.66 44.21 49.41 72.49 93.39 46.44 11.12 36.40 74.14 44.03 49.58 72.08 93.19 47.03 10.54 36.87 76.02 43.21 49.86 71.76 92.66 50.58 10.43
2.42 42.40 17.78 20.59 37.14 40.36 2.15 41.74 21.62 22.93 40.90 41.36 2.33 40.43 15.09 19.84 37.82 41.41 2.36 41.47 17.28 22.23 34.67 37.52 2.36 43.42 20.85 19.31 40.89 43.67 2.09 44.79 16.58 21.33 47.27 42.21
5.05 4.36 4.91 4.75 4.54 4.72
0 37 42 HyCoCLIP-B clampOff 0 37 42
44.39 89.01 59.42 55.08 74.36 94.94 55.21 15.60 10.20 3.98 52.38 22.17 26.76 41.14 43.32 42.53 87.94 58.63 54.44 77.43 94.12 56.57 17.59 9.11 3.86 51.01 24.56 25.80 41.80 46.98 44.06 89.04 59.27 55.04 76.93 95.06 59.65 18.37 9.80 3.28 53.28 26.06 24.84 39.89 49.62 43.71 87.84 57.11 54.47 75.80 94.33 57.46 16.29 10.64 3.09 51.83 26.39 26.38 42.13 42.92 42.38 88.26 57.93 54.09 75.22 93.82 55.14 15.63 10.08 3.56 49.77 24.97 22.98 32.28 44.28 43.83 88.66 59.20 54.32 76.56 94.16 55.67 14.27 11.01 3.99 52.02 26.80 27.07 40.70 45.60
5.57 5.45 5.81 5.98 5.92 5.41
PHyCLIP-B baseline
43.38 88.68 59.48 55.41 77.69 94.10 58.72 16.19 9.67 3.63 53.00 25.36 25.90 41.59 46.92 43.63 88.16 58.45 55.28 75.07 95.10 58.51 17.79 10.29 3.16 52.68 23.81 27.93 47.18 45.27 44.25 88.70 60.74 55.92 77.31 95.19 59.20 15.86 10.55 3.17 52.99 25.29 27.61 40.17 48.66 43.56 88.17 58.90 55.14 76.45 93.99 55.10 14.64 9.21 4.23 54.13 24.26 25.64 38.01 46.24 43.55 87.08 58.18 54.50 75.67 94.94 55.90 14.69 9.26 3.45 50.82 26.25 26.17 39.39 45.19 43.27 88.16 58.34 53.77 77.67 95.00 57.40 15.47 10.59 2.75 50.40 23.35 27.18 41.87 46.45
5.36 5.60 5.63 5.37 5.68 5.73
MERU-B clampOff
HyCoCLIP-B baseline
PHyCLIP-B clampOff
0 37 42 0 37 42
7.22 6.84 7.07 6.96 6.21 6.43
Table 12: Per-seed absolute compositional (hard-negative) accuracy (points). VL-CheckList-Object: Location (Center/Mid/Margin), Size (Large/Medium/Small). SugarCrepe: Replace (Obj/Att/Rel), Swap (Obj/Att), Add (Obj/Att), overall. Current-GRIT baseline and clampOff, seeds 0, 37, 42. VL-CheckList-Object
SugarCrepe
Model / setting
seed Loc-C Loc-M Loc-Mg Sz-L Sz-M Sz-S Rep-O Rep-A Rep-R Swp-O Swp-A Add-O Add-A SC-All
MERU-B baseline
0 37 42 0 37 42
62.90 62.10 54.40 60.80 64.60 57.50
60.30 58.10 53.20 60.50 62.80 56.30
61.60 57.10 54.00 61.80 61.00 55.70
64.00 60.60 58.80 62.90 57.60 55.50 56.10 54.10 50.70 63.40 61.10 55.20 65.30 62.50 59.30 59.50 56.60 56.00
90.68 88.92 89.95 89.53 89.65 90.13
79.57 81.09 79.19 79.95 79.57 79.44
69.84 69.77 70.06 69.20 69.49 69.35
57.14 60.82 55.10 57.96 60.41 53.06
63.96 66.52 64.41 67.42 67.72 62.01
80.02 80.80 78.90 80.21 81.96 80.75
71.68 72.69 75.29 72.83 71.10 73.27
77.47 77.89 77.31 77.63 78.10 77.29
0 37 42 HyCoCLIP-B clampOff 0 37 42
67.50 75.00 76.30 69.00 67.80 69.80
68.10 74.30 73.80 68.70 67.60 68.70
69.30 70.90 71.30 70.60 67.90 68.00
69.60 65.00 66.40 75.70 71.00 71.20 77.10 71.30 71.50 70.80 66.80 69.50 67.90 64.90 69.90 70.90 65.20 68.10
90.80 91.16 90.92 90.62 91.40 90.86
77.66 80.71 78.81 81.60 80.84 81.60
65.79 65.58 65.22 69.56 67.78 67.85
58.37 58.78 57.14 60.82 59.18 53.06
64.41 64.56 63.66 65.17 69.07 65.17
82.06 81.57 81.28 82.25 83.37 83.27
73.41 70.81 74.42 74.13 72.98 72.69
77.34 77.35 77.15 78.68 78.94 78.31
PHyCLIP-B baseline
76.60 73.30 77.60 78.20 76.20 77.20
73.00 70.80 74.80 75.80 72.60 75.40
73.10 68.80 70.80 75.10 71.70 75.40
76.70 70.70 72.00 73.80 67.70 66.90 79.50 70.00 70.80 77.80 75.00 74.30 75.90 70.70 71.20 78.10 73.20 74.60
91.10 90.98 91.16 90.74 91.16 91.22
80.33 80.33 79.70 82.11 80.20 80.20
65.50 66.64 68.14 68.92 67.71 66.71
60.41 60.82 60.41 60.41 59.18 56.73
64.86 66.52 66.82 65.77 62.91 62.91
81.09 82.10 81.04 83.75 83.46 82.69
74.42 70.09 72.69 74.86 75.00 75.14
77.57 77.79 78.01 79.16 78.47 78.02
MERU-B clampOff
HyCoCLIP-B baseline
PHyCLIP-B clampOff
0 37 42 0 37 42
31
B.4
Cone aperture and entailment violation diagnostics
Table 13 reports the full cone diagnostics across families, sizes, released checkpoints, and clampOff, for all three seeds. The patterns match Section 5.2: the i→t rate is ≈ 100% everywhere (text never enters the active image cone), and the small t→i rate reflects embedding compression, not order. From opposite baseline image-cone regimes—MERU narrow (≈ 0.74 rad), HyCoCLIP wide (≈ 1.45 rad)—both families saturate to π/2 under clampOff at ViT-B. The exception is MERU-S clampOff (0.75 → 1.175 rad, not fully saturated). In all cases the clampOff cone provides no operational hierarchy.
Table 13: Full cone aperture and entailment-violation diagnostics across families, sizes, released checkpoints, and the clampOff intervention. Image-aperture median (radians) with the fraction of samples/factors saturated at π/2 (ω ≥ π/2 − 0.01) in parentheses. Text-side apertures saturate at π/2 for every model. Violation is reported in both directions: i→t tests text embeddings against the image-side cone (parent = image), t→i image embeddings against the text-side cone (parent = text), and bt→bi the box-text→box-image cone for the box-supervised families. Violations are measured at the exact cone boundary (η=1.0: angle exceeds the aperture), independent of the softer η=1.2 training margin, so the diagnostic reports true geometric containment rather than trained slack (the η intervention is in Appendix C.3). PHyCLIP values are factor-wise means across 64 subspaces. Cones are measured on a 256-sample GRIT batch (processed shards 00000– 00001), the same evaluation distribution as the geometry in Table 14. ViT-B configurations are reported for three seeds (ViT-S has a single trained seed). MERU-B clampOff is seed-consistent (all three seeds saturate just below the 2K edge; Section 7.3). The released PHyCLIP-B/L bt→bi cells are carried from earlier runs. The others are re-measured here, drawing one box per sample under a fixed seed. Every bt→bi rate is near-trivial, so it corroborates rather than carries the conclusion. Model
Setting
c
img aper (rad, sat%)
text aper (rad, sat%)
i→t viol %
t→i viol %
bt→bi viol %
MERU-S MERU-B MERU-L HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
released‡ released‡ released‡ released released released released
0.100 0.100 0.100 0.100 0.100 0.100 0.100
0.875 (0%) 0.852 (0%) 0.802 (0%) 1.475 (30.5%) 1.456 (26.2%) 1.219 (24.6%) 1.185 (21.4%)
1.571 (98.8%) 1.571 (98.0%) 1.571 (97.7%) 1.571 (100%) 1.571 (100%) 1.569 (99.4%) 1.569 (99.5%)
100.0 100.0 100.0 100.0 100.0 100.0 100.0
48.4 64.8 70.7 0.39 0.39 8.5 6.3
– – – 0.00 0.00 6.3 5.4
MERU-S MERU-S MERU-B MERU-B MERU-B MERU-B MERU-B MERU-B HyCoCLIP-S HyCoCLIP-S HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B PHyCLIP-S PHyCLIP-S PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-B
baseline clampOff baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42) baseline clampOff baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42) baseline clampOff baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42)
0.100 0.042 0.100 0.100 0.100 0.029 0.028 0.029 0.100 0.011 0.100 0.100 0.100 0.009 0.009 0.009 0.100 0.015 0.100 0.100 0.100 0.014 0.014 0.015
0.751 (0%) 1.175 (0%) 0.732 (0%) 0.726 (0%) 0.733 (0%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.446 (15.6%) 1.571 (100%) 1.422 (10.9%) 1.420 (12.9%) 1.423 (10.5%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.417 (44.7%) 1.569 (95.5%) 1.373 (41.7%) 1.372 (40.9%) 1.366 (41.4%) 1.569 (95.6%) 1.569 (95.5%) 1.569 (95.7%)
1.571 (99.2%) 1.571 (100%) 1.571 (99.6%) 1.571 (99.2%) 1.571 (99.2%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (100%) 1.571 (99.5%) 1.569 (100%) 1.569 (99.7%) 1.569 (99.7%) 1.569 (99.7%) 1.569 (100%) 1.569 (100%) 1.569 (100%)
100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0
5.1 5.5 3.1 2.7 3.1 1.6 1.2 3.1 0.0 0.0 0.0 0.0 0.4 0.0 0.0 0.0 13.8 0.20 9.5 9.3 9.1 0.05 0.10 0.09
– – – – – – – – 0.8 0.0 0.0 0.4 0.4 0.0 0.0 0.0 6.5 0.05 6.1 6.4 5.6 0.04 0.04 0.05
‡ Released MERU is RedCaps-trained rather than GRIT-trained, so its GRIT-evaluated t→i rate (48–71%) partly reflects
domain shift, not learned order. Its i→t rate is still ≈ 100%.
32
B.5
Current-GRIT clampOff geometry by modality
Table 14: Full modality-wise geometry for current-GRIT baseline and clampOff interventions. The baseline/clampOff setting is folded into the model name. Statistics use 500 GRIT samples (processed shards 00000–00001), the same evaluation distribution as the cone diagnostics in Table 4. This is the GRITevaluated counterpart to the ImageNet/COCO-evaluated Table 2. MERU has no box modality. Coneaperture and entailment-violation diagnostics are reported separately in Table 13. PHyCLIP is√omitted from the present table because its 64 independently-curved subspaces have no single per-modality cρ, and its factor-wise cone geometry is given there. All models remain near-Euclidean (nonlocal-regime fraction 0% for every model and modality). Rows √ are reported at the representative seed (seed 0). All cells are seedstable, including MERU-B clampOff ( cρ = 0.176–0.180 across seeds, just below the 2K=0.2 saturation edge; Table 2). √ Model Modality c ρ cρ H(u) MERU-S baseline MERU-S baseline MERU-S clampOff MERU-S clampOff MERU-B baseline MERU-B baseline MERU-B clampOff MERU-B clampOff
image text image text image text image text
0.100 0.100 0.042 0.042 0.100 0.100 0.029 0.029
0.929 0.591 1.061 0.773 0.947 0.595 1.060 0.759
0.294 0.187 0.217 0.158 0.299 0.188 0.182 0.130
1.0145 1.0059 1.0079 1.0042 1.0150 1.0059 1.0055 1.0028
HyCoCLIP-S baseline HyCoCLIP-S baseline HyCoCLIP-S baseline HyCoCLIP-S baseline HyCoCLIP-S clampOff HyCoCLIP-S clampOff HyCoCLIP-S clampOff HyCoCLIP-S clampOff
full image box image full text box text full image box image full text box text
0.100 0.100 0.100 0.100 0.011 0.011 0.011 0.011
0.637 0.631 0.392 0.313 1.848 1.853 1.350 1.266
0.201 0.200 0.124 0.099 0.195 0.196 0.143 0.134
1.0068 1.0067 1.0026 1.0016 1.0064 1.0064 1.0034 1.0030
HyCoCLIP-B baseline HyCoCLIP-B baseline HyCoCLIP-B baseline HyCoCLIP-B baseline HyCoCLIP-B clampOff HyCoCLIP-B clampOff HyCoCLIP-B clampOff HyCoCLIP-B clampOff
full image box image full text box text full image box image full text box text
0.100 0.100 0.100 0.100 0.009 0.009 0.009 0.009
0.639 0.632 0.383 0.320 2.058 2.065 1.515 1.337
0.202 0.200 0.121 0.101 0.195 0.196 0.144 0.127
1.0068 1.0067 1.0024 1.0017 1.0064 1.0064 1.0035 1.0027
33
Table 15: Per-seed absolute image-side geometry (500 ImageNet and 500 COCO val images), backing the summary in Table 2. ViT-B rows are reported for all three seeds (0, 37, 42). ViT-S rows are single-seed (seed √ 0). The nonlocal-regime fraction is 0% in every row. MERU-B clampOff is seed-consistent: cρ = 0.176– 0.180 across seeds, below 2K=0.2 and saturated. The GRIT per-modality / aperture breakdown is Table 14. √ Model / setting seed c ρ (img) cρ H(u) MERU-S baseline MERU-S clampOff
0 0
0.1000 0.0419
0.9206 1.0523
0.2911 0.2154
1.0142 1.0078
MERU-B baseline
0 37 42 0 37 42
0.1000 0.1000 0.1000 0.0294 0.0281 0.0291
0.9362 0.9430 0.9371 1.0506 1.0515 1.0451
0.2961 0.2982 0.2964 0.1800 0.1762 0.1784
1.0147 1.0149 1.0147 1.0054 1.0052 1.0053
0 0
0.1000 0.0112
0.6348 1.8390
0.2007 1.0067 0.1944 1.0063
0 37 42 HyCoCLIP-B clampOff 0 37 42
0.1000 0.1000 0.1000 0.0090 0.0094 0.0095
0.6379 0.6377 0.6374 2.0531 2.0022 2.0017
0.2017 0.2017 0.2016 0.1949 0.1945 0.1947
1.0068 1.0068 1.0068 1.0063 1.0063 1.0063
PHyCLIP-S baseline PHyCLIP-S clampOff
0 0
0.1000 0.0155
0.6440 1.5084
0.2036 0.1877
1.0070 1.0059
PHyCLIP-B baseline
0 37 42 0 37 42
0.1000 0.1000 0.1000 0.0143 0.0144 0.0145
0.6501 0.6500 0.6507 1.5700 1.5612 1.5539
0.2056 0.2055 0.2058 0.1876 0.1874 0.1873
1.0071 1.0071 1.0071 1.0059 1.0059 1.0059
MERU-B clampOff
HyCoCLIP-S baseline HyCoCLIP-S clampOff HyCoCLIP-B baseline
PHyCLIP-B clampOff
B.6
PHyCLIP product-geometry convention
PHyCLIP uses a product-space implementation whose reported scalar curvature convention differs from the global-curvature convention used in MERU and HyCoCLIP. For this reason, PHyCLIP product-space diagnostics are reported factor-wise where applicable, and we avoid comparing a single scalar curvature across families unless the implementation convention is matched.
C
GRIT Snapshot and Training Details
C.1
Current-GRIT snapshot
A “parent-box annotation” is counted as one parent{NNN}.txt grounding entry per sample (the box_text source consumed by GroundedDatasetTarMapper), matching the per-sample box unit used elsewhere in this paper. The documented 35.9M figure is reproduced from the original release and may aggregate box annotations under a slightly different convention.
34
Table 16: Statistics of the fixed current-GRIT snapshot used for controlled interventions. GRIT is URL-derived. The effective corpus is smaller than the original release due to dead links. All 2,051 shards were enumerated. The per-shard breakdown is provided in the released artifact (data_provenance/grit_snapshot_count_per_shard.csv). Quantity
Value
Download / crawl date Total shards (.tar) Processed corpus size Image–text pairs Parent-box annotations Mean parent boxes per pair Pairs per shard Effective epochs (training) Mapper Train image transform Decoding failures
C.2
2026-04-13 (file mtime) 2,051 (all enumerated) 731 GB documented 20.5M / obtained 13,064,747 documented 35.9M / obtained 25,018,042 1.915 (25,018,042/13,064,747) min 5,611, mean 6,370, max 6,540 ≈29 (384M exposures / 13.06M pairs) GroundedDatasetTarMapper RandomResizedCrop(224, scale=(0.5,1.0)) + ToTensor webdataset warn_and_continue
Training hyperparameters
Table 17: Training hyperparameters for current-GRIT interventions. All causal comparisons are within the same current-GRIT snapshot and change only the listed intervention variable. The PHyCLIP clampOff variants use the same floor-relaxation protocol ([0.1, 10] → [0.001, 10]) in the product-space configuration. Family
Size
Variant
Curv. clamp
Changed variable
MERU MERU MERU HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP PHyCLIP PHyCLIP
S/B S/B B S/B S/B S/B B S/B S/B
baseline clampOff λe =0 baseline clampOff intra=0.7 λe =0 baseline clampOff
[0.1, 10] [0.001, 10] [0.1, 10] [0.1, 10] [0.001, 10] [0.1, 10] [0.1, 10] [0.1, 10] [0.001, 10]
none curvature floor entailment weight (→ 0) none curvature floor entailment threshold entailment weight (→ 0) none curvature floor
The baseline and clampOff variants are matched on the same current-GRIT snapshot with global batch size 768, 500,000 iterations, AdamW (lr = 5 × 10−4 , β = (0.9, 0.98), weight decay 0.2), linear warmup (4000 steps) followed by cosine decay, no gradient accumulation (each iteration is a single optimizer step over 768 samples), and AMP enabled (fp16 forward, fp32 hyperbolic operations via explicit .float() casts in phyclip/lorentz.py).5 Hardware was determined by model size: ViT-S variants were trained on a single H100, and all ViT-B runs (MERU, PHyCLIP, and HyCoCLIP) on 2×H200. MERU-B clampOff seeds 37/42 were retrained on this unified 2×H200 setup after preliminary single-GPU runs showed hardware-dependent geometry. The retrained seeds converge with seed 0. The global batch size of 768 was held fixed across all configurations and devices. ViT-L is not trained here. It is audited only at the released-checkpoint level (Table 8). 5 The remaining variants share this recipe. The η=0.7 intervention runs use the full 500,000-iteration budget (seed 0, ViT-S and ViT-B). Only the curvature-collapse probes are short: the λe =0 probes run 11.0–11.2k steps (seeds 0/37/42, both families; the MERU-B seed-0 probe was additionally logged to 38.1k), the seed-0 run in each family is extended to 40k (Section 7.4), and the dedicated λe =0.2 collapse probes (seeds 23/37/42) run ∼10k steps.
35
C.3
Threshold-activation stress test (η intervention)
This auxiliary single-seed intervention complements the entailment-off ablation of Section 7.4: rather than removing the entailment term, we activate it more strongly and find it likewise fails to install hierarchy. Whereas the main-text interventions (Section 7) add or modify a curvature path, this stress test stays within the published objective, lowering the intra-modal entailment threshold η from the default 1.2 to 0.7. This is the most direct test of whether more active entailment cones induce hierarchy. We run it as a matched current-GRIT intervention on HyCoCLIP at both ViT-S and ViT-B (seed 0), changing only η.
Table 18: Lowering the entailment threshold η from 1.2 to 0.7 (HyCoCLIP, current-GRIT, seed 0) leaves the operating geometry unchanged, activates √ cones only marginally, and shifts behavior in a scale-dependent √ direction. Geometry rows (c, cρ, H, % cρ>1) are identical to two decimals. Cone activation rises but stays sub-π/2. Traversal and downstream effects differ by scale and, at ViT-B, worsen. Single-seed comparison, read as decoupling evidence rather than a performance claim. The ZSC mean is over the zero-shot suite for this η study and is not directly comparable in absolute terms to the 16-task suite in Table 11. ViT-B ViT-S η=1.2
η=0.7
∆
η=1.2
η=0.7
∆
Operating geometry (unchanged) c√ 0.100 cρ (image) 0.202 H(u) 1.007 √ % cρ>1 0%
0.100 0.201 1.007 0%
0 −0.001 0 0
0.100 0.201 1.007 0%
0.100 0.201 1.007 0%
0 0 0 0
Cone activation (marginal, sub-π/2) image aper. (rad) 1.422 1.469 image sat. % 10.9 25.8 t→i violation % 0.0 1.2
+0.05 +14.9 +1.2
1.446 15.6 0.0
1.491 32.8 0.8
+0.05 +17.2 +0.8
Behavior (scale-dependent) traversal mono. (INet) terminal collapse (INet) COCO t2i R@5 ZSC mean
−8.3 +98.5 −3.4 −2.0
15.2 10.0 51.5 38.0
17.4 7.0 52.9 37.2
+2.2 −3.0 +1.4 −0.8
Quantity
16.5 1.5 57.4 41.4
8.2 100 54.0 39.4
The result separates cleanly into three layers (Table 18). First, the operating geometry is unchanged: √ curvature stays pinned at the floor (c = 0.100), the dimensionless radius cρ holds at ≈ 0.20, the local distortion factor stays at 1.007, and no embedding enters the nonlocal regime—at either scale. Second, the cones become more active, but only marginally and without operational benefit: image-side aperture saturation rises (HyCoCLIP-B 10.9 → 25.8%, HyCoCLIP-S 15.6 → 32.8%) and the text→image violation rate ticks up from zero to ∼ 1%, yet the median image-side aperture stays below π/2 and the geometry regime is unmoved. Third, the behavioral effects are scale-dependent and, where they move, point the wrong way for hierarchy: at ViT-B, traversal monotonicity on ImageNet falls (16.5 → 8.2%) and terminal retrieval collapses (1.5 → 100%, i.e. every trajectory ends at a single caption), while downstream retrieval and zeroshot classification degrade (∆COCO t2i −3.4, ∆ZSC −2.0). At ViT-S the same metrics move slightly the other way (∆COCO +1.4, traversal +2.2). This scale-dependence mirrors the downstream pattern noted in Section 5.4. The interpretation is consistent with the gradient analysis. Lowering η activates the entailment term, but the entailment term itself contains the low-curvature shortcut (Section 7.3): activating it more strongly does not build radial hierarchy, it deepens the regularization pressure that keeps curvature at the floor. The natural threshold intervention thus behaves like a regularizer, not like a mechanism for installing hierarchy. 36
C.4
Intervention scope
Current-GRIT ViT-B interventions are three-seed structural comparisons (ViT-S and η interventions are single-seed), reported as seed means without a formal significance test. Their role is to test whether changing a geometric constraint changes the operating geometry, hierarchy diagnostics, or downstream behavior under matched training conditions, not to claim downstream improvements.
D
Full Traversal Results
D.1
Traversal summary across interventions
Table 19 aggregates the traversal diagnostics for MERU and HyCoCLIP—on the released checkpoints and on the current-GRIT baseline and curvature-unclamped runs across scales. All 200 trajectories collapse to a single shared terminal caption (from a 25,014-caption pool) in every released and current-GRIT setting. PHyCLIP is evaluated separately across all 64 product factors: the per-factor breakdown is in Table 21, and the aggregate product-distance result (monotonicity 8–19%, no perfect traversal) is in Section 5. Per-run current-GRIT values are in Table 20. Table 19: Traversal diagnostics show mode collapse rather than operational hierarchy, for the MERU and HyCoCLIP released checkpoints (the audit target) and the current-GRIT baselines and their curvatureunclamped variants. Terminal retrievals collapse to a single caption in every setting on the COCO val2017 pool (distinct terminals = 1/200). Monotonicity is the min–max range across the ImageNet, COCO, and Flickr30k pools (all rows, including clampOff). Released rows are each a single published checkpoint (n = 1; no seed ensemble exists), so the single-checkpoint caveat applies a fortiori.
D.2
Model
Setting
Norm monotonicity
Perfect traversal
Collapse signal
MERU-S MERU-S MERU-S MERU-B MERU-B MERU-B MERU-L
released baseline clampOff released baseline clampOff released
15.1–20.8% 13.7–15.7% 15.4–18.4% 13.6–17.7% 9.5–14.5% 13.1–17.2% 13.5–16.5%
0% 0% 0% 0% 0% 0% 0%
1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions
HyCoCLIP-S HyCoCLIP-S HyCoCLIP-S HyCoCLIP-B HyCoCLIP-B HyCoCLIP-B
released baseline clampOff released baseline clampOff
14.5–18.7% 14.2–18.0% 14.1–18.8% 14.9–21.1% 14.9–18.9% 12.2–18.1%
0% 0% 0% 0% 0% 0%
1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions
Traversal by retrieval pool
Table 20: Per-pool norm monotonicity for MERU and HyCoCLIP current-GRIT trained baselines (η = 1.2). Perfect traversal is zero in all evaluated settings. PHyCLIP is reported separately by factor in Table 21. Its aggregate product-distance monotonicity is reported in Section 5. Model ImageNet COCO Flickr30k Perfect Collapse signal MERU-S MERU-B HyCoCLIP-S HyCoCLIP-B
15.7% 12.6% 15.2% 16.5%
15.2% 14.5% 18.0% 18.9%
13.7% 9.5% 14.2% 14.9%
37
0% 0% 0% 0%
1/200 terminal captions 1/200 terminal captions 1/200 terminal captions 1/200 terminal captions
Table 21: PHyCLIP factor-wise traversal across 64 product factors, on the released PHyCLIP-B and PHyCLIP-L checkpoints (ViT-L is audited only at the released level). No factor exceeds chance-level monotonicity and no factor achieves a perfect traversal across all evaluated datasets. Model Dataset Mean mono. Range Factors > 0.5 Perfect factors PHyCLIP-B PHyCLIP-B PHyCLIP-B PHyCLIP-L PHyCLIP-L PHyCLIP-L
ImageNet COCO Flickr30k ImageNet COCO Flickr30k
0.214 0.256 0.251 0.213 0.253 0.247
[0.178, 0.258] [0.217, 0.307] [0.205, 0.290] [0.181, 0.259] [0.222, 0.292] [0.208, 0.298]
0/64 0/64 0/64 0/64 0/64 0/64
0/64 0/64 0/64 0/64 0/64 0/64
PHyCLIP factor-wise traversal is evaluated by isolating each product factor and repeating the same traversal protocol used for the aggregate representation. This test addresses the possibility that hierarchy is hidden in individual factors rather than visible in the full product distance. The result is negative across both PHyCLIP-B and PHyCLIP-L: every factor remains below the chance-level monotonicity reference, and no factor yields a perfect traversal.
D.3
Terminal retrieval collapse
We record terminal nearest-neighbor captions at the end of traversal trajectories. In representative MERU and HyCoCLIP settings, all 200 traversal trajectories collapse to a single shared terminal caption, drawn from a pool of 25,014 candidate captions (COCO val2017; distinct terminals = 1/200). This supports the quantitative monotonicity result: traversal does not provide an operational path from general to specific. To make this concrete, Tables 22 and 23 show step-by-step retrieved captions for representative trajectories on the released checkpoints and on the current-GRIT seed-42 runs (the Figure 4 checkpoints), respectively. Selection and screening rules are given in the captions. Both sources contained plausible-early trajectories, so no random fallback was drawn. Both settings show the same mirror pattern: early steps retrieve relevant paraphrases (source cosine 0.65–0.85)—the regime that published qualitative interpolation demos showcase— while continuation collapses to an arbitrary terminal hub shared by all 200 trajectories. The released and current-GRIT runs collapse identically. For MERU-B the released run even reaches the same terminal hub caption (“I do not know what this is supposed to be..”) as the current-GRIT run, underscoring that the collapse recurs across independently trained and released artifacts, not just a single run. The collapse is not specific to the COCO pool, and in particular not an artifact of evaluating on COCO rather than the data behind published interpolation demos. On the original grounded Flickr30k pool of Pal et al. (2025)—the dataset used by their qualitative interpolation figure, here re-encoded in full to 747,260 image, box-crop, and caption items—200 image→root traversals on the released HyCoCLIP-B reach only 13 distinct terminal captions, with 79% converging onto just two degenerate hubs (e.g. “hosiery”, “The”), and all 200 trajectories reaching [ROOT] (Table 24; step-by-step trajectories in Table 25). Both free-form caption pools therefore collapse—COCO to 1/200 and grounded Flickr30k to 13/200 (two hubs covering 79%)—whereas a structured relation pool retains diverse terminals (the native box→full NG2 walk yields 706–756/1000 distinct endpoints; Table 31). The collapse magnitude tracks how free-form the retrieval pool is, consistent with a high-dimensional retrieval-hubness effect rather than a pool-specific quirk—which is why terminal collapse is read as a pool-dependent symptom, with the load-bearing evidence being the controlled graded readouts on the native relation (Table 32; Section 5.3) and per-step monotonicity corroborating. 38
Table 22: Step-by-step retrieved captions on the released MERU-B and HyCoCLIP-B checkpoints, illustrating the early-plausible / late-hub-collapse pattern behind the 0% perfect traversal of Table 19. Selection rule: the first two COCO val2017 source images by index. Captions truncated to 80 characters and PII/junkscreened (0 redactions). “cos” is cosine to the source image. The onset row marks where the trajectory converges to the hub. Across all 200 source trajectories the terminal is a single shared hub caption. Model
Steps
Retrieved caption (≤ 80 chars)
Source A: COCO 000000465718.jpg (idx 0) MERU-B 0–4 This workstation features three desktop monitors with a single keyboard. . . 5–8 a laptop computer a keyboard and two monitors 9–17 Not the biggest workspace in the world, but it works 18–20 I do not know what this is supposed to be.. (onset/hub) HyCoCLIP-B
0–4 5 6 7–14 15–19 20
Two computer screens that are sitting on a desk. A desk has two computer monitors, a keyboard, and a laptop that is all connected. a desk with a monitor and a keyboard a desk with a computer a laptop and monitor a table that has some computers on it Picture of living room with modern furniture and decor (onset/hub)
Source B: COCO 000000014888.jpg (idx 1) MERU-B 0–2 A young calf drinks from its mother’s udders 3–6 A dairy cow is being milked by machine.. 7–8 A cow is being milked by a machine. 9–20 I do not know what this is supposed to be.. (onset/hub) HyCoCLIP-B
0–6 7–13 14–18 19 20
A dairy cow is being milked by machine.. A dairy cow suffering in confinement hooked up to a milking machine. a small calf nursing a cow in a pasture a number of people in a bod of water Picture of living room with modern furniture and decor (onset/hub)
Norm
Cos
0.571
0.790
0.496 0.396 0.384
0.783 0.749 0.679
0.382 0.369
0.839 0.842
0.344 0.327 0.320 0.318
0.848 0.849 0.827 0.532
0.613 0.525 0.488 0.384
0.752 0.742 0.727 0.682
0.391 0.365
0.828 0.826
0.329 0.318 0.318
0.730 0.568 0.417
Table 24: Terminal-retrieval collapse is pool-dependent. Distinct terminal captions for 200 image→root traversals on the released HyCoCLIP-B across two free-form caption pools, compared with the structured native box→full NG2 radial walk (1000 box/full pairs, Table 31). Free-form caption pools collapse onto one or two degenerate hubs. The structured relation pool retains diverse terminals. Collapse magnitude tracks pool free-formness—consistent with a high-dimensional retrieval-hubness effect rather than a geometry signal—so terminal collapse is a pool-dependent symptom, read alongside the load-bearing controlled graded readouts (Section 5.3, Table 32). Retrieval pool
Pool type
COCO val2017 grounded Flickr30k native NG2 box→full
free-form captions free-form captions structured relations
Items
Distinct terminals
25,014 747,260 1,000
1/200 13/200 706–756/1000
39
Top-hub share 100% 48% (top-2 79%) ≤2.9%
Table 23: Companion to Table 22 on the current-GRIT seed-42 checkpoints (the Figure 4 runs): step-bystep retrieved captions illustrating the same early-plausible / late-hub-collapse pattern. Selection rule: the first two COCO val2017 source images by index. Captions truncated to 80 characters and PII/junk-screened (0 redactions). “cos” is cosine to the source image. The onset row marks where the trajectory converges to the hub. Across all 200 source trajectories the terminal is a single shared hub caption. Model
Steps
Retrieved caption (≤ 80 chars)
Norm
Cos
0.595
0.766
0.564 0.474 0.456
0.737 0.543 0.485
0.359
0.837
0.323 0.318 0.306
0.836 0.768 0.433
Source B: COCO 000000014888.jpg (idx 1) MERU-B 0–1 A dirty white and black dog next to a bottle of soda. 2–8 Cow looking into the bottom of a machine. 9–11 A young calf drinks from its mother’s udders 12–20 I do not know what this is supposed to be.. (onset/hub)
0.623 0.602 0.560 0.456
0.706 0.704 0.674 0.513
HyCoCLIP-B
0.361
0.737
0.337 0.306
0.649 0.471
Source A: COCO 000000465718.jpg (idx 0) MERU-B 0–10 This workstation features three desktop monitors with a single keyboard. . . 11–13 A desk set up as a workstation with a laptop 14–15 I am unable to see an image above. 16–20 I do not know what this is supposed to be.. (onset/hub) HyCoCLIP-B
0–7 8–17 18 19–20
0–14 15 16–20
A desk with both a laptop computer and a desktop computer. a desk with a laptop and a desktop computer a laptop computer with a keyboard set upon it a big bowl with some mix inside of it (onset/hub)
A dairy cow suffering in confinement hooked up to a milking machine. a small calf nursing a cow in a pasture a big bowl with some mix inside of it (onset/hub)
Table 25: Step-by-step retrieved items on the released HyCoCLIP-B for image→root traversal over the original grounded Flickr30k pool (the data behind the published HyCoCLIP interpolation figure; 747,260 items), using 21 interpolation steps (0–20) as in Table 22. Selection rule: the first two full-image sources by index. Captions truncated to 80 characters. “Norm” is the retrieved item’s space-norm, “Cos” its cosine to the source image. The onset row marks convergence to the degenerate hub. Both trajectories pass from plausible related captions to the shared degenerate hub “The” before [ROOT]—the same early-plausible / late-hub-collapse pattern as on COCO. Steps
Retrieved item (≤ 80 chars)
Source A: grounded Flickr30k 3359636318.jpg (idx 517) 0–6 ⟨source image⟩ 7–8 People walk down a city street past a record store. 9 A picture of a storefront with a few people passing by. 10 passersby stare 11–12 Passersby interact 13–16 The (onset/hub) 17–20 [ROOT] Source B: grounded Flickr30k 6959556104.jpg (idx 527) 0–7 ⟨source image⟩ 8–9 Protesters with a sign reading “ASTI, Save Our Schools” march outside. 10–13 a protest 14–15 The (onset/hub) 16–20 [ROOT]
40
Norm
Cos
0.641 0.396 0.370 0.294 0.272 0.179 0.000
1.000 0.868 0.853 0.816 0.806 0.712 0.000
0.641 0.346
1.000 0.825
0.260 0.179 0.000
0.792 0.687 0.000
E
Taxonomy Mapping and Robustness
E.1
CIFAR-100 to WordNet mapping
Table 26: Manually disambiguated CIFAR-100 to WordNet synset mapping used in the main taxonomy analysis (Table 5). For 13 classes, the naive top-3-synset heuristic (described in Table 27) includes semantically incorrect senses—in several cases excluding the correct sense entirely—so we manually pin each to the single sense below. The remaining 87/100 CIFAR classes retain their naive candidate sets. CIFAR label
Manual synset
Sense / reason
seal ray turtle skunk whale dolphin sweet pepper plate maple tree oak tree palm tree pine tree willow tree
seal.n.09 ray.n.07 turtle.n.02 skunk.n.04 whale.n.02 dolphin.n.02 bell_pepper.n.02 plate.n.04 maple.n.02 oak.n.02 palm.n.03 pine.n.01 willow.n.01
marine mammal, not sealing wax cartilaginous fish, not light beam aquatic reptile, not sweater musteline mammal, not pejorative person cetacean, not fictional giant toothed whale, not dolphinfish bell pepper vegetable, not spice dinner dish, not baseball home plate Acer tree, not generic tree Quercus tree, not generic tree Palmae tree, not generic tree coniferous tree, not generic tree Salix tree, not generic tree
41
Table 27: Robustness of the taxonomy decomposition to the WordNet mapping choice. Naive last-token mapping (top-3 WordNet noun synsets of each label’s final token; minimum shortest-path distance over candidate pairs) vs. manually disambiguated mapping (Table 26) on the same 100 CIFAR classes. ViT-B current-GRIT rows are seed means over 0/37/42. Manual disambiguation raises Pearson r by 0.01–0.10 across all 22 models. The broad hyperbolic > Euclidean pattern is preserved and the gap widens slightly. 2 The angular/radial decomposition is unchanged under either mapping: ∆Rnorm ≤ 0.003 for all models, and Mantel permutation tests do not detect a significant norm contribution at α = 0.05 in any of the 44 model×mapping combinations (minimum pperm = 0.125). Per-seed, naive-mapping false positives occur at the expected rate (1 of 36 ViT-B seed×mapping cells), while the manual mapping is non-significant throughout. ∆r is computed from unrounded values (the rounded column difference may differ by ±0.001). Model
Setting
naive r
manual r
∆r
pperm (naive)
pperm (manual)
CLIP-S CLIP-B CLIP-L
released released released
0.359 0.358 0.336
0.377 0.370 0.363
+0.018 +0.012 +0.027
0.343 0.257 0.379
0.308 0.245 0.330
MERU-S MERU-B MERU-L HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
released released released released released released released
0.448 0.428 0.409 0.376 0.400 0.384 0.405
0.488 0.489 0.446 0.449 0.465 0.456 0.456
+0.040 +0.062 +0.037 +0.073 +0.065 +0.072 +0.051
0.125 0.978 0.467 0.694 0.677 0.682 0.910
0.724 0.340 0.977 0.531 0.501 0.735 0.499
MERU-S MERU-S HyCoCLIP-S HyCoCLIP-S PHyCLIP-S PHyCLIP-S
baseline clampOff baseline clampOff baseline clampOff
0.432 0.452 0.397 0.395 0.411 0.416
0.492 0.508 0.448 0.491 0.465 0.493
+0.060 +0.056 +0.051 +0.096 +0.054 +0.077
0.559 0.888 0.374 0.413 0.894 0.812
0.666 0.385 0.701 0.301 0.538 0.913
MERU-B MERU-B HyCoCLIP-B HyCoCLIP-B PHyCLIP-B PHyCLIP-B
baseline clampOff baseline clampOff baseline clampOff
0.398 0.422 0.397 0.432 0.395 0.416
0.459 0.485 0.462 0.508 0.471 0.479
+0.061 +0.063 +0.065 +0.076 +0.076 +0.064
0.506 0.660 0.325 0.480 0.404 0.642
0.366 0.362 0.437 0.436 0.425 0.676
42
E.2
Full taxonomy-distance correlation results
Table 28: CIFAR-100 WordNet path-distance correlation and its angular/radial decomposition across all evaluated models, using the manually disambiguated mapping (Table 26). Columns match Table 5: Pearson 2 2 2 2 2 r, the cosine-only Rcos , the norm-only Rnorm-only , the incremental ∆Rnorm = Rcos +norm − Rcos , and its Mantel pperm . Hyperbolic checkpoints obtain higher r than Euclidean CLIP across all sizes, but the radial 2 increment ∆Rnorm is non-significant (pperm > 0.05) for every model in every family and intervention: norm adds no detectable structure beyond cosine anywhere. The largest increment among the hyperbolic/box2 trained models (MERU-B baseline, ∆Rnorm = 0.0029) equals the Euclidean CLIP noise floor (CLIP-B, 2 0.0029). ViT-B current-GRIT rows are seed means over 0/37/42, over which ∆Rnorm scatters at noise level (MERU-B baseline, 0.0001–0.0054). Other rows are single checkpoints. The hyperbolic-vs-Euclidean r separation is preserved under the naive last-token mapping, with within-group rankings shifting (Table 27).
E.3
Model
Setting
Taxonomy r
2 Rcos
2 Rnorm-only
2 ∆Rnorm
pperm
CLIP-S CLIP-B CLIP-L
released released released
0.377 0.370 0.363
0.142 0.137 0.132
0.001 0.001 0.002
+0.0024 +0.0029 +0.0024
0.308 0.245 0.330
MERU-S MERU-B MERU-L HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
released released released released released released released
0.488 0.489 0.446 0.449 0.465 0.456 0.456
0.238 0.239 0.199 0.201 0.217 0.208 0.208
0.001 0.001 0.001 0.002 0.005 0.002 0.003
+0.0002 +0.0013 +0.0000 +0.0007 +0.0010 +0.0002 +0.0009
0.724 0.340 0.977 0.531 0.501 0.735 0.499
MERU-S MERU-S HyCoCLIP-S HyCoCLIP-S PHyCLIP-S PHyCLIP-S
baseline clampOff baseline clampOff baseline clampOff
0.492 0.508 0.448 0.491 0.465 0.493
0.242 0.258 0.200 0.241 0.216 0.243
0.002 0.000 0.000 0.000 0.003 0.002
+0.0003 +0.0013 +0.0002 +0.0019 +0.0007 +0.0000
0.666 0.385 0.701 0.301 0.538 0.913
MERU-B MERU-B HyCoCLIP-B HyCoCLIP-B PHyCLIP-B PHyCLIP-B
baseline clampOff baseline clampOff baseline clampOff
0.459 0.485 0.462 0.508 0.471 0.479
0.210 0.236 0.214 0.258 0.222 0.230
0.002 0.000 0.001 0.001 0.003 0.002
+0.0029 +0.0021 +0.0018 +0.0012 +0.0016 +0.0006
0.366 0.362 0.437 0.436 0.425 0.676
Radial parent-child ordering diagnostics
Table 29 reports the full radial parent-child ordering results, providing the numerical values behind Figure 3 and extending them to ViT-S and ViT-L and to the HyCoCLIP current-GRIT interventions. Directed pairs use the CIFAR-100 fine→coarse label hierarchy, not the WordNet synset mapping in Table 26, which is reserved for the Section 6.1 taxonomy-distance correlation. No model shows a stable positive excursion above its shuffle null at any scale or under curvature unclamping: no three-seed mean crosses the threshold, and the single-seed crossings that occur are seed-unstable—the sign of z flips across the three current-GRIT seeds (+0.70, −1.73, −1.97 for HyCoCLIP baseline) despite comparable downstream performance—consistent with a seed-unstable near-null statistic rather than a stable learned feature. The ViT-S intervention rows are single-seed and exploratory. Model-level intervention conclusions rest on the three-seed ViT-B conditions. In particular, the single largest positive excursion in the table (PHyCLIP-S clampOff, z = +2.73) is one such unreplicated ViT-S cell.
43
Table 29: Full radial parent-child ordering diagnostics. “Radial consistency” is the fraction of 100 directed CIFAR-100 fine-to-superclass pairs with child norm > parent norm. Under the radial-depth hypothesis this exceeds the shuffle null (z = (real − null)/s.d. over nperm = 10,000 permutations), read one-sided against the |z| ≥ 1.6 threshold (calibrated in Appendix F.1; per-test false-positive rate ≈ 6%). Individual seeds cross ±1.6 in both directions at this rate, but none survives multiple-comparison correction (Bonferroni for the ≈ 24 current-GRIT runs is |z| ≈ 2.9, reached by no cell). Tellingly, the Euclidean CLIP-B baseline, which encodes no radial hierarchy, itself crosses (z = +1.73), confirming that a crossing reflects the near-unit-norm marginal artifact rather than radial depth. No model passes the operative graded test: under the controlled geodesic readout the largest correlation belongs to the Euclidean baseline, and strict monotonicity is 0% (Tables 32, 31)—so a directed bit, where present, is at most directionally significant, never traversable.
Model
Setting
Scale
Radial cons.
Shuffle null
z
Interpretation
Euclidean CLIP Euclidean CLIP Euclidean CLIP MERU MERU MERU HyCoCLIP HyCoCLIP PHyCLIP PHyCLIP
released released released released released released released released released released
ViT-S ViT-B ViT-L ViT-S ViT-B ViT-L ViT-S ViT-B ViT-B ViT-L
35% 22% 10% 63% 61% 62% 56% 35% 42% 18%
35.5 ± 2.1 18.0 ± 2.3 9.5 ± 1.6 62.1 ± 3.5 61.0 ± 3.7 62.3 ± 3.5 55.9 ± 2.8 37.1 ± 3.3 44.2 ± 3.3 22.9 ± 3.1
−0.25 +1.73 +0.32 +0.25 −0.00 −0.09 +0.04 −0.64 −0.67 −1.59
near chance above null near chance near chance near chance near chance near chance near chance near chance near chance
MERU MERU MERU MERU MERU MERU MERU MERU HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP HyCoCLIP PHyCLIP PHyCLIP PHyCLIP PHyCLIP PHyCLIP PHyCLIP PHyCLIP PHyCLIP
baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42) baseline (s0) clampOff (s0) baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42) baseline (s0) clampOff (s0) baseline (s0) baseline (s37) baseline (s42) clampOff (s0) clampOff (s37) clampOff (s42) baseline (s0) clampOff (s0)
ViT-B ViT-B ViT-B ViT-B ViT-B ViT-B ViT-S ViT-S ViT-B ViT-B ViT-B ViT-B ViT-B ViT-B ViT-S ViT-S ViT-B ViT-B ViT-B ViT-B ViT-B ViT-B ViT-S ViT-S
82% 84% 85% 80% 82% 83% 77% 69% 38% 34% 22% 74% 67% 63% 47% 59% 17% 17% 24% 63% 62% 64% 41% 69%
78.1 ± 3.0 77.1 ± 3.5 83.1 ± 3.0 74.9 ± 3.4 79.6 ± 3.1 80.5 ± 3.5 73.6 ± 3.4 69.5 ± 3.5 35.9 ± 3.1 39.4 ± 3.1 28.2 ± 3.2 75.8 ± 3.5 70.9 ± 3.4 61.1 ± 3.7 48.9 ± 3.3 62.4 ± 3.5 23.7 ± 2.9 24.2 ± 3.6 29.9 ± 3.1 58.5 ± 2.9 56.6 ± 3.3 62.8 ± 3.3 39.0 ± 3.3 59.5 ± 3.5
+1.31 +1.99 +0.66 +1.50 +0.77 +0.72 +0.98 −0.14 +0.70 −1.73 −1.97 −0.52 −1.17 +0.51 −0.58 −0.97 −2.34 −2.00 −1.93 +1.57 +1.65 +0.37 +0.62 +2.73
near chance above null near chance near chance near chance near chance near chance near chance near chance below null below null near chance near chance near chance near chance near chance below null below null below null near chance above null near chance near chance above null
44
E.4
Mantel permutation and prompt robustness
We use Mantel-style permutations because pairwise class distances are not independent. For the native directed test we additionally report a length-matched subset and a length-residualized variant (Section 5.1, Appendix E.5). Across those controls the norm contribution beyond cosine remains negligible. E.5
Model-native directed radial test (GRIT box-caption)
Table 30 reports the native-relation radial test of Section 5.1, with the same pairing (full-image caption = more specific, predicted larger norm; box captions = more general, predicted smaller), estimand, and shufflenull design. Column definitions are in the caption. Because the native box→caption relation has no tree metric, the decomposition regresses a binary same-sample-pair indicator (y=1 iff a (full, box) pair belongs to 2 the same GRIT sample) on the same cosine-distance and norm-difference predictors as Section 6.1. ∆Rnorm 2 is the norm term’s incremental R over cosine. All directed z are positive (consistent with the hypothesized direction), but the cosine-controlled increment is negligible on released checkpoints and small even under clampOff. Angular distance is the dominant predictor throughout. Table 30: Model-native GRIT box-caption directed radial test, with the current-GRIT clampOff rows reported per training seed (0/37/42; released checkpoints “—”). Directed consistency (ρfull > ρbox ) vs. sample-level shuffle null; z on the full set and the length-matched subset (|∆tok| ≤ 3, n=157); incremental 2 beyond cosine with its Mantel p. The directed z measures the one-bit full>box norm contribution ∆Rnorm ordering (a shuffle-controlled residual, not a graded-axis detection) and is seed-robust. The graded-axis 2 , seed-unstable for HyCoCLIP-B clampOff and small for PHyCLIP-B clampOff (per-seed evidence is ∆Rnorm 2 ≈ 0.013. CLIP-B significance in the p column). Every value stays far below the detection threshold ∆Rnorm is the Euclidean, non-box-trained baseline. The directed z is computed against the re-pairing shuffle null (which preserves the marginal norm distributions), not the 50% precheck reference, so a model can show z > 0 while its raw ρfull > ρbox fraction is below 50%—as for CLIP-B (precheck 46.6% yet z = +2.97), whose 2 = 0.00001 confirms no usable radial structure. ∆Rnorm Model
seed
z (full, n=1000)
z (len-matched)
2 ∆Rnorm
Mantel p
HyCoCLIP-B released HyCoCLIP-B clampOff
— 0 37 42 — 0 37 42 — —
+9.71 +9.66 +9.82 +10.00 +8.41 +7.63 +8.27 +7.88 +9.35 +2.97
+12.31 +13.00 +12.84 +13.63 +12.81 +11.83 +10.78 +11.55 +11.06 +4.40
0.00015 0.00170 0.00225 0.00000 0.00000 0.00024 0.00035 0.00016 0.00040 0.00001
< 0.001 < 0.001 < 0.001 0.285 0.41 < 0.001 < 0.001 < 0.001 < 0.001 0.02
PHyCLIP-B released PHyCLIP-B clampOff
MERU-B released CLIP-B (Euclidean)
Traversal along the native radial direction. Table 31 reports the operational counterpart to the directed test: we traverse each model’s representation along the same radial (norm) coordinate the NG2 ordering occupies, interpolating from the general (box) end outward to the specific (full) end (10 steps, n = 1000 pairs), and ask whether retrieved captions move monotonically from general to specific (column definitions in the caption). Strict monotonicity is 0% for every box-trained model despite the strong directed z-scores in Table 30: the native radial direction is directionally significant but not traversable. Terminal diversity is high (706–756 distinct endpoints out of 1000, top-1 share ≤ 2.9%), so the failure is non-monotonic, many-to-diverse mapping rather than collapse onto a few hubs—the radial walk neither orders captions by specificity nor funnels them to a single attractor. Endpoint cosine is higher under clampOff (0.92 vs. 0.72– 0.76 released), but monotonicity remains 0%—endpoint proximity is not step-wise traversability. CLIP-B fails the direction precheck (46.6%, at chance) and is not traversed. 45
Table 31: Traversal along the native NG2 radial direction. Direction precheck = fraction with ρfull > ρbox ; strict monotonicity is the fraction of trajectories increasing in specificity at every one of the 10 steps (the 50% figure is the per-step direction chance; the all-steps random baseline is far lower, so 0% is well below chance either way); endpoint cosine with the specific target F ; terminal diversity as distinct terminals / 1000 and the most-common-terminal share. Strong directed z (Table 30) coexists with 0% monotonicity: the direction is significant but not walkable. CLIP-B fails the precheck and is not traversed (–). Model HyCoCLIP-B released PHyCLIP-B released MERU-B released HyCoCLIP-B clampOff PHyCLIP-B clampOff CLIP-B (Euclidean)
Dir. precheck
Strict mono.
Endpoint cos(·, F )
Distinct term.
Top-1 share
91.4% 92.2% 77.7% 92.0% 91.4% 46.6%
0% 0% 0% 0% 0% –
0.72 0.76 0.72 0.92 0.92 –
756/1000 755/1000 724/1000 706/1000 743/1000 –
1.0% 1.2% 1.7% 2.4% 2.9% –
Graded readouts and controls. Strict monotonicity is a knife-edge statistic, so we also report a graded step-vs-specificity rank correlation ρ (Spearman of retrieved-caption specificity against step index, K = 10, n = 1000; specificity proxy cos(retrieved, F ) − cos(retrieved, B)) under two interpolation readouts, with a shuffle-step null band (200 reshuffles, 95%) and two controls (Table 32). Under the radial readout (angular component fixed at uF , norm interpolated) cosine retrieval is invariant to norm scaling, so the same caption is returned at every step up to ties and ρ ≈ 0 (mean |ρ| ≤ 0.017, residual being tie-breaking noise)— formalizing why the radial axis is not exposed by any cosine-based retrieval. Under the geodesic readout (linear xB → xF in the ambient embedding, which approximates the hyperbolic geodesic in the nearEuclidean regime H(u) ≈ 1) ρ is positive (0.56–0.65 for the hyperbolic models), but two controls show this reflects generic ambient interpolation, not hyperbolic hierarchy. (i) A mismatched-target control (xB → xF ′ , F ′ a different sample’s full caption) stays well above the null band (ρ = 0.35–0.58), so endpoint identity is not the dominant source. The true-pair excess ∆pair = ρtrue − ρmismatched is only 0.13–0.22, smaller than the generic mismatched component. (ii) The Euclidean CLIP-B baseline—with no usable native radial direction—produces the largest geodesic ρ of all seven models (0.71), above every box-trained hyperbolic checkpoint. Across all three readouts the NG2 direction is therefore non-operative: ≈ 0 under radial, positive-but-generic under geodesic, and never strictly monotonic.
Table 32: Graded NG2 traversal under two readouts, with controls. ρ is the mean per-pair Spearman correlation of retrieved-caption specificity against step index. Radial readout is ≈ 0 (cosine is norm-invariant up to ties). Geodesic ρ is positive but the true-vs-mismatched excess ∆pair is small, and the Euclidean CLIPB baseline (no usable native radial direction) attains the largest geodesic ρ—so geodesic positivity is generic to ambient interpolation, not a hyperbolic-hierarchy signature. All geodesic and mismatched values lie above the shuffle-step null band. The radial readout was computed for the three released ViT-B checkpoints and the HyCoCLIP-B and PHyCLIP-B clampOff variants. “–” marks models where it was not run (MERU-B clampOff), and CLIP-B is omitted from the radial column (fails the precheck). ρtrue , ρmism , and ∆pair are rounded independently to two decimals, so the displayed ∆pair may differ from the column difference by ±0.01. Model Radial ρ Geodesic ρtrue Geodesic ρmism ∆pair CLIP-B (Euclidean) MERU-B released MERU-B clampOff HyCoCLIP-B clampOff PHyCLIP-B released PHyCLIP-B clampOff HyCoCLIP-B released
– +0.006 – +0.017 +0.003 −0.010 −0.013
0.71 0.65 0.60 0.57 0.56 0.56 0.56
46
0.58 0.46 0.46 0.35 0.37 0.36 0.38
0.13 0.18 0.14 0.22 0.19 0.21 0.18
F
Diagnostic Positive Controls
We include a synthetic positive control to verify that our diagnostics can detect hierarchy when radial and pair-specific structure is present. We construct a balanced tree with branching factor 4 and depth 6, yielding 5,461 nodes and 5,460 parent–child edges. Nodes are embedded in the Poincaré disk with radius increasing with depth, and angular sectors are assigned so that each child lies inside its parent’s sector. Table 33 shows that the diagnostics recover the planted hierarchy. Parent–child radial ordering, depth– radius correlation, and chain monotonicity are all perfect in the synthetic tree. When radii are shuffled, radial ordering collapses to chance-level behavior and the depth–radius correlation disappears. Similarly, the planted angular-sector containment signal is nearly perfect, whereas angle-shuffled controls remove the pair-specific sector signal. The radial shuffle null is high in this construction because shuffled children remain marginally deeper than parents. This illustrates the same point as our main directed-norm diagnostics: raw radial order alone is insufficient, and pair-specific structure should be evaluated against a shuffle-null gap. This sanity check verifies only the narrower claim needed for this paper: the radial and shuffle-controlled diagnostics can detect the specific radial/sector hierarchy mechanism that hyperbolic VLMs claim to instantiate. Table 33: Positive-control diagnostic sanity check. The same diagnostics recover planted radial and angular hierarchy in a synthetic hyperbolic tree, while radius- and angle-shuffled controls collapse the corresponding signals. Pair gap is real radial order minus its shuffled-pairing null. Sector gap is real sector containment minus its shuffled-pairing null. Embedding Synthetic radial tree Radius-shuffled control Angle-shuffled control
F.1
Radial order
Pair gap
Depth–radius Spearman
Chain mono.
Sector gap
2 ∆Rnorm
1.000 0.498 –
0.200 – –
1.000 -0.008 –
1.000 0.528 –
0.996 – 0.000
0.232 – –
Sensitivity of the shuffle-controlled diagnostics
The control above establishes non-blindness at a single, perfectly ordered point. To answer the stronger question—what is the smallest signal the diagnostics would have detected—we extend it to a graded series and report a minimum-detectable-effect (MDE). Both diagnostics test a shuffle-controlled excess, not a raw rate: for the radial test, ∆pair = Orderreal − Eshuffle [Ordershuffle ]; for the taxonomy test, the incremental 2 ∆Rnorm beyond cosine. A purely marginal effect—a uniform shift of all child norms, or a norm signal that merely re-encodes the coarse partition already carried by cosine—lies in the null space of these statistics by construction and yields zero excess. This is a desirable property, not a limitation: a global norm offset carries no information about which child belongs to which parent, and the radial mechanism is a pair-specific claim (a child sits deeper than its own parent). Radial directed ordering (Part A). We plant a pair-specific radial signal using matched edge margins on the same 100-pair, 20-parent design and the same coarse-permutation null as Section 5.1: each child is placed a margin m above its own parent’s anchor radius, with additive noise, and m is swept while the noise level is tuned so the signal-zero detection rate matches the nominal α (δ=0 power ≈ 0.05–0.10). Detection power (|z| ≥ 1.6, R≥200 draws) reaches 80% once the planted shuffle-controlled excess reaches ∆pair ≈ 13 pp. As a consistency check, the normal approximation from the observed null SDs (≈ 2.6–3.7 pp) predicts an analytic 80% MDE of ≈ 8–9 pp. The simulated value is slightly larger, i.e. the simulation is conservative. The curve is non-monotonic at very large m: once every child exceeds every parent, the real and shuffled rates both approach 100% and ∆pair returns toward zero—directly visualizing why a marginal mean shift is undetectable. 2 Taxonomy radial contribution (Part B). We repeat the analysis for ∆Rnorm using the real CIFAR100 cosine distances and WordNet tree distances, so the Mantel null dispersion is inherited from the data rather than assumed (median real null SD ≈ 0.0028, matched by the synthetic background to within 7%).
47
A radial contribution that is correlated with tree distance beyond what cosine already explains is detected 2 at 80% power once it reaches ∆Rnorm ≈ 0.013. The largest seed-mean increment across all audited models (MERU-B baseline, 0.0029, at the Euclidean CLIP noise floor; per-seed values reach 0.0054, still 2.4× below; Table 28) is more than four times below this threshold and non-significant at every seed. The same setup also makes the mechanism explicit: a norm signal that re-encodes only the coarse taxonomic partition saturates 2 at ∆Rnorm ≈ 0.0017 (z < 0.2) at any planted strength, because the real cosine distances already capture 2 that partition (Rcos ≈ 0.16). The near-zero observed increment is therefore consistent with norm–cosine redundancy, not with an absence of norm structure in general. Interpretation. Across both diagnostics, “near chance” means the audited models lack detectable pairspecific radial structure beyond the angular taxonomy—at a sensitivity quantified above, which the observed values fall below by roughly 2.6–4.5× (the radial test by ≈ 2.6×, worst case PHyCLIP-L 5pp vs. the 13pp MDE; the taxonomy test by ≈ 4.5×). The tests do not, and are not intended to, rule out a purely marginal norm separation between all fine and all coarse prompts, which can arise from prompt wording or level marginals rather than learned parent–child relations. They rule out, at the stated sensitivity, a shufflesurviving pair-specific radial alignment beyond those marginals.
G
Multi-Granularity Retrieval Details
G.1
Query construction
For each WordNet depth level, we construct queries from internal ImageNet ancestors that have at least two descendant leaves and at most 500 descendant leaves. Only levels with at least one qualifying ancestor appear in Table 35 (the odd depths 1–13 for the ImageNet leaf set). Query features are computed by averaging descendant leaf class embeddings. Averaging shrinks norms, so coarse centroid queries mechanically sit nearer the origin, and the multi-granularity comparison is accordingly read as a supervision probe, not a radialgeometry readout. Retrieval is performed over ImageNet validation images, and we report mean average precision over queries at each depth. Table 34: Multi-granularity retrieval separates supervision effects from hyperbolic geometry, and the separation is scale-robust. Across ViT-S/B/L, the hyperbolic-but-non-box MERU stays close to Euclidean CLIP at coarse depths (gap ≤ 0.02 at every scale), while the box/compositional models HyCoCLIP and PHyCLIP improve substantially (+0.14–0.15 coarse AP over CLIP). Coarse = mean of depths 1, 3, 5; Fine = mean of depths 11, 13 (unweighted, computed from unrounded depth-wise AP); depth-wise AP in Table 35. No released HyCoCLIP-L or PHyCLIP-S checkpoint exists. Model Coarse (d ≤ 5) Fine (d ≥ 11) CLIP-S CLIP-B CLIP-L
0.160 0.194 0.194
0.529 0.567 0.575
MERU-S MERU-B MERU-L
0.180 0.179 0.203
0.534 0.579 0.583
HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
0.314 0.333 0.338 0.346
0.567 0.602 0.603 0.631
48
Table 35: Depth-wise multi-granularity retrieval AP across released ViT-S/B/L. Depths 1/3/5/7/9/11/13 have 2/9/21/83/54/46/25 qualifying queries (up to 50 sampled per depth). Queries are internal WordNet ancestors of ImageNet leaves. Retrieval is over ImageNet validation images. Model 1 3 5 7 9 11 13 CLIP-S CLIP-B CLIP-L
0.045 0.073 0.057
0.169 0.210 0.224
0.265 0.300 0.303
0.221 0.245 0.255
0.448 0.472 0.488
0.568 0.609 0.635
0.491 0.525 0.514
MERU-S MERU-B MERU-L
0.063 0.048 0.070
0.182 0.186 0.227
0.294 0.304 0.311
0.236 0.262 0.268
0.464 0.497 0.511
0.575 0.626 0.637
0.493 0.532 0.530
HyCoCLIP-S HyCoCLIP-B PHyCLIP-B PHyCLIP-L
0.116 0.124 0.114 0.128
0.396 0.417 0.441 0.438
0.431 0.459 0.460 0.470
0.346 0.379 0.365 0.391
0.550 0.582 0.571 0.600
0.585 0.628 0.615 0.656
0.549 0.576 0.591 0.606
H
Depth-Supervision Diagnostics
H.1
Pairwise depth ranking has no direct curvature gradient
For a pairwise norm-ranking loss Ldepth = max(0, m − (ρc − ρp )), the curvature parameter c does not appear explicitly. Therefore ∂Ldepth = 0. ∂c This explains why pairwise depth ranking cannot identify curvature by itself. H.2
Full c-gradient diagnostic and multi-seed collapse-phase signs
The per-loss curvature-gradient trajectory is reported in the main text (Figure 4). We do not duplicate it here. The single-batch c-only implementation check below complements it by isolating the depth-loss path. Multi-seed collapse-phase sign pattern. Table 36 reports the collapse-phase sign analysis for three seeds (42, 37, 23) of the curvature-collapse probe in each model. The collapse phase is defined per run as the probe steps before c reaches the curvature floor (c ≤ 0.1001). Probes are taken every 250 steps. Across seeds the floor is reached at 7750 steps (MERU-B, identical in all three seeds) and 7250–7500 steps (HyCoCLIP-B), and the entailment gradient reaches its minimum within one to four probe steps of the floor. MERU-B carries no box-level depth term (depth column “none”) yet exhibits the same 100% entailment-downward collapsephase sign pattern in all three seeds, indicating the curvature shortcut does not depend on depth supervision. The contrastive gradient is consistently c-upward only while it carries signal (100% of high-signal steps, c > 0.5, for four of six runs; 85% for the two lowest-signal MERU seeds) and decays to sign-unstable noise (∼10−3 , c-up fraction 56–74%) once c floors—matching its framing as the vanishing, non-load-bearing term.
49
Table 36: Collapse-phase per-loss gradient signs across three seeds (curvature-collapse probe, ViT-B). Sign convention: positive ∂L/∂ log c = c-down. “contr. c-up (all / c>0.5)” splits all collapse-phase steps from the high-signal sub-phase where the contrastive gradient is above its ∼10−3 noise floor. These signs are measured on the full-objective (λe =0.2) collapse trajectory and are not a general property of the contrastive objective. Seed Steps Floor Entail c-down Contr. c-up (all / >0.5) Entail>contr. mag. Depth term
Model
MERU-B 42 MERU-B 37 MERU-B 23 HyCoCLIP-B 42 HyCoCLIP-B 37 HyCoCLIP-B 23
30 30 30 29 29 28
7750 7750 7750 7500 7500 7250
100% 100% 100% 100% 100% 100%
70% / 100% 70% / 85% 67% / 85% 79% / 100% 90% / 100% 82% / 100%
100% 96.7% 96.7% 100% 100% 100%
none none none yes yes yes
Table 37: The depth-loss (c-only) path, measured at two scales. ViT-S: deterministic eight-draw probe on the HyCoCLIP-S baseline checkpoint (16-sample GRIT batches, eval mode, all RNGs fixed per draw, gradients on). ViT-B: per-loss gradients logged on the HyCoCLIP-B collapse-probe trajectories (three seeds; collapse-phase means; the same runs as Figure 4 and Table 36). Gradients are with respect to log curvature. A positive gradient induces a c-down update. The per-loss gradients sum to the objective-weighted total (entailment scaled by λe = 0.2, depth by λdepth ), so the decomposition is complete. At both scales the entailment term is c-down and overwhelms the depth term—by 89–4500× (ViT-S; range and median of the eight per-draw ratios) and ∼300× (ViT-B). The depth term is c-up on the baseline checkpoint but flips to c-down on about a third of collapse-phase steps at ViT-B (batch dependence; Section 7.3). Either way it is far too small to counteract the shortcut. Absolute magnitudes vary across draws, so the probe is read for sign and relative scale, not as a population-level magnitude estimate. Per-seed ViT-S values and batch manifests, and ViT-B per-step gradients, are in the reproducibility bundle (raw_json/c_only_impl_check.json, raw_json/gradient_probe/).
I
Quantity
ViT-S (8 draws, baseline)
ViT-B (3 seeds, collapse)
Entailment gradient sign |∂Le /∂ log c| (raw entailment) Depth gradient sign |∂Ldepth /∂ log c| (raw depth) |entailment| / |depth|
8/8 c-down 0.315–0.551, med. 0.315 8/8 c-up 1.2×10−4 –3.5×10−3 , med. 2.3×10−3 89–4500×, med. 157×
c-down, 100% steps 0.35–0.38 (mean) net c-up; 34–37% steps c-down 1.1–1.3×10−3 (mean) 289–330×
Reproducibility Notes
All diagnostics in this paper are produced by an audit suite provided as supplementary material: the measurement scripts (geometry, cones, radial and native directed tests, traversal, taxonomy decomposition, and the curvature-gradient probe), the raw per-run JSON outputs behind every diagnostic table, and √ convention documents that trace, per model family, the radial coordinate ρ, the curvature transform to cρ, and the cone-aperture formula to the released models’ own Lorentz operations. The release also records SHA-256 hashes of every released checkpoint and all from-scratch current-GRIT runs (including the λe =0 ablations), the current-GRIT training configurations and final curvature values, a load-sanity check reproducing the reported reference metrics, and toy analytic-versus-code checks for the radius and aperture computations. Evaluation sets (ImageNet/COCO for per-modality geometry, GRIT shards 00000–00001 for cones and native diagnostics) are specified in each table caption. I.1
Lorentz fp32 numerical patch
During reproduction we found that the Lorentz operations in our training code did not cast to fp32, causing numerical instability—contrastive loss collapse—under fp16 (AMP) training. The instability was most acute in PHyCLIP’s product-space batched operations (*_batch in phyclip/lorentz.py), but we added fp32 casts uniformly to all Lorentz operations, including the scalar-curvature versions used by MERU and HyCoCLIP. 50
This affects the reproducibility of the from-scratch current-GRIT training runs across all three families but is not used as primary evidence for our hierarchy claims. I.2
MERU current-GRIT configuration compatibility
MERU ViT-B checkpoints trained on one server saved a configuration containing the keyword argument curv_min. We added backward-compatible optional keyword handling for curv_min and curv_max during evaluation. This change affects configuration loading only and does not alter model computations. I.3
ATMG implementation caveat
We do not include ATMG as a primary audited family. ATMG is an angle-based approach and is complementary to our diagnosis, but we could not unambiguously match the publicly released implementation to the paper’s stated objective, so audit findings on the public checkpoints could not be attributed to the method as published. We therefore discuss ATMG as related work rather than as a direct checkpoint audit target.
51