LatentMD: Benchmarking Markdown Boundary Failures in LLM-Generated Text∗ †
arXiv:2609.06993v1 [cs.SE] 7 Sep 2026
Sungjune Lee Seoul National University [email protected]
Myungjoo Kang Seoul National University [email protected]
Abstract Large language models (LLMs) increasingly generate Markdown that is consumed by renderers, agents, code extractors, and structured downstream pipelines. Yet existing evaluations often conflate content quality with format adherence, leaving Markdown boundary failures under-measured. We introduce LatentMD, a benchmark and evaluation protocol for diagnosing CommonMark-level fence-boundary failures in LLM-generated Markdown. LatentMD separates content correctness from boundary correctness, enabling detection of outputs that are content-correct but boundary-broken. The benchmark contains 4,179 prompts and a CLI for scoring arbitrary model outputs. Across 9 LLMs and roughly 37,600 generations, we find that Markdown boundary failures are widespread: 38.0% of valid maingrid outputs are content-correct but boundary-broken, with substantial boundary breakage under unspecified prompts and in a small human-authored validation set. Ablations show that failures are driven primarily by same-family symmetricdelimiter collisions rather than nesting alone, are only partially mitigated by prompt hints, and generalize to Python triple-quote docstrings while JSON remains robust as an asymmetric-delimiter control. LatentMD provides a reproducible diagnostic target for parser-sensitive LLM evaluation.
1
Introduction
LLM Markdown output is increasingly read by automated tools, not just human eyes — chat UIs render it on screen, agentic systems parse it to decide actions, and search-and-extraction systems pull code blocks out of it for reuse [20]. These tools rely on the implicit assumption that the output’s structure is correct; when the structure breaks, the tools break with it. This breaks regularly in the wild. A user asks a chat model to write a Markdown tutorial about Python generators with two code examples; the model wraps its reply in a fenced block tagged markdown and embeds the two Python examples as inner fenced blocks. In many chat UIs and developer tools, this output visibly breaks: the inner fences close the outer wrapper at the first length match, and the trailing content leaks outside the intended block as raw text or stripped fence characters. In other cases the breakage is invisible at render time but still corrupts any tool that reads the document’s spec-level structure. This is not a contrived test case: the same fence-boundary failure has been filed across major LLM and developer tools, including OpenAI ChatKit, Google Gemini CLI, Microsoft VS Code, and open-source agentic projects [19, 11, 16, 2, 6, 17]. These incidents show that nested-fence breakage is a recurring deployment failure rather than a benchmark artefact. ∗ Code and scoring tool: https://anonymous.4open.science/r/LatentMD-code-B51D. †
Corresponding author.
Preprint.
Yet existing LLM-output benchmarks do not measure this failure. They score content correctness and surface format adherence [24, 5, 8, 25, 14], but none isolate whether the model can correctly track nested fence boundaries. A generation that is content-correct but boundary-broken would therefore be scored as passing by these benchmarks. The same output would still corrupt any specstrict downstream parser or renderer. We close this gap with LatentMD, a 4,179-prompt rule-based benchmark that adds an explicit Boundary axis orthogonal to Content. This isolates a content-correct but boundary-broken failure mode, the latent failure. It supports paired causal and trend claims that aggregate-only benchmarks cannot. Contributions. We make the following contributions. (a) Benchmark. We introduce LatentMD, the first systematic benchmark of LLM Markdown fence-boundary failures, comprising 4,179 prompts that systematically vary outer wrapping, inner nesting, and delimiter properties. (b) Diagnostic methodology. We define an error-type taxonomy (premature closure, unclosed fence, latent delimiter collision) and a four-cell decomposition that separates content correctness from CommonMarklevel boundary correctness, surfacing latent failures missed by content/format-only evaluations. (c) Empirical study. Across 9 LLMs we show that Markdown boundary failures are widespread, driven primarily by same-family symmetric-delimiter collision, only partially mitigated by prompting, and echoed in Python but not JSON. (d) Reproducible artifact. We release prompts, generation scripts, and a stand-alone scoring tool for applying LatentMD to arbitrary model outputs.
2
Background & Related Work
Existing structured-output benchmarks Table 1: Failure mode coverage. ✓: directly measured; measure adjacent but non-overlapping partial: indirectly detectable, no cause-level diagnosis; properties. StructEval [24] and concept: precursor without implementation; –: not covered. StructTest [4] score AST-level struc- Per-cell: App. A.4. tural matching across JSON, XML, and Markdown; MDEval [5] evaluFailure Type StructEval FormatBias IFEval LatentMD ates Markdown content quality; FormatE1 Premature Closure partial – – ✓ Bias [8] argues for separating format E2 Unclosed Fence partial – – ✓ E3 Latent Collision – – – ✓ from content; IFEval [25] and FollowF1 Instr. Compliance – – partial ✓ Bench [14] measure format-instruction F2 Fence Balance partial – – ✓ compliance; FMBench [22] benchmarks F3 Safe Length – – – ✓ format-following; STED [21] addresses 4-cell Decomposition – concept – ✓ structured-output reliability. None diLatent Failure Detection – – – ✓ rectly tests fence-balance or safe-length selection inside the same delimiter family — the precise mechanism that creates latent failures. Formal-language work [7] has shown that autoregressive models struggle with balanced-bracket tracking on synthetic formal-language tasks. Our benchmark extends this observation to LLM Markdown generation, where nested fences impose the same balanced-bracket structure under realistic deployment conditions. Table 1 gives the per-failure-type coverage matrix.
3
Benchmark Design
Three Roles of Symmetric Delimiters. In Markdown fenced code blocks, the same character sequence (triple backticks ``` or tildes ~~~) serves three distinct roles within a single generation: (1) Wrapper — the outer fence enclosing the entire response; (2) Content — inner fences that form the document’s structure (e.g., code blocks within a tutorial); (3) Literal — fence characters appearing as verbatim text (e.g., when explaining Markdown syntax). When an autoregressive model fails to disambiguate these three roles during generation, the output incurs a boundary-state tracking failure, violating the delimiter nesting rules of CommonMark Section 4.5 [15]. The benchmark axes below stratify generations by how each role is constrained. Operational Basis. We adopt CommonMark Spec 0.31.2 [15] as the operational reference for boundary failure; this also applies to GFM [10], which inherits CommonMark’s fenced-code rules unchanged. The spec serves as an externally specified reference, ensuring boundary-failure detection is reproducible rather than renderer-dependent. 2
3.1
Experimental Axes
The benchmark is organized around three orthogonal axes, each isolating a distinct source of boundarystate tracking failure. The A-axis manipulates whether the model must itself produce an outer fence, isolating wrapper-induced collision. The B-axis varies the inner nesting load, isolating literal-fence interference inside the response. The D-axis ablates delimiter properties (family, length, cross-family) to identify which property of symmetric delimiters drives the failure. A visual overview of all three axes is in Appendix B.1. A-axis (Outer Constraint) controls wrapper pressure. A1 (Direct) forbids an outer fence, giving a no-wrapper baseline for failures that do not require wrapping. A2 (Wrapped) requires a full-response fenced block, serving as the stress condition for wrapper-induced collision. A3 (Unspecified) gives no wrapper instruction and measures realistic base-rate self-wrapping. A1 is a negated instruction; LLM under-compliance [13] is isolated in Table 5 (Step 3). B-axis (Inner Nesting) controls same-character delimiter load. B1 (Single) has one code block and no nesting pressure. B2 (Nested Single) adds one literal Markdown rendering of that block, forcing content-vs-literal fence disambiguation; B3 (Nested Multiple) repeats this with two or more literal fence examples, increasing opportunities for premature closure. B4 (Mixed Elements) adds tables, blockquotes, or lists as non-fence distractors while retaining code-block structure. D-axis (Delimiter Ablations) isolates which delimiter property drives failure. D-Family swaps the outer delimiter between backticks and tildes; D-Outer-Length fixes the outer length to 3, 4, or 5; D-Inner-Run fixes an inner same-family run of length 3, 4, or 5 to test a-priori F3 safe-length prediction. D-Cross-Family uses different outer/inner delimiter families, removing same-family collision by design and distinguishing collision from generic nesting. Full A×B Crossing. Each prompt instantiates one A condition and one B condition. The 3 A levels and 4 B levels are fully crossed, giving 12 A×B cells per (TASK, LANG) pair. This full crossing supports within-task paired comparisons across treatments (§3.2). 3.2
Slot-Based Prompt Generation
Each prompt is constructed from a structural template whose two slots, LANG and TASK, are filled from external sources. Drawing slots from an external dataset (McEval [3]) rather than model-generated content ensures the test distribution is independent of the evaluated model. LANG comprises 9 programming languages (popularity consensus, Appendix B.3) and 3 structured formats (Markdown, JSON, HTML), all restricted to languages with McEval task coverage. TASK provides 15 task instances per LANG (180 in total), mechanically extracted from McEval instruction fields by stratified sampling that selects 5 easy, 5 middle, and 5 hard tasks under a fixed random seed. The benchmark totals 4,179 unique prompts. The breakdown is: 2,160 from the main A×B grid, 1,440 from D-axis delimiter ablations, 432 from hint ablations, 27 from inline code-span probes, 90 from cross-format probes, and 30 human-authored natural-prompt validation prompts. Released Artifact. All prompts, templates, and the generation script are released as part Write a Markdown document that addresses the following task. Include one <LANG> of the artifact. We additionally release a standcode example. In the document, also show the raw Markdown source for that code alone command-line tool that scores any model example, including the opening fence, on the benchmark with the metrics defined in §4. the language tag, the code, and the closing fence. Wrap your entire response The tool ingests a JSONL file of (prompt_id, in a fenced markdown code block. response) pairs and emits per-record labels and Task: aggregate failure rates with Wilson 95% confi<TASK> dence intervals. External users can therefore score arbitrary models without re-implementing the metrics. Figure 1 shows one instantiated Figure 1: Slot-based prompt template (A2×B2). prompt to anchor the slot-template structure; the Per-condition examples: App. B.4. slots <LANG> and <TASK> are filled per (LANG, TASK) pair. 3
4
Evaluation Metrics
We score each response along two orthogonal axes: boundary correctness (fence balance and safe outer-fence length) and content correctness (lightweight structural checks ensuring the response is a meaningful document). Let r ∈ Σ∗ denote a response and p ∈ Pmain a prompt — a combination of experimental axis settings and slot values, with full domain definitions in Appendix A.1. Predicates that do not apply to a given prompt return NA (Not Available), which is ignored in the boolean conjunctions used below (x ∧ NA = x). 4.1
Format Metrics (F1–F3)
F1 (Instruction Compliance) measures whether the model follows the A-axis wrapping directive: pass for A1 if the response does not contain an outer wrapper, pass for A2 if it does, NA for A3; it extends prior instruction-compliance metrics [25] to the wrapper presence/absence case. F2 (Fence Balance) checks CommonMark [15] fence balance with a stack discipline: opening fences push their family and length, and only bare same-family fences of at least the opening length can close the current block; the response passes iff the stack is empty at EOF, and the check is implemented on top of markdown-it-py [9] with full pseudocode in Appendix A.2. F3 (Safe Length) checks whether the outer fence is long enough to survive any same-family run inside its content (NA when no outer fence exists), reported in two variants: a-priori (used in the D-inner-run ablation, requiring an outer fence one character longer than the specified inner run) measures predictive compliance, and post-hoc (used in the main A×B analysis, with safe length recomputed from realized inner content) measures internal consistency. 4.2
Content Metrics (C1–C3)
Three content metrics ensure the response is a meaningful document rather than structurally empty output, in the spirit of FormatBias [8] and prior structural checks [24, 5]. C1 (Heading Structure) passes when the response contains at least two headings. C2 (Code Language Match) passes when at least one code block’s info string matches the requested LANG (with alias map; fails if no code blocks are present). C3 (Block Compliance) is a B-condition-aware check that the response contains the expected number of literal-syntax fenced blocks: zero for B1 and B4 (no nested literal fences), one for B2 (single nested fence), and two for B3 (multiple nested fences). 4.3
Boundary Violation Error Types (E0–E3)
Every generation receives one of four mutually exclusive and exhaustive error labels — E0 (Correct), E1 (Premature Closure), E2 (Unclosed Fence), or E3 (Latent Collision) — with conditions and interpretations given in Table 2, and a minimal worked example for each failure type shown in Figure 2. E1 fires when F2 fails and an interior bare same-family fence of length at least the outer-fence length is present; E2 fires when F2 fails without such an interior bare match.
Table 2: Boundary violation error types (E0–E3). Label E0 (Correct) E1 (Premature Closure) E2 (Unclosed Fence) E3 (Latent Collision)
Condition
Interpretation
F2(r) ∧ F3(r) ¬F2(r) ∧ ∃ bare match ¬F2(r) ∧ ∄ bare match F2(r) ∧ ¬F3(r)
boundary-correct outer terminates early at a matching bare fence outer stays open at EOF (no matching bare fence) length collision masked by balanced match
Per the CommonMark spec, a closing fence must contain only fence characters plus optional trailing whitespace, so a fence with an info string (e.g., ```python) can never close an outer fence; premature closure arises when the model writes bare triple backticks (```) inside an inner block whose length matches the outer. When the response wraps an outer fence, the structural fix in all three failure cases is to satisfy F3, i.e., to make the outer fence strictly longer than the longest inner same-family run. 4
1 ```markdown 2 # Tutorial 3 ``` 4 Leaked text 5 ```
(a) E1 source # Tutorial Leaked text
(d) E1 rendered
1 ```markdown 2 # Tutorial 3 ```python 4 print() 5 ``` 6 ```
1 ```markdown 2 # Tutorial 3 ```python 4 print() 5 End of doc.
(b) E2 source # Tutorial ```python print() End of doc.
(c) E3 source # Tutorial ```python print()
(e) E2 rendered
(f) E3 rendered
Figure 2: Minimal source examples for E1/E2/E3 under the LatentMD boundary classifier. Top row: source. Bottom row: how it renders to a reader. (a, d) E1 (Premature Closure): the bare fence at line 3 matches the outer and prematurely closes it; line 5 opens a new bare-fence block that never closes, so F2 fails with an interior bare match. (b, e) E2 (Unclosed Fence): the outer fence never closes, so all content stays trapped; F2 fails without an interior bare match. (c, f) E3 (Latent Collision): inner and outer fences nest cleanly so F2 passes, yet the inner ```python shares the outer’s 3-backtick length, so F3 fails. 4.4
Four-Cell Decomposition and Latent Failure
Every generation is decomposed along two orthogonal axes. Formally: Boundary(r) := F2(r) ∧ F3(r), Content(r, p) := C1(r) ∧ C2(r, p) ∧ C3(r, p).
(1) (2)
The four cells (Figure 3) are mutually excluBoundary ✓ Boundary × sive and exhaustive over valid generations. The Both correct Format-only Content starred cell — outputs that are content-correct Spec-OK, on-topic Bdry OK, content fail ✓ but boundary-wrong — is the latent failure, de⋆ Latent failure Both wrong Content fined formally in Eq. (3) below. Prior contentContent OK, bdry broken Bdry + content fail × only or format-only benchmarks judge such outputs as passing, even though they violate the Figure 3: Four-cell decomposition. spec at the boundary level. Latent failures may render plausibly in lenient chat interfaces yet corrupt downstream consumers (agents, code extractors, structured parsers) that depend on spec-compliant boundary parsing; Appendix A.5 shows a concrete instance, and Appendix A.3 details which Markdown constructs are in scope. For a model M and a prompt subset P ′ ⊆ P, let V ⊆ P ′ denote the valid subset — prompts whose response was not externally truncated by hitting the token limit (see §4.5). We define latent failure and the latent failure rate (LFR) as LF(r, p) := ¬Boundary(r) ∧ Content(r, p), X [ M (P ′ ) = 1 LFR 1{LF(M (p), p)}, |V|
(3) (4)
p∈V
reported with the Wilson 95% score interval [23]. This is the headline metric our benchmark surfaces, and that the research questions in §5 ultimately decompose. 4.5
Validity Filtering
Each model in our evaluation reports a finish_reason field on every generation indicating why decoding stopped. We exclude generations whose finish_reason = length (the model hit the max_tokens limit), as these are token-truncation artefacts rather than complete free-form outputs. No generation in our data triggered the safety content_filter flag, so this case did not arise empirically. Valid N (after exclusion) is the denominator for all reported rates. Per-model and per-LANG truncation statistics are in Appendix D.1. 5
5
Experiments
We evaluate 9 LLMs on LatentMD, following a mechanism-to-prevalence arc: trigger (wrapping × nesting), cause (same-family collision), mitigation limits (safe-length planning and hints), crossformat generalization, and latent-failure distribution. We structure the results around 5 RQs. (RQ1) How does explicit outer-container wrapping interact with inner nesting complexity to trigger boundarystate tracking failures? (RQ2) Are boundary failures driven by same-family delimiter collision rather than nesting complexity alone? (RQ3) Can LLMs select a safe outer-fence length when innerdelimiter pressure is explicit, and do prompt-level mitigations close the gap? (RQ4) Does the boundary-tracking failure generalize beyond Markdown fenced blocks to other structured-output formats? (RQ5) How prevalent are latent failures, and how are they distributed across models, tiers, and languages? Section 5.5 validates A1/A3/A2 stratification and human-authored prompts. 5.1
Setup
Large
Mid
Small
We evaluate 9 models spanning 3 tiers Table 3: Models evaluated. “Open”: open-weight (Small, Mid, Large/API) and both open- (vLLM/HF-served); “Closed”: closed-source API. Tiers weight and closed-source API categories reflect open-weight parameter count; closed-source (Table 3). Open-weight models are served counts are unavailable (Large/API group). locally via vLLM with HuggingFace Transformers as fallback for CUDA-constrained Tier Model Provider Params Type nodes; closed-source models — GPTQwen3-7B Alibaba 7B Open 4o [18], Claude-Sonnet-4 [1], Gemini-2.5Gemma-2-9B Google 9B Open Flash [12] — are accessed through their Llama-3.1-8B Meta 8B Open APIs. We use a common decoding configQwen3-32B Alibaba 32B Open uration wherever supported: temperature= Gemma-3-27B Google 27B Open 0 for greedy decoding, top_p= 0.9, and Llama-3.1-70B Meta 70B Open max_tokens= 7168 (the largest completion GPT-4o OpenAI – Closed budget compatible with Gemma-2-9B’s 8K Gemini-2.5-Flash Google – Closed context window after prompt overhead). Claude-Sonnet-4 Anthropic – Closed Open-weight runs fix seed= 42; reasoning modes are disabled so outputs are pure Markdown; provider-specific top_k and disable flags are in Appendix C.1. The full benchmark generates 4,179 unique prompts × 9 models ≈ 37,600 generations; we record finish_reason for truncation filtering (§4.5, Appendix D.1). Main results report Wilson 95% confidence intervals (CIs) over valid N (Appendix E.2) as descriptive binomial uncertainty over prompt instances. 5.2
Wrapping × Nesting Triggers Boundary Failure (RQ1)
6
E3 (Latent Collision) Rate (%)
A-axis (outer container)
Figure 4 shows the pooled E3 pattern across the main 100 A×B grid. We order A by experimental severity 8.9% 11.7% 19.3% A1 (Direct) 17.8% 80 rather than condition ID: A1 (no-wrap baseline), A3 60 (unspecified, realistic anchor), and A2 (forced-wrap 15.1% 25.7% 37.8% A3 (Unspecified) 38.6% 40 stress test). In the pooled E3 view, all A1 cells remain below 20%. A3 (realistic anchor) self-wraps 37% 20 48.7% 60.9% 72.4% A2 (Wrapped) 62.7% of responses (Step 1 of Table 5) and pools to 40–55% 0 boundary failure across B conditions (Appendix G.5). B1 B2 B3 B4 B-axis (inner nesting complexity) Total boundary failure under A2 reaches 100% in many model×B cells (Wilson 95% CI [97.9, 100], Figure 4: E3 (latent collision) rate across n=180 prompts per cell; Appendix E.1), confirmA×B, averaged over 9 models and LANG. ing that A2 is a stress condition for testing whether Full per-model total boundary-failure rates wrapper-induced collision can drive failure to a ceil(E1 ∪ E2 ∪ E3) in Appendix E.1. ing, not a deployment prevalence estimate. The E3 mass concentrates on the A2 row (49–72%) and peaks at A2×B4, showing that wrapping itself triggers same-family delimiter collision even when inner content is structurally simple. The lower E3 rate at A2×B2 should not be read as robustness: Appendix G.5 shows substantial mass shifts into E1/E2 while total boundary failure remains high. Together, outer wrapping is the dominant trigger, while inner nesting modulates the E1/E2/E3 failure mix; the full A×B×E breakdown is in Appendix G.5.
5.3
Same-Family Collision, Not Nesting Alone, Drives Failure (RQ2)
If boundary failure were caused primarily by Table 4: D-axis ablation: boundary-correct rate by generic nesting complexity, changing delimiter delimiter condition. Boundary Correct = F2 ∧ F3; family while preserving nesting structure should higher is better. Per-group max in bold. not yield a selective improvement; if failures are driven by same-family symmetric-delimiter Condition Value Bdry OK (%) F2 (%) F3 (%) collision, manipulations that reduce or remove Backtick 7.2 70.0 3.8 Family same-family collision should help most. SwitchTilde 37.0 83.2 36.7 ing the outer fence from backtick to tilde raises outer= 3 10.3 69.8 7.0 Outer length outer= 4 32.4 69.5 30.7 the boundary-correct rate from 7.2% to 37.0% outer= 5 28.4 62.5 25.9 (Table 4 Family rows). Explicit outer-length inBt outer 22.2 73.3 20.3 structions raise the boundary-correct rate from Cross-Family Ti outer 35.3 74.0 35.3 10.3% at outer=3 to 32.4% at outer=4, with diminishing returns at outer=5 (28.4%). The most direct diagnostic contrast comes from cross-family nesting: using different delimiter families for outer and inner fences structurally removes same-family collision while preserving nesting, raising the boundary-correct rate to 22.2% for backtick-outer/tildeinner and 35.3% for tilde-outer/backtick-inner. Residual failures remain high, indicating that delimiter collision is not the only failure source; nevertheless, the selective improvement supports same-family symmetric collision as a dominant driver rather than generic nesting alone. 5.4
Safe-Length Selection Failure and Prompt Mitigations (RQ3)
Fraction of responses (%)
The diagnosis of why mitigation is necessary in the outer=3 outer=4 outer=5 outer 6 first place lies in the safe-length selection failure: 100 even when the prompt explicitly states the innerfence run length (D-inner-run ablation, n=144 per 80 cell), 8 of 9 models achieve negligible F3 a-priori pass rates, with most outputs still using an unsafe 60 3-backtick outer fence; only GPT-4o adapts non40 trivially (run=3: 13.2%, run=5: 41.7%). Figure 5 shows the per-model distribution of outer20 fence backtick lengths chosen. Stacked colors are red (outer= 3, fails all three inner-run conditions), 0 r=3 r=4 r=5 r=3 r=4 r=5 r=3 r=4 r=5 r=3 r=4 r=5 orange (outer= 4), light green (outer= 5), dark Llama-3.1-70B GPT-4o Gemini-2.5-FlashClaude-Sonnet-4 green (outer ≥ 6); within each model, three subbars correspond to inner-run conditions r= 3, 4, 5. Figure 5: Outer-fence length distribution for Among these, Llama-3.1-70B, GPT-4o, and Gemini- the 4 models with visible outer-length varia2.5-Flash partially respond by matching the inner tion; full 9-model table in Appendix E.3. length (outer = run) but rarely exceed it, so these matched-length responses still fail the strict > criterion; Claude-Sonnet-4 shows minor variation but little systematic safe-length planning. The remaining 5 models (omitted from the figure) use outer = 3 across nearly all runs, with F3 pass rates ≤ 0.7%. Three prompting interventions then test whether the failure can be mitigated without delimiter changes: H1 (boundary-awareness rule: “outer fences must use more characters than inner fences”), H2 (one-shot example of correct nesting), H3 (length-specific instruction: “use at least 4 backticks for the outermost fence”). H3 is the strongest mitigation, dropping mean failure by 30.9 percentage points (96% → 65%); H1 helps moderately (−23pp); H2 (one-shot) is the weakest (−8pp). H3 helps because it supplies a concrete length cue, but it does not solve the general safe-length problem: “at least 4” remains unsafe whenever inner same-family runs are length 4 or longer. The per-model breakdown (Figure 8, Appendix E.4) reveals strong model-level heterogeneity: the two Mid models barely respond to any hint, while the Large/API models split, with some dropping below 10% under H3. Thus, prompt hints reduce but do not close the safe-length gap.
7
5.5
Manipulation Check and Natural Prompt Validation
Manipulation check. The A-axis design assumes A1, Table 5: A1 four-step decomposition (9A3, and A2 stratify wrapper pressure and boundary- model average). Low Step 3 + high Step 4 failure risk (baseline < realistic < stress). We verify isolates wrapper-induced boundary failure. this with a four-step decomposition (Table 5): Step 1 Per-model: App. F.1. A3 wrap rate 36.9% (models self-wrap roughly one third of the time without any instruction); Step 2 A1 Step Metric Description Rate non-compliance 19.0% (responses that ignored “do not 1 A3 wrap rate Base-rate wrap (no instr.) 36.9% wrap”); Step 3 compliance-conditional A1 residual 2 A1 non-comp. “Do NOT wrap” ignored (F1) 19.0% 3 A1 residual Obeyed no-wrap; bdry fail 8.1% 8.1% (boundary failure among the responses that did 4 A2 bdry fail Boundary fail under wrap 95.5% obey “do not wrap” — i.e. the failure that remains when no outer fence exists); Step 4 A2 boundary failure 95.5%. The 8.1% → 95.5% gap (Step 3 vs Step 4) confirms the manipulation is well-separated and the A2 ceiling is not a generic instruction-following artefact. Two models (Gemma-3-27B, Llama-3.1-70B) exhibit non-trivial Step 3 residual (34.5%, 40.7%), indicating a heterogeneous sub-population whose failures are not fully explained by outer wrapping alone. Paired McNemar evidence (Appendix I.4) confirms the A2 increase is not due to task composition. The 30 human-authored prompts (Appendix D.3) provide a second ecological check: high-elicitation prompts reach up to 56% failure, broadly overlapping the A3 realistic-anchor range (§5.2); the small set serves as a sanity check, not a population-level prevalence estimate. 5.6
Failure Generalizes to Python but Not JSON (RQ4)
To test whether symmetric-delimiter boundary failure is specific to Markdown or a more general property of LLM generation, we replicate the design in two contrasting target formats: Python (symmetric """ triple-quote docstrings) and JSON (asymmetric {} brace nesting). JSON achieves 100% parse-validity across every model and every nesting depth (asymmetric-delimiter control), while Python triple-quote conditions vary widely: the three Python conditions escalate triple-quote pressure (K1: literal """ inside docstring; K2: both """+''' styles; K3: nested """), and JSON’s K1/K2/K3 vary brace-nesting depth (1–2, 3–4, or ≥5 levels). The tier-aggregated view (Table 6) inverts Table 6: Cross-format parse-validity (%) per tier (Smallthe usual “larger is safer” expectation: un- /Mid/Large/API = 3/2/4 models). Per-model: App. J.1. der Python K1 the Small tier averages 97.8% pass rate, the Mid tier 56.7%, and Python JSON K1 K2 K3 K1 K2 K3 Tier the Large/API tier only 31.7%, with the K1 drop driven primarily by GPT-4o and Small 97.8 93.3 91.1 100.0 100.0 100.0 Mid 56.7 73.3 66.7 100.0 100.0 100.0 Gemini-2.5-Flash, both hitting 0% as litLarge/API 31.7 68.3 55.0 100.0 100.0 100.0 eral delimiter placement terminates the docstring prematurely. An analogous symmetric-delimiter boundary failure observed in Markdown thus reappears under Python triple-quote, indicating that scale or frontier-model status alone does not eliminate boundary-tracking failures; the per-model breakdown (Appendix J.1) shows that exact failure modes and model ordering are format- and condition-dependent. Inline code spans provide a related but noisier delimiter-handling probe: pooled run=1/2 failures are high, but Appendix J.2 shows the aggregate mixes interpretation defaults with genuine boundary-collision errors, so we treat it only as complementary evidence. 5.7
38% of Generations Are Latent Failures (RQ5)
The four-cell decomposition reveals that 38.0% of main-grid generations are latent failures — contentcorrect outputs with spec-level boundary violations invisible to prior format-adherence benchmarks (Table 7) — and among content-correct generations 57.2% fail boundary metrics (38.0/66.4), so any scoring that ignores fence-level boundary state misses this entire mass. Per-model variation is large: Gemma-3-27B pairs the highest latent failure (73.8%) with high content correctness (76.0%), while Gemma-2-9B has only 1.4% latent failure but the lowest content correctness (15.5%), so its low latent rate reflects content failure rather than boundary success. The per-language stratification (Appendix G.7) exposes a counterintuitive ordering — code-heavy languages (Shell 46.2%, C# 45.1%, C++ 42.9%) carry the highest latent failure rate while Markdown carries the lowest (9.2%), an artefact of Markdown’s lowest content correctness (21.6%) pre-empting latent classification rather than a genuine boundary success. The Cochran–Armitage trend test shows a significant population8
Model
Valid N Content Correct (%) Latent Failure (%)
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
2,141 2,157 2,146
69.3 [67.3, 71.2] 15.5 [14.0, 17.1] 48.2 [46.1, 50.3]
48.1 [46.0, 50.2] 1.4 [ 1.0, 2.0] 14.5 [13.1, 16.1]
Mid
Qwen3-32B Gemma-3-27B
2,152 2,158
71.6 [69.7, 73.5] 76.0 [74.2, 77.8]
46.1 [44.0, 48.2] 73.8 [71.9, 75.6]
Large/API
Table 7: Per-model content correctness and latent failure rates with Wilson 95% CIs (valid N on main A×B grid, post-truncation). Content Correct = C1–C3 all pass; Latent Failure = content-correct ∧ ¬(F2 ∧ F3) (either fence balance or safe length fails). Each cell: rate [lower, upper].
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
2,159 2,160 1,895 2,160
69.1 [67.1, 71.0] 80.0 [78.3, 81.7] 76.1 [74.1, 78.0] 93.0 [91.9, 94.0]
35.7 [33.7, 37.8] 43.3 [41.3, 45.4] 49.9 [47.6, 52.1] 30.6 [28.7, 32.6]
Overall
19,128
66.4 [65.7, 67.1]
38.0 [37.3, 38.7]
level monotonic trend in E3 (latent collision) rate with inner nesting complexity (pooled Z = +6.17, p < 10−3 ; Appendix I.2). Together, these results establish latent failure as a prevalent (38.0%), heterogeneously distributed (extremes 1.4% to 73.8%), language-pervasive, and dose-responsive failure mode that cannot be detected by content-only or format-only scoring. Robustness to content-metric choice. The headline Latent Failure rate is 38.0% under the primary Content definition (C1–C3) versus 33.1% under Full-C (adding C4 keyword-based Task Completion), with every model showing a lower but directionally consistent rate under Full-C; the finding does not invert under stricter content gating (Appendix F.2).
6
Limitations and Future Work
LatentMD is a diagnostic benchmark for Markdown fence-boundary failures, not a universal ranking of LLM reliability. Its boundary judgments are defined against CommonMark 0.31.2, but deployment impact depends on the downstream parser or renderer: some failures may be harmless in lenient chat interfaces while corrupting spec-strict tools such as agents, code extractors, or parsers. The benchmark focuses primarily on Markdown fenced code blocks, with smaller cross-format probes for Python triple-quote docstrings and JSON; our natural-prompt validation set is small and serves as an ecological sanity check rather than a field prevalence estimate. Future work should extend this boundary-evaluation framework to other structured-text formats (LATEX verbatim environments, YAML block scalars, HTML <pre> regions), expand the natural-prompt component with larger and more diverse user-derived prompts, and use LatentMD to track model progress and evaluate delimiter-aware generation or decoding-time mitigations.
7
Conclusion
We presented LatentMD, a benchmark and evaluation protocol for diagnosing Markdown boundary failures in LLM-generated outputs. By separating content correctness from CommonMark-level boundary correctness, LatentMD reveals failures that can be missed by content-only or surface format evaluations. Across 9 LLMs, we find that boundary failures are widespread, driven primarily by same-family symmetric-delimiter collisions, and only partially mitigated by prompting; 38.0% of main-grid generations are content-correct but boundary-broken. An analogous failure appears in Python triple-quote docstrings while JSON remains robust as an asymmetric-delimiter control; scale or frontier-model status alone does not eliminate symmetric-delimiter boundary-tracking failures. These results suggest that structured-text evaluation should measure parser-level boundary correctness, not only visible rendering, instruction compliance, or content quality. LatentMD provides a reproducible protocol, prompt suite, diagnostic taxonomy, and scoring tool for evaluating whether LLM-generated Markdown is safe for downstream tools that depend on spec-compliant structure. 9
References [1] Anthropic. Claude Sonnet 4. https://www.anthropic.com/news/claude-4, 2025. Accessed: 2026-04-01. [2] Block. Desktop markdown rendering bug: fenced code blocks can break message layout. https://github.com/block/goose/issues/8290, 2026. Goose AI agent issue. [3] Linzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin, Ke Jin, Jiaheng Liu, Tao Sun, Ge Zhang, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang, Xianjie Wu, Bing Wang, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang, and Zhoujun Li. McEval: Massively multilingual code evaluation. In The Thirteenth International Conference on Learning Representations (ICLR), 2025. URL https://openreview.net/forum?id=UunCPtPOlZ. [4] Hailin Chen, Fangkai Jiao, Mathieu Ravaut, Nawshad Farruque, Xuan Phi Nguyen, Chengwei Qin, Manan Dey, Bosheng Ding, Caiming Xiong, Shafiq Joty, and Yingbo Zhou. StructTest: Benchmarking LLMs’ reasoning through compositional structured outputs. arXiv preprint arXiv:2412.18011, 2024. [5] Zhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang, Xingyan Chen, Yu Zhao, and Fuji Ren. MDEval: Evaluating and enhancing markdown awareness in large language models. In Proceedings of the ACM on Web Conference 2025 (WWW ’25), pages 2981–2991. Association for Computing Machinery, 2025. doi: 10.1145/3696410.3714674. [6] Continue. When LLM generate markdown with code start with three backquote, usually it will be broken. https://github.com/continuedev/continue/issues/6059, 2025. Continue IDE extension issue. [7] Grégoire Délétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Lars Buesing, Elliot Catt, Marcus Catt, Tim Mattern, Marcus Hutter, Shane Legg, and Pedro A. Ortega. Neural networks and the Chomsky hierarchy. In International Conference on Learning Representations (ICLR), 2023. [8] Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, and Min-Yen Kan. LLMs are biased towards output formats! systematically evaluating and mitigating output format bias of LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 299–330. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.naacl-long.15. [9] ExecutableBookProject. markdown-it-py: Markdown parser in Python. https://github. com/executablebooks/markdown-it-py, 2024. Version 4.0.0 (released 2024-08-11); CommonMark 0.31.2 compliant. [10] GitHub. GitHub Flavored Markdown spec. https://github.github.com/gfm/, 2019. Accessed: 2026-04-01. [11] Google. CLI Markdown rendering is broken for triple-backtick code blocks. https://github. com/google-gemini/gemini-cli/issues/10515, 2025. Google Gemini CLI issue. [12] Google DeepMind. Gemini 2.5 flash. https://deepmind.google/technologies/ gemini/flash/, 2025. Accessed: 2026-04-01. [13] Joel Jang, Seonghyeon Ye, and Minjoon Seo. Can large language models truly understand prompts? A case study with negated prompts. In Proceedings of the 1st Transfer Learning for Natural Language Processing Workshop, PMLR 203, 2023. [14] Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang. FollowBench: A multi-level fine-grained constraints following benchmark for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 4667–4688, 2024. [15] John MacFarlane et al. CommonMark spec version 0.31.2. https://spec.commonmark. org/0.31.2/, 2024. Accessed: 2026-04-01. 10
[16] Microsoft. create_file tool wraps content in extra backtick fences when file contains markdown code blocks. https://github.com/microsoft/vscode/issues/295126, 2026. VS Code Copilot Chat issue. [17] Open WebUI. Incorrect rendering of nested code blocks in Markdown. https://github. com/open-webui/open-webui/issues/5016, 2024. Open WebUI issue. [18] OpenAI. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/, 2024. Accessed: 2026-04-01. [19] OpenAI. Nested triple-backtick code blocks break when generating Markdown files in responses. https://github.com/openai/chatkit-js/issues/89, 2025. OpenAI official rendering library issue. [20] Jiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang, Weipeng Chen, and Ji-Rong Wen. HtmlRAG: HTML is better than plain text for modeling retrieved knowledge in RAG systems. In Proceedings of the ACM on Web Conference 2025 (WWW ’25), pages 1733–1746. Association for Computing Machinery, 2025. doi: 10.1145/3696410.3714546. [21] Guanghui Wang, Jinze Yu, Xing Zhang, Dayuan Jiang, Yin Song, Peiyang He, Xuefeng Liu, and Tomal Deb. STED and consistency scoring: A framework for evaluating LLM structured output reliability. In NeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM), 2025. URL https://openreview.net/forum?id=rSCV1hTZvF. [22] Yaoting Wang, Yun Zhou, and Henghui Ding. FMBench: Adaptive large language model output formatting. arXiv preprint arXiv:2602.06384, 2026. [23] Edwin B. Wilson. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158):209–212, 1927. [24] Jialin Yang, Dongfu Jiang, Tony He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, and Wenhu Chen. StructEval: Benchmarking LLMs’ capabilities to generate structural outputs. Transactions on Machine Learning Research, 2026. URL https://openreview.net/forum?id=buDwV7LUA7. [25] Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023.
11
A
Metric and Algorithm Details
A.1
Format Metrics: Full Definitions
Notation, recap. Σ∗ is the response (Markdown text) space; Pmain = A × B × L × T is the main prompt space (§4), and a prompt p = (a, b, ℓ, t) ∈ Pmain is a tuple of an A-axis condition a, a B-axis condition b, a LANG ℓ, and a TASK t. The four slot domains are: A = {A1, A2, A3} (A-axis levels, §3.1); B = {B1, B2, B3, B4} (B-axis levels, §3.1); L contains the 12 LANG values (nine programming languages plus three structured formats, §3.2); and T contains the 15 TASK instances per LANG (180 in total, §3.2). A model M : P → Σ∗ under greedy decoding maps each prompt to a single response. We use the indicator 1{·} ∈ {0, 1}, the boolean conjunction ∧ extended to {0, 1, NA} via x ∧ NA := x (so NA is the identity), and the negation ¬ defined on {0, 1} only. Format Metric Definitions. For r ∈ Σ∗ , p = (a, b, ℓ, t) ∈ Pmain : 1{¬HasOuterFence(r)} if a = A1, F1(r, p) = 1{HasOuterFence(r)} if a = A2, NA if a = A3. F2(r) = 1{ S TACK BALANCE(r) = ∅ } (CommonMark Spec, Section 4.5) 1{ outerLen(r) > maxinner run(r) } if HasOuterFence(r), F3post (r) = NA otherwise.
(5) (6) (7)
F3apr (r, p) = 1{ outerLen(r) > innerRun(p) } (D-inner-run only) (8) Here S TACK BALANCE is the parser pseudocoded in Appendix A.2, returning the residual stack contents at end-of-text; outerLen(r) is the length of the outermost opening fence; run(r) scans inner content for the longest same-family backtick/tilde run. Content Metric Definitions. C1(r) = 1{ |{h ∈ r : h is an ATX or Setext heading}| ≥ 2 } (9) C2(r, p) = 1{ ∃ code block b ∈ r : infoStr(b) ∈ alias(ℓ) } (10) C3(r, p) = 1{ nliteral (r, p) ≥ k(b-condition of p) } (11) where nliteral (r, p) counts fenced blocks that serve as literal Markdown-source examples under the B-condition templates; k(·) encodes the prompt-specified literal-block count — k(B1) = k(B4) = 0 (no inner block), k(B2) = 1 (single inner block), k(B3) = 2 (multiple inner blocks) — and alias(ℓ) is the LANG-alias map released with the artefact. C2 fails when the response contains no code blocks (matching the released scoring CLI). Cells, Indicator Functions, and Aggregate Estimators. The four-cell decomposition (§4.4) defines four indicators on (r, p): both_correct(r, p) := Boundary(r) ∧ Content(r, p) (12) LF(r, p) := ¬Boundary(r) ∧ Content(r, p) (latent failure) (13) format_only(r, p) := Boundary(r) ∧ ¬Content(r, p) (14) both_wrong(r, p) := ¬Boundary(r) ∧ ¬Content(r, p) (15) For Boundary := F2 ∧ F3post and Content := C1 ∧ C2 ∧ C3 with the NA-as-identity convention applied to F3post when no outer fence is present (the only NA case in the framework), the four indicators sum to 1 on every valid response (mutual exclusion and exhaustiveness). Aggregate P ′ [ M (P ′ ) = |V|−1 rates are LFR p∈V LF(M (p), p), with V ⊆ P the valid (non-truncated) subset; analogous estimators apply to the other three cells. Deleted Metrics. The initial metric design included four additional format metrics that were removed during the final framework consolidation: F3-old (CommonMark Parse Success) was subsumed by the pipeline pre-filter; F4-old (AST Structural Match) was replaced by descriptive block-count statistics; F5-old (Embedded Block Preservation) was decomposed into C2 (Code Language Match) and C3 (Block Compliance); F6-old (Render Equivalence) was superseded by the latent failure definition (§4.4), which uses the four-cell decomposition instead of a rendering-based proxy. 12
Two Compliance Layers: Prompt vs Specification. The F-metrics span two distinct compliance layers that are easy to conflate. F1 (Instruction Compliance) is prompt-level: it asks “did the model follow the explicit A-axis directive (wrap / do-not-wrap)?” and is judged against the prompt itself. F2 (Fence Balance) and F3 (Safe Length, post-hoc variant) are specification-level: they ask “does the produced text conform to CommonMark closing-fence rules?” and are judged against the spec without reference to the prompt. F3 additionally carries an a-priori variant used in the D-inner-run ablation that is again prompt-level: it tests whether the model picked a safe outer length given an explicit inner-run constraint. Keeping these layers explicit clarifies what each metric isolates: F1 measures intent-following, F2/F3 measure spec-conformant production, and the post-hoc/a-priori variants of F3 separate reactive from predictive safety. A.2
Auto-Detection Algorithms
Stack-Based Fence Balance Checker (F2). The pseudocode tracks opening fences via a LIFO stack and pops on each matching closing fence: function check_fence_balance(text): stack = [] for each line in text: if line matches OPENING_FENCE_PATTERN: extract (family, length, info_str) push (family, length) onto stack elif line matches CLOSING_FENCE_PATTERN: extract (family, length) # closing fence must carry no info string if stack is not empty: (ofam, olen) = stack.top() if family == ofam and length >= olen: pop stack return stack is empty # True = balanced
Safe Length Calculator (F3). The pseudocode scans interior content for the longest run of the outer fence character and verifies the strict-inequality safety condition: function check_safe_length(text, ofam, olen): max_run = 0 for each line in text (excluding outer fences): for each run of ofam characters in line: max_run = max(max_run, run_length) return olen > max_run # strict inequality
Additional pseudo-code for E0–E3 classification and the 4-cell decomposition algorithm is provided in the code release. A.3
Scope Justification
Our analysis focuses on paired delimiters that simultaneously serve as opening and closing boundaries: fenced code blocks (backtick, tilde) and inline code spans (backtick). These constructs exhibit the symmetric-delimiter problem: the same character sequence must be interpreted as wrapper, content, or literal depending on generation state. Single-marker constructs (blockquotes >, lists -, headings #) do not form paired boundaries and thus do not create delimiter-collision opportunities. Inline emphasis (*, **) operates within a single line and does not expose the multi-line boundary-state tracking challenge we investigate. YAML front matter is excluded because --- markers are not defined as paired delimiters in CommonMark 0.31.2 and their semantics vary across Markdown extensions. A.4
Coverage Matrix: Cell-Level Justification
StructEval “partial” for E1/E2. StructEval includes syntactic-validity and structural-element checks for structured outputs, so boundary failures that visibly alter parseability or required structural 13
elements can be indirectly reflected in its scores. However, StructEval does not inspect CommonMark fence-closing semantics directly, distinguish E1 from E2, or compare outer and inner fence lengths. Hence “partial”: some consequential symptoms may be detectable, but cause-level fence-boundary diagnosis is not. StructEval “–” for E3. StructEval does not perform fence length comparisons and does not include the opening/closing fence length relationship in its metric definitions. E3 (Latent Collision) is a fence-length-specific check and is fundamentally undetectable by StructEval. StructEval “partial” for F2. StructEval’s syntax-validity checks can indirectly penalize malformed structured outputs when fence imbalance affects parseability or required elements. However, it does not implement a stack-based CommonMark fence-balance checker or diagnose the specific cause of imbalance. IFEval “partial” for F1. IFEval evaluates general instruction-following with verifiable constraints, which conceptually overlaps with our F1 wrapping-instruction compliance. However, it does not target the specific wrapper-presence/absence manipulation used here, nor does it condition downstream CommonMark boundary metrics on compliance. FormatBias “concept” for 4-cell. FormatBias argues that format and content should be evaluated separately, which conceptually aligns with our 4-cell decomposition. However, FormatBias does not implement a boundary-correct/wrong axis and does not perform an actual 2×2 decomposition. It is a conceptual precursor, hence “concept.” Remaining “–” cells. For each remaining “–” cell, we reviewed the respective benchmark’s metric definitions and confirmed that no metric is capable of detecting the corresponding failure type. Benchmarks omitted from Table 1. MDEval [5] and FMBench [22] are excluded from the matrix columns because their metric definitions do not directly cover our fence-boundary failure types. MDEval evaluates Markdown awareness, readability, and content-structure quality without inspecting CommonMark fence-closing semantics or outer/inner fence-length safety. FMBench evaluates adaptive Markdown output formatting under diverse structural and layout constraints, including mixed content and code-block formatting, but does not perform per-fence stack-balance checks, fence-length comparisons, or same-family nesting diagnosis. Including all-– columns would inflate the matrix without informational content, so we restrict it to benchmarks with at least one partial or positive coverage cell. A.5
Latent Failure Downstream Example
This example illustrates how a latent failure—a generation that is content-correct but boundary-wrong (F2 or F3 fail)—produces different outcomes depending on the downstream consumer. A naive regex-based extractor may either truncate at the first apparent closing fence or, in balanced-but-unsafe cases, miss the latent length collision entirely. LatentMD’s F3 check identifies unsafe same-family fence-length collisions (E3), which can cause spec-strict downstream parsers or extractors to mis-nest or truncate the document. This demonstrates why spec-level boundary checking is necessary even when visual rendering appears correct. Anonymized real-world incident. Beyond engineered prompts, the same failure mode has been observed organically by end users on production chat platforms. In one incident witnessed by the authors during a routine usage of a major chat assistant, the model produced a Markdown tutorial about Python that visually rendered as expected inside the chat UI; when the user copied the output into a CommonMark-strict Markdown editor for archival, all code blocks beyond the first silently merged with the surrounding prose, corrupting the document’s structure. The phenomenology aligns precisely with the latent failure cell quantified in our slot-based experiments (Table 7): visually normal output, structurally broken at the spec level, surfaced only when downstream tooling diverges from the lenient renderer used for the initial display. We do not identify the chat platform here; the public bug-tracker references in §1 provide independently verifiable instances of the same pattern across multiple platforms. 14
Table 8: Concrete latent-failure example. Model GPT-4o on prompt main_A2_B1_python_Python_35 (A-cond=A2, B-cond=B1, task=Python/35). The response passes all content checks (C1∧C2∧C3) but the inner code-fence run length collides with the outer wrapper, so F3 fails (E3 latent collision). A naive regex extractor or visual rendering would flag this as “correct”; only spec-level CommonMark parsing surfaces the boundary violation. Model output (excerpt) 1 ```markdown 2 # Calculating the Number of Apples Tao Tao Can Reach 3 4 In this document, we will discuss how to calculate the number of apples that Tao Tao can
reach. Tao Tao can reach a certain height with her hand stretched upwards, and she can also use a stool to extend her reach by an additional 30 centimeters. We will write a Python function to determine how many apples are within her reach.
5 6 ## Problem Statement 7 8 Given a list of apple heights and Tao Tao's maximum reach, we need to calculate how many
apples she can reach. The function should consider both her natural reach and the additional height provided by a stool.
9 10 ### Function Signature 11 12 ```python 13 def count_apples_reachable(apple_heights: List[int], tao_reach: int) -> int: 14 ``` 15 16 ### Parameters 17 18 [... abbreviated ...]
Diagnosis: C1/C2/C3 all pass (heading present, language-tag matches, code block of expected type present), so the response is judged Content-correct. However, an inner code block uses a backtick run that matches the outer wrapper, breaking F3 (safe-length) — the CommonMark spec terminates the outer fence at the first matching closing run, so the document parses with mis-nested blocks (E3).
B
Benchmark Design
B.1
Visual Overview of Axes
A-axis Outer constraint
B-axis Inner structure
D-axis Delimiter ablations
A2 Wrapped
A1 Direct B1 Single
A3 Unspecified
B2 Nested single
B3 Nested multi
B4 Mixed
Family
Length
Inner run
Cross-fam
backtick / tilde
outer 3 / 4 / 5
3/4/5
mixed pair
Slot-based generation: LANG (12) × TASK (15 per LANG, from McEval) = 180 base pairs Main: A (3) × B (4) × 180 base pairs = 2,160 prompts Main 2,160 + D-axis 1,440 + Hint 432 + Inline 27 + Cross-format 90 + Natural 30 = 4,179 unique prompts × 9 models ≈ 37,600 generations (T = 0)
Figure 6: Visual overview of the three benchmark axes (cross-reference for §3.1). A controls outer constraint (3 levels), B controls inner nesting (4 levels), and D enumerates delimiter ablations. Total: 4,179 prompts (breakdown in §3.2). 15
B.2
Evaluation Pipeline Models
Axis
Metrics
Errors
(main)
A, B
Prompt
Response
F, C
E
(ablation)
D
Prompt
Response
F, C
E
Figure 7: Evaluation pipeline overview. Each row represents one experimental track. The main track conditions on A and B values, while the ablation track conditions on D values. In both tracks, the axis values are rendered into a Prompt; one of nine Models produces a Response; the Response is scored by Format (F1–F3) and Content (C1–C3) metrics; the metric outcomes determine one of four boundary-violation labels (E0–E3). The dashed wrapper marks the model-interaction stage, the only stochastic component when temperature > 0 (we use T = 0 throughout). B.3
LANG Selection: Programming Stratum Consensus
The nine programming languages in the LANG pool are selected by a popularity consensus rule applied to three independent rankings published in 2024-2025: the TIOBE Programming Community Index (December 2025 snapshot), the GitHub Octoverse 2024 most-used languages list, and the Stack Overflow 2025 Developer Survey “most popular technologies” table. A language is included in the programming stratum if and only if it appears in at least two of the three rankings within the top thirty entries. This two-of-three consensus reduces the influence of any single ranking’s idiosyncrasies (e.g., TIOBE’s search-engine bias, GitHub’s repository-creation bias, or Stack Overflow’s questionnaire selection bias) while still admitting any language that two independent sources agree is widely used. After applying the consensus rule, we further restrict to languages that have at least one instructiontuned task in the McEval [3] benchmark, which is required by our slot-based template (see §3.2). The final nine languages are: Python, JavaScript, TypeScript, Java, C++, C, C#, Go, and Shell. The structured-format stratum (three additional values — Markdown, JSON, and HTML) is then drawn from the McEval non-programming task set, since these formats are well-defined by their respective specifications and do not need a popularity ranking. B.4
Prompt Templates: Per-Condition Examples
The full per-condition prompt set (9 YAML files, 4,179 prompts) is released as part of the LatentMD dataset; the examples below show one representative prompt per axis condition, with the two filled-in slots written as <LANG> and <TASK> (cf. Figure 1). A2 × B2 Example. Write a Markdown document that addresses the following task. Include one <LANG> code example. In the document, also show the raw Markdown source for that code example, including the opening fence, the language tag, the code, and the closing fence. Wrap your entire response in a fenced markdown code block. Task: <TASK>
16
A3 × B2 Example (no wrap directive — realistic anchor). Write a Markdown document that addresses the following task. Include one <LANG> code example. In the document, also show the raw Markdown source for that code example, including the opening fence, the language tag, the code, and the closing fence. Task: <TASK>
A2 × B1 Example (simple, no nested fence). Write a Markdown document that addresses the following task. Include exactly one <LANG> code example. Wrap your entire response in a fenced markdown code block. Task: <TASK>
A2 × B3 Example (multiple nested fences). Write a Markdown document that addresses the following task. Include <LANG> code examples. In at least two places in the document, show the raw Markdown source for the code examples, including the opening fence, the language tag, the code, and the closing fence. Wrap your entire response in a fenced markdown code block. Task: <TASK>
A2 × B4 Example (mixed elements). Write a Markdown document that addresses the following task. Include <LANG> code examples, a comparison table, a blockquote with a citation, and a numbered list of steps. Wrap your entire response in a fenced markdown code block. Task: <TASK>
D-Family (Tilde) Example. [A2 x B2 template as above] + Use tilde fences (~~~) instead of backtick fences for all code blocks in your response.
D-Inner-Run Example. [A2 x B4 template as above] + When writing inner code blocks, use exactly 4 backticks for the opening fence.
D-Length Example. [A2 x B2 template as above] + When writing fenced code blocks, use exactly 5 backticks for the outermost fence.
D-Cross-Family Example. [A2 x B2 template as above] + Use backtick fences for your outer wrapper, and tilde fences for any inner code block examples.
17
Hint H1 (Boundary-Awareness Rule) Example. [A2 x B2 template as above] + Hint: Maintain correct nesting of all code fences. Outer fences must use MORE fence characters than any same-type fence appearing inside the content.
Hint H2 (One-Shot) Example. [A2 x B2 template as above] + Hint: Here is an example of correct nesting: ````markdown # Title ```python print('hello') ``` ```` Now write yours following this pattern.
Hint H3 (Length-Specific) Example. [A2 x B2 template as above] + Hint: Use at least 4 backticks for your outermost fence.
Cross-Format Python K1 Example. Write a Python function that addresses the following task. The function must have a docstring that mentions the triple-quote delimiter """ as part of the documentation text. Task: <TASK>
Cross-Format JSON K1 Example (asymmetric-delimiter control). Write a JSON document that addresses the following task. Use 1 to 2 levels of nested objects or arrays. Task: <TASK>
Inline I1, run=1 Example. Write one sentence in Markdown that contains a single inline code span whose displayed content is the text: `foo`.
C
Experimental Setup
C.1
Hyperparameter Table Table 9: Full hyperparameter configuration for all experiments. Parameter
Value
Notes
Temperature top_p top_k max_tokens seed Serving (open) Serving (closed)
0 0.9 40 7168 42 vLLM / HF Transformers Official APIs
Greedy; sweep in App. J.3 Cross-API compatible Gemini/Anthropic; ignored by OpenAI Fits Gemma-2 8K context Open-weight only See Table 3 OpenAI, Google, Anthropic
Open-weight models were served on NVIDIA A100 40GB, RTX 6000 Ada 48GB, H100 SXM 80GB, and RTX 3090 24GB GPUs, via vLLM (HuggingFace Transformers for A100 due to CUDA version constraints). Closed-source models were queried from a desktop workstation via official APIs. Under 18
temperature-0 greedy decoding, open-weight outputs are deterministic up to floating-point precision and therefore independent of the specific GPU used. Approximate runtime. Representative wall-clock measurements include ∼8 hours for Llama3.1-8B on a single A100 and ∼52 hours for Llama-3.1-70B on 5 A100 GPUs (full main grid via HuggingFace Transformers). Open-weight ablation generations ran on a 2×H100 SXM cloud node within ∼32 wall-clock hours using vLLM. Together with the remaining open-weight runs, aggregate open-weight compute was on the order of several hundred GPU-hours, dominated by Llama-3.170B main-grid generation. Closed-source API generations were queried in parallel from a single workstation, with wall-clock time governed by provider latency and rate limits; aggregate API spend is itemized in Appendix C.2. Reasoning-mode disable flags. Three models in our suite ship with optional “thinking” / extendedreasoning modes: Qwen3 (both 7B and 32B), Gemini-2.5-Flash, and Claude-Sonnet-4. We disable all of them so that every response is a single Markdown answer rather than a reasoning trace followed by an answer. Per-model flags: Qwen3 via enable_thinking=False, Gemini-2.5-Flash via thinkingBudget=0, and Claude-Sonnet-4 via its default-off extended-thinking setting (no thinking field in the request). The remaining six models (Gemma-2-9B, Gemma-3-27B, Llama3.1-8B/70B, GPT-4o) do not expose a reasoning-mode toggle and were used as-is. Closed-source group classification. GPT-4o, Claude-Sonnet-4, and Gemini-2.5-Flash do not publish parameter counts, so they cannot be placed in a parameter-defined size tier in the strict sense. We group them with the largest open-weight model (Llama-3.1-70B) under a unified “Large/API” label in Table 3, treating “Tier” as an operational rather than a parameter-strict grouping. The Small and Mid tiers are defined by parameter count among open-weight models only. This grouping affects descriptive aggregations (e.g., the tier averages in Table 6) but does not enter any model-level statistical claim, all of which are reported per individual model. C.2
Closed-source API Usage and Cost
Closed-source API calls were issued from a workstation, with approximate usage costs of $67 (Anthropic Claude-Sonnet-4), $25 (OpenAI GPT-4o), and $30 (Google Gemini-2.5-Flash plus a small Flash-Lite probe), totalling ≈ $122 across the three closed APIs for the full benchmark (4,179 prompts × 3 closed models, covering both main grid and ablations). These figures are based on provider billing records at the time of the experiments and are intended only as a reproducibility reference; actual costs may vary with provider pricing, rate limits, batching, and response length. Table 10: Closed-API inference cost for the full benchmark (approximate, in USD). API
Cost
Coverage
Anthropic Claude-Sonnet-4 OpenAI GPT-4o Google Gemini-2.5-Flash/Lite
≈ $67 ≈ $25 ≈ $30
Main + Ablation Main + Ablation Main + Ablation (Flash + Lite probe)
Total
≈ $122
3 closed APIs, full benchmark
D
Truncation Filtering and Natural-Prompt Validation
D.1
Truncation Statistics
Responses with finish_reason=“length” are excluded from all boundary metrics and reduce the effective sample size (valid N ) for the affected model. Under identical conditions (max_tokens=7168, temperature= 0), Gemini-2.5-Flash exhibits a truncation rate of ∼12–13%, while GPT-4o and Claude-Sonnet-4 show 0%. This discrepancy arises because Gemini generates 2–3× longer outputs for the same prompts and occasionally enters degenerate repetition loops (e.g., repeating a single character until the token limit). This pattern is consistent with the observed repetition loops in our generations. No post-processing (e.g., trailing whitespace removal) is applied to any model’s output; all responses are evaluated as-is. 19
Table 11: Truncation statistics per model (main + ablation experiments combined). Generations with finish_reason = “length” are excluded from all metric computations. Valid N is used as the denominator for Wilson confidence intervals. Total N
Truncated
Valid N
Qwen3-7B Gemma-2-9B Llama-3.1-8B Qwen3-32B Gemma-3-27B Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
4179 4179 4179 4179 4179 4179 4179 4179 4179
30 6 34 22 8 4 0 503 1
4149 4173 4145 4157 4171 4175 4179 3676 4178
Total
37611
608
37003
Model
Excluded model (Gemini-2.5-Flash-Lite). We additionally ran Gemini-2.5-Flash-Lite on the full 4,179-prompt suite as a within-Google-family probe. Its responses are even more verbose than Flash’s (avg ∼18K characters per response vs Flash’s ∼14K), yielding 803/4,179 = 19.2% truncation — markedly above Flash’s 12.0%, producing an effective sample size (3,376) that diverges from the 9 paper models (≥ 3,676) widely enough to substantially weaken per-cell comparability across stratified analyses. We therefore exclude Flash-Lite from the main analysis cohort; per-record results are released alongside the 9-model data for downstream verification. D.2
Per-Language Truncation Distribution
The aggregate truncation report above pools across languages. Stratified by language (Table 12), Gemini-2.5-Flash’s 12.3% overall truncation rate concentrates in Go (26.1%), JavaScript (22.2%), TypeScript (22.2%), and C++ (17.2%) and is much lower in Markdown (3.9%) and Shell (2.2%). Other models show < 1% truncation across all languages. The pattern is consistent with the verboseoutput hypothesis (Gemini generates 2–3× longer responses for verbose-syntax languages and exhausts the 7168-token budget).
D.3
Jav aSc
Ty peS cri
Sh ell
N
C#
C
Go
HT
JSO
Ma rkd ow n
Small
2.8 0.0 3.9
0.0 0.0 0.0
0.0 2.2 0.6 0.0 0.0 0.0 0.0 0.0 0.0
0.6 0.0 0.0
0.6 1.1 0.0
1.7 0.0 0.0
1.7 0.0 0.0
0.6 0.0 1.1
0.0 0.6 0.0
0.0 0.0 2.8
0.9 0.1 0.6
Mid
Qwen3-32B Gemma-3-27B
1.7 0.0
0.0 0.0
0.0 1.1 0.0 0.0 0.0 0.0
0.0 0.0
1.1 0.0
0.0 0.0
0.6 0.6
0.0 0.0
0.0 0.0
0.0 0.6
0.4 0.1
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
0.0 0.0 7.8 0.0
0.0 0.0 22.2 0.0
0.0 0.0 0.0 0.0 0.0 0.0 5.0 17.2 6.1 0.0 0.0 0.0
0.0 0.0 22.2 0.0
0.0 0.0 0.0 0.0 0.0 0.0 2.2 13.9 26.1 0.0 0.0 0.0
0.6 0.0 0.0 0.0 7.2 13.3 0.0 0.0
0.0 0.0 3.9 0.0
0.0 0.0 12.3 0.0
Overall
1.8
2.5
0.6 2.3 0.7
2.5
0.6
1.0
0.8
1.6
+ C+
Jav a
ML
Py
Total
Qwen3-7B Gemma-2-9B Llama-3.1-8B
rip t
tho n
Model
Large
pt
Table 12: Truncation rate (%) per model and language. Truncation = finish_reason == length (max_tokens budget exhausted). Each cell is computed on N ≈45 prompts (3 A × 4 B × ∼3.75 TASK per LANG, before truncation filtering). Bold marks model rows with >5% overall truncation.
1.7
3.2
1.5
Natural Prompt Validation
Thirty human-authored prompts—phrased as a developer might naturally write them—serve as an external validity check. Prompts span three elicitation levels: 10 high (explicitly request nested code blocks or raw Markdown syntax), 10 medium (request code examples without mentioning nesting), and 10 low (technical topics where code blocks are likely but not explicitly requested). With N = 30, we report descriptive failure-rate ranges and treat this set as an ecological sanity check rather than a population-level prevalence estimate. 20
Table 13: Natural prompt validation results. 30 human-authored prompts × 9 models across three elicitation levels. Fail Rate = boundary-incorrect responses / total responses observed for that prompt (truncated responses excluded). #
Prompt (abbreviated)
High elicitation 1 Tutorial on fenced code blocks 2 Display raw Markdown source in a document 3 Markdown syntax cheat sheet (raw + rendered) 4 Tutorial on nesting code blocks 5 Static site generator docs: fenced examples 6 Reference card: literal triple backticks 7 Blog post on Markdown pitfalls 8 Markdown linter docs: correct/incorrect 9 GitHub README: raw Markdown source 10 CommonMark spec: fence lengths and languages Medium elicitation 11 Python README for data pipeline 12 JavaScript async/await guide 13 REST API docs with curl examples 14 README.md for Rust CLI tool 15 Overview of sorting algorithms (C++) 16 API reference for Go package 17 TypeScript React getting-started guide 18 Python list comprehensions vs for loops 19 SQL joins quick reference 20 JS to TypeScript migration guide Low elicitation 21 Python virtual environment setup 22 Debugging memory leaks in Node.js 23 JWT authentication implementation 24 Git rebase vs merge 25 PostgreSQL database setup 26 Docker environment variables 27 Optimizing slow SQL queries 28 Webhooks: how they work and setup 29 Nginx reverse proxy configuration 30 Automated testing in CI/CD pipeline
E
Main Results: Full Tables
E.1
Full A×B×Model Results
Level
Fail Rate
High High High High High High High High High High
11.1% 55.6% 28.6% 22.2% 22.2% 33.3% 44.4% 37.5% 22.2% 55.6%
Medium Medium Medium Medium Medium Medium Medium Medium Medium Medium
0.0% 0.0% 11.1% 0.0% 0.0% 11.1% 0.0% 33.3% 0.0% 0.0%
Low Low Low Low Low Low Low Low Low Low
0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0% 0.0%
Table 14: Boundary failure rate (%) per model across A×B conditions (= E1∪E2∪E3, i.e. F2 or F3 fail). Columns are ordered by experimental severity rather than condition ID: A1 (no-wrap baseline) → A3 (unspecified, realistic anchor) → A2 (forced-wrap stress test). Per-cell mechanism breakdown (E1/E2/E3) is in Appendix G.6.
Small
A2 (Wrapped) B2 B3 B4
Qwen3-7B Gemma-2-9B Llama-3.1-8B
5.6 23.5 44.1 0.0 91.0 98.3 92.7 77.6 90.5 97.8 98.3 92.2 4.4 7.8 5.6 0.0 1.7 31.7 18.3 0.0 60.6 97.8 93.3 87.7 0.0 7.2 2.8 0.6 0.6 52.2 26.8 1.7 93.3 91.1 96.1 86.4
Mid
A3 (Unspecified) B1 B2 B3 B4
B1
Qwen3-32B Gemma-3-27B
52.2 18.9 18.9 75.8 71.7 41.1 37.4 46.9 96.1 97.8 99.4 98.9 88.8 75.0 84.4 93.3 95.0 93.9 96.1 92.8 95.0 100.0 100.0 96.1
Large
A1 (Direct) B2 B3 B4
Model
Llama-3.1-70B 0.0 83.9 78.9 0.0 0.0 80.6 76.7 0.0 93.3 100.0 100.0 95.0 GPT-4o 16.7 1.7 1.1 11.1 83.3 2.2 10.6 93.3 93.3 100.0 98.9 97.8 Gemini-2.5-Flash 15.3 22.5 20.0 2.9 75.5 98.8 96.4 51.5 93.2 100.0 100.0 99.3 Claude-Sonnet-4 0.0 0.0 0.6 0.0 0.0 0.0 0.0 0.0 99.4 100.0 100.0 100.0
B1
B1=single, B2=nested single, B3=nested multiple, B4=mixed. A3 (Unspecified) reflects realistic prompts: models self-wrap ∼37% of responses (Step 1 of Table 5); their boundary fail rate sits between A1 (low) and A2 (stress-test ceiling).
Table 14 reports the per-model boundary failure rate (E1∪E2∪E3) for every A×B cell; the body §5.2 cites the pooled view (Figure 4) and the per-condition severity ordering. Complete per-metric breakdowns (F1–F3, C1–C4, E0–E3) for every cell and model are also provided in the supplementary materials as structured JSONL files. 21
E.2
Main Results with Wilson 95% Confidence Intervals
Table 15: Boundary failure rate (%) per model under A1 (Direct) prompts, with Wilson 95% confidence intervals over valid N (after truncation filtering). Each cell shows rate [CI lower, CI upper]. Wilson’s small-sample correction yields tight intervals near the [0, 100] extremes. Model
23.5 [17.8,30.2] 44.1 [37.0,51.4] 7.8 [4.7,12.6] 5.6 [3.0,9.9] 7.2 [4.3,12.0] 2.8 [1.2,6.4]
B4
Small
B3
Mid
B2
5.6 [3.1,10.0] 4.4 [2.3,8.5] 0.0 [0.0,2.1]
Qwen3-32B Gemma-3-27B
52.2 [45.0,59.4] 18.9 [13.8,25.2] 18.9 [13.8,25.2] 75.8 [69.0,81.5] 88.8 [83.4,92.7] 75.0 [68.2,80.8] 84.4 [78.4,89.0] 93.3 [88.7,96.2]
Large
B1
Qwen3-7B Gemma-2-9B Llama-3.1-8B
0.0 [0.0,2.1] 0.0 [0.0,2.1] 0.6 [0.1,3.1]
Llama-3.1-70B 0.0 [0.0,2.1] 83.9 [77.8,88.5] 78.9 [72.4,84.2] GPT-4o 16.7 [11.9,22.8] 1.7 [0.6,4.8] 1.1 [0.3,4.0] Gemini-2.5-Flash 15.3 [10.6,21.7] 22.5 [16.8,29.3] 20.0 [14.6,26.7] Claude-Sonnet-4 0.0 [0.0,2.1] 0.0 [0.0,2.1] 0.6 [0.1,3.1]
0.0 [0.0,2.1] 11.1 [7.3,16.5] 2.9 [1.1,7.3] 0.0 [0.0,2.1]
Table 16: Boundary failure rate (%) per model under A2 (Wrapped) prompts, with Wilson 95% confidence intervals over valid N (after truncation filtering). Each cell shows rate [CI lower, CI upper]. B1
B2
B3
B4
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
90.5 [85.3,94.0] 60.6 [53.3,67.4] 93.3 [88.6,96.1]
97.8 [94.4,99.1] 97.8 [94.4,99.1] 91.1 [86.1,94.5]
98.3 [95.1,99.4] 93.3 [88.7,96.2] 96.1 [92.1,98.1]
92.2 [87.3,95.3] 87.7 [82.1,91.7] 86.4 [80.6,90.7]
Mid
Qwen3-32B Gemma-3-27B
96.1 [92.2,98.1] 97.8 [94.4,99.1] 99.4 [96.9,99.9] 95.0 [90.8,97.4] 100.0 [97.9,100.0] 100.0 [97.9,100.0]
98.9 [96.0,99.7] 96.1 [92.2,98.1]
Large
Model
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
93.3 [88.7,96.2] 93.3 [88.7,96.2] 93.2 [88.2,96.2] 99.4 [96.9,99.9]
100.0 [97.9,100.0] 100.0 [97.9,100.0] 95.0 [90.8,97.4] 100.0 [97.9,100.0] 98.9 [96.0,99.7] 97.8 [94.4,99.1] 100.0 [97.7,100.0] 100.0 [97.6,100.0] 99.3 [96.3,99.9] 100.0 [97.9,100.0] 100.0 [97.9,100.0] 100.0 [97.9,100.0]
This table is the per-cell counterpart of Table 14 (body), augmented with Wilson 95% score intervals for the boundary failure rate. Why a CI under T = 0? Greedy decoding is intended to suppress the generation-level variance of M (p) to (approximately) zero, so repeated sampling is not used to estimate within-prompt stochasticity. The reported intervals instead quantify a different source of uncertainty: the promptsampling variance induced by drawing 15 TASKs per LANG (5 Easy + 5 Mid + 5 Hard at seed=42) from a larger McEval-restricted pool. A different seed would yield a different 180-prompt subset, and hence a slightly different observed pass rate; the Wilson interval provides an approximate 95% binomial uncertainty interval over the sampled prompt instances. Within-prompt comparisons (A1 vs A2 on the same TASK × LANG) are handled by paired McNemar tests (Appendix I.4), which are insensitive to this between-prompt sampling variance and therefore complement, rather than duplicate, the Wilson CI. Why Wilson rather than Wald. The Wilson interval is preferred over the normal approximation because failure rates near 0% and 100% are common in our data, where the normal (Wald) interval would yield negative or super-unit bounds, or collapse to a degenerate point at the extremes (e.g., pb = 1.0, n = 180 gives Wald = [1.0, 1.0], asserting certainty). Wilson’s small-sample correction keeps both endpoints in [0, 100] while preserving nominal coverage; for the same case it gives the asymmetric [97.9%, 100%], properly reflecting upper-bound uncertainty. Inspection by row shows that the dominant per-condition failure rates (≥ 90% for A2 cells across most models) have intervals of width ≤ 8pp, supporting the qualitative claim that A2 failures are not borderline. E.3
Outer-Fence Length Distribution (All Nine Models)
This subsection provides the full per-model breakdown of outer-fence length choices on D-inner-run prompts, complementing the four-model body figure in §5.4. For each (model, inner-run condition r), Table 17 reports the percentage of responses choosing each outer-fence backtick length (out of n=144 records per cell) and the resulting F3 a-priori pass rate (sum of cells with outer > r). 22
Model
r
outer=3
outer=4
outer=5
outer≥6
F3 pass
Qwen3-7B
3 4 5
100.0 100.0 100.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
Gemma-2-9B
3 4 5
100.0 100.0 100.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
Llama-3.1-8B
3 4 5
100.0 100.0 100.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
Qwen3-32B
3 4 5
100.0 99.3 99.3
0.0 0.7 0.0
0.0 0.0 0.0
0.0 0.0 0.7
0.0 0.7 0.7
Gemma-3-27B
3 4 5
100.0 100.0 100.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
Llama-3.1-70B
3 4 5
100.0 31.9 95.1
0.0 68.1 4.2
0.0 0.0 0.7
0.0 0.0 0.0
0.0 0.0 0.0
GPT-4o
3 4 5
86.8 36.1 28.5
13.2 63.9 25.7
0.0 0.0 4.2
0.0 0.0 41.7
13.2 0.0 41.7
Gemini-2.5-Flash
3 4 5
100.0 60.4 73.6
0.0 39.6 7.6
0.0 0.0 18.8
0.0 0.0 0.0
0.0 0.0 0.0
Claude-Sonnet-4
3 4 5
100.0 83.3 99.3
0.0 16.7 0.7
0.0 0.0 0.0
0.0 0.0 0.0
0.0 0.0 0.0
Table 17: Outer-fence length distribution and F3 a-priori pass rate per (model, inner-run condition r) on D-inner-run prompts. Each row sums to 100% across the four outer-length columns (n=144). The F3 pass column equals the sum of cells with outer > r. Most models default to outer = 3 regardless of r; a few (Llama-3.1-70B, GPT-4o, Gemini-2.5-Flash, Claude-Sonnet-4) partially adapt by matching the inner length (outer = r), but only GPT-4o non-trivially exceeds it.
Hint Ablation: Per-Model Heatmap
40
80
t-4
Mean (%): 96 72 87 65
on ne
Cla
ud
e-S
-2. mi ni Ge
60
Boundary failure (%)
100 5 79 7
h
98 79 93 33 5-F las
T-4 o GP
70 ma -3
98 23 48 38
Lla
-3mm a
.1-
27
2B 20
Ge
Qw
en
3-3
.10
97 82 94 56 B
98 97 98 99 B
98 99 94 95
Lla
Ge
mm a
ma -3
-2-
3-7 en Qw
92 90 94 83 8B
85 81 89 79 9B
95 95 95 92
None H1 H2 H3
B
Hint
E.4
100
Figure 8: Per-model boundary failure rate (%) across hint conditions (rows: None / H1 / H2 / H3; columns: 9 models in tier order). Cell color: deeper red indicates higher failure rate. Right column: row means across models. Body summary: §5.4.
F
Manipulation Checks and Robustness Diagnostics
F.1
A1 Per-Model Decomposition
Per-model breakdown of the four-step A1 decomposition (Table 5 in §5.5 shows the 9-model averages; Table 18 below shows the dispersion behind each step). 23
Table 18: Per-model A1 decomposition. Step 1: A3 wrap rate; Step 2: A1 non-compliance (F1); Step 3: compliance-conditional A1 residual boundary failure; Step 4: A2 boundary failure rate.
F.2
Model
Step 1
Step 2
Step 3
Step 4
Qwen3-7B Gemma-2-9B Llama-3.1-8B Qwen3-32B Gemma-3-27B Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
77.6% 0.1% 0.1% 46.2% 83.8% 0.0% 47.9% 82.5% 0.0%
18.2% 5.8% 0.4% 43.3% 79.4% 0.0% 7.2% 16.0% 0.0%
1.5% 3.1% 2.4% 1.5% 34.5% 40.7% 0.7% 3.2% 0.1%
94.7% 85.7% 91.7% 98.0% 97.9% 97.1% 98.1% 98.4% 99.9%
Full-C Sensitivity Analysis
Our pre-specified escalation rule (§4.2) demotes C4 (keyword-based Task Completion) when the Kendall’s τ between Full-C (C1∧C2∧(C4≥0.7)) and Structural-C (C1∧C2) model rankings on the “content-correct, boundary-wrong” cell falls below 0.8. We use C1∧C2 for the ranking-stability trigger because C3 is a template-structure check; the final primary Content definition retains C3 as specified in §4.2. On our data, τ = 0.667 < 0.8, so C4 is demoted and the primary Content judgment is C1∧C2∧C3 throughout the main text. For transparency we retain the Full-C (with C4) four-cell decomposition here; the headline Latent Failure rate under Full-C is 33.1% (overall), versus 38.0% under the primary definition without C4. The direction of model rankings is preserved (all 9 models exhibit higher Latent Failure under the primary definition than under Full-C), and the most extreme cases—Gemma-3-27B (73.8% vs 57.0%) and Llama-3.1-8B (14.5% vs 10.8%)—are shifted by the content-filter relaxation, consistent with the view that Full-C’s keyword gate was removing structurally correct responses from the Content-Correct side; the primary C1–C3 definition is therefore the more conservative choice for boundary-failure attribution. Full per-model numbers under both definitions are provided as LatentMD/results/tables/four_cell.json in the supplementary materials (primary_struct_only_c and sensitivity_full_c entries).
G
Stratified Error-Type Analyses
This section breaks down the error-type distribution (E0 correct / E1 premature / E2 unclosed / E3 latent collision) along four orthogonal stratifications: language (12 LANG values), model size tier (Small/Mid/Large), inner-nesting condition (B1–B4), and McEval task difficulty (easy/middle/hard). All numbers are pooled across the 9 main-experiment models and exclude truncated responses. G.1
Per-Language Distribution
Table 19: Error type (E0–E3) distribution per LANG, pooled across all 9 models. N = number of valid records per language. LANG
N
E0 (%)
E1 (%)
E2 (%)
E3 (%)
Python JavaScript Java C++ C# TypeScript Shell C Go HTML JSON Markdown
1591 1580 1611 1583 1608 1579 1611 1592 1568 1603 1595 1607
43.6 45.0 43.3 43.6 41.0 42.6 37.6 41.9 43.7 44.4 38.6 64.8
11.0 8.9 10.2 8.6 13.1 8.4 9.9 8.0 9.2 6.7 10.6 4.2
12.3 12.0 16.1 11.6 12.8 13.7 12.2 13.9 14.5 10.0 10.7 3.3
33.1 34.1 30.4 36.2 33.1 35.3 40.2 36.2 32.5 39.0 40.2 27.7
Markdown stands out as the easiest LANG (E0 = 64.8%) while JSON, Shell, and HTML show the highest E3 rates (≥39%); the spread is consistent with the view that E3 (latent collision) is driven 24
by the syntactic similarity between code-fence delimiters and inner content tokens, not by language semantics. G.2
Per-Model-Size-Tier Distribution Table 20: Error type distribution per model size tier (Small=3, Mid=2, Large=4 models). Tier
N
E0 (%)
E1 (%)
E2 (%)
E3 (%)
Small Mid Large
6444 4310 8374
53.4 22.3 48.3
10.8 11.9 6.3
17.5 7.0 10.2
18.3 58.8 35.3
The ordering is non-monotonic in tier — Mid models (Qwen3-32B and Gemma-3-27B) have the lowest E0 rate (22.3%) and highest E3 (58.8%), driven by Gemma-3-27B’s outlier behavior; Small and Large tiers are closer to each other (53.4% vs 48.3% on E0). G.3
Per-B-Condition Distribution
Table 21: Error type distribution per B-condition (inner nesting complexity), pooled across all 9 models. E3 (latent collision) climbs sharply with B complexity. B-cond
N
E0 (%)
E1 (%)
E2 (%)
E3 (%)
B1 B2 B3 B4
4800 4812 4793 4723
47.7 40.1 41.1 48.0
3.0 19.4 10.7 3.0
9.7 16.4 15.7 5.7
39.6 24.1 32.5 43.3
E1 (premature closure) peaks at B2 (19.4%) where the response must include a code example plus a literal Markdown explanation — the premature-closure failure mode is most strongly activated by exactly one layer of literal-fence content. E3 (latent collision) is high at B1 (39.6%) and B4 (43.3%) where the inner code block competes most directly with the outer wrapper run length. G.4
Per-Difficulty Distribution
Table 22: Error type distribution per task difficulty, pooled across all 9 models. Difficulty levels (easy/middle/hard) are taken from McEval task metadata (LatentMD/src/prompt_gen/_mceval_ tasks.py in the released code). Difficulty
N
E0 (%)
E1 (%)
E2 (%)
E3 (%)
Easy Middle Hard
6437 6374 6317
43.6 45.2 43.7
8.7 9.0 9.5
11.4 11.8 12.6
36.3 33.9 34.2
The distribution is essentially flat across McEval task difficulty levels (E0 43.6/45.2/43.7%), suggesting that boundary-state tracking failures are not primarily explained by McEval semantic task difficulty. G.5
A×B Joint Distribution (Pooled)
The body Table 14 reports the boundary failure rate (E1∪E2∪E3) per cell. Table 23 pools the full E0/E1/E2/E3 distribution across all 9 models to surface which boundary mechanism each (A, B) cell preferentially triggers (A3 included). Two patterns merit attention. (i) The A2 row is E3-dominant across all B conditions (49–72% latent collision); explicit wrapping makes E3 the dominant failure mode, while inner nesting modulates the E1/E2/E3 mix (e.g., A2×B2 still carries substantial E1 mass). (ii) A1×B2 and A1×B3 are the only A1 cells where E2 (unclosed fence) becomes the dominant failure mode (12.8% and 11.6%); these are the conditions where the model must produce literal-syntax inner blocks without an outer wrapper, and a non-trivial fraction of responses leave an inner fence open. A3 cells fall between A1 25
Table 23: Error-type composition per A×B cell, averaged across all 9 models. Each row sums to 100% (rounding aside); E0 = correct (F2∧F3 pass), E1 = premature outer-fence closure, E2 = unclosed outer fence, E3 = latent collision (F2 pass but F3 fail). Bold marks the dominant error type per row (excluding E0), exposing which boundary mechanism each (A, B) cell preferentially triggers. A
B
E0 (%)
E1 (%)
E2 (%)
E3 (%)
A1
B1 B2 B3 B4
79.7 73.3 71.5 79.6
1.2 5.1 5.2 0.3
1.4 12.8 11.6 0.8
17.8 8.9 11.7 19.3
A2
B1 B2 B3 B4
9.5 1.7 1.6 5.2
5.7 37.1 20.9 8.1
22.1 12.4 16.6 14.3
62.7 48.7 60.9 72.4
A3
B1 B2 B3 B4
53.5 44.6 49.5 59.6
2.2 16.7 6.2 0.9
5.7 23.5 18.6 1.8
38.6 15.1 25.7 37.8
and A2 with E3 as the dominant failure type (15–39%), reflecting the base-rate wrapping behavior measured in Step 1 of the A1 decomposition (Table 5). G.6
A×B Joint Distribution (Per-Model)
Table 24 disaggregates the same A×B×E breakdown to each of the 9 models. The per-model view reveals mechanism fingerprints that the pooled summary obscures: some Large/API models, especially Claude-Sonnet-4 and GPT-4o, are E3-dominant under A2 (latent collision ≥92% of the cell mass), while small models such as Gemma-2-9B and Llama-3.1-8B exhibit substantial E2 (unclosed fence) mass under A2, indicating that smaller models additionally leave fences unclosed rather than failing only by length-safety.
G.7
Model
B1
A1 (Direct) B2 B3
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
2/1/2 2/3/0 0/0/0
19/1/3 14/4/26 0/0/0 11/5/75 74/6/18 22/5/72 9/0/83 2/6/0 2/4/0 0/0/0 9/48/3 66/21/12 33/45/15 2/79/7 1/6/0 0/3/0 0/1/0 6/64/23 19/38/33 13/65/18 6/31/49
Mid
Qwen3-32B Gemma-3-27B
0/1/51 2/2/15 2/1/16 0/1/75 6/5/86 46/7/44 38/2/60 31/2/66 6/7/77 16/12/47 27/4/53 3/6/85 2/22/72 22/13/64 31/11/58 2/2/92
Large
Table 24: Per-model error mechanism breakdown across A×B conditions. Each cell shows E1 / E2 / E3 percentages (rounded to integers); E0 (correct) is implied as 100 − sum, and the body Table 14 boundary failure rate equals the sum. Bold marks the dominant mechanism per cell. Pooled (model-averaged) version: Table 23.
Llama-3.1-70B 0/0/0 0/84/0 GPT-4o 0/0/17 0/2/0 Gemini-2.5-Flash 1/1/13 7/2/14 Claude-Sonnet-4 0/0/0 0/0/0
B4
B1
A2 (Wrapped) B2 B3
B4
0/79/0 0/0/0 8/53/32 18/27/56 28/23/49 8/14/72 0/1/0 0/0/11 2/0/92 2/0/98 2/0/97 0/0/98 2/8/10 0/1/2 9/1/83 47/0/53 19/0/81 15/0/84 0/1/0 0/0/0 0/0/99 41/0/59 1/0/99 0/0/100
Per-Language Latent Failure Distribution
Among code-heavy languages the latent failure rate is uniformly high (Shell 46.2%, C# 45.1%, C++ 42.9%, TypeScript 42.8%, JavaScript 42.2%, Python 41.5%, C 41.5%, Go 40.1%). Markup languages have lower latent failure (HTML 28.9%, JSON 39.0%, Java 36.8%); Markdown is the lowest (9.2%) but only because its content correctness is also lowest (21.6%) — latent failure requires content-correct = True, so when the model fails to produce a valid Markdown document it cannot enter the latent-failure cell at all. The substantive finding is that wherever the model can produce a meaningful document, the document is roughly equally likely to be boundary-broken across all programming languages. 26
Table 25: Latent failure rate per language (pooled across all 9 models) with Wilson 95% confidence intervals (computed on valid N after truncation filtering). Latent Failure = content-correct ∧ boundarywrong (F2/F3 fail). Languages are sorted by latent failure rate (descending). Each cell shows rate [CI lower, CI upper]. Code-heavy languages (Shell, C#, C++, TypeScript) have the highest latent failure rate (43–46%): models produce structurally correct content but break the outer fence boundary. Markdown is the lowest (9%) only because its content correctness is also lowest (22%) — latent failure requires content-correct = True, so when the model fails to produce a valid Markdown document it can never enter the latent-failure cell. LANG
Valid N
Content Correct (%)
Latent Failure (%)
Shell C# C++ TypeScript JavaScript Python C Go JSON Java HTML Markdown
1,611 1,608 1,583 1,579 1,580 1,591 1,592 1,568 1,595 1,611 1,603 1,607
74.7 [72.5, 76.7] 75.9 [73.7, 77.9] 74.0 [71.8, 76.1] 74.9 [72.7, 76.9] 75.0 [72.8, 77.1] 73.7 [71.5, 75.8] 70.5 [68.3, 72.7] 72.4 [70.1, 74.5] 63.3 [60.9, 65.7] 68.5 [66.2, 70.7] 53.2 [50.8, 55.6] 21.6 [19.7, 23.7]
46.2 [43.8, 48.6] 45.1 [42.7, 47.5] 42.9 [40.5, 45.3] 42.8 [40.4, 45.3] 42.2 [39.7, 44.6] 41.5 [39.1, 44.0] 41.5 [39.1, 44.0] 40.1 [37.7, 42.6] 39.0 [36.6, 41.4] 36.8 [34.5, 39.2] 28.9 [26.8, 31.2] 9.2 [ 7.9, 10.7]
Overall
19,128
66.4 [65.7, 67.1]
38.0 [37.3, 38.7]
H
Joint and Cascade Failure Analyses
H.1
F1×F2 Joint Failure Analysis
F1+F2+
F1+F2−
F1−F2+
F1−F2−
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
49.2 35.4 43.2
10.4 10.9 4.9
24.8 33.6 28.9
15.6 20.1 23.0
Mid
Qwen3-32B Gemma-3-27B
40.5 29.1
10.7 7.0
45.4 47.0
3.3 16.8
Large
Table 26: F1×F2 joint distribution per model (% of valid main experiment responses, N ≈2160 each). F1 = Instruction Compliance (A1+A2 directive followed), F2 = Fence Balance (closing fence present and length-matched). Cells sum to 100% per row (rounding aside). Model
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
38.1 63.6 52.9 63.1
18.9 0.6 8.6 3.5
20.2 35.5 27.1 33.3
22.7 0.3 11.5 0.0
The four cells decompose responses by whether the model followed the A1/A2 instruction (F1) and whether the closing fence is present and length-matched (F2). The F1+F2− cell isolates instructioncompliant but boundary-broken outputs—the locus of the wrapper-induced collisions discussed in §5.7. The F1−F2+ cell, in contrast, captures responses that ignored the wrap directive yet still produced a balanced inner block; F1 and F2 are not independent in either direction. H.2
Block Compliance by Boundary Condition
C3 (Block Compliance) measures whether the response contains the number of literal Markdownsource fenced blocks required by the B-condition, independent of fence balance (F2) or length safety (F3). Because C3 requires zero literal-source blocks for B1 and B4, these cells are near-saturated by construction; the informative contrast is B2 (one literal-source block required) and especially B3 (two required), where small models drop to single digits while large models retain ≥50%. This stratification motivates separating Content (C1∧C2∧C3) from Boundary (F2∧F3): a B3 response can pass Content yet still fail at the fence boundary. 27
Model
B1
B2
B3
B4
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
100.0 100.0 100.0
90.5 40.6 58.7
41.9 5.9 5.6
100.0 100.0 100.0
Mid
Qwen3-32B Gemma-3-27B
100.0 100.0
88.1 94.2
84.8 85.2
100.0 100.0
Large
Table 27: C3 (Block Compliance) pass rate (%) per model, stratified by B-axis (inner nesting complexity). C3 requires the response to contain a code block matching the requested structural type. B1=single code block, B2=nested single, B3=nested multiple, B4=mixed elements. Computed on main experiment responses (N ≈540 per (model, B) cell).
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
100.0 100.0 100.0 100.0
92.0 92.8 85.1 95.9
36.7 49.6 64.6 85.0
100.0 100.0 100.0 100.0
Table 28: Decomposition of E2 (Unclosed Fence) responses by re-parsing the raw response text and inspecting the last fence-line that follows the outer opening. Sub-patterns: No close (no fence-line at all ⇒ model ended in prose); Short close (last fence is same family but shorter than the outer, so CommonMark §4.5 ignored it as a close); Wrong family (last fence is the opposite family, e.g. tilde when outer is backtick); Imbalanced mid (last fence matches outer length+family but the stack stayed non-empty, indicating internally inconsistent nesting earlier); No outer (no opening fence detected, mostly A1 responses without wrapper). Percentages are of total responses per model; the five sub-pattern rates per row sum to the E2 total.
H.3
Model
E2 (%) No close Short close Wrong family Imbal. mid No outer
Qwen3-7B Gemma-2-9B Llama-3.1-8B Qwen3-32B Gemma-3-27B Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
6.9 21.4 24.0 3.5 10.5 36.5 0.5 2.8 0.0
0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.1 0.0
0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
6.9 21.4 24.0 3.5 10.5 36.5 0.5 2.7 0.0
0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
Overall
11.9
0.0
0.0
0.0
11.9
0.0
E2 Unclosed-Fence Sub-Patterns
We re-parse the raw response text of every E2 record with a stack-based CommonMark fence walker to characterize how the outer fence ended up unclosed. Across all 9 models, nearly every E2 response falls into the Imbalanced mid sub-pattern: the model successfully wrote at least one matching inner-block close but the cumulative stack ended non-empty. Pure No close (model wrote no fence-line at all after the outer opening) accounts for < 0.2% of total responses, even on the worst model. Short close attempts (a closing fence shorter than the outer) are also negligible because CommonMark requires fence runs of length ≥ 3, so any plausible close attempt would already match a 3-backtick outer. The takeaway: E2 is not a “model gave up writing” failure mode but rather a nesting-bookkeeping failure—models produce structurally complex output where the cumulative open/close balance silently leaves the outer wrapper unclosed.
H.4
Error Cascade Analysis
The conditional gap between P (content-correct | E1) and P (content-correct | ¬E1) is small in aggregate (70.0% vs 66.1%), suggesting that premature-closure failures do not systematically truncate or omit content; content correctness and fence-tracking failure are not tightly coupled in this diagnostic. Per-model rows show high variance (e.g., GPT-4o: 50.0/80.2%, Claude-Sonnet-4: 100.0/92.8%); small NE1 in some rows (GPT-4o n=10, Claude n=75) limits the per-model interpretation. 28
Table 29: Error cascade: probability the response is content-correct conditional on whether E1 (premature closure) fired. A large gap between the two columns indicates that early-closure failures tend to truncate or omit content (cascade effect); a small gap means boundary failures and content failures are decoupled. NE1=true
P(C | E1) (%)
NE1=false
P(C | ¬E1) (%)
Qwen3-7B Gemma-2-9B Llama-3.1-8B Qwen3-32B Gemma-3-27B Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
407 207 83 226 288 113 10 326 75
91.2 5.8 45.8 54.0 87.2 77.0 50.0 77.6 100.0
1734 1950 2063 1926 1870 2046 2150 1569 2085
64.1 16.5 48.3 73.7 74.3 68.6 80.2 75.8 92.8
Overall
1735
70.0
17393
66.1
Model
Table 30: Statistical test plan. All families of multiple comparisons use Holm–Bonferroni correction where applicable. Comparison
Test
Effect Size
Between models (failure rate) Within model, between conditions Outer length × failure C3 sensitivity (model ordering)
Chi-squared / Fisher’s exact McNemar’s test (paired) Cochran-Armitage trend Kendall’s τ
Cramér’s V Odds ratio – –
Confidence intervals: 95% Wilson score intervals (binomial proportion, valid N denominator). Stability: CV and ICC(1,1) from 30 prompts × 5 runs.
I
Statistical Tests
I.1
Inter-Model Agreement (Fleiss’ Kappa)
Table 31: Fleiss’ kappa for inter-model agreement on the binary boundary-correct judgment (F2∧F3). Computed per tier (Small/Mid/Large) and overall across the 9 models. Higher kappa indicates that models tend to fail (or pass) on the same prompts; values near 0 indicate near-independent failure modes. Group Small Mid Large Overall (all 9)
Models
Prompts N
Fleiss’ κ
3 2 4 9
2125 2150 1895 1868
0.433 0.008 0.392 0.349
Aggregate Fleiss’ κ = 0.349 across the 9 models indicates moderate but far-from-perfect agreement on which prompts cause boundary failures. The Mid tier’s near-zero κ (+0.008) reflects Gemma-327B’s anomalously high failure rate combined with Qwen3-32B’s much milder profile — the two models show little agreement on which prompts fail, indicating that boundary-failure susceptibility is not a single “hard prompt” axis but at least partially model-specific. I.2
Cochran–Armitage Trend Test
The pooled Cochran–Armitage Z-statistic is +6.17 (p < 10−3 ), providing evidence for a positive ordered trend in E3 (latent collision) rate with B-condition complexity, despite non-monotone per-cell profiles in some models. At the per-model level, 5/9 models show statistically significant positive trends; the four exceptions (Gemma-2-9B, Qwen3-32B, GPT-4o, Claude-Sonnet-4) either have very low E3 rates throughout (Gemma-2-9B) or show U-shaped rather than monotone profiles (Qwen3-32B, GPT-4o have lower E3 at B2/B3 than at B1/B4). 29
Table 32: Cochran–Armitage trend test for E3 (latent collision) rate as a function of B-condition complexity (B1→B2→B3→B4 treated as equally-spaced dose levels). A positive Z indicates that E3 rate increases monotonically with nesting complexity. All p-values two-sided. Model
B1 (%)
B2 (%)
B3 (%)
B4 (%)
Z
p
Qwen3-7B Gemma-2-9B Llama-3.1-8B Qwen3-32B Gemma-3-27B Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
53.7 1.1 7.6 67.0 73.7 10.6 63.5 46.7 33.1
13.2 3.9 11.1 30.1 56.4 18.5 33.5 31.3 19.8
47.0 5.0 6.0 36.1 58.5 16.3 35.6 57.5 33.0
52.6 2.2 16.3 62.1 86.1 24.1 67.0 45.5 33.3
+3.20 +1.36 +3.56 -0.95 +4.43 +5.27 +1.31 +2.32 +1.56
0.00137 0.174 < 10−3 0.343 < 10−3 < 10−3 0.191 0.0205 0.12
Overall (pooled)
39.6
24.1
32.5
43.3
+6.17
< 10−3
Table 33: Chi-squared pairwise model comparisons of boundary-correct judgments on the main experiment (N = 2,160 prompts per model). Cells show Cramér’s V effect size; significance markers reflect Holm-Bonferroni-corrected p-values: ∗ p < 0.05, ∗∗ p < 0.01, ∗∗∗ p < 0.001. Lower triangle only; diagonal blank. G2-9B L3.1-8B Q32B G3-27B L3.1-70B GPT-4o Gemini-2.5-Flash Claude-S-4
I.3
Q7B
G2-9B
L3.1-8B
0.33∗∗∗ 0.29∗∗∗ 0.05∗∗ 0.31∗∗∗ 0.09∗∗∗ 0.17∗∗∗ 0.02 0.34∗∗∗
Q32B
0.04∗∗ 0.29∗∗∗ 0.61∗∗∗ 0.25∗∗∗ 0.17∗∗∗ 0.31∗∗∗ 0.01
0.25∗∗∗ 0.57∗∗∗ 0.36∗∗∗ 0.21∗∗∗ 0.04∗ 0.13∗∗∗ 0.12∗∗∗ 0.27∗∗∗ 0.02 0.05∗∗ 0.30∗∗∗
G3-27B L3.1-70B GPT-4o Gemini-2.5-Flash
0.39∗∗∗ 0.46∗∗∗ 0.08∗∗∗ 0.34∗∗∗ 0.06∗∗∗ 0.14∗∗∗ 0.61∗∗∗ 0.26∗∗∗ 0.18∗∗∗
0.32∗∗∗
Chi-Squared Pairwise Comparisons
Pairwise comparison of boundary-correct judgments between every pair of the 9 models. Cramér’s V in the lower triangle ranges from 0.01 (Claude vs Gemma-2-9B, negligible effect size) to 0.61 (Claude vs Gemma-3-27B, very large effect). Effect sizes correlate with the per-model Latent Failure rates reported in Table 7: model pairs with similar latent-failure profiles cluster together with low V , while pairs spanning the catastrophic and proficient extremes show V ≥ 0.5. The four highest-V pairs all involve Gemma-3-27B (the dominant outlier) or Claude-Sonnet-4 (the dominant proficient). I.4
McNemar Paired Test (A1 vs A2)
Table 34: McNemar paired test for boundary-correct judgments under A1 (Direct) versus A2 (Wrapped), with each prompt observed under both conditions. b = discordant pairs A1=pass / A2=fail; c = discordant pairs A1=fail / A2=pass. The odds ratio b/c quantifies the asymmetry: A2 is dramatically worse than A1 in nearly every discordant case. Quantity
Value
Discordant pairs b (A1=pass, A2=fail) Discordant pairs c (A1=fail, A2=pass) Odds ratio b/c McNemar χ2 statistic p-value (two-sided)
4,509 4 1127.25 4495.02 < 10−3
The McNemar paired test asks: of prompts where A1 and A2 outcomes differ, how asymmetric is the difference? The answer is decisively asymmetric (b/c ≈ 1127): in 4,509 paired prompts the model passes under A1 (no wrapper) but fails under A2 (wrapper required), while only 4 prompts show the reverse pattern. This is the within-prompt analogue of the marginal ∼95% A2 failure observed in Table 14: the wrapper directive does not just shift the marginal failure rate—it flips the per-prompt outcome in nearly every discordant case. 30
J
Cross-Format and Auxiliary Probes
J.1
Cross-Format Analysis
Model
Python (triple-quote collision) K1 K2 K3
JSON (nesting depth) K1 K2 K3
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
93.3 100.0 100.0
93.3 93.3 93.3
80.0 100.0 93.3
100.0 100.0 100.0
100.0 100.0 100.0
100.0 100.0 100.0
Mid
Qwen3-32B Gemma-3-27B
40.0 73.3
60.0 86.7
86.7 46.7
100.0 100.0
100.0 100.0
100.0 100.0
Large
Table 35: Cross-format parse-validity rate (%) per model. Python conditions test escalating triplequote (""" / ''') collision pressure inside docstrings (K1: literal mention; K2: both delimiter styles; K3: nested triple-quote in docstring code example). JSON conditions test escalating brace-nesting depth (K1: 1–2 levels; K2: 3–4 levels; K3: ≥5 levels). Pass = the extracted code block parses successfully via Python ast.parse or json.loads; failure = syntax/parse error. JSON serves as a partial negative control: its asymmetric braces do not exhibit the same length-matching collision as Markdown fences, so we expect higher pass rates regardless of nesting depth.
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
46.7 0.0 0.0 80.0
86.7 93.3 73.3 20.0
26.7 80.0 20.0 93.3
100.0 100.0 100.0 100.0
100.0 100.0 100.0 100.0
100.0 100.0 100.0 100.0
We test whether the symmetric-delimiter collision phenomenon extends beyond Markdown by replicating the experiment in two contrasting target formats: Python (symmetric triple-quote docstrings) and JSON (asymmetric braces). Each prompt asks the model to produce a syntactically valid Python function or JSON document under escalating delimiter-collision pressure (K1→K3); we evaluate by attempting to parse the extracted code block with Python ast.parse or json.loads. The pattern in Table 35 mirrors the Markdown finding: JSON, the asymmetric control, achieves 100% parse-validity across every model and every nesting depth, while Python triple-quote conditions exhibit large per-model failure rates, with several Large models hitting 0% on the seemingly easiest condition (K1: “mention the triple-quote delimiter in the docstring text”), where literally placing the delimiter inside the docstring terminates it prematurely. The K2 condition (asking for both """ and ''' styles) often improves the pass rate over K1 because the prompt itself nudges the model toward delimiter alternation as a workaround. These results support the claim that symmetric-delimiter boundary tracking can fail beyond Markdown, while JSON remains robust as an asymmetric-delimiter control. J.2
Inline Code Span Results
Model
Run=0
Run=1
Run=2
Small
Qwen3-7B Gemma-2-9B Llama-3.1-8B
0.0 0.0 11.1
100.0 100.0 100.0
100.0 100.0 100.0
Mid
Qwen3-32B Gemma-3-27B
0.0 11.1
100.0 100.0
100.0 100.0
Large
Table 36: Inline code span failure rate (%) per (model, run). Failure = the model fails to produce any inline span whose CommonMark-rendered content exactly equals the keyword text specified in the prompt (inline_match). Run = N denotes N literal backticks placed around the keyword content (so safe wrapper length > N ). Cells are computed via markdown-it-py CommonMark parsing of the raw response and exact match against the prompt’s verbatim target. Each cell is over 9 records (3 keyword groups × 3 sentence templates) per (model, run).
Llama-3.1-70B GPT-4o Gemini-2.5-Flash Claude-Sonnet-4
0.0 44.4 0.0 0.0
100.0 100.0 100.0 66.7
100.0 100.0 100.0 55.6
Overall
7.4
96.3
95.1
31
Setup. Inline code span experiments use 27 unique prompts (3 sentence templates × 3 keyword groups × 3 runs) × 9 models = 243 generations. The prompt asks the model to display a verbatim text (foo, `foo`, or ``foo`` for run=0/1/2) inside a single inline code span. Run=0 has no collision pressure (the keyword text contains no backticks); run=1 places one backtick on each side; run=2 places two backticks on each side. The safe wrapper length follows the CommonMark inline rule: wrapper backtick run length must be strictly greater than the longest same-family run inside the span content (run=1 needs wrapper length ≥ 2, run=2 needs wrapper length ≥ 3, with appropriate space-padding so the literal backticks display). Evaluator. The legacy block-fallback metric (any fenced code block in the response) under-counted real inline failures because it never inspected the inline span’s rendered content. We re-evaluate every inline response with a CommonMark parser (markdown-it-py): for each response we extract all code_inline tokens, take their rendered content, and pass the response if and only if at least one token’s content exactly matches the prompt’s verbatim target. Wrapper length is recovered post hoc from the raw text and an auxiliary inline_F3_safe metric (wrapper > inner max same-family run) is computed for the matched span. Findings. Table 36 reports the verbatim-match failure rate. At run=0 (no collision) the pooled failure is only 7.4%, dominated by a single model (GPT-4o, 44.4%) that produces an inline span enveloping an entire introductory phrase rather than the keyword alone. Under inline collision pressure (run=1 / run=2) the picture inverts: 8 of 9 models fail 100% of cases at both runs. Only Claude-Sonnet-4 handles the collision partially (66.7% fail at run=1, 55.6% at run=2), and even Claude’s non-failures use exactly run+1 backticks for the wrapper without any further safety margin. Note on interpretation: the prompt’s verbatim target (`foo` or ``foo``) appears as inline code in the prompt source, so the dominant failure mode (Qwen3, Gemma, Llama; Backtick omission below) reflects models defaulting to wrap the bare keyword in a single backtick — structurally safe (inline_F3_safe = True) but verbatim-mismatched. The genuine collision-tracking failures concentrate in the asymmetric-close and length-matched-collision modes; the four-cell decomposition (§4.4) classifies the dominant mode as format_only rather than latent failure, and no inline record is classified as latent (a wrapper too short to span backtick-bearing content cannot also produce a verbatim-correct rendering). Per-model fail-mode taxonomy. The 96%+ pooled fail rate decomposes into four qualitatively distinct error modes, which we identified by inspecting the parsed spans against the prompt’s target: • Backtick omission (Qwen3, Gemma, Llama at run=1/2): the model writes the keyword without the surrounding literal backticks, producing a span like `foo` instead of the requested ``foo``. This is a verbatim-display mismatch rather than a boundary-safety failure: the produced span is usually structurally safe, but its rendered content does not match the prompt’s literal target. • Whole-sentence wrap (GPT-4o run=1, run=2): the model wraps the entire introductory sentence (“Here is some inline code: `foo`.”) in one inline span instead of isolating the keyword. Wrapper length is structurally safe (2-bt for run=1) but semantic content is wrong. • Asymmetric/short close (Claude run=2 partial): the model opens the span with ≥3 backticks but closes with a shorter run, producing parses like ``foo` (mismatched) or breaking into multiple disjoint spans. • Length-matched collision (Llama-3.1-70B, Gemini under high collision): the model uses a wrapper length equal to the keyword’s internal backtick run, so the parser silently splits the response into multiple inline spans (the inner literal backticks act as premature closing). Phenomenologically identical to the E1 premature closure observed at the block level. Synthesis. The block-level boundary-state tracking failure (§5.2) is not an artefact of fence-line semantics: the asymmetric-close and length-matched-collision modes documented above reproduce, for inline spans, the same default-shortest-wrapper heuristic that drives E3 latent collision in fenced blocks. The aggregate 96%+ verbatim-match failure additionally absorbs an interpretation default (Backtick omission), which inflates the headline number without itself constituting a collisiontracking failure. Even after netting out interpretation defaults, the structural failures observed in the asymmetric-close and length-matched-collision modes support the view that boundary-state tracking under symmetric, length-matched delimiters is a generation-time bottleneck rather than a 32
Markdown-fence-specific quirk; inline therefore serves as a complementary probe to the block-level evidence rather than an independent generalization. J.3
Temperature Sensitivity (Future Work)
Temperature sensitivity analysis—testing whether boundary failure patterns persist across the full temperature range 0.0–1.0—is left to future work. Our main experiments use temperature= 0 (greedy decoding), which ensures deterministic outputs but leaves open the question of whether higher temperatures exacerbate or mitigate boundary failures.
K
Practical Evidence: Full List
Table 37: Complete list of independently reported boundary-state tracking failures across LLM platforms and developer tools. URLs verified at the time of submission. Platform
Source
OpenAI
chatkit-js #89
Title and URL Nested triple-backtick code blocks break https://github.com/openai/chatkit-js/issues/89
OpenAI
Community (Bugs)
Use longer Markdown fences so code blocks don’t break https://community.openai.com/t/chatgpt- use- longer- markdownfences-so-code-blocks-don-t-break/1357890
OpenAI
Community (Bugs)
Minor code block bug with nested code blocks https : / / community . openai . com / t / bug - minor - code - block - bug with-nested-code-blocks/1018718
OpenAI
Community (Bugs)
Missing triple backquote in Markdown code block https://community.openai.com/t/missing- triple- backquote- inmarkdown-code-block/1234674
OpenAI
Community (API)
Markdown formatting issues with GPT-5 https : / / community . openai . com / t / markdown - formatting - issues with-gpt-5/1337570
gemini-cli #10515
CLI Markdown rendering is broken https://github.com/google-gemini/gemini-cli/issues/10515
Microsoft
vscode #295126
create_file tool wraps extra backtick fences https://github.com/microsoft/vscode/issues/295126
JetBrains
YouTrack LLM-1623 Markdown code snippets generated broken https://youtrack.jetbrains.com/issue/LLM- 1623/Markdown- codesnippets-generated-broken
Open WebUI #5016
Incorrect rendering of nested code blocks https://github.com/open-webui/open-webui/issues/5016
LangChain4j #1446
JSON wrapped in backticks extraction failure https://github.com/langchain4j/langchain4j/issues/1446
Obsidian
Forum #60565
Allow nested code blocks / allow triple backticks in code blocks rendering result https://forum.obsidian.md/t/allow- nested- code- blocks- allowtriple-backticks-in-code-blocks-rendering-result/60565
Block
goose #8290
Fenced code blocks break message layout https://github.com/block/goose/issues/8290
Streamdown #473
Fenced code blocks don’t render incrementally during streaming https://github.com/vercel/streamdown/issues/473
Continue
#6059
LLM generate markdown with broken backquotes https://github.com/continuedev/continue/issues/6059
Medium
Cultman Sachs
Why Can’t AI Models Output Clean Markdown? A Technical Mess That Still Isn’t Fixed https://medium.com/@CultmanSachs/why- cant- ai- models- outputclean - markdown - a - technical - mess - that - still - isn - t - fixed 1dc70ff366a3
Medium
Daniel Olshansky
Escaping Backticks in your LLM System Prompt https://olshansky.medium.com/escaping- backticks- in- your- llmsystem-prompt-6507a25b7bc8
33