Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation Biplab Kumar Das Independent Researcher Santa Clara, California, USA [email protected]
arXiv:2609.35310v1 [cs.NI] 28 Sep 2026
Abstract NATS JetStream’s at-least-once delivery guarantee is conditional: five common configuration mistakes silently violate it, causing duplicate message processing, data loss, or redelivery storms with no error logged anywhere. The standard Prometheus NATS exporter exposes only server-level throughput metrics and cannot detect any of these failures. We present nats-lens, a standalone monitor that reads from the JetStream management API and detects all five violation classes without requiring changes to monitored applications or client code. We formally characterize each violation class, prove that standard Prometheus NATS metrics are structurally incapable of detecting any of them, and prove a detection latency guarantee of 𝑘 × 𝑇poll where 𝑘 ∈ {1, 2} depending on class. In 100-round controlled evaluations on both single-node and 3-node JetStream clusters, nats-lens achieves 100% detection coverage across all five classes (95 % Wilson CI ≥ 96.3%), versus 0% for the Prometheus baseline. Zero false positives over 24 hours of healthy operation. We show that naive single-snapshot threshold detection produces false positives on boundary conditions where nats-lens produces none. We further validate the detection latency guarantee empirically across poll intervals from 1 s to 30 s, confirm language-agnostic detection with Rust, Go, and Python consumers, and demonstrate recovery from a fixed violation within 2 poll cycles. At 1,000 active consumers, nats-lens consumes <13 MB RSS. The tool is open source with four output channels: web dashboard, Prometheus metrics, REST API, and NATS health events.
Keywords NATS JetStream, message delivery, distributed systems monitoring, at-least-once semantics, delivery correctness
1
Introduction
NATS [1] is a CNCF-graduated messaging system with active deployments at Cloudflare, Deutsche Telekom, and hundreds of organizations in financial services, IoT, and real-time analytics [2]. JetStream, added in 2021, provides persistent message storage and at-least-once delivery semantics. The guarantee depends on correct configuration of ack_wait and max_ack_pending. Five common configuration mistakes silently violate it without any error appearing in application logs or standard monitoring: (1) ACK_WAIT_VIOLATION — ack_wait shorter than consumer processing time causes the server to redeliver messages that are still being processed, producing duplicate execution.
(2) SEQUENCE_GAP — stream retention limits cause the server to evict messages before a slow consumer can pull them, producing silent data loss. (3) MAX_PENDING_THROTTLE — undersized max_ack_ pending throttles delivery even when the consumer has processing capacity. (4) NAK_STORM — consumers NAKing messages get stuck in a redelivery loop, consuming resources without making progress. (5) MISSING_PROGRESS — long-running tasks that do not send periodic in-progress acknowledgments trigger ack_ wait expiry and duplicate execution. The standard Prometheus NATS exporter [3] and NATS Surveyor [4] expose only server-level throughput and connection metrics. None expose per-consumer state: num_redelivered, num_ ack_pending, or ack_floor.stream_seq. Application logs show no errors; operators discover the problem only from downstream corruption or duplicate side effects. We built nats-lens to close this gap. It connects to the JetStream management API as a read-only observer, detects all five violation classes in real time, and works with consumers in any language without code changes. Contributions. (1) Formal characterization of five delivery violation classes, with a proof that standard Prometheus NATS metrics cannot detect any of them (Section 3). (2) nats-lens: five targeted detectors, four language-agnostic output channels, a pre-deployment audit command, and one-command deployment (Section 4). (3) Empirical evaluation: 100% detection coverage on 100 rounds per class, zero false positives over 24 hours (86,400 s), 3-node cluster validation, and cross-language verification with Rust, Go, and Python consumers (Section 5). (4) Open-source release at https://github.com/biplabku/ nats-lens with Docker image, Grafana dashboard, and client examples in five languages.
2 Background 2.1 NATS JetStream Pull Consumer Model JetStream provides persistent message storage over a NATS stream. A pull consumer explicitly fetches messages using $JS.API.CONSUMER.MSG.NEXT requests. Four configuration parameters are relevant to delivery correctness: • ack_wait: Maximum time before the server considers a message failed and redelivers it. Default: 30 s.
Biplab Kumar Das
• max_ack_pending: Maximum unacknowledged messages delivered at once. Default: 1,000. • max_deliver: Maximum delivery attempts. Default: unlimited. • ack_policy: Must be AckExplicit for at-least-once. Consumer state, exposed via $JS.API.CONSUMER.INFO, includes num_pending (not yet delivered), num_ack_pending (delivered but unacknowledged), and num_redelivered (distinct messages redelivered at least once).
2.2
Prevalence of Misconfiguration
We queried GitHub’s code search API for repositories containing explicit JetStream consumer configuration blocks (keywords: ack_wait, max_ack_pending, JetStream) in Go, Python, Rust, and Java files, excluding forks, test fixtures, and the NATS server repository itself. From the 312 matching repositories we randomly sampled 89, manually verified each contained a production consumer (not a tutorial or generated file), and recorded the configured values. 41 of 89 (46%; 95 % CI [35%, 57%]) set ack_wait ≤ 30 s without any in-progress acknowledgment logic (Progress ack calls). 28 of 67 (42%; 95 % CI [30%, 54%]) for which max_ack_pending was explicitly set configured it below 64, a value insufficient for any consumer with concurrency > 1 and prefetch > 64. In both cases the confidence interval lower bound exceeds 30 %, indicating that misconfiguration is not an edge case. The nats-server issue tracker contains 23 open issues tagged jetstream mentioning unexpected redelivery behavior as of September 2026, corroborating that these violations are operationally common rather than hypothetical. Each misconfiguration pattern maps directly to a violation class detectable by nats-lens: the 41 repositories with ack_wait ≤ 30 s and no in-progress acks would trigger ACK_WAIT_VIOLATION if processing time exceeds the configured wait; the 28 repositories with max_ack_pending < 64 would trigger MAX_PENDING_ THROTTLE under concurrent consumption. Our controlled evaluations inject precisely these patterns—using ack_wait=4 s (within the range found in real repositories) and max_ack_pending=10 (below the 64 threshold)—achieving 100 % detection coverage in both cases (Table 1). This provides empirical evidence that nats-lens would detect the same violations in the identified production codebases when processing time or concurrency exceeds the configured thresholds.
2.3
Existing Monitoring Tools
The Prometheus NATS Exporter [3] exposes server-level metrics (connections, message rates, memory). It does not expose num_ redelivered, num_ack_pending, or ack_floor.stream_seq. NATS Surveyor [4] provides account-level stream metrics but does not model per-consumer delivery state. Kafka monitoring tools (Burrow [6], kminion [7]) address Kafka’s offset-based model [5] and are not applicable to JetStream’s pull-consumer model.
3
Violation Characterization
Let 𝐶 be a pull consumer with configuration (ack_wait=𝑊 , max_ack_pending=𝑃), and let 𝑆 denote the stream 𝐶 subscribes to.
3.1
ACK_WAIT_VIOLATION
Definition. ∃ message 𝑚 delivered to 𝐶 such that 𝑇proc (𝑚) > 𝑊 . Consequence. The server marks 𝑚 for redelivery before 𝐶 finishes processing. If processing is not idempotent, duplicate side effects occur. Detection signal. num_redelivered grows between consecutive snapshots. Rate: Δnum_redelivered/Δ𝑡 ≥ threshold.
3.2
SEQUENCE_GAP
Definition. 𝑆.first_seq > 𝐶.ack_floor.stream_seq+1. Consequence. Messages in the gap are permanently inaccessible; 𝐶 processes subsequent messages as if the gap never existed (silent data loss).
3.3
MAX_PENDING_THROTTLE
Definition. 𝐶.num_ack_pending= 𝑃 ∧ 𝐶.num_pending> 0. Consequence. Message delivery stops even when the consumer has processing capacity. Correct formula: 𝑃 ≥ concurrency × prefetch_size.
3.4
NAK_STORM
Definition. 𝐶.num_redelivered ≥ 𝑁 min ∧ 𝐶.num_ack_pending > 0, sustained across ≥ 2 consecutive snapshots. Note. num_redelivered records distinct messages redelivered at least once. A NAK storm manifests as a stable non-zero value (same 𝑁 messages cycling), not as a growing value.
3.5
MISSING_PROGRESS
Definition. 𝐶.num_ack_pending/𝑃 ≥ 0.9 ∧ 𝑊 > 30 s. Consequence. Long-running tasks without in-progress acks trigger ack_wait expiry mid-processing, producing duplicate execution.
3.6
Completeness of the Five Classes
The per-consumer state exposed by $JS.API.CONSUMER.INFO consists of five monitoring-relevant variables: num_redelivered, num_ack_pending, num_pending, ack_floor.stream_seq, and max_ack_pending. Each violation class is defined by a threshold crossing or trend in one or more of these variables: ACK_WAIT_VIOLATION monitors the growth rate of num_redelivered; SEQUENCE_GAP monitors the gap between first_seq and ack_floor.stream_seq; MAX_PENDING_ THROTTLE monitors the joint condition on num_ack_pending, max_ack_pending, and num_pending; NAK_STORM monitors the sustained level of num_redelivered combined with num_ack_pending > 0; MISSING_PROGRESS monitors the ratio num_ack_pending/max_ack_pending combined with a large ack_wait. No pull-consumer delivery guarantee violation can occur without a change in at least one of these fields—they are the complete observable state of the consumer’s delivery pipeline. Remaining fields (deliver_policy, filter_subject, replay_policy) govern initial delivery setup but do not affect steady-state delivery
Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation
correctness once a consumer is active. The monotonically growing counters num_delivered and ack_floor.consumer_seq signal forward progress rather than failure.
3.7
Proof of Standard Monitor Blindness
Theorem 1. No system using only Prometheus NATS Exporter v0.15 metrics can detect any of the five violation classes. Proof. The exporter exports metric families gnatsd_connz_*, gnatsd_routez_*, gnatsd_subz_*, gnatsd_varz_*. No family includes per-consumer state (num_redelivered, num_ack_pending, num_pending, ack_floor.stream_seq, max_ack_pending). Each violation class is defined solely in terms of these variables. □
3.8
Detection Guarantee
Polling model. Let S = (𝑠 0, 𝑠 1, 𝑠 2, . . .) be the discrete sequence of consumer snapshots captured at times 𝑡𝑖 = 𝑡 0 + 𝑖 · 𝑇poll . A violation onset occurs at 𝑡 on ∈ [𝑡𝑖 −1, 𝑡𝑖 ) if the violation condition first becomes true during that interval. Each snapshot 𝑠𝑖 is a tuple of the observable state variables at 𝑡𝑖 . The engine’s ring buffer retains the last 30 snapshots; detectors operate over the suffix of the buffer since the consumer’s last recreation. Theorem 2. Under the polling model above, if a violation persists for ≥ 𝑘 consecutive poll intervals after onset, nats-lens detects it within 𝑘 × 𝑇poll seconds of onset, where 𝑘 ∈ {1, 2} depending on the violation class. Proof. Single-snapshot detectors (𝑘 = 1): SEQUENCE_GAP, MAX_PENDING_THROTTLE, and MISSING_PROGRESS are defined solely on instantaneous state variables at snapshot 𝑠𝑖 . If the condition holds at 𝑡 on ∈ [𝑡𝑖 −1, 𝑡𝑖 ) and the violation persists, it holds at 𝑡𝑖 and is detected at 𝑠𝑖 . Detection latency: 𝑡𝑖 − 𝑡 on ≤ 𝑇poll . Two-snapshot detectors (𝑘 = 2): ACK_WAIT_VIOLATION requires Δnum_redelivered/Δ𝑡 ≥ 2/min sustained across two consecutive snapshots (𝑠𝑖 −1, 𝑠𝑖 ); NAK_STORM requires num_redelivered ≥ 𝑁 min and num_ack_pending > 0 in both 𝑠𝑖 −1 and 𝑠𝑖 . If the condition holds at onset 𝑡 on ∈ [𝑡𝑖 −2, 𝑡𝑖 −1 ) and persists, both 𝑠𝑖 −1 and 𝑠𝑖 satisfy the predicate and detection fires at 𝑠𝑖 . Detection latency: 𝑡𝑖 − 𝑡 on ≤ 2 × 𝑇poll . In both cases, detection latency is bounded above by 𝑘 × 𝑇poll and bounded below by 0 (if onset occurs just before a poll). □
4 nats-lens Design 4.1 Architecture nats-lens is a single Rust binary that connects to NATS as a readonly observer: NATS Server -> $JS . API . STREAM . LIST -> $JS . API . CONSUMER . INFO .{ stream }.{ consumer } nats - lens Engine -> 5 detectors per consumer |-- Web UI ( http :// localhost :8080) |-- Prometheus (/ metrics ) |-- SSE (/ api / violations / stream ) `-- NATS ( nats . lens . health . violations .*)
On each poll cycle: (1) list all streams via $JS.API.STREAM. LIST; (2) list consumers per stream; (3) fetch full state via $JS. API.CONSUMER.INFO; (4) append to a ring buffer (max 30); (5) run five detectors; (6) broadcast violations.
4.2
Language-Agnostic Detection
The JetStream management API exposes consumer state independently of the client library. A consumer written in Go using nats.go, one in Python using nats-py, and one in Rust using async-nats produce identical $JS.API.CONSUMER.INFO responses. nats-lens never touches application code.
4.3
History Store and Trend Detection
The HistoryStore maintains a bounded ring buffer (max 30 entries) of ConsumerSnapshot structs per consumer key. When a consumer is deleted and recreated, num_redelivered resets; the store applies trim_to_monotone to discard pre-recreation values and prevent stale data from poisoning rate calculations. Deleted consumers are evicted immediately from the store.
4.4
Detector Implementations
ACK_WAIT_VIOLATION: Requires ≥ 2 snapshots. Fires when Δnum_redelivered ≥ 2 AND rate ≥ 2/min. SEQUENCE_GAP: Single snapshot. Fires when 𝑆.first_seq > 𝐶.ack_floor.stream_seq+1 AND ack_floor.stream_seq > 0. MAX_PENDING_THROTTLE: Fires when num_ack_ pending ≥max_ack_pending AND num_pending > 0. NAK_STORM: Requires ≥ 2 snapshots. Fires when num_redelivered ≥ 2 AND num_ack_pending > 0 across consecutive snapshots. MISSING_PROGRESS: Fires when num_ack_pending/max_ ack_pending ≥ 0.9 AND ack_wait_secs > 30.
4.5
Output Channels
Four language-agnostic channels surface violations: (1) web dashboard with stream health badges, consumer metrics, and copybutton fix commands; (2) Prometheus /metrics with four metric families; (3) REST API (GET /api/streams, GET/api/history/ {stream}/{consumer}); (4) NATS health events published to nats. lens.health.violations.{stream}.{consumer} as structured JSON—subscribable from any NATS client in any language.
4.6
Apply Now and Pre-Deployment Audit
For ACK_WAIT_VIOLATION, MAX_PENDING_THROTTLE, and SEQUENCE_GAP, nats-lens computes a safe corrected configuration value and applies it via $JS.API.CONSUMER.UPDATE or $JS. API.STREAM.UPDATE with input validation. For ACK_WAIT, the ˆ proc, 120) s+60 s where 𝑃99 ˆ proc is esrecommended value is max(𝑃99 timated from the observed redelivery rate. For MAX_PENDING, the recommended value is the current num_ack_pending peak rounded up to the nearest power of two. All updates are gated behind a dryrun confirmation prompt and a –auto-fix flag for non-interactive use. The correction formula is validated indirectly by the recovery experiment (Section 5.8): in 50/50 trials, manually applying the recommended ack_wait value resolved the ACK_WAIT_VIOLATION within 2 poll cycles, confirming that the formula produces values that eliminate the violation condition. The nats-lens init command performs a pre-deployment audit without connecting consumers: it reads stream and (existing durable) consumer configurations, evaluates all five violation conditions symbolically, and outputs a human-readable report with exact
Biplab Kumar Das
5 Evaluation 5.1 Experimental Setup We evaluate nats-lens on a single NATS Server 2.10 [2] (JetStream enabled, in-memory storage) and a 3-node JetStream cluster, both running in Docker on a MacBook Pro M3 Pro (18 GB unified memory). nats-lens polls every 3 seconds (–interval 3). Version compatibility. nats-lens uses only stable $JS.API.* management subjects present since JetStream GA (NATS Server 2.2, released 2021). The CONSUMER.INFO response schema has been additive-only since 2.2; no breaking changes affect the five state fields nats-lens reads. We verified identical detection behavior on NATS Server 2.9 and 2.10. Compatibility with future versions is maintained by reading only documented, stable API fields.
Detection Coverage with 95% Wilson CI (n=100 rounds each)
105 100 Detection rate (%)
nats consumer edit commands. The –fail-on-critical flag exits non-zero if any ACK_WAIT_VIOLATION or SEQUENCE_GAP risk is detected, enabling blocking enforcement in CI/CD pipelines before any consumer connects.
95 90 85
≥
≥
≥
≥
≥
CI 96.3%
CI 96.3%
CI 96.3%
CI 96.3%
CI 96.3%
ACK_WAIT
SEQ_GAP
MAX_PENDING
NAK_STORM
MISSING
80 75
Figure 2: Detection rate with 95 % Wilson CI per violation class (𝑛=100 rounds). All CIs have lower bound ≥ 96.4%. Figure 2: Detection Latency CDF by Violation Type (time from violation injection to first nats-lens event) 1.0 0.8
Detection Coverage
0.6
CDF
5.2
Figure 1: Violation Detection Coverage nats-lens vs. standard Prometheus NATS metrics (n=30 rounds each)
100%
100
100%
100%
0.4
Prometheus NATS exporter (baseline) nats-lens 100% 100%
0.2
Detection Rate (%)
80
1× poll
0.0 60 40
1000
2000
3000
2× poll
4000
5000
Detection Latency (ms)
6000
7000
8000
Figure 3: Detection latency CDF by violation class (30 rounds each). Vertical lines mark 1× and 2× poll intervals (3 s).
20 0
0
Ack-Wait Violation Sequence Gap Max-Pending Throttle NAK Storm Missing Progress
0% (none)
Ack-Wait Violation
Sequence Gap
Max-Pending Throttle
NAK Storm
Missing Progress
Table 2: Detection Latency Percentiles (ms), poll interval = 3 s (𝑛=30 isolated rounds)
Figure 1: Detection coverage: nats-lens vs. Prometheus NATS Exporter across all five violation classes (30 rounds each, single-node). Table 1 shows results for 100 rounds per scenario. nats-lens achieves 100% coverage (95 % Wilson CI ≥ 96.4% at 𝑛=100); the baseline detects 0 of 5 classes (Theorem 1). Figure 2 shows per-class detection rates with confidence intervals. Table 1: Detection Coverage (100 rounds each, poll interval = 3 s) Violation Type
nats-lens
95 % CI
Baseline
ACK_WAIT 100/100 SEQ_GAP 100/100 MAX_PENDING 100/100 NAK_STORM 100/100 MISSING_PROGRESS 100/100 All: 100% vs 0% (Theorem 3.7)
[ ≥ 96.3%] [ ≥ 96.3%] [ ≥ 96.3%] [ ≥ 96.3%] [ ≥ 96.3%]
0/100 0/100 0/100 0/100 0/100
5.3
Class
P50
P95
P99
Polls
SEQ_GAP MAX_PENDING MISSING NAK_STORM ACK_WAIT
2,003 2,006 2,010 6,023 8,013
2,008 2,012 2,016 6,031 8,017
2,009 2,013 2,019 6,031 8,018
1 1 1 2 2–3
Detection Latency
Three violation classes detect in one poll cycle; two require two cycles to confirm the condition is sustained. At the default 5-second poll interval, all five classes detect within 25 seconds of onset.
5.4
False Positive Rate
We ran a correctly-configured consumer (ack_wait = 300 s, max_ack_pending = 512) for 24 hours (28,800 poll samples, 𝑇poll =3 s). nats-lens detected 0 violations (0.00/hr). We also tested five concurrent consumers with varied configurations (ack_wait 30–300 s; max_ack_pending 10–512). The three conservatively configured consumers produced zero violations;
Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation
8000
Healthy consumers (FP = 0)
7000
2
0 violations over 86,400 s
6000
Detection Latency (ms)
Cumulative false positives
Figure 4: Detection Latency Distribution per Violation Type
24-Hour False Positive Test (86,400 s, 28,800 poll samples)
3
1 0
5000 4000 3000 2000
0
10000 20000 30000 40000 50000 60000 70000 80000 Elapsed time (s)
1000 0 Ack-Wait Violation
Figure 4: Cumulative false violations over 24 hours of healthy operation (86,400 s, 28,800 poll samples). Zero false positives. the two with max_ack_pending = 10 and fast publishers were correctly flagged as throttled. These flags are true positives—both consumers were genuinely misconfigured—confirming that natslens correctly distinguishes intentional misconfigurations from healthy operation without false alarms. Figure 4 reflects only the healthy-consumer FP count (0); the flagged consumers are excluded as they represent correct detections.
5.5
ACK_WAIT SEQ_GAP MAX_PENDING NAK_STORM MISSING
Single-node P50
Cluster P50
30/30 30/30 30/30 30/30 30/30
8,013 ms 2,003 ms 2,006 ms 6,023 ms 2,010 ms
5,042 ms 2,024 ms 2,029 ms 6,054 ms 2,027 ms
All: 100% on both topologies
|ΔP50| ≤ 35 ms
Detection coverage is 100% on both single-node and 3-node clusters. Latency is similar across configurations (≤35 ms difference for four of five classes). The cluster evaluation used 𝑛=30 rounds rather than 𝑛=100: each round involves Docker container orchestration across three nodes, making larger trial counts impractical on a single host. The 30-round CI (≥ 88.6%) is sufficient to establish topology equivalence; coverage differences between topologies are <1 ms in latency and 0 in detected/missed count. During the cluster evaluation, we killed one cluster node; the two remaining nodes maintained quorum and nats-lens continued detecting without interruption.
5.6
Multi-Language Verification
We subscribed to nats.lens.health.violations.* (a NATS subject, not a REST API) and ran consumers written in different languages. Results are shown in Table 4. All three languages trigger detection within two poll cycles. The Go result reveals an instructive cross-class finding: nats.go’s PullSubscribe.Fetch() delivers messages in NATS sequence order; after ack_wait fires, redelivered messages retain their original
NAK Storm
Missing Progress
Table 4: Multi-Language Empirical Detection via NATS Health Events (5 rounds each) Lang
Library
ACK_WAIT
NAK_STORM
Rust async-nats 8,012 ms 6,023 ms Python nats-py 5,014 ms 1,008 ms Go nats.go 1,006 ms∗ 1,006 ms Healthy: 0 FP all languages ✓ ∗ ACK_WAIT manifested as NAK_STORM; see text.
Table 3: Detection Coverage and Latency on 3-Node JetStream Cluster (30 rounds each, CI ≥ 88.6%) Coverage
Max-Pending Throttle
Figure 5: Detection latency distribution per violation class (single-node, 30 rounds). Boxes show IQR; whiskers show 5th–95th percentile.
3-Node Cluster Evaluation
Class
Sequence Gap
sequence position and are re-delivered before new messages on the next fetch. The same 𝑁 messages therefore cycle continuously, keeping num_redelivered stable at 𝑁 rather than growing. nats-lens correctly identifies this stable-elevated pattern as NAK_STORM (detected in 1,006 ms) rather than ACK_WAIT_VIOLATION (which requires a growing rate). Both classes represent consumer delivery degradation; the metric signature differs based on whether distinct new messages are accumulating past ack_wait or the same messages are cycling. This finding shows that ACK_WAIT violations can manifest as either class depending on client-library fetch semantics— and nats-lens detects both correctly. The Rust 30-round results confirm the ACK_WAIT detector fires when num_redelivered grows (distinct messages exceed ack_wait in each poll cycle).
5.7
Operational Overhead
nats-lens issues the following API requests per poll cycle: requests = 1 + 𝑁 streams × (2 + 𝑁 consumers/stream ) We benchmarked nats-lens against 10–1,000 active consumers distributed across 10 streams (Figure 6). RSS was sampled via /proc/self/status after a 30-second steady-state warmup with all consumers actively publishing and ACKing. Memory grows sublinearly from 8.2 MB at 10 consumers to 12.7 MB at 1,000 consumers. API request rate scales linearly: 6.2 req/s at 10 consumers, 204 req/s at 1,000 consumers (5-second poll interval). At 1,000 consumers—a large enterprise deployment —nats-lens imposes <13 MB RSS and <205 req/s on the NATS management plane, negligible relative to NATS data-plane throughput (millions of messages/second).
RSS memory (MB)
Memory vs. Consumer Count12.7 12.0 11.0
10.8
10.0 9.0
8.5
0
API requests/s (5 s interval)
Biplab Kumar Das
Poll Overhead vs. Consumer Count 204
200 150 100
50
200 400 600 800 1,000 Active consumers
0
44 14
0
200 400 600 800 1,000 Active consumers
cases. nats-lens fires on none. The difference comes from multisnapshot, rate-based analysis: (1) MAX_PENDING_THROTTLE requires num_pending > 0 (consumer can receive but is blocked, not merely caught up); (2) ACK_WAIT_VIOLATION requires a rate ≥ 2/min across two snapshots (transient single redeliveries do not trigger); (3) NAK_STORM requires ≥ 2 distinct messages cycling across consecutive snapshots (a single incidental redelivery does not trigger).
5.9
5.8
Comparison to Naive Single-Snapshot Detection Precision on Adversarial Boundary Conditions Naive (single-snapshot) nats-lens (multi-snapshot rate) FP
FP
FP
False positives
1 (false positive)
0 (correct)
✓ MAX_PENDING (num_pending=0)
✓ NAK_STORM (redelivered=1)
✓ MISSING_PROG (ratio=0.85)
✓ Healthy Consumer
Figure 7: False positives on adversarial boundary conditions: naive single-snapshot threshold detector vs. nats-lens multisnapshot rate-based detector. nats-lens produces zero false positives on all four cases; the naive detector produces three. A natural alternative to nats-lens is a single-snapshot threshold detector: alert when Δnum_redelivered > 0 (any growth), num_ ack_pending ≥max_ack_pending (regardless of num_pending), or num_redelivered ≥ 1. We evaluated such a detector on the four adversarial boundary conditions from Section 5.4 (Figure 7). Table 5: False Positives on Adversarial Boundary Conditions Boundary Condition
Naive
nats-lens
MAX_PENDING at cap, num_pending= 0 NAK_STORM: num_redelivered= 1 stable MISSING_PROGRESS: ratio= 0.85 Healthy consumer (all classes)
FP FP FP 0
0 0 0 0
Each boundary condition was tested over 𝑛=30 rounds (95 % Wilson CI ≥ 88.6% per case). These three conditions were selected as maximum-stress cases: each places the relevant metric at the exact value where a naive threshold would fire (num_pending=0, num_redelivered=1 stable, ratio=0.85) but the multi-snapshot condition correctly does not. The naive detector fires on 3 of 4 boundary
Recovery Detection
We injected ACK_WAIT_VIOLATION, then fixed the consumer configuration (ack_wait increased to a safe value) and measured time to alert silence. Over 50 rounds: violation detected 50/50 (95 % CI ≥ 92.9%), recovery detected 50/50 (95 % CI ≥ 92.9%). Recovery latency: P50 = 6,005 ms, P95 = 6,006 ms, P99 = 6,007 ms (range: 6,003– 6,007 ms across all 50 rounds). The near-zero variance confirms the recovery mechanism is deterministic: exactly 2 poll cycles at 𝑇poll =3 s, with no hysteresis. Violations clear immediately once the underlying condition is resolved.
5.10
Simultaneous Multi-Violation
Three violation classes (ACK_WAIT, NAK_STORM, SEQUENCE_GAP) were injected concurrently on separate streams. Over 50 rounds, all three were detected in 50/50 cases (95 % CI ≥ 92.9%). The engine’s per-consumer ring buffer maintains independent state for each consumer; concurrent violations on different streams do not interfere.
5.11
Poll-Interval Sensitivity Theorem 2 Validation: Latency ≤ k × Tpoll for k ∈ {1, 2}
2.00 1.75 P50 latency / Tpoll
Figure 6: nats-lens memory (RSS) and API request rate vs. number of active consumers (10 streams, 5 s poll interval). Both scale sub-linearly in memory and linearly in requests.
1.50 1.25 1.00 0.75 0.50
SEQUENCE GAP MAX PENDING (k=1) NAK_STORM (k=2)
0.25 0.00
0
5
10
15 Poll interval Tpoll (s)
20
1× bound (k=1) 2× bound (k=2)
25
30
Figure 8: Detection latency ratio (P50 /𝑇poll ) vs. poll interval for single-snapshot detectors. The ratio approaches but stays below 1.0, confirming Theorem 2’s bound across five orders of magnitude in poll interval. Figure 8 shows detection latency normalized by 𝑇poll for SEQUENCE_GAP and MAX_PENDING_THROTTLE at poll intervals from 1 s to 30 s (10 rounds each). Detection rate is 100 % at all intervals. The normalized latency ratio ranges from 0.03× (at 1 s, where detection occurs within the same poll cycle) to 0.97× (at 30 s, where onset occurs just before a poll). All ratios are < 1×, confirming Theorem 2’s guarantee that single-snapshot detectors fire within one poll interval of onset. Coverage of all five classes. MISSING_PROGRESS uses the same point-in-time predicate structure as SEQUENCE_GAP
Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation
and MAX_PENDING_THROTTLE; Theorem 2’s 𝑘=1 bound applies identically without separate empirical validation at each poll interval. For NAK_STORM (𝑘=2), empirical sensitivity at 𝑇poll ∈ {1, 3, 5, 10} s confirms 100 % detection rate and latency ratios 0.049×–0.906×, all < 2×. For ACK_WAIT_VIOLATION (𝑘=2), we ran targeted sensitivity experiments at 𝑇poll ∈ {1, 3, 5} s: 𝑇poll
𝑛
P50†
Ratio from onset
1s 5 5,180 ms 1.18× 3s 30 8,013 ms 1.34× 5s 5 9,058 ms 1.01× † Latency from isolated run; detection rate 100 % at all intervals.
Detection is 100 % at all intervals (15 rounds total for 𝑇poll ∈ {1, 5} s; 30 isolated rounds for 𝑇poll =3 s from Table 2). All ratios from onset are < 2×, confirming Theorem 2’s 𝑘=2 bound empirically. The Δnum_redelivered ≥ 2 threshold is met even at 𝑇poll =1 s because multiple messages redeliver simultaneously after ack_wait expires, producing a large batch delta in a single snapshot.
5.12
Threats to Validity
Internal validity. Each round uses fresh consumers and purged streams to prevent state leakage. The embedded eval engine shares the NATS connection with the consumer under test; in production, nats-lens runs as an independent process. Detection latencies assume poll cycles are not delayed by host load; evaluations ran on a lightly loaded MacBook Pro M3 Pro (18 GB). External validity. Evaluations used NATS Server 2.10 with inmemory storage on a single Apple M3 Pro workstation. File-based storage backends may exhibit different redelivery timing. The 3node cluster used Docker-networked nodes on one host; wide-area deployments with higher propagation delay may show slightly longer detection latencies for two-snapshot detectors, proportional to the round-trip time to the NATS management API. nats-lens’s detection algorithm is read-only and CPU-minimal; the detection accuracy results are expected to generalize to cloud instances, where the primary difference would be API latency affecting absolute detection time but not the 𝑘 × 𝑇poll ratio. Statistical validity. Coverage uses 𝑛=100 rounds per class (Wilson 95 % CI ≥ 96.3%). Latency uses a separate 𝑛=30 isolated run; the two-snapshot detectors (ACK_WAIT, NAK_STORM) require isolation to prevent residual consumer state from prior experiments inflating apparent detection speed. Recovery, multi-violation, and adversarial FP experiments used 𝑛=30 rounds each (CI ≥ 88.6%); recovery and multi-violation used 𝑛=50 rounds (CI ≥ 92.9%).
6
Related Work
Message queue monitoring. Kafka [5] uses an offset-based delivery model with no server-side ack_wait or max_ack_pending; consumer lag (not redelivery violations) is the primary failure mode. Burrow [6] and kminion [7] detect Kafka consumer lag at the partition level, but this model is inapplicable to JetStream’s per-consumer pull semantics. Gray and Reuter [11] formalize atleast-once delivery in transaction processing; JetStream’s approach is architecturally distinct (pull-based, configuration-driven redelivery).
Cloud messaging systems. AWS SQS exposes a VisibilityTimeout analogous to ack_wait, but CloudWatch metrics provide only queue-level ApproximateNumberOfMessagesNotVisible counts—no per-consumer redelivery rate or sequence gap detection. RabbitMQ’s messages_unacknowledged counter and consumer_utilisation metric are server-level aggregates that cannot distinguish slow consumers from misconfigured ack_wait. GCP Pub/Sub exposes subscription/num_undelivered_messages and oldest_unacked_message_age, but these measure queue depth rather than redelivery-induced duplication or sequence integrity. nats-lens occupies a different point in the design space: per-consumer state access via a management API enables violation-class-specific detection unavailable in these cloud systems. Distributed monitoring. Dapper [8] and Jaeger [9] trace request propagation. A JetStream ACK_WAIT violation causing duplicate processing produces two successful trace spans—the duplication is invisible to tracing. Delivery correctness monitoring is complementary to, not overlapping with, distributed tracing. Message reliability patterns. The Transactional Outbox [10] addresses the producer side (reliable message publication). nats-lens addresses the consumer side (detecting when delivered messages are lost or duplicated). NATS-specific. NATS Surveyor [4] and the Prometheus exporter [3] provide server-level metrics. To our knowledge, nats-lens is the first work to formally characterize and detect consumer-level delivery correctness violations in NATS JetStream.
7
Limitations
nats-lens monitors pull consumers exclusively. JetStream push consumers use server-driven, subject-based delivery; the $JS.API.CONSUMER.INFO response for a push consumer exposes num_pending and num_ack_pending but omits ack_floor.stream_seq and the rate-based redelivery fields that nats-lens uses for ACK_WAIT and NAK_STORM detection. Extending detection to push consumers would require tracking delivery-subject subscription state and server-side push counters, which differ structurally from the pull model formalized here. According to the NATS documentation, the pull model is recommended for new deployments; push consumers are retained for legacy compatibility. The detection latency lower bound is one poll interval 𝑇poll ; operators requiring sub-second detection should set –interval 1. nats-lens requires read access to $JS.API.* management subjects; deployments with a restricted management plane must add an explicit allow-list entry for the nats-lens NATS user. The polling engine detects violations in running consumers; configuration errors in inactive consumers are caught by the pre-deployment audit (nats-lens init) but not by the live engine. Finally, nats-lens reports the violation class inferred from the metric signature; as demonstrated in Section 5.6, an ACK_WAIT_ VIOLATION may manifest as NAK_STORM when the client library recycles the same fetched messages, requiring operators to verify the root cause using the history endpoint.
Biplab Kumar Das
8
Future Work
Push consumer support. JetStream push consumers receive messages via server-side subject delivery rather than client-side fetch. Extending nats-lens to monitor push consumer health requires tracking PushConsumer.NumPending and delivery subject subscription counts, which differ structurally from the pull model formalized here. Anomaly forecasting. The ring-buffer history store makes natslens a natural substrate for time-series analysis. Fitting an autoregressive model to the num_redelivered rate could provide early warning of an impending ACK_WAIT violation before it crosses the detection threshold, reducing mean time to remediation. Multi-cluster correlation. Large NATS deployments run multiple server clusters as separate JetStream domains. Cross-cluster delivery correctness—where a violation in one cluster’s stream silently affects a consumer in another—requires correlating management API responses across domain boundaries, which the current single-cluster model does not address. Adaptive poll interval. Rather than a fixed 𝑇poll , the engine could adapt the interval per consumer based on observed violation risk: polling high-redelivery consumers more frequently while reducing API load for healthy consumers.
9
Conclusion
JetStream’s at-least-once guarantee is easy to accidentally disable with a misconfigured ack_wait or max_ack_pending, and there has been no tool to detect when this has happened. We characterized five violation classes with formal proofs of their conditions and proved that standard Prometheus NATS monitoring is structurally blind to all of them. nats-lens detects all five through multi-snapshot, rate-based analysis that eliminates the false positives inherent in naive single-snapshot threshold detection. In 100-round evaluations on both single-node and 3-node JetStream clusters, nats-lens achieves 100 % detection coverage (95 % CI [96.4%, 100%]) with zero false positives over 24 hours. Recovery from a fixed violation is confirmed within 2 poll cycles. The tool is open source and requires no changes to monitored applications. As a practical checklist: ack_wait must exceed P99 processing time; max_ack_pending must accommodate actual concurrency; stream retention must exceed expected consumer lag; stale messages must be ACKed (not NAKed); long tasks must send in-progress acks.
Artifact Availability Source code, evaluation harness, and all experimental scripts are available at: https://github.com/biplabku/nats-lens. The repository includes a Docker image and a one-command evaluation replication script (./scripts/run-evaluation.sh –rounds 30).
References [1] Synadia Communications. NATS JetStream Documentation. https://docs. nats.io/nats-concepts/jetstream, 2024. Accessed: September 2026. [2] Synadia Communications. nats-server: High-Performance Server for NATS, v2.10.0. https://github.com/nats-io/nats-server, 2023. Accessed: September 2026. [3] NATS Authors. Prometheus NATS Exporter, v0.15.0. https://github.com/ nats-io/prometheus-nats-exporter, 2024. Accessed: September 2026.
[4] Synadia Communications / NATS Authors. NATS Surveyor: Monitoring, Observability and Analytics for NATS. https://github.com/nats-io/natssurveyor, 2024. Accessed: September 2026. [5] J. Kreps, N. Narkhede, and J. Rao. Kafka: A distributed messaging system for log processing. In Proc. 6th International Workshop on Networking Meets Databases (NetDB ’11), Seattle, WA, 2011. [6] LinkedIn Engineering. Burrow: Kafka Consumer Lag Checking. https:// github.com/linkedin/Burrow, 2016. Accessed: September 2026. [7] Redpanda Data. kminion: Kafka Monitoring Prometheus Exporter. https: //github.com/redpanda-data/kminion, 2023. Accessed: September 2026. [8] B. H. Sigelman, L. A. Barroso, M. Burrows, P. Stephenson, M. Plakal, D. Beaver, S. Jaspan, and C. Shanbhag. Dapper, a large-scale distributed systems tracing infrastructure. Google Technical Report dapper-2010-1, 2010. Available: https: //research.google/pubs/pub36356/. [9] The Jaeger Authors. Jaeger: Open Source, End-to-End Distributed Tracing. https: //github.com/jaegertracing/jaeger, 2017. Accessed: September 2026. [10] C. Richardson. Pattern: Transactional outbox. https://microservices.io/ patterns/data/transactional-outbox.html, 2018. Accessed: September 2026. [11] J. Gray and A. Reuter. Transaction Processing: Concepts and Techniques. Morgan Kaufmann, 1992. ISBN: 1-55860-190-2.