Conceptio › Archive › arXiv CS
arXiv CSopen access

Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

1

Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC

arXiv:2609.05249v1 [cs.AR] 4 Sep 2026

Saad Memon

, Rafal Graczyk , Jan Swakoń , Leszek Grzanka and Mike Papadakis , Member, IEEE

Abstract: As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commercial, black-box architectures severely restricts the development of verifiable radiation mitigation strategies. Open-source, registertransfer level (RTL)-accessible accelerators resolve this limitation by enabling user-defined instrumentation, yet few have empirical radiation-response baselines. This work establishes a foundational system-level proton-irradiation baseline for an unmitigated open-source Tensil NN accelerator deployed on a Zynq UltraScale+ SoC executing ResNet-20 inference. Under 20– 58-MeV proton irradiation, we delivered 4.29×1010 p/cm2 within monitored operational windows. Seven workload interruptions required two restarts of the notebook process, four reboots or board resets, and one power-cycle sequence. Two outputcorruption events returned incorrect CIFAR-10 classes without loss of service. In the longer event, the accelerator returned a class absent from the ten-image CIFAR-10 pool for 39 consecutive inputs at normal cadence. The process remained alive, while the kernel log, limited memory test, and sampled power showed no anomaly. Observation of the stuck-class sequence ended with scheduled bitstream reconfiguration. All nine onsets occurred under the nominal 4-cm beam, which exposed the SoC, LPDDR4, and additional board circuitry; none occurred under the 2cm SoC-centered field. This pattern shows a field association but does not establish LPDDR4 as the cause because field size was confounded with run order and dose. Linux-managed accelerators require end-to-end content checks and recovery that reaches the state in which corruption can persist. This baseline documents availability loss and silent output corruption, supporting future software hardening of COTS FPGA-SoCs for neural-network inference in space systems. Index Terms: Commercial off-the-shelf (COTS), FPGA-SoC, Linux, neural-network accelerator, proton irradiation, silent data This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. This work was supported in part by the European Union’s Horizon 2020 research and innovation programme under Grant Agreement No. 101008126 (RADNEXT), the Horizon Europe research and innovation programme under Grant Agreement No. 101057511 (EURO-LABS), and the Luxembourg National Research Fund (FNR) under Grant CS20/IS/14689454 (HERA). For the purpose of open access, the authors have applied a Creative Commons Attribution 4.0 (CC BY 4.0) license to any Accepted Manuscript version arising from this submission. Saad Memon (Corresponding author) and Mike Papadakis are with the Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg, L-1855 Luxembourg, Luxembourg (e-mail: [email protected]; [email protected]). Rafal Graczyk was with the Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg, L-1855 Luxembourg, Luxembourg. Jan Swakoń, Leszek Grzanka, and Sebastian Kusyk are with the Henryk Niewodniczański Institute of Nuclear Physics, Polish Academy of Sciences (IFJ PAN), 31-342 Kraków, Poland (e-mail: [email protected]; [email protected]; [email protected]).

, Sebastian Kusyk

,

corruption, single-event functional interrupt.

I. I NTRODUCTION N inference service on a commercial off-the-shelf (COTS) field-programmable gate array (FPGA) systemon-chip (SoC) is not the accelerator alone. In a typical deployment, the processing system (PS) runs Linux and the application, the programmable logic (PL) implements the accelerator, and external dynamic random-access memory (DRAM) holds the model and data in flight. Every inference crosses all three. Radiation can disturb any of them, and at the workload level the disturbance surfaces in one of two ways. Either the service stops (a process terminates, Linux hangs, a reset is needed), or it keeps running and delivers a wrong answer. Operationally, these failure modes are distinct. A liveness watchdog can detect a stopped workload but cannot verify a completed inference. A content checker can detect a wrong result while Linux remains responsive. Grouping both under one generic “failure” hides which monitor would have caught it, which recovery it needs, and which exposure belongs in its rate denominator. Campaigns on Linux-managed Zynq platforms have recorded crashes and corrupted inferences in the same beam session [1], [2], but their reporting conventions differed. Agiakatsikas et al. classified runs by outcome [1]; Sabogal et al. grouped consecutive erroneous outputs as one error [2]; and Stirk et al. power-cycled after each Linux-benchmark failure [3]. None reported how long a corrupted state persisted or what the standard monitors recorded during the event. Device-level cross sections for configuration memory, block RAM, and processor resources [3], [4] do not resolve this system-level question because utilization, activation, and masking stand between an upset and a wrong answer. The open question is not whether a Linux-managed accelerator can silently produce wrong outputs; prior work shows that it can. Instead, we ask what happens when corrupted state persists through a block of inputs: how long it lasts, how it ends, and what the available operational monitors report during the event. To investigate this question, we conducted a protonirradiation campaign on an Avnet Ultra96-V2, a Zynq UltraScale+ XCZU3EG with adjacent LPDDR4, running PYNQ Linux and the open-source Tensil Tensor Compute Unit (TCU) in the PL [5]. The TCU executed a ResNet-20 classifier on a fixed pool of ten CIFAR-10 images [6], [7] in blocks

A

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

1

2

3

ship 11

ship 21

13

ship 22

32

frog

frog

dog

dog

63

bird 72

bird 81

bird 82

bird 91

bird

bird

bird

bird 85

bird 94

bird

bird

bird

bird

bird

bird

bird 90

bird 99

bird

bird 80

89

98

bird

bird

bird

bird

horse 70

79

88

97

bird

bird

bird 87

96

69

78

truck 60

frog

horse

bird 77

86

95

bird

bird 76

ship

horse

bird 100

bird

II. R ELATED W ORK AND R ESEARCH G AP A. Device, Platform, and Accelerator Campaigns

50

59

68

truck 40

frog

horse

dog 67

ship

49

58

airplane

denominator to the monitor that observes its endpoint and state that monitor’s blind spots. Third, record the recovery that restored operation without treating it as evidence of the fault’s origin.

30

39

ship

dog

airplane

bird

truck

48

57

66

75

84

93

bird

bird

bird

bird

65

74

83

92

bird

bird 73

dog

ship 64

29

38

ship 20

dog

airplane

dog

frog 56

19

28

47

10

ship

airplane

truck

ship

ship 55

horse

37

46

9

18

27

truck

dog

airplane 54

airplane

airplane

36

45

8

17

26

dog

horse

frog

airplane 62

71

frog 53

truck

35

44

ship 16

25

airplane

truck

7

ship

ship

34

43

52

61

dog

6

15

24

33

42

51

horse

dog

horse 41

ship 14

23

horse

airplane 31

5

frog

horse 12

4

2

bird

Fig. 1. Run 5, block 10 (Section IV-B), in execution order: 61 correct predictions, then 39 consecutive predictions of bird, a class absent from the ten-image pool, while every availability indicator stayed nominal. Composite of CIFAR-10 test-set images 10–19 [7].

of 100 inferences, and the PL was reloaded at each block boundary. We retained every predicted class in execution order and monitored workload progress, the kernel log, network reachability, sampled power rails, and a userspace memory test. In one block at 58 MeV, the first 61 predictions were correct and the next 39 were all bird, a class absent from the pool (Fig. 1). These outputs arrived at the usual cadence while every availability indicator remained nominal, until the scheduled reload ended the observation. The workload was also interrupted seven times, with recovery ranging from a process restart to a power cycle. We used two beam fields. One nominally covered the SoC, whereas the wider field also covered the LPDDR4 package and surrounding board area. This geometry enabled an exploratory comparison between SoC-focused and wider-system exposure. We pose two questions. Q1: When corrupted state persists through a block, how long does it last, how does it end, and what do the Linux liveness, timing, memory-test, and sampled-power indicators report? Q2: Does exposing board circuitry beyond the SoC footprint change the rate of Linuxlevel workload interruptions at matched beam energies? From this monitored PS–PL inference stack, we report a silent constant-class corruption and its complete monitor record, seven Linux-level interruptions and their recovery tiers, and condition-specific cross sections for both endpoints. We also report a matched-energy field comparison, but field is confounded with run order and dose. We neither identify a physical mechanism nor evaluate a mitigation. Instead, we apply and recommend three reporting practices. First, count events by onset so consecutive wrong outputs from one corrupted state form one event. Second, match each

Device-level irradiation of Zynq UltraScale+ and related SRAM-based FPGA families has quantified susceptibility in configuration memory, block RAM, flip-flops, and digitalsignal-processing (DSP) blocks [4]. Platform-level campaigns have added processor-side cache and translation-lookasidebuffer (TLB) cross sections [1], [3]. Stirk et al. compared neutron-test methods for an Ultra96 and reported process/hang, kernel, and application-processingunit reset failures under Linux [3]. Agiakatsikas et al. irradiated a Linux-managed ZCU102 running a Vitis deeplearning processing unit (DPU) and observed both crashes and numerical corruption [1]. Sabogal et al. irradiated a Linuxmanaged convolutional-neural-network (CNN) accelerator on Zynq-7020 and ZU3EG boards with wide-spectrum neutrons [2]. They recorded erroneous and hung executions and grouped consecutive erroneous outputs as one error, but did not report an episode’s duration or its monitor record. Running under an operating system also changes which failures appear: Santini et al. found that Linux raised the functional-interrupt rate of an embedded SoC under neutrons without a comparable change in its silent-corruption rate [8]. Radiation and fault-injection studies of FPGA accelerators have examined quantization, architectural parallelism, redundancy, and scrubbing [9]–[13]. The reported outcome may be a changed class, a numerical deviation, a hang, or an accelerator reset. It depends on the checker: a top-1 oracle observes only class-changing corruption, whereas a score-vector checker detects numerical changes that preserve the winning class. The protocol matters as well. Reloading the PL after every input ends any persistent error that another campaign might observe across many inputs. For example, a Linux-managed Jetson SoC showed persistent output errors when its main memory was inside the beam [14]. FINN-generated designs have beam baselines [12], [13]; we found no published beam data for Tensil, and beam data for open accelerators exercised through a Linux-managed PS–PL path remain scarce. We treat the complete inference path as the service. This path extends from PL configuration through direct-memory-access (DMA) control and the Advanced eXtensible Interface (AXI) interconnect to external memory. The open register-transfer-level (RTL) source permits future rebuilding and instrumentation. Because this campaign added neither instrumentation nor mitigation, its results provide an unhardened baseline for that work. B. Attribution Under Whole-System Exposure External DRAM creates a recurring attribution problem. A wide beam exposes the SoC, memory package, traces, clocks, and power-delivery circuitry at the same time. Systemlevel observations alone cannot identify which component contained the initiating disturbance. Guertin and Cui noted

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

this difficulty for a SoC tested with its supporting LPDDR memory [15]. Our proton tests of two Linux-managed LPDDR platforms also could not attribute memory-test mismatches to a component under whole-system exposure [16]. Those tests used 20–58 MeV at the device from a nominal 60-MeV beam. Following system-level single-event-effect (SEE) guidance, we use field geometry only to test association at matched energies. We do not infer a physical fault location from a workload symptom alone [17], [18]. III. E XPERIMENTAL M ETHOD A. Hardware and Software Stack The device under test (DUT) was an Avnet Ultra96-V2 populated with an XCZU3EG-1SBVA484I Zynq UltraScale+ multiprocessor SoC (MPSoC). The device uses a 16-nm FinFET process, speed grade −1, and an industrial temperature rating. The board contained 2 GB of Micron MT53D512M32D2DS053 LPDDR4 on the PS DDR controller [19] in a 32-bit configuration without error-correcting code (ECC), which the single x32 device cannot provide [20]. A 16-GB microSD card held the boot image and root filesystem. The device lot and date codes were not recorded. A Rigol DP821 supplied 8.000 V, within the board’s 7– 14 V input range [19], with remote sensing. Overvoltage and overcurrent limits were 8.4 V and 3.5 A, respectively. The heat sink was removed to expose the package, and the SoC surface was oriented approximately normal to the beam. The DUT operated in room-temperature air. PYNQ Linux v2.7, based on Ubuntu 20.04 with kernel 5.4.0-xilinx-v2020.2, ran on the PS. A laptop outside the irradiation room controlled the board through USB-gadget Ethernet. It launched notebook blocks and collected outputs and kernel messages. It also sent Internet Control Message Protocol (ICMP) echo requests to assess reachability. All Linux evidence reached the laptop through this link, which limits what service loss alone can establish (Section V-D). The platform used no experiment-specific watchdog, configuration readback, Soft Error Mitigation (SEM) controller, or redundant inference checker [21]. The inference workload was a ResNet-20 classifier for CIFAR-10. The design was adapted from the public Tensil Ultra96-V2 flow [5]. The TCU used a 16×16 systolic array at 100 MHz and the 16-bit fixed-point FP16BP8 format throughout. This format covered the array, local memory, accumulator memory, and single-instruction, multiple-data (SIMD) datapath. The TCU contained 20,480 local-memory vectors and 4,096 accumulator-memory vectors. For the compiled model used in this campaign, Tensil reported 27 layers, 101,840 ninebyte instructions, 568,474 constant scalars, and 97.2% constant utilization. Two 128-bit AXI4 master interfaces moved weights, inputs, intermediate data, and outputs between the TCU and external LPDDR4. A 128-bit AXI direct-memory-access (DMA) path supplied TCU instructions. Local and accumulator memories resided in PL block RAM. Each inference exercised PS software, PS–PL interfaces, external LPDDR4, DMA and interconnect logic, configuration state, and accelerator-local state.

3

The bitstream used in the campaign was implemented with Vivado 2021.2 using the original RTL, target device, and board definition. The routed design met timing with +0.523 ns worst-case slack. It used 24,382 lookup tables (34.6%), 18,907 registers (13.4%), 180 block-RAM tiles (83.3%), and 273 DSP48E2 blocks (75.8%). Vivado’s essential-bit mask for this design marks Bess = 13,222,948 of 30,876,800 positions (42.8%) as essential. We report this count for future comparison but do not use it for normalization. B. Inputs, Oracle, and Observability The input pool contained zero-based indices 10–19 of the canonical CIFAR-10 test set. These ten images span six classes: truck, dog, horse, and ship each appear twice, while airplane and frog each appear once. Automobile, bird, cat, and deer are absent. A pre-irradiation validation classified all ten images correctly, so the dataset labels served as the online reference. Each inference selected one image uniformly with replacement. No fixed pseudorandom seed was used, so the realized order cannot be regenerated from the sampling rule alone. The log retained the selected input index, reference class, and predicted class for every retained record. The checker retained only the predicted class, defined by the largest output score. It did not retain the complete ten-element score vector, intermediate tensors, or accelerator state. It can detect a class change but not a numerical perturbation that leaves the winning class unchanged. We considered a block verified when its class summary was retained, with or without the per-inference order. Of the 92 launched blocks, 62 met this criterion. All availability analyses use a separate operationalwindow denominator that does not assume output verification for missing blocks. C. Irradiation Conditions The AIC-144 cyclotron at IFJ PAN delivered nominal 20-, 40-, and 58-MeV protons in air [22]. The beam was normal to the exposed board surface. Facility dosimetry followed the station procedure, with 3% reported uncertainty for fluence and flux. Dose-to-fluence conversion used SRIM2013 stopping powers [23]. We report nominal incident energy because component-specific energy was not measured. The three energies are settings below the facility’s nominal 60MeV proton beam. Two collimated fields changed the exposed board area; Table I labels them by the operator’s nominal designation. The operator record describes the nominally 2-cm field as primarily covering the SoC. The facility record identifies a 25-mm collimator, which we use as the authoritative aperture. Because this exceeds the 19-mm SBVA484 package, the edge of the adjacent LPDDR4 package may also have been exposed. The nominally 4-cm field covered the SoC, the LPDDR4 package, and additional board area. Both fields were used at 20 and 40 MeV. Only the wider field was used at 58 MeV. Fig. 2 shows the mounted DUT. Table I lists the six monitored irradiation intervals. The monitored intervals accumulated 6.19 × 1010 p/cm2 .

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

4

After a sustained loss of progress, recovery began with the least disruptive available action. Operators escalated through process restart, Linux reboot, board reset, and power cycle. When several actions changed together or observation ended, we report a lower bound or the full recovery sequence rather than assign causality to one action. E. Event Definitions and Statistical Analysis

Fig. 2. Ultra96-V2 DUT at the AIC-144 station with the heat sink removed and the SoC facing the beam collimator (right). TABLE I P ROTON B EAM C ONDITIONS FOR M ONITORED RUNS Run

t (s)

Ea (MeV)

Flux (p/cm2 /s)

Φtotal (p/cm2 )

Dw (Gy)

Field (cm)

1 2 3 4a 4b 5

1126.5 1076.4 1016.4 781.4 388.2 1191.8

20 40 20 40 40 58

1.14 × 107 1.20 × 107 1.01 × 107 1.20 × 107 1.28 × 107 9.61 × 106

1.29 × 1010 1.29 × 1010 1.02 × 1010 9.37 × 109 4.98 × 109 1.15 × 1010

54.7 31.2 43.4 22.6 12.0 20.6

2 2 4 4 4 4

Total

5580.7

6.19 × 1010

184.5

a

t is irradiation time, E is nominal incident energy, and Dw is absorbed dose to water. Facility dose, fluence, and flux uncertainty is 3%. Run 4 contains intervals 4a and 4b; unmonitored exposures are excluded.

D. Block Protocol and Recovery One block corresponded to one notebook execution. At block start, the notebook issued two PL configuration downloads, ran one memtester 10M 1 pass (a stuck-address test and 17 pattern tests over a 10-MB region, 0.5% of the memory), and sampled the board’s power rails [24]. The notebook then executed and logged 100 inferences, sampled the power rails again, recorded a class summary, and performed a final PL configuration download. The 100-inference sequence required about 2.34 s. Individual output records were typically 23–24 ms apart. The userspace memory test typically required 15–20 s. Successive block starts were separated by a median of 40 s, with a 23– 127 s range. The final reconfiguration rewrote PL configuration, reinitialized block RAM, reset the TCU and DMA, and reloaded the model, which rewrote the model constants in their LPDDR4 buffer. It did not reset Linux, the PS, on-chip memory (OCM), the remaining LPDDR4 contents, or the Linux page cache. Operators continuously followed block summaries, the control terminal, kernel messages, and network reachability. A continuous supply-current trace is available only for run 1; the other runs contain two power-rail samples per block.

Individual particle interactions were not observed directly. We count an operational event onset as the first abnormal observation after a recorded normal state or recovery action. Consecutive abnormal outputs, log records, or failed relaunches without an intervening normal state form one cluster. A recorded process restart, reboot, reset, or power cycle ends the cluster, so a later failure begins a new onset. One persistent corruption is therefore counted once, regardless of how many predictions it affects. Clustering is performed separately for each endpoint. Thus, an output event and the interruption that follows it are counted under their respective endpoints. This rule affects run 4, whose record is ambiguous. The primary count treats a recorded reboot as a cluster boundary although no block completed before the next failure; Section IV-E reports alternative readings. We retain the label numbering of the campaign’s event taxonomy and define only the categories used here. F3 denotes multiple incorrect predictions followed by block completion, and F4 denotes output corruption still present when observation ended. F6 denotes an isolated accelerator or DMA timeout; F7, an isolated PS–PL interface failure; F8, failure of the prescribed user process; and F9, a Linux- or systemlevel failure. The labels describe observed behavior, not fault locations. Recovery levels identify the least disruptive successful action: R5 is a process restart, R6 or R7 a reboot or board reset, and R8 a power cycle. Following the workload-conditioned definition of our prior study, we classify every F8 or F9 onset that prevented completion of the prescribed notebook workload as a Linux-level single-event functional interrupt (Linux-SEFI) [25]. Its tiers correspond to the process, kernel, and reset classes of Stirk et al. [3] and the symptom classes of Esquer et al. [26]. The term describes an observed symptom: a workload that stopped and required at least a process restart. It does not identify a location or mechanism. The definition includes both process failure while Linux remains reachable and any stall observed as a loss of progress (Section IV-B). The count can therefore include interruptions of any origin, so each Linux-SEFI cross section is an upper bound on the radiation-induced rate. Unlike the JEDEC SEFI definition [27], our definition includes the run 4 episode that ended in a power cycle. Table V lists this episode separately so it can be excluded. We report two primary endpoints separately: every F3/F4 inference-output onset and every F8/F9 Linux-SEFI onset. Corrected-OCM reports (OCM ECC correctable-error reports from the Linux driver) are telemetry rather than workload failures and form a separate endpoint. By itself, such a report establishes neither radiation origin nor the software object at that address. Every onset began while the beam was on, so we

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

call them radiation-associated. This criterion is temporal. The error-free pre-beam blocks and the 44 completed small-field blocks, including 31 with verified outputs, weigh against but do not exclude a beam-independent cause. The Linux-SEFI denominator is operational-window fluence, denoted Φop . For each contiguous monitored window, we integrate the facility-reported flux from the first block launch until the final block ends or failure recovery begins. Because the beam remained on between blocks, Φop includes inter-block intervals; it is not the fluence of the 2.34-s inference phases alone. Classification evidence was retained for only 62 of 92 launched blocks. For output events, we therefore allocate each run’s operational fluence in proportion to the fraction of blocks with retained output verification:

5

TABLE II RUN S UMMARY Run 1 2 3 4 5 Σ

E Field Φop Blocks Err. Event Op. (MeV) (cm) (1010 p/cm2 ) launch. compl. verif. outputs onsets time (%) 20 40 20 40 58

2 2 4 4 4

1.19 1.17 0.78 0.29 0.86

23 21 20 7 21

23 21 19 3 17

23 8 19 3 9

0 0 0 3 39

0 0 1 3 5

93 91 76 21 75

4.29

92

83

62

42

9

70

Blocks are launched, completed, and verified (retained class records); Err. outputs are incorrect predictions (affected outputs, not event counts). The nine primary onsets are seven Linux-SEFIs and two output events. Output-event cross sections use Φver from the 62 blocks with retained class evidence. Op. time is top /t, where top = Φop /flux is the operational-window duration and t is the run duration in Table I.

(r)

(r) Φ(r) ver = Φop (r)

nver (r)

,

(1)

nblk

(r)

where nver and nblk are verified and launched block counts. Equation (1) is an exposure-allocation approximation: it does not assert that verified and unverified blocks share the same latent event history. For endpoint k under condition g, the observed system-level event cross section is X Nk,r r∈g

σk,g = X

[cm2 /system],

(2)

Φk,r

r∈g

where Nk,r is the number of clustered onsets and Φk,r the endpoint-specific denominator; the subscripts SEFI and FE denote the Linux-SEFI and the F3/F4 output-event endpoints. We write cm2 /system for a cross section per DUT under the stated field; because the exposed area differs between fields, the values are condition-specific. For nonzero counts we report two-sided 95% Garwood intervals under a Poisson counting model, and for zero observed events the one-sided 95% upper limit 2.996 95%,UL σk,g =X , (3) Φk,r r∈g

where 2.996 = − ln 0.05. These intervals quantify counting uncertainty only and assume independent onsets at a constant rate within each window. Clustering or a rate that changes with accumulated dose would widen the intervals. When a window ends at the failure that terminated the run, as in run 4, the ratio N/Φop is biased upward for small counts. The intervals do not include this bias. The facility assigns 3% uncertainty to fluence and flux. Minute-resolution operator timestamps add reconstruction bounds of ±90 s for runs 1–2 and ±150 s for runs 3–5. These bounds equal 8.6%, 9.2%, 19%, 62%, and 17% of the reconstructed Φop for runs 1–5. Because these are bounds rather than independent random errors, we do not combine them in quadrature with dosimetry uncertainty. Instead, we repeat the field comparison after lengthening every wide-field window and shortening every small-field window by its bound. This adjustment weakens the observed contrast.

For two conditions modeled as independent Poisson counts, conditioning on the total count under equal rates per fluence yields a binomial allocation. If conditions 1 and 2 have exposures Φ1 and Φ2 , the null probability that an event falls in condition 1 is Φ1 . (4) p0 = Φ1 + Φ 2 For the matched-energy comparison, we condition on the total count within each energy and use the total wide-field count as the test statistic. When every event falls in the wider field, the one-sided exact p-value is the product of the per-energy conditional probabilities. The analysis was not preregistered, so these p-values provide exploratory rate evidence rather than confirmatory hypothesis tests. We apply no multiplicity adjustment. IV. R ESULTS A. Exposure and Data Completeness The monitored runs delivered 6.19 × 1010 p/cm2 . Reconstructed operational windows account for 4.29 × 1010 p/cm2 ; the remainder was delivered before the first monitored block, after the last, or during recovery. Fig. 3 places each window, recovery interval, and onset in campaign order. Table II gives the per-run accounting. Ninety-two blocks were launched and 83 completed, yielding 8300 inferences. Class records were retained for 6200 inferences in 62 blocks. Record retention varied by energy: 42 of 43 launched blocks at 20 MeV retained their records, compared with 11 of 28 at 40 MeV and 9 of 21 at 58 MeV. Thus, output verification was thinnest at the energies where output events occurred. In run 5, the eight completed blocks without class records were also the eight blocks with corrected-OCM reports (Fig. 3). We do not know why these records were lost. If the loss correlated with abnormal blocks, (1) allocates too much exposure to those runs and underestimates the output-event cross sections. The zero-event result for the 40-MeV small field is based on only eight verified blocks. Nine launched blocks did not complete, two more than the seven Linux-SEFI onsets, because failed relaunches within one cluster are not separate onsets (Section III-E).

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

block OK (100/100) block logged, status blank

6

corrected OCM ECC report output corruption (F3/F4)

block failed (hang/crash) beam on (shade = run energy)

Run 1: 20 MeV 2 cm, 1126 s OCM ECC (corrected)

Run 2: 40 MeV 2 cm, 1076 s Run 3: 20 MeV 4 cm, 1016 s

61/100 correct, then 39 x bird

process restart

OCM ECC (corrected)

power cycle; boot fails until beam off

Run 4: 40 MeV 4 cm, 1170 s

3/100 wrong

Run 5: 58 MeV 4 cm, 1192 s reboot under beam

0

15

30

45

60

75

90

105

120

hang + oops; reboot

135

Time from start of the first monitored run (min) Fig. 3. Monitored campaign timeline. Open circles mark logged blocks without a recorded class status; a corrected-OCM marker replaces a block’s status marker; run 4 comprises intervals 4a and 4b. Annotations reproduce operator notes and event order; they do not establish causality. All seven Linux-SEFI and both output-event onsets began under the wider field.

TABLE III W HAT THE M ONITORS R ECORDED D URING THE C ONSTANT-C LASS B LOCK Indicator

Check and timing

Record for run 5, block 10

Notebook process

alive, no exception; continuous interval between records; every inference kernel-reported errors; continuous over the link pattern mismatches; once, before inference sampled voltages and currents; twice per block ECC-corrected OCM reads; logged on read

alive, no exception

Output cadence Kernel log Memory test (10 MB) Power rails Corrected-OCM telemetry Top-1 oraclea

predicted class versus reference; every inference

23–24 ms throughout (each instruction stream completed, no DMA error reported) no related message (the userspace-driven accelerator path emits none) passed within campaign range reports in adjacent blocks; none attributable to this block in the record outputs 62–100 wrong (39 consecutive, one absent class)

outputs 1–61 were correct and the wrong class did not depend on the input, the corrupted state became effective between inferences 61 and 62. This transition occurred within one 23– 24 ms interval of the inference phase, after model load and PL configuration at block start, and in state not refreshed between inferences. Every availability indicator stayed nominal (Table III). The notebook process remained alive and raised no exception. Outputs kept their 23–24 ms cadence, showing that each instruction stream completed without a reported DMA error. The kernel log was silent, as expected for corruption inside a userspace-driven accelerator. The block’s memory test had passed, and its power-rail samples were in range. A watchdog using only these signals would not have detected the 39 incorrect outputs. The record does not show whether a correctedOCM report was logged inside this block; the same address was reported in the blocks around it (Section IV-C).

B. Silent Output Corruption

The error was still present at the block boundary. It did not clear during the 39 affected predictions, about 0.9 s at nominal cadence. Nothing in the protocol acted until the scheduled PL reload, which rewrote the configuration, reinitialized the PL memories, and reloaded the model. The observation therefore ended with the reload. The protocol, rather than the fault, rightcensored the episode at 39 predictions.

In run 5, block 10, the first 61 predictions matched their references and predictions 62–100 all read bird (Figs. 1 and 4). No bird image exists in the pool, so a stale but valid input cannot explain the sequence. The constant class (index 2) differs from the last correct output (dog, index 5), excluding a frozen output buffer. It also differs from the first-maximum argmax of an all-equal score vector (index 0), excluding a zeroed output vector. A corrupted input or output path remains possible. The same class persisted across 39 changing inputs, which rules out a single mislabeled input and establishes a sustained class-output error in the observed sequence. Because

The next launched block did not complete (Fig. 3), so the record does not show whether the reload ended the underlying abnormal state. If the subsequent failure continued the same condition, the output event and one run 5 Linux-SEFI would be one physical episode counted under two endpoints. The run 4 output event was also followed by a failed launch: the F3 block at 89.4 min, then a failure at 90.4 min (Fig. 3). The record cannot determine whether one condition produced both symptoms. Section IV-D therefore includes a merged reading. We report one onset and 39 affected outputs, separating the event frequency from its persistence. Treating all 39 outputs

a

All rows except the last are availability or telemetry indicators and do not use a content oracle; none examines what the accelerator computed. Added for the experiment; retained for 62 of 92 blocks.

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

correct (matches label)

7

incorrect

truck

Predicted class

ship horse frog

first incorrect result: inference 62 (1.44 s after first logged result)

dog

every output = bird until block teardown (39 consecutive; event duration right-censored)

deer cat bird automobile airplane

0

20

40

60

80

100

Inference index within the 100-inference block Fig. 4. Run 5, block 10: 61 correct outputs followed by 39 constant bird outputs. Observation ended at the scheduled block-end reconfiguration.

TABLE IV T HE N INE P RIMARY O NSETS IN C AMPAIGN O RDER Run Block time (min) Endpoint 3

66.4

4

89.4

4

90.4, 91.4

4

97.4, 104.4a

5

123.4

5

124.4 to 136.4b

TABLE V R AW O BSERVED O UTCOMES BY N OMINAL E NERGY

Evidence and recovery

F8 Linux-SEFI

process termination, no retained signal or kernel message; process restart (R5) F3 output three wrong outputs, order not retained; block completed; next launch failed F9 Linux-SEFI two failed launches; loss of progress and service; reboot (R6/R7) F9 Linux-SEFI failed relaunch after the reboot, then boot failures under beam; beam off and power cycle (R8) F4 output 61 correct, then 39 consecutive bird; censored at the reload; next launch failed F8 + 3 F9 Linux-SEFIs one segmentation fault (R5); three reboot-tier interruptions, one with a photographed oops, one recovery-censored (R6 lower bound)

Times are block launches in Fig. 3. a The 104.4-min launch lies in interval 4b, inside the recovery sequence. b Failed launches at 124.4, 127.4, 132.4, and 136.4 min; the timeline places the kernel oops at the final launch; the archive does not map the other launches to onsets.

as independent events would overcount the episode by a factor of 39. The second output event, in run 4 at nominal 40 MeV, survives only as a block summary: three incorrect predictions in a block that completed (F3). Its output order was not retained, so the record supports one to three distinct onsets; the primary analysis uses one, and Table V carries the uncertainty. No isolated accelerator or DMA timeout (F6) or PS–PL interface failure (F7) was identified. The instrumentation could not distinguish either condition because the driver waits for the accelerator without a timeout. A stalled accelerator or DMA would therefore appear as a loss of workload progress and be classified as a Linux-SEFI. C. Availability Loss: Linux-SEFIs, Recovery, and Telemetry The seven Linux-SEFIs span three recovery levels (Tables V and VI). Two F8 events stopped the prescribed workload but cleared with a process restart while Linux remained reachable. The run 3 record contains no signal, exit code, or kernel

20 MeV 40 MeV 58 MeV Primary operational onsetsa multi-output deviation (F3)b sustained constant-class output (F4) process-restart Linux-SEFI (F8, R5) reboot/reset-tier Linux-SEFI (F9)c power-cycle-sequence Linux-SEFI (F9, R8) Corrected-OCM telemetry episodes Observed incorrect inference outputs Launched blocks not completed a b c

1 0 0 1 0 0 1 0 1

3 1 0 0 1 1 1 3 4

5 0 1 1 3 0 1 39 4

F3/F4 output events plus F8/F9 Linux-SEFIs. Output order was not retained; one onset is the minimum supported count. Includes one recovery-censored run 5 onset. Energy columns mix field conditions and are not an energy-response experiment.

TABLE VI O BSERVED L INUX -SEFI E VIDENCE AND R ECOVERY Recovery category

N

Runs and available evidence

Process restart (R5)

2

Reboot/reset (R6/R7)

4

Power-cycle sequence (R8)

1

Run 3: Jupyter/IPython process termination. Run 5: segmentation fault with a core dump of the notebook server process. One event in run 4 and three in run 5. Evidence included workload-progress loss, notebook or secure-shell service loss, and a photographed null-pointer kernel oops with a call trace in the clone/fork path while ipython was active. One run 5 onset was recovery-censored and is assigned only a lower-bound R6 tier. During run 4, boot attempts did not complete while irradiation continued. Operation returned after a sequence that included beam termination and power cycling.

N counts clustered operational onsets. Recovery is the least disruptive successful action when isolated; censored or multi-action episodes are reported as lower bounds or complete sequences.

message for the termination, so a software-only cause cannot be excluded for that event. Of the four reboot-tier F9 events, one is documented by a photographed console showing a null-pointer kernel oops. Its call trace passes through the clone/fork path while ipython was active (Fig. 5(c)). The other three are evidenced by loss of workload progress and

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

8

(a) Corrected-OCM report

(b) Crash-handler failures and EXT4 errors

(c) Kernel oops and call trace

Fig. 5. Three archived Linux/OCM records: a corrected-OCM report, filesystem errors, and a kernel oops. Images are contrast-enhanced and checked against the originals.

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

9

TABLE VII O N -C HIP -M EMORY C ORRECTABLE -E RROR T ELEMETRY Run

E (MeV)

Field (cm)

1 2 5

20 40 58

2 2 4

Address

Reporting pattern

0xFFFEB8A0 0xFFFF6B50 0xFFFEA4A0

single report single report repeated across nine blocks; observed across one reboot

All addresses lie in documented OCM [20]. Repeated run 5 reports at one address are grouped as one telemetry episode; that episode brackets the run 5 F4 block and one Linux-SEFI block (Fig. 3).

notebook or secure-shell service, subject to the USB-gadget link limitation in Section V-D. Table IV lists the nine onsets in campaign order with the block times shown in Fig. 3. The interpretation of run 4 affects the field comparison (Section IV-E). Launches at 90.4 and 91.4 min failed and were followed by a reboot. The relaunch at 97.4 min then failed and began the boot-failure sequence. During interval 4b, which lasted 388 s under beam, another launch failed within that sequence before the beam was stopped and the board was power-cycled. We count two onsets in this record. Φop for run 4 covers 87.4 to 91.4 min, so the second onset lies outside the denominator window. Crediting the minute around that relaunch to Φop would move the matchedenergy p-value from 0.016 to about 0.02. It would also reduce the 40-MeV wide-field cross section from 6.9 × 10−10 to 5.5 × 10−10 cm2 /system. Both changes fall within the applied timing bound. We retain the defined window because the relaunch failed within the timeline’s resolution. Table X also reports readings that merge the two onsets or count every failed relaunch and credit interval 4b. The second run 4 episode led to power cycling. Repeated boot attempts did not complete while irradiation continued, and operation returned after a sequence that included beam termination and a power cycle. At the wide-field onset rates in Table VIII and each run’s flux, the mean interval between onsets is 2 to 13 min, or 2 to 11 min at the run 4 flux. This interval is comparable to a Linux boot of about one minute at the 40-MeV rate, although the episode itself contributes one of the two onsets. If the rate during boot matched the workload rate, the onset rate alone makes repeated boot failure under irradiation plausible. Beam termination may therefore have been sufficient. Similar behavior has been reported on a Linuxmanaged ZCU102 [1]. Because beam exposure and power state changed together, the evidence supports the recorded R8 sequence but does not show that power cycling was necessary. With only two power-rail samples per block in runs 2–5, the record also cannot exclude a single-event latch-up or another high-current condition during the episode. Fig. 5 presents three archived Linux/OCM images that support the telemetry, filesystem, and kernel summaries. The EXT4 errors in panel (b) show similarly malformed entries in two directory blocks of two directories on the microSD root filesystem. Both entries have the same offset and inode field and a zero-length name. This repeated pattern is more consistent with a storage-path or shared-state mechanism than

TABLE VIII C ONDITION -S TRATIFIED L INUX -SEFI C ROSS S ECTIONS E Field N (MeV) (cm)

Φop (p/cm2 )

σSEFI (cm2 /system)

95% interval or UL (cm2 /system)

20 20

2 4

0 1.19 × 1010 – < 2.5 × 10−10a 1 0.78 × 1010 1.3 × 10−10 [3.2 × 10−12 , 7.1 × 10−10 ]

40 40

2 4

0 1.17 × 1010 – < 2.6 × 10−10a 2 0.29 × 1010 6.9 × 10−10b [8.4 × 10−11 , 2.5 × 10−9 ]

58

4

4 0.86 × 1010 4.7 × 10−10

[1.3 × 10−10 , 1.2 × 10−9 ]

N is the clustered Linux-SEFI count; cross sections follow (2); – marks an undefined point estimate for zero events. a One-sided 95% Poisson upper limit from (3). Nonzero intervals are two-sided exact Garwood intervals. Counting, dosimetry, and bounded timing uncertainties are reported separately. b The second onset lies outside Φop (Section IV-C); with the relaunch minute credited, 5.5 × 10−10 [6.7 × 10−11 , 2.0 × 10−9 ].

TABLE IX O UTPUT-E VENT C ROSS S ECTIONS BY F IELD AND BY N OMINAL E NERGY Condition

N

Φver σFE (1010 p/cm2 ) (cm2 /system)

95% interval or UL (cm2 /system)

Small field (runs 1, 2) 0 Wide field (runs 3–5) 2

1.64 1.23

– < 1.8 × 10−10a 1.6 × 10−10 [2.0 × 10−11 , 5.9 × 10−10 ]

20 MeV (both fields) 40 MeV (both fields) 58 MeV (wide field)

1.93 0.57 0.37

– < 1.6 × 10−10a 1.8 × 10−10 [4.4 × 10−12 , 9.8 × 10−10 ] 2.7 × 10−10 [6.9 × 10−12 , 1.5 × 10−9 ]

a

0 1 1

Φver follows (1) with the verified block counts of Table II; the F3 event counts as one onset (its record supports one to three). One-sided 95% upper limit.

with independent bit upsets. The record does not show whether the corruption persisted across later reboots. Panel (b) also records crash-handler (apport) failures, but the archive does not place either record relative to a specific onset. Corrected-OCM telemetry appeared under both fields (Table VII). Runs 1 and 2 each logged a single report at a distinct address under the small field. This is consistent with the small field reaching the SoC, although no workload event followed over 2.36 × 1010 p/cm2 . Run 5 logged the same address, 0xFFFEA4A0, in nine blocks and across one reboot. We group these reports as one episode, a deliberate exception to the onset rule. OCM ECC corrects the data returned by a read without rewriting the stored word [20], so one stored upset is reported on every read until that word is rewritten. Of the nine blocks, the eight that completed retained no class record. The ninth was either the constant-class block or a block that did not complete; the record does not identify which. The three addresses lie in the upper OCM region that holds the Arm Trusted Firmware image in the default boot flow. A reboot rewrites this region, so recurrence after reboot is consistent with either a persistent cell or repeated upsets; the record cannot distinguish them. The 97 complete memtester records, comprising 1746 pattern results with no mismatch, do not map one-to-one onto the 92 blocks. Each shows only that a selected 10-MB region passed the 18 tests before inference. They neither establish LPDDR4 integrity during an output event nor localize one.

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

D. Cross Sections Table VIII reports the Linux-SEFI cross sections by energy and field. The small-field conditions yield one-sided upper limits, whereas the wide-field point estimates have broad intervals because the counts are small. Pooled over the wide field, seven onsets in 1.93×1010 p/cm2 give 3.6×10−10 cm2 /system [1.5 × 10−10 , 7.5 × 10−10 ]. Table IX reports the output-event cross sections with the Φver denominator of (1), stratified by field and energy. Each nonzero energy-specific estimate is based on one event. The output endpoint does not distinguish the two fields because the small-field upper limit exceeds the wide-field estimate. Operational availability provides a complementary view of the Linux-SEFI results. The smallfield runs kept the workload operational for 91–93% of the beam interval; the shortfall is beam time outside the block sequence, not downtime. Runs 3 and 5 reached 75–76%, and run 4 reached 21% (Table II). The run 5 Linux-SEFI count depends on three interpretations of the record. The alternatives merge the ambiguous Linux sequence near recovery into an adjacent episode, exclude the recovery-censored onset, or treat the failed block after the constant-class block as part of that episode. Each interpretation alone reduces the 58-MeV Linux-SEFI count from four to three and its point estimate from 4.7 × 10−10 to 3.5 × 10−10 cm2 /system. Applied together, they leave one or two events because the archive does not establish whether the first two interpretations concern the same block (Table IV). None changes the output-event count. The 40-MeV widefield value in Table VIII also halves when the run 4 events are merged. The same numbers result if the failed launch at 90.4 min is treated as a continuation of the F3 event, as in the second row of Table X. Treating the run 4 F3 record as three onsets would triple the 40-MeV output-event cross section without altering the existence of class-output corruption. E. Field Contrast Q2 asks whether exposing board circuitry beyond the SoC footprint changed the Linux-SEFI rate at matched energies, not where any fault occurred. At 20 MeV the matched exposures are 1.19 × 1010 p/cm2 (small field) and 0.78 × 1010 p/cm2 (wide), with counts 0 and 1. At 40 MeV, the exposures are 1.17 × 1010 and 0.29 × 1010 p/cm2 , with counts 0 and 2. Under equal rates within each energy pair, (4) gives 0.78/(1.19 + 0.78) = 0.396 for the 20-MeV event and [0.29/(1.17 + 0.29)]2 = (0.199)2 for the two 40-MeV events. Because every matched-energy onset fell in the wider field, the one-sided exact p-value is their product: pmatched = 0.396 (0.199)2 ≈ 1.6 × 10−2 .

(5)

With only three events, this is the smallest value the test can produce. Table X shows its sensitivity to alternative readings. Two of the three events lie in the run 4 wide-field window, which lasted about four minutes and has a ±150 s timing bound, equal to 62% of its fluence. Depending on how the recovery record is interpreted, this window contains one, two, or three onsets. Across the reported readings and timing bound, p ranges from 0.016 to 0.46. The largest value excludes both

10

TABLE X S ENSITIVITY OF THE M ATCHED -E NERGY F IELD T EST Reading of the record

Events 20/40 MeV

p

p (timing)

1/2

0.016

0.043

1/1

0.079

0.14

1/3

0.026

0.050

1/0

0.40

0.46

0/2

0.040

0.094

0/1 3 of 3

0.20 0.030

0.31 0.061

Primary: inclusive definition; recorded reboot ends a cluster Second run 4 onset merged; equivalently, the power-cycle episode excluded (JEDEC-consistent) Every failed relaunch after a recorded recovery counted; interval 4b credited to the run 4 denominator (0.79 × 1010 p/cm2 ) Failed launch at 90.4 min read as part of the F3 episode (alone, as the second row) and the power-cycle episode excluded Reboot-tier and power-cycle events only Reboot-tier events only Runs 1 to 4 pooled without energy stratification

One-sided exact conditional p-values; the timing bound lengthens every wide-field window and shortens every small-field window by its reconstruction bound. Under the primary reading the two-sided exact 95% interval for the wide-to-small rate ratio is [1.25, ∞), [0.76, ∞) with the timing bound, and [0.47, ∞) with the second run 4 onset merged. Crediting all run 4 exposure outside the operational window as well gives 0.066 (0.10).

run 4 interruptions from the Linux-SEFI count, leaving one event in the test. V. D ISCUSSION A. Alive, on Schedule, and Wrong Run 5 shows what a reset-on-error protocol misses: the behavior of corrupted state before the reset. The accelerator returned the same wrong class for the remaining 39 inferences of the block, about 0.9 s at its usual cadence. The scheduled reload ended the observed sequence, but the next block did not complete. The record therefore cannot establish whether the reload cleared the underlying condition. Although the same availability indicators recorded every availability failure in the campaign, none flagged this event. They did not inspect output content, and the memory test and power-rail samples were taken outside the inference phase. The behavior is consistent with numerical or data-path corruption that leaves control flow intact. Several fault classes could produce an input-independent constant class at the usual cadence. These include a corrupted constant or final-layer bias in the LPDDR4 model buffer, a stuck accumulator or SIMD path, or an instruction stream that bypasses layers. The class index excludes a stale output buffer (Section IV-B). The blockend model reload would refresh the corrupted constant or bias. It would also refresh an instruction-stream corruption if it lay in the program buffer rather than the decode logic. PL reconfiguration would refresh the logic involved in all three candidates. The retained class sequence does not distinguish among these candidates or exclude corruption in a transfer, configuration path, or PS-side driver state. These observations motivate content checks alongside availability monitors. Each check should specify which state it validates. The top-1 oracle used here cannot detect corruption that preserves the winning class, and the stock deployment

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

had no content checker. Periodic known-answer inputs could bound an episode to the probe interval if their state and expected outputs were protected. A class-histogram check would have flagged the absent class at its first occurrence. Checksums of static model buffers before each reload, together with a bounded wait and DMA status check, could test the candidates above at low cost. None of these measures was evaluated here. B. Persistent State and Recovery Scope The campaign recorded output corruption, a repeatedly reported corrected-error address, and malformed directory entries. Only the output events were silent. The longer output event lasted for the rest of its block without an alert from any availability indicator. It may have continued as the subsequent hang, a sequence observed after both output events. If so, the condition resided in state that the block-end reload did not refresh (Section III-D). This would exclude the three candidates in Section V-A and point instead to state beyond the reach of a process restart or PL reload. The correctederror address was reported in nine blocks spanning a reboot, consistent with a persistent cell or repeated upsets. However, ECC corrected every read and telemetry logged every report, so the corrupted value was never delivered to software. The malformed entries appeared in two root-filesystem directory blocks at one time. They replicate one pattern but do not show persistence, and their origin is unknown. Each recovery action refreshes different state. A process restart recreates userspace state but leaves Linux, drivers, PS state, external memory, and PL configuration unchanged. In this protocol, however, every recovery was followed by a new block that re-downloaded the PL and reloaded the model. The protocol therefore refreshed more state than the recovery action alone. On this platform, a Linux reboot performs a system reset that clears PL configuration and reruns the boot firmware [20]. PYNQ loads no bitstream at boot, so the next block start re-downloaded the PL. A reboot still cannot remove a persistent hardware condition; the run 5 OCM address returned after one. PL reconfiguration with a model reload restores the implemented logic, initialized PL memories, and model buffers, but not Linux, OCM, other LPDDR4 contents, or the page cache. Power cycling resets the broadest set of state at the cost of the longest interruption. SEM scrubbing, had it been used, would have covered PL configuration memory but not block-RAM contents, flip-flops, PS registers, OCM, LPDDR4, or software state [21]. Recovery should therefore escalate according to the state that may remain corrupted. A process supervisor can restart a terminated notebook. A failed known-answer check can trigger a model reload or PL reconfiguration, followed by further escalation only if the refreshed system fails again. This guidance follows from the refresh analysis, not from a comparison of recovery methods; no recovery strategies were tested against each other. C. What the Field Contrast Supports The field comparison supports an association, not component-level attribution. All nine primary onsets and all

11

TABLE XI S ELECTED L INUX -M ANAGED Z YNQ I RRADIATION C AMPAIGNS Study and platform

Monitoring scope and reported outcomes

This work; protons; Ultra96-V2, ZU3EG; PYNQ Linux + Tensil

Workload progress/recovery, Linux records, reachability, sampled rails, and retained top-1 classes (62 of 92 blocks). Seven Linux-SEFI and two class-changing output-event onsets under the clustering rule; class-preserving changes were outside the oracle. Reported 68 unexpected SoC failures, each followed by a power cycle: 13 process-failure/hang events, 17 kernel failures, and 38 application-processing-unit resets. Its uniform power-cycle recovery differs from the staged recovery used here. Heartbeat liveness and golden checksums per output. For the simplex accelerator (scrubbed static region, DDR ECC), 25 erroneous and one hung execution in 75,527 executions at 3.49 × 1011 n/cm2 ; consecutive erroneous outputs were counted as one error. Class sequences and Linux failure classes were not reported. Among 5985 ResNet-50 runs, reported 2964 correct, 89 crashes, 46 critical silent data corruptions (SDCs) associated with misclassification, and 2886 tolerable SDCs without a final-class change. Its numerical-output checker detects effects that a top-1-only checker cannot observe.

Stirk et al. [3]; neutrons; Ultra96; Linux + Dhrystone

Sabogal et al. [2]; neutrons; UltraZed-EG, ZU3EG; PetaLinux + ReCoN CNN

Agiakatsikas et al. [1]; neutrons; ZCU102, XCZU9EG; Linux + Vitis DPU

Counts are not susceptibility rankings. Particle spectra, devices, workloads, monitor coverage, clustering, recovery rules, and denominators differ.

three matched-energy Linux-SEFIs occurred under the wider field, which received 45.0% of the operational exposure. Under the small field, corrected-OCM reports were consistent with exposure of the SoC (Section IV-C), yet no LinuxSEFI occurred in 44 completed blocks and no output event occurred in the 31 verified blocks. This pattern is consistent with a contribution from circuitry outside the SoC footprint and supports treating the stack, rather than the chip, as the observed system. It does not identify the responsible circuitry. Every wide-field run followed both small-field runs, so field is confounded with run order and accumulated dose. The 58MeV run also has no small-field counterpart. Table XI summarizes the monitoring scope and reported outcomes of selected Linux-managed Zynq irradiation campaigns.

D. Limitations and Threats to Validity The campaign evaluated a specific hardware–software configuration: a Tensil accelerator implemented on an Ultra96-V2 and running a compiled ResNet-20 workload on a fixed set of ten CIFAR-10 images. The resulting cross sections apply only to this system and protocol; they do not quantify device-todevice variation, susceptibility across neural-network models, CIFAR-10 accuracy under irradiation, or flight qualification. The two beam configurations also did not isolate individual components. The edge of the nominal 2-cm field may have reached the LPDDR4 package, whereas the nominal 4-cm field exposed additional board resources. Run order was not randomized, and every wide-field run followed the smallfield runs. Five unmonitored same-day exposures added dose outside the reported denominators, and 58 MeV was tested only under the wide field. Consequently, the association between wide-field exposure and the observed onsets cannot be attributed specifically to LPDDR4. The field comparison

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

remains exploratory, and the energy-stratified cross sections in Tables VIII and IX do not establish an energy dependence. The experiment did not retain logits, intermediate tensors, model-memory checksums, DMA transactions, physical LPDDR4 addresses, or PL configuration readback. It therefore cannot reconstruct the internal propagation path of either output event or localize any onset. Because Linux evidence reached the laptop through the USB-gadget link, loss of notebook or secure-shell service cannot be distinguished from loss of the link itself without a local console. A link loss would have appeared as a workload interruption and could affect both the onset count and the assigned tier. The userspace memory test covered only 10 MB of LPDDR4 and did not retain physical-page mappings, limiting both memory coverage and fault localization. VI. C ONCLUSION Proton irradiation exposed two failure modes in this Linuxmanaged inference stack. Seven Linux-SEFIs stopped the workload. Two output-corruption events instead returned incorrect classes while execution continued. The longer event showed a stuck-class pattern: 39 consecutive CIFAR-10 inputs were assigned the same class, which was absent from the tenimage pool. Outputs continued at the usual 23–24 ms interval, and the availability monitors reported no anomaly. The stuckclass sequence was still present when the scheduled block-end bitstream reconfiguration began. The next launch failed, so the record cannot determine whether reconfiguration cleared the underlying condition. The other output event was also followed by a failed launch. The kernel also reported corrected OCM ECC errors at the same address in nine blocks spanning one reboot. ECC corrected every read, and the record does not link these reports to the output corruption. All nine event onsets in the primary analysis occurred under the wider beam field. This pattern is consistent with a contribution from circuitry outside the SoC footprint. However, field size was confounded with run order and accumulated dose. The results therefore do not localize the failures or identify LPDDR4 as their source. These results show that liveness and nominal timing do not establish inference correctness. Linux-managed accelerators need end-to-end, content-aware checks and staged recovery that can refresh the layers where corrupted state may persist. Radiation campaigns should count events by onset, match each cross-section denominator to the monitor that can observe its endpoint, and state that monitor’s blind spots. Recovery actions should be reported without using them to infer the fault’s physical origin. Future tests should retain complete score vectors, add a local console, randomize the beam-field order, and include long runs without scheduled bitstream reloads. Such runs would allow fault persistence to be measured beyond one block. This study provides a system-level baseline for future software hardening of COTS FPGA-SoCs used for neural-network inference in modern space systems. ACKNOWLEDGMENT The authors thank the operations staff of the Department of Radiation Research and Proton Radiotherapy at IFJ PAN for

12

beam time and dosimetry support. OpenAI ChatGPT (GPT5.6 Pro, September 2026) and Anthropic Claude (Fable 5.1, September 2026) were used for language editing, structural revision, and consistency checking of the abstract and Sections I–VI. They did not generate experimental observations. The authors checked all numerical values, equations, citations, interpretations, figures, and final wording against the campaign records and accept full responsibility for the manuscript. R EFERENCES [1] D. Agiakatsikas et al., “Single event effects assessment of UltraScale+ MPSoC systems under atmospheric radiation,” IEEE Trans. Rel., vol. 73, no. 1, pp. 771–783, Mar. 2024, doi: 10.1109/TR.2023.3312548. [2] S. Sabogal, A. D. George, and G. Crum, “ReCoN: A reconfigurable CNN acceleration framework for hybrid semantic segmentation on hybrid SoCs for space applications,” in Proc. IEEE Space Comput. Conf. (SCC), 2019, pp. 41–52, doi: 10.1109/SpaceComp.2019.00010. [3] W. Stirk, E. Poff, J. Smith, J. Goeders, and M. Wirthlin, “Comparison of neutron radiation testing approaches for a complex SoC,” IEEE Trans. Nucl. Sci., vol. 70, no. 4, pp. 505–514, Apr. 2023, doi: 10.1109/TNS.2023.3237080. [4] D. M. Hiemstra, V. Kirischian, and J. Brelski, “Single event upset characterization of the Zynq UltraScale+ MPSoC using proton irradiation,” in Proc. IEEE Radiat. Effects Data Workshop (REDW), New Orleans, LA, USA, Jul. 2017, pp. 1–4, doi: 10.1109/NSREC.2017.8115448. [5] Tensil AI, “Tensil: Open source machine learning accelerators,” ver. 1.0.15, 2022 (repository archived Oct. 2025). Accessed: Sep. 4, 2026. [Online]. Available: https://github.com/tensil-ai/tensil [6] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90. [7] A. Krizhevsky, “Learning multiple layers of features from tiny images,” Univ. Toronto, Toronto, ON, Canada, Tech. Rep., Apr. 2009. [8] T. Santini, L. Carro, F. R. Wagner, and P. Rech, “Reliability analysis of operating systems and software stack for embedded systems,” IEEE Trans. Nucl. Sci., vol. 63, no. 4, pp. 2225–2232, Aug. 2016, doi: 10.1109/TNS.2015.2513384. [9] P. Rech, “Artificial neural networks for space and safety-critical applications: Reliability issues and potential solutions,” IEEE Trans. Nucl. Sci., vol. 71, no. 4, pp. 377–404, Apr. 2024, doi: 10.1109/TNS.2024.3349956. [10] F. Libano, B. Wilson, M. Wirthlin, P. Rech, and J. Brunhaver, “Understanding the impact of quantization, accuracy, and radiation on the reliability of convolutional neural networks on FPGAs,” IEEE Trans. Nucl. Sci., vol. 67, no. 7, pp. 1478–1484, Jul. 2020, doi: 10.1109/TNS.2020.2983662. [11] P. Maillard et al., “Radiation-tolerant deep learning processor unit (DPU)-based platform using Xilinx 20-nm Kintex UltraScale FPGA,” IEEE Trans. Nucl. Sci., vol. 70, no. 4, pp. 714–721, Apr. 2023, doi: 10.1109/TNS.2022.3216360. [12] I. Souvatzoglou et al., “Assessing the reliability of FPGA-based quantized neural networks under neutron irradiation,” IEEE Trans. Nucl. Sci., vol. 71, no. 12, pp. 2565–2577, Dec. 2024, doi: 10.1109/TNS.2024.3491503. [13] F. Benevenuti et al., “Reliability of FINN-generated CNN accelerators for image classification on SRAM-based FPGAs under heavy-ioninduced faults,” IEEE Trans. Nucl. Sci., vol. 72, no. 8, pp. 2830–2838, Aug. 2025, doi: 10.1109/TNS.2025.3588691. [14] J. M. Badia et al., “Reliability of vision transformers and CNNs on edge AI systems under neutron radiation,” IEEE Trans. Nucl. Sci., vol. 72, no. 8, pp. 2706–2716, Aug. 2025, doi: 10.1109/TNS.2025.3536519. [15] S. M. Guertin and M. Cui, “SEE test results for the Snapdragon 820,” in Proc. IEEE Radiat. Effects Data Workshop (REDW), New Orleans, LA, USA, Jul. 2017, pp. 1–6, doi: 10.1109/NSREC.2017.8115452. [16] S. Memon, R. Graczyk, T. Rajkowski, J. Swakoń, and M. Papadakis, “Patterns that break memory: SEU characterization of COTS LPDDR2 and LPDDR4 SDRAM via stress testing under 60 MeV proton beam,” in Proc. 25th Int. Conf. Softw. Qual., Rel. Secur. (QRS), Hangzhou, China, Jul. 2025, pp. 462–472, doi: 10.1109/QRS65678.2025.00053. [17] S. M. Guertin, “Guideline for single-event effect (SEE) testing of system on a chip (SOC) devices,” Jet Propulsion Lab., California Inst. Technol., Pasadena, CA, USA, Rep. JPL-PUB-18-2, Feb. 2018. [Online]. Available: https://ntrs.nasa.gov/citations/20190002148

PREPRINT. SUBMITTED TO IEEE TRANSACTIONS ON NUCLEAR SCIENCE

[18] H. Quinn, “Challenges in testing complex systems,” IEEE Trans. Nucl. Sci., vol. 61, no. 2, pp. 766–786, Apr. 2014, doi: 10.1109/TNS.2014.2302432. [19] Ultra96-V2 Single Board Computer Hardware User’s Guide, ver. 1.3, Avnet, Phoenix, AZ, USA, Jun. 2021. [20] Zynq UltraScale+ Device Technical Reference Manual, UG1085 (v2.5), AMD, Santa Clara, CA, USA, Mar. 2025. [21] UltraScale Architecture Soft Error Mitigation Controller v3.1 LogiCORE IP Product Guide, PG187 (v3.1), Xilinx/AMD, San Jose, CA, USA, Nov. 2022. [22] J. Swakoń et al., “Facility for proton radiotherapy of eye cancer at IFJ PAN in Kraków,” Radiat. Meas., vol. 45, no. 10, pp. 1469–1471, Dec. 2010, doi: 10.1016/j.radmeas.2010.06.020. [23] J. F. Ziegler, M. D. Ziegler, and J. P. Biersack, “SRIM: The stopping and range of ions in matter (2010),” Nucl. Instrum. Methods Phys. Res. B, vol. 268, no. 11–12, pp. 1818–1823, Jun. 2010, doi: 10.1016/j.nimb.2010.02.091. [24] C. Cazabon, “memtester: A userspace memory-test utility,” ver. 4.7.1. Accessed: Sep. 4, 2026. [Online]. Available: https://pyropus.ca/softwar e/memtester/ [25] S. Memon et al., “Where Linux breaks under radiation: A cross-architecture kernel-level characterization of proton-induced failures in COTS SoCs,” 2026, arXiv:2503.03722v4, doi: 10.48550/arXiv.2503.03722. [26] S. Esquer, B. D. Sierawski, A. F. Witulski, R. D. Schrimpf, G. Karsai, and M. Turowski, “Single event functional interrupt (SEFI) sensitivities of a multicore microprocessor,” in Proc. IEEE Aerosp. Conf., 2024, pp. 1–11, doi: 10.1109/AERO58975.2024.10521240. [27] Measurement and Reporting of Alpha Particle and Terrestrial Cosmic Ray-Induced Soft Errors in Semiconductor Devices, JEDEC Standard JESD89A, JEDEC Solid State Technol. Assoc., Arlington, VA, USA, Oct. 2006.

13

Record · ID 660807 · SHA-256 921d16223ba2730f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.