Date of publication xxxx 00, 0000, date of current version xxxx 00, 0000. Digital Object Identifier 10.1109/ACCESS.2017.DOI
Trust-Aware Output Management for Physical Neural Network in Cloud-Continuum Systems MALIHEH HARIRI1
arXiv:2609.22443v1 [cs.DC] 18 Sep 2026
1, 2
, STEFAN FISCHER2
, (Member, IEEE)
University of Lübeck, Institute of Telematics, 23562 Lübeck, Germany
Corresponding author: Stefan Fischer (e-mail: [email protected]).
ABSTRACT Physical Neural Networks (PNNs) introduce new opportunities for cloud continuum computing, but their outputs may be affected by noise, drift, delay, and incomplete reliability information. Existing substrate-management approaches mainly focus on discovery, invocation, and monitoring, while the reliability of the returned output is often left unaddressed. This paper proposes a trustaware output management framework for heterogeneous PNNs. Each output is represented with quality and context information, and a lightweight edge-level trust score decides whether it should be accepted, rejected, or forwarded to the fog. At the fog layer, compatible outputs are checked for disagreement and combined using reliability- and uncertainty-aware fusion. Historical trust is also tracked to detect sustained degradation and support recalibration requests. The framework is evaluated using controlled and randomized PNN output models. Across 20 random seeds, the proposed full-trust policy reduces unsafe acceptance from approximately 61.5% for raw output handling to about 5.0%, while accepted-output MAE decreases from 0.504 to 0.242. Risk coverage analysis shows that this improvement is not explained only by lower local acceptance. The proposed fog fusion method also achieves the lowest aggregate mean error among the evaluated methods, with a small but consistent advantage over strong uncertainty-aware baselines. The prototype adds about 2 µs of edge processing per evidence record. These results show that post-invocation reliability management can improve the safe use of PNN outputs across edge–fog–cloud systems. INDEX TERMS Physical Neural Networks, Trust-Aware Output Management, Post-invocation Reliability, Physical AI, Cloud Continuum
I. INTRODUCTION
Cloud continuum systems distribute computation across edge, fog, and cloud resources instead of relying only on centralized infrastructure. This supports low-latency processing and closer interaction with the physical environment. As these systems evolve, they are also beginning to include non-conventional computing resources in addition to digital processors and accelerators. Physical Neural Networks (PNNs) are one example. They use the dynamics of physical substrates, including chemical, biological, memristive, photonic, and other material systems, as part of the computation itself [1]. Previous work has shown that nonlinear responses, memory effects, and device dynamics can provide useful computational behavior [2, 3]. These properties make VOLUME 4, 2016
PNNs relevant for edge and extreme-edge environments, but they also make their outputs sensitive to physical conditions. A PNN output may be affected by noise, drift, delay, calibration state, aging, or environmental changes. Its result may also take different forms, such as spike activity, optical intensity, conductance, or chemical concentration. Therefore, successful execution alone does not guarantee that the returned output is reliable enough to use in a distributed decision process. Several systems support the programming and access of physical computing substrates. NIR and EdgeMap improve portability and deployment for neuromorphic systems [4, 5], while BioCRNpyler supports compilation for chemical reaction networks [6]. Wetware platforms 1
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
such as DishBrain, NeuroPlatform, and the Cortical Labs API provide software-facing access to biological neural systems [7, 8, 9]. More recently, CP2 N2 provides a substrate-aware control plane for discovering, invoking, and monitoring heterogeneous PNN resources across the cloud continuum [10]. Related ideas also appear in the Model Context Protocol, Web of Things, and digitaltwin systems [11, 12, 13]. These approaches mainly focus on how a substrate is described, selected, invoked, or monitored. They do not explicitly decide whether the output returned after invocation should be trusted. This becomes important for PNNs because a correctly invoked substrate may still produce an untrustable result. The problem considered in this paper is therefore the reliability of the output after successful invocation. To address this problem, we introduce a trust-aware output management framework for PNNs in the cloud continuum. The main contributions are: • A common evidence representation and edge-level trust model that considers confidence, uncertainty, drift, and freshness when deciding whether an output should be accepted, rejected, or forwarded. • A fog-level mechanism that checks compatibility and source disagreement before applying reliabilityand uncertainty-aware fusion to eligible PNN outputs. • An adaptive trust mechanism that tracks reliability over time and can trigger recalibration requests when persistent degradation is detected. The remainder of the paper is organized as follows. Section II reviews related work. Section III presents the system model and assumptions. Section IV describes the proposed framework. Section V presents the evaluation, and Section VI concludes the paper. II. RELATED WORK
The reliability of Physical Neural Network (PNN) outputs is related to several areas, including physical neural computing, edge intelligence, uncertainty-aware inference, sensor fusion, trustworthy edge computing, and adaptive system management. These areas provide useful methods for handling heterogeneous computation and uncertain data, but none directly addresses the post-invocation question considered in this work: whether a returned PNN output is reliable enough to be accepted, rejected, or further processed. Fischer et al. [1] review physical neural computing across memristive, photonic, chemical, mechanical, and other substrates, showing both the computational potential and the strong dependence of these systems on their physical dynamics. Tanaka et al. [2] and Marković et al. [3] similarly show how physical dynamics can provide efficient computation, but also make performance dependent on device and environmental conditions; these works study the computing mechanisms 2
rather than how their outputs should be trusted after execution. Several works instead focus on making heterogeneous substrates easier to program or access. Pedersen et al. [4] introduce a common intermediate representation for neuromorphic systems, while Xue et al. [5] optimize the mapping of spiking neural networks to edge hardware. These approaches improve portability and deployment, but assume that a successfully executed model produces an output that can be used directly. Similar abstractions exist for other physical substrates. BioCRNpyler [6] and ChemComp [14] provide compilation methods for chemical reaction networks, reducing the gap between high-level models and physical implementations. Their focus is on translating and executing computations, not on judging the reliability of the physical result returned after execution. Wetware platforms make this issue even more visible. Kagan et al. [7] demonstrated closed-loop interaction with cultured neural cells, while Jordan et al. [8] and Hogan et al. [9] provide remote and software-facing access to biological neural systems. These platforms show that physical neural computation can be exposed through software interfaces, but output quality remains closely tied to biological state and experimental conditions and is not handled through a common postinvocation trust mechanism. Shi et al. [15] study communication-efficient Edge AI and show how distributing learning and inference closer to data sources can reduce latency and communication cost. Xu et al. [16] provide a broader view of edge intelligence in which data, models, computation, and decisions are distributed across the network; both approaches motivate local decision-making but mainly treat model outputs as conventional digital results. Murshed et al. [17] review machine-learning deployment at the network edge, including model compression, hardware support, and resource constraints. Baccour et al. [18] focus on resource-efficient distributed AI for IoT systems and consider computation, communication, and energy jointly. These works provide the systems basis for edge–cloud execution, but do not evaluate the reliability of a physical output after the underlying computation has completed. Qendro et al. [19] introduce a lightweight uncertaintyaware sensing approach suitable for resource-constrained edge devices. Their method makes prediction uncertainty available with limited overhead, but mainly captures uncertainty of a digital learning model rather than physical effects. This distinction matters for PNNs because uncertainty can originate from both computation and the physical substrate itself. A returned value may therefore appear confident at the model level while still being affected by changing physical conditions, which motiVOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
vates combining several reliability indicators rather than relying on predictive confidence alone. Gruber et al. [20] show that uncertainty-aware sensor fusion can reduce the influence of unreliable measurements compared with treating all sensor values equally. Their later work [21] applies this idea to physical sensor networks and connects uncertainty information with digital representations of sensing entities. These methods provide a useful basis for weighted fusion, but they mainly consider measurements from conventional sensors rather than computational outputs produced by heterogeneous physical neural substrates. Vedurmudi et al. [22] review automated uncertainty handling in sensor-network metrology, including middleware, fusion, machine learning, and agent-based techniques. This work treats uncertainty as part of the dataprocessing pipeline, but generally assumes sensing systems with defined measurement and calibration models; PNN outputs may additionally depend on the internal dynamics of the computing substrate. Wang et al. [23] review trustworthy edge intelligence from the perspectives of reliability, security, transparency, and sustainability. Their work provides a broad view of trust in distributed AI systems, but does not define an output-level decision process for noisy, drifting, or stale physical neural computations. Hallyburton et al. [24] use adaptive trust estimation in sensor fusion, allowing the influence of a source to change as new observations become available. This supports the idea that reliability should evolve over time, although their setting mainly addresses security and faulty sensing rather than degradation caused by the physical computing substrate itself. Tan and Matta [25] study the synchronization problem between physical systems and their digital twins and formalize when twin state should be updated. Kamburjan et al. [13] address lifecycle management of digital twins through explicit models of system state and evolution. These approaches are useful for longterm monitoring and recovery, but they do not provide a mechanism for deciding whether an individual PNN output should be trusted immediately after invocation. We summarize all research directions in Table 1 which are mainly address conventional digital AI models, sensor measurements, security-aware fusion, or substratelevel control. None of the proposed papers directly address the post-invocation reliability problem of heterogeneous PNN outputs.
agement layer maintains longer-term information such as calibration history, historical trust, and digital-twin state. PNN outputs are not assumed to be fully reliable digital values. Depending on the substrate and operating conditions, an output may be affected by noise, drift, delay, calibration error, or changes in the physical state. Each returned output is therefore represented as an evidence record containing its value together with available quality and context information, including confidence, uncertainty, drift, timestamp, modality, and provenance. These indicators are obtained from substrate adapters, runtime telemetry, calibration records, repeated observations, or digital-twin information rather than being manually supplied by the user. The edge layer acts as the first reliability decision point after a PNN invocation. It performs lightweight trust estimation and decides whether an output should be accepted locally, rejected, or forwarded to the fog. When required quality information is missing, the output is treated conservatively and can be forwarded for further assessment. The fog layer may receive outputs from several PNN sources referring to the same task or observation. Fusion is performed only when the evidence is compatible in task, modality, time, and decision context. Compatible outputs are checked for disagreement before numerical fusion, so strongly inconsistent evidence is not combined automatically. The cloud or management layer keeps longer-term reliability information and may receive recalibration requests when sustained degradation is detected. The current prototype evaluates trust tracking and recalibration triggering, but does not directly modify physical calibration parameters or automatically synchronize a deployed digital twin. Communication between continuum layers is assumed to have non-zero delay and limited bandwidth. Network optimization is outside the scope of this work; communication delay is reflected through evidence freshness, while compact evidence records are preferred over transferring raw substrate traces. The framework is intended for heterogeneous PNN technologies, including wetware, chemical, memristive, photonic, and other physical computing substrates, without assuming a common internal implementation.
III. SYSTEM MODEL A. ASSUMPTIONS
Given a PNN output generated by a physical substrate, the objective is to decide whether the output is reliable enough for local use, should be rejected, or should be forwarded to the fog layer for additional fusion. Each output is represented as an evidence record containing the returned value and its quality indicators. The decision problem is therefore defined over an evidence
B. PROBLEM FORMULATION
We consider a cloud-continuum architecture with edge, fog, and cloud layers. PNN substrates are located close to the physical environment, mainly at the extreme edge or edge layer. Fog nodes provide local coordination among multiple substrates, while the cloud or manVOLUME 4, 2016
3
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
Table 1: Comparison of the proposed framework with related research directions. Research direction Physical neural computing and reservoir systems [1, 2, 3] Substrate programming and deployment [4, 5, 6, 14] Wetware access and biological computing [7, 8, 9] Edge intelligence [15, 16, 17, 18] Uncertainty-aware edge inference [19] Uncertainty-aware sensor fusion and metrology [20, 21, 22] Trustworthy edge AI and adaptive trust [23, 24] Digital-twin synchronization and lifecycle management [25, 13] CP2 N2 / substrate control [10, 11] Proposed framework
Edge/Fog –
Uncertainty Partial
Fusion/Trust –
PNN-aware ✓
Post-invocation output trust –
Partial
–
–
✓
–
Partial
–
–
✓
–
✓ ✓ ✓
– ✓ ✓
– – ✓
– – –
– – –
✓
✓
✓
–
–
Partial
Partial
–
–
–
✓ ✓
Partial ✓
Partial ✓
✓ ✓
– ✓
record ri , where the system must select an action Di ∈ {Accept, Reject, Forward}. The objective is to reduce unsafe acceptance of unreliable physical outputs while preserving useful local decision-making for reliable outputs. This creates a trade-off between reliability and local availability. Accepting too many outputs at the edge may propagate noisy, stale, or drifting physical results, whereas a more conservative policy reduces local coverage and increases the number of outputs forwarded for additional processing. The proposed framework addresses this trade-off through trust-aware edge scoring and fog-level compatibility checking, disagreement assessment, and fusion. The trust score Ti therefore serves as a compact reliability estimate derived from the quality indicators associated with ri . Rather than being treated as an optimization objective itself, Ti provides the basis for the threshold-based decision policy introduced in Section IV-C. The acceptance and rejection thresholds determine how conservatively the system operates and can be selected according to the reliability requirements and risk tolerance of the target application. IV. METHODOLOGY
The proposed methodology consists of five main stages: evidence construction and parameter estimation, edgelevel trust computation, trust-based decision-making, fog-level compatibility checking and fusion, and adaptive feedback for long-term reliability management. The following subsections describe each stage in detail. A. EVIDENCE CONSTRUCTION
When a PNN substrate produces an output, the first step is to convert the raw physical signal into a structured evidence record. This step is required because PNNs may return different types of outputs depending on the substrate. Although these outputs are physically different, the proposed framework represents them through a common evidence structure: ri = (vi , ci , σi , di , ti , mi , ρi ), 4
(1)
where vi denotes the output value returned by substrate i; ci ∈ [0, 1] is its confidence score; σi ∈ [0, 1] represents the normalized noise or uncertainty level; di ∈ [0, 1] represents the normalized drift from calibrated behavior; ti is the output-generation timestamp; mi identifies the output modality; and ρi contains provenance metadata, such as the substrate, adapter, calibration, location, or digital-twin identifier. The evidence record therefore captures both the physical output and the information required to assess its reliability. Outputs with similar values may consequently receive different trust assessments depending on their confidence, noise, drift, freshness, and provenance. B. PARAMETER ESTIMATION AND NORMALIZATION
The quality parameters in the evidence record (Formula 1 ) are not assumed to be manually provided by the user. They are estimated from telemetry, repeated measurements, adapter metadata, calibration records, timestamps, and digital-twin state. This makes the proposed method suitable for integration with substrateaware control layers that already expose runtime and life cycle information. The confidence value (ci ) in formula 1 represents how reliable the output appears from the viewpoint of the substrate or adapter. It can be derived from signal strength, backend health, convergence quality, viability status, classification margin, or the stability of repeated responses. For example, a spike-based output with a clear and repeatable response pattern receives a higher confidence value than a weak or inconsistent response. The confidence score is normalized to the range ([0,1]), where 1 indicates high confidence, and 0 indicates very low confidence. The uncertainty or noise value (σi ) captures shortterm instability in the output. If repeated measurements or samples are available, it can be estimated using the standard deviation of the returned values: v u K u1 X raw (vik − µi )2 , (2) σi =t K k=1
VOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
where (vik ) is the (k)-th repeated output from substrate (i), (K) is the number of samples, and (µi ) is their mean. The raw uncertainty is then normalized as raw σi ,1 , (3) σi = min σimax where (σimax ) is the maximum acceptable uncertainty for the corresponding task or substrate class. A lower value of (σi ) means that the output is more stable. The drift value (di ) measures how far the current behavior of the substrate has moved away from its calibrated or expected behavior. This can be estimated by comparing the current response with a calibration baseline or with a digital-twin prediction: current ∆i ,1 , (4) di = min ∆max i where (∆current ) is the deviation between the current i response and the reference behavior, and (∆max ) is the i maximum tolerable deviation. A small drift value means that the substrate is still close to its calibrated behavior, while a high drift value indicates that recalibration or reduced trust may be required. Freshness is computed from the timestamp of the output. An output that was produced recently is usually more relevant than an old output, especially in timesensitive edge applications. The freshness value is defined as fi = e−∆ti /τ ,
(5)
where (∆ti = tnow − ti ) is the age of the output and (τ ) is a task-dependent time constant. A larger (τ ) can be used for slow physical processes, while a smaller (τ ) is suitable for fast edge decisions. If required reliability telemetry is missing, the framework applies a conservative decision override. Diagnostic estimates may still be constructed internally to preserve a complete evidence representation, but an output with required quality information missing is not accepted locally solely on the basis of these estimates. Instead, the evidence is marked as incomplete and forwarded for additional processing. This policy prevents an apparently high numerical trust score from masking the absence of information required to verify the reliability of the output. C. EDGE-LEVEL TRUST ESTIMATION
After construction and normalization of evidence, the edge layer computes a trust score for each PNN output. The trust score summarizes the output quality into a single value that can be used for fast decision-making: Ti = αci + β(1 − σi ) + γ(1 − di ) + δfi ,
(6)
where (Ti ) is the trust score of output (i), and (α), (β), (γ), and (δ) are non-negative weights. These weights VOLUME 4, 2016
control the importance of confidence, noise, drift, and freshness. They satisfy α + β + γ + δ = 1.
(7)
The weighted additive form in Eq. (6) is intentionally selected for edge-level decision-making. Edge nodes may operate under limited computational resources and cannot always execute expensive probabilistic inference after each physical invocation. The proposed score therefore provides a bounded, monotonic, and interpretable trust estimate that can be computed in constant time when the required telemetry values are available. Boundedness. If ci , σi , di , fi ∈ [0, 1], and if α, β, γ, δ ≥ 0 with α + β + γ + δ = 1, then the trust score Ti is bounded in the interval [0, 1]. This follows because the terms ci , 1 − σi , 1 − di , and fi are all in [0, 1]. Since Ti is a convex combination of these bounded terms, Ti must also lie in [0, 1]. Monotonicity. The score is monotonic with respect to the quality indicators. Increasing confidence ci or freshness fi increases the trust score, while increasing noise σi or drift di decreases the trust score. This behavior is desirable for PNN outputs because the trust score should reward stable, fresh, and confident physical evidence while penalizing uncertainty and substrate degradation. D. EDGE-LEVEL DECISION POLICY
The edge layer uses the trust score to decide how the output should be handled. The decision rule is defined as: Accept, Di = Forward to Fog, Reject,
Ti ≥ θaccept , θreject ≤ Ti < θaccept , Ti < θreject ,
(8)
where (θaccept ) and (θreject ) are trust thresholds. If the output has high trust, it is accepted and can be used locally. If the trust score is very low, the output is rejected early to avoid propagating unreliable evidence. If the output is uncertain, it is forwarded to the fog layer. This middle decision region is important because uncertain outputs may still be useful when compared or combined with outputs from other PNNs or edge nodes. The three-way policy exposes a direct trade-off between local availability and reliability. High-trust outputs can be used immediately at the edge, while clearly unreliable outputs are rejected and intermediate cases are escalated for additional evidence processing. The framework therefore avoids unconditional forwarding, but does not assume that fog escalation is rare. Under challenging uncertainty conditions, a large fraction of outputs may be forwarded intentionally because the objective is to avoid forcing unreliable local decisions. Communication cost itself is not optimized in this work 5
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
and depends on the selected trust thresholds and operating conditions. E. FOG-LEVEL COMPATIBILITY CHECKING AND FUSION
where ci denotes the confidence of source i, di denotes its normalized drift, and fi denotes its freshness. The non-negative fog-level reliability weights satisfy αf + γf + δf = 1.
The fog layer receives uncertain evidence records that have been forwarded by edge nodes because their edgelevel trust scores fall within the intermediate decision region. Rather than forcing an independent decision from each uncertain output, the fog layer compares evidence from multiple PNN sources and performs fusion only when the evidence is sufficiently compatible and mutually consistent. Before fusion, the fog layer performs compatibility checking. In the current framework, two or more evidence records are considered compatible when they refer to the same task and output modality. Their timestamps must also fall within a configured temporal window. When decision-context or observation-window identifiers are available in the provenance metadata, these identifiers must also be consistent across the evidence records. Therefore, compatibility is determined using the task identifier, output modality, timestamp, and available contextual provenance information. Evidence records that fail these checks are not fused. For a set of N compatible numerical outputs {v1 , . . . , vN }, the fog layer next evaluates the degree of disagreement among the sources. To reduce sensitivity to an extreme individual value, the median of the compatible outputs is first used as a robust reference value:
The fog-level reliability factor Ri is distinct from the edge-level trust score Ti defined in Eq. (6). The edge-level trust score combines confidence, uncertainty, drift, and freshness and is used to select among local acceptance, rejection, and forwarding. In contrast, Ri is used only during fog-level fusion. The uncertainty term is deliberately excluded from Ri because uncertainty is incorporated separately through the fusion weight. This avoids penalizing the same uncertainty estimate once inside the reliability factor and again through inverseuncertainty weighting. The final fusion weight assigned to source i is defined as wi =
Ri 2
max (σi , σmin ) + ϵ
(9)
max |vi − ṽ| ∆fog =
i
max (|ṽ|, 1)
.
(15)
wi
i=1
.
(10)
The denominator prevents numerical instability when the reference value is close to zero. Fusion is permitted only when ∆fog ≤ θdis ,
(11)
where θdis denotes the maximum acceptable disagreement threshold. If the disagreement exceeds this threshold, the fog layer does not force a fused output. Instead, the current evidence set is treated as insufficiently consistent, and the system can request an additional measurement or initiate an appropriate recovery action, such as recalibration. When compatibility and disagreement conditions are satisfied, the fog layer computes a reliability factor for each source. The fog-level reliability factor is defined as Ri = αf ci + γf (1 − di ) + δf fi , 6
(14)
w i vi
ŷ = i=1 N X
The normalized disagreement is then defined as
,
where σi denotes the normalized uncertainty associated with source i, σmin > 0 is a minimum uncertainty floor, and ϵ > 0 is a small constant for numerical stability. The uncertainty floor prevents a source with a near-zero estimated uncertainty from receiving an excessively large fusion weight. The final numerical output is then computed as the normalized weighted average N X
ṽ = median (v1 , . . . , vN ) .
(13)
(12)
Consequently, a source receives greater influence when it has high confidence, low drift, good freshness, and low estimated uncertainty. Confidence, drift, and freshness determine the fog-level reliability factor Ri , while uncertainty affects the source separately through the inverseuncertainty component of wi . The complete fog-level procedure therefore consists of three stages. First, the system filters evidence according to compatibility. Second, it evaluates whether the compatible sources agree sufficiently according to Eq. (10). Third, only evidence that passes both checks is combined using Eqs. (12)–(15). Evidence that is incompatible or exhibits excessive disagreement is not forced into a fused decision and is instead handled through additional measurement or recovery mechanisms. F. ADAPTIVE RELIABILITY TRACKING AND RECALIBRATION TRIGGERING
Physical neural substrates may change over time because of drift, aging, environmental variation, biological VOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
variability, or calibration loss. Evaluating every invocation independently may therefore fail to detect gradual or persistent degradation. To provide temporal context, the framework maintains a historical trust estimate for each substrate. Let T̄i (t) denote the historical trust estimate of substrate i after invocation t. It is updated using an exponential moving average: T̄i (t) = λT̄i (t − 1) + (1 − λ)Ti (t),
(16)
where λ ∈ [0, 1] controls the balance between historical and recent observations. A larger λ produces a smoother long-term estimate, whereas a smaller value makes the historical trust more responsive to recent changes. The current implementation combines historical trust degradation with a counter of consecutive low-trust observations. A recalibration request is generated when persistent low current trust or sufficiently low historical trust indicates sustained degradation. This mechanism helps distinguish temporary reliability fluctuations from longer-term changes in substrate behavior. The resulting trigger can be passed to an external substratemanagement, calibration, or digital-twin component. The present evaluation validates historical trust tracking and recalibration triggering; it does not perform physical recalibration of a real PNN substrate or automatically update a deployed digital twin. Algorithm 1 summarizes the proposed trust-aware output management process. G. COMPUTATIONAL COMPLEXITY
The proposed framework keeps the latency-sensitive edge processing lightweight. Evidence construction, missing-telemetry checking, freshness computation, trust-score evaluation, and the three-way decision rule each require O(1) time per output when the required quality indicators are already available. If uncertainty is estimated from K repeated measurements, the corresponding edge-level cost becomes O(K) per source. The adaptive trust update also requires only O(1) time per invocation. At the fog layer, compatibility checking and numerical fusion over N evidence records require O(N ) time, while the median-based disagreement analysis requires O(N log N ) in the current implementation. Therefore, the overall complexity is dominated by disagreement checking and is O(N log N ) when quality indicators are precomputed. When uncertainty estimation from K repeated measurements is included for all N sources, the end-to-end worst-case complexity becomes O(N K + N log N ).
Algorithm 1 Trust-Aware Management of PNN Outputs Require: PNN output vi , confidence ci , uncertainty σi , drift di , timestamp ti , modality mi , provenance ρi Ensure: Edge-level acceptance or rejection, fog-level fused result, or uncertain decision 1: Build evidence record ri ← (vi , ci , σi , di , ti , mi , ρi ) 2: Normalize confidence, uncertainty, drift, and freshness-related values 3: if required reliability telemetry is missing then 4: Mark the evidence as incomplete 5: Forward the evidence record ri to the fog layer 6: else 7: Compute freshness fi using Eq. (5) 8: Compute edge-level trust score Ti using Eq. (6) 9: if Ti ≥ θaccept then 10: Accept the PNN output at the edge 11: Update historical trust for substrate i 12: return vi 13: else if Ti < θreject then 14: Reject the PNN output as unreliable 15: Update historical trust for substrate i 16: return rejection decision 17: else 18: Forward the evidence record ri to the fog layer 19: end if 20: end if 21: Identify evidence records compatible in task identifier, modality, timestamp window, and available contextual provenance 22: if no compatible evidence set exists then 23: Request additional evidence or measurement 24: Update historical trust for the involved substrate(s) 25: return uncertain decision 26: end if 27: Compute disagreement ∆fog using Eq. (10) 28: if ∆fog > θdis then 29: Request additional measurement or initiate recovery action 30: Update historical trust for the involved substrate(s) 31: return uncertain decision 32: end if 33: for each compatible evidence record rj do 34: Compute fog-level reliability factor Rj using Eq. (12) 35: Compute fusion weight wj using Eq. (14) 36: end for 37: Compute fog-level fused result ŷ using Eq. (15) 38: Update historical trust for the involved substrate(s) 39: return ŷ
V. EVALUATION & ANALYSIS VOLUME 4, 2016
7
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
A. EXPERIMENTAL SETUP AND METRICS
The evaluation examines whether explicit postinvocation reliability management improves the handling of PNN outputs after successful substrate invocation. The proposed framework is implemented as an extension of the CP2 N2 [10] prototype. The underlying CP2 N2 workflow remains responsible for substrate discovery, matching, invocation, and telemetry access, whereas the proposed layer operates on the returned output and its associated reliability information. The experiments use CP2 N2 -compatible numerical PNN output models under controlled and randomized uncertainty conditions. Hidden physical noise and drift are used to generate corrupted outputs, while the trust layer receives imperfect estimates of these quantities rather than their exact ground-truth values. Confidence is generated from an independent latent backend-quality component together with smaller contributions from estimated noise and drift. Delay influences the physical error through simulated state evolution and independently affects the freshness component of the trust score. The default edge configuration uses equal trust weights and decision thresholds of θaccept = 0.75 and θreject = 0.35. Equal weights are used as a neutral reference configuration rather than as an assumed optimal setting. Separate sensitivity experiments examine alternative weight priorities and different levels of decisionpolicy strictness. The edge-level evaluation compares four methods: direct use of the raw CP2 N2 output, confidence-only selection, freshness-only selection, and the proposed full trust score. The controlled evaluation varies physical noise, drift, output delay, and the freshness time constant τ . A separate randomized stress experiment samples a broader range of uncertainty conditions. The main controlled, randomized, and fog-fusion experiments are repeated across 20 independent random seeds to evaluate statistical stability. We distinguish several reliability quantities that capture different aspects of decision quality. Coverage is defined as P (Accept). The conditional false acceptance rate (FAR) is defined as P (Accept | Bad), while the unsafe acceptance frequency is defined as P (Accept ∩ Bad). The accepted-output contamination rate is defined as P (Bad | Accept). These metrics have different denominators and therefore should not be interpreted interchangeably. The conditional false rejection rate is similarly defined as P (Reject | Good). Accepted-output mean absolute error (MAE) measures the numerical error only among outputs that are accepted locally. Because a lower acceptance rate can itself reduce accepted-output error, a separate risk–coverage experiment varies the acceptance threshold while keeping the rejection threshold fixed. This experiment evaluates whether the proposed method continues to achieve lower 8
selective risk than confidence-only filtering at comparable coverage levels. At the fog layer, seven numerical fusion methods are evaluated: simple averaging, confidence weighting, inverse-uncertainty weighting, trust-only weighting, median fusion, the original Ti /σi2 formulation, and the proposed orthogonal Ri /σi2 formulation. Compatibility and normalized disagreement are assessed separately from numerical fusion. Deterministic functional tests additionally verify the rejection of evidence with mismatched task identifiers, modalities, decision contexts, observation windows, or timestamps. Additional experiments evaluate trust-component ablation, missing reliability telemetry, adaptive historical trust, and computational scalability. Unless otherwise stated, confidence intervals in the multi-seed evaluation represent 95% confidence intervals computed over seedlevel aggregate results. B. RESULTS AND DISCUSSION
Figure 1 examines accepted-output quality as delay increases. CP2 N2 and confidence-only continue to accept outputs even at long delays because neither policy explicitly prevents stale evidence from being used. Freshness-only stops accepting once the freshness value becomes too low, while the full trust method also stops local acceptance when the combined evidence is insufficiently reliable. Consequently, their curves end because accepted-output MAE is undefined when no outputs are accepted locally; the missing points therefore represent selective forwarding or rejection, not zero error or missing experimental data. The non-monotonic MAE of the remaining methods is expected because MAE is computed only over the subset accepted at each delay and the physical error also depends on the simulated state evolution. Figure 2 examines whether the lower accepted-output error of the full trust method results only from accepting fewer outputs. Over the overlapping coverage range, the full-trust curve remains below the confidence-only curve, showing that the multi-factor trust score ranks reliable outputs more effectively than confidence alone. CP2 N2 appears as a single reference point rather than a curve because it is an accept-all baseline: it has no postinvocation acceptance threshold to vary and therefore operates only at full coverage. The result consequently reflects a reliability–coverage trade-off rather than a comparison based only on one fixed threshold. Figure 3 shows the availability cost of responding to substrate degradation. CP2 N2 maintains full local acceptance because it performs no reliability filtering. Confidence-only gradually reduces acceptance, whereas full trust reduces it more strongly because increasing drift directly lowers the trust score. The falling fulltrust curve therefore represents deliberate escalation of degraded evidence rather than loss of functionality. VOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
Figure 1: Accepted-output MAE as delay increases. Shorter curves indicate operating regions in which a policy accepts no outputs locally, making acceptedoutput MAE undefined.
Figure 3: Local acceptance under increasing substrate drift. Full trust progressively shifts uncertain outputs away from local acceptance as degradation becomes stronger.
Figure 2: Reliability–coverage trade-off for confidenceonly and full-trust selection. CP2 N2 is shown as a single accept-all reference point because it has no postinvocation reliability threshold to sweep.
Figure 4: Local acceptance under increasing physical noise. The full trust policy becomes more selective as output uncertainty increases.
Figure 4 shows a similar reliability–availability response to physical noise. CP2 N2 continues accepting all outputs, while confidence-only becomes more selective and full trust reduces local acceptance more rapidly. This behavior is consistent with the proposed policy: strong uncertainty is handled through forwarding or rejection instead of preserving edge coverage at any cost. Figure 5 shows the safety consequence of these decisions. Unsafe acceptance increases strongly for raw CP2 N2 as drift grows, whereas confidence-only remains substantially lower. Full trust maintains the lowest unsafe acceptance and approaches zero under severe drift. This reduction must be interpreted together with Figure 3: part of the safety gain results from deliberately VOLUME 4, 2016
withholding unreliable outputs from local acceptance. Figure 6 shows the same pattern under physical noise. Raw CP2 N2 increasingly propagates unreliable outputs, while confidence-only reduces the risk and full trust suppresses it further. At high noise levels, the nearzero unsafe acceptance of full trust coincides with very low local coverage, demonstrating the intended safety– availability trade-off rather than claiming that highly corrupted outputs become accurate. Figure 7 shows how the individual trust components affect accepted-output quality under increasing noise. Removing uncertainty has the clearest effect at severe noise because the policy loses the component that directly represents short-term physical instability. Removing drift also degrades accepted-output quality 9
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
Figure 5: Safety and local availability under increasing physical uncertainty at τ = 5 s.
Figure 7: Accepted-output MAE under increasing noise for the complete trust model and each ablated variant.
Figure 6: Safety and local availability under increasing physical uncertainty at τ = 5 s.
Figure 8: False acceptance under increasing substrate drift for the full trust model and its component ablations.
because the noise bins still contain scenarios with different substrate conditions. Some curves terminate at higher noise levels because the corresponding variants accept no outputs in those bins; their missing MAE values therefore indicate selective behavior rather than zero error. Figure 8 isolates the role of drift information. The complete method increasingly suppresses unsafe local acceptance as degradation becomes stronger. When drift is removed, unsafe acceptance no longer falls in the same way and remains present across degraded conditions because the policy cannot directly recognize movement away from calibrated behavior. Other components can still reduce acceptance indirectly, but they do not replace an explicit drift signal. Figure 9 demonstrates the distinct role of freshness. Variants that retain freshness rapidly stop unsafe local acceptance as evidence ages. Removing freshness leaves 10
a persistent unsafe-acceptance level even at long delays, because confidence, uncertainty, and drift contain no direct information about the age of the evidence. Freshness therefore contributes temporal validity that cannot be fully replaced by the other trust components. Figure 10 compares numerical fusion as the number of PNN sources increases. All evaluated methods show lower mean error with additional sources, indicating the benefit of combining more independent evidence. Simple averaging, confidence weighting, and median fusion remain weaker because they do not jointly account for physical uncertainty and reliability. The uncertainty-aware methods form the strongest group. The proposed Ri /σi2 formulation remains very close to the uncertainty-only and original trust–uncertainty approaches and achieves the lowest aggregate mean error across the 20-seed evaluation, although the numerical VOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
Figure 9: False acceptance under increasing output delay, highlighting the role of freshness in preventing staleoutput acceptance.
Figure 11: Fusion-allowed rate as the number of PNN sources increases under the fixed disagreement threshold.
Figure 10: Mean fusion error as the number of PNN sources increases for the seven evaluated fusion strategies.
Figure 12: Adaptive current and historical trust during healthy operation, sustained degradation, and simulated recovery.
advantage is small. The main contribution of the fog layer should therefore be interpreted as the combination of compatibility checking, disagreement gating, and reliability-aware fusion rather than a large improvement from weighting alone. Figure 11 shows that disagreement gating becomes more selective as the evidence set grows. With more sources, the probability that at least one value exceeds the maximum allowed deviation from the median increases, so fewer evidence sets pass the fixed disagreement threshold. This reduction in fusion availability is therefore an expected consequence of the gating rule rather than evidence that numerical fusion becomes less accurate with additional sources. Figure 12 illustrates how the adaptive mechanism responds to persistent changes in substrate reliability.
Current trust reacts quickly to individual observations, whereas the exponential moving average changes more slowly and filters short-term variation. During sustained degradation, repeated low-trust observations eventually satisfy the persistence condition and generate a recalibration request. During the simulated recovery phase, current trust improves first while historical trust follows more gradually. The experiment demonstrates the intended temporal behavior of the trigger, but does not represent physical recalibration or automatic digitaltwin synchronization.
VOLUME 4, 2016
C. SCALABILITY AND LIMITATIONS
Figure 13 shows that the total edge-level processing time increases approximately proportionally to the number of evidence records. The corresponding per record mea11
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
Figure 13: Edge-level trust scalability. Total runtime grows approximately proportionally with the number of processed evidence records.
Figure 14: Fog-level processing scalability for numerical fusion and compatibility/disagreement assessment as the number of PNN sources increases.
surements remain close to 2 µs across the tested workload sizes, which is consistent with the O(1) trust-score and decision cost per record when quality indicators are already available. Figure 14 shows that both numerical fusion and compatibility/disagreement assessment remain lightweight for the evaluated source counts. Their measured runtime increases smoothly with the number of sources. Numerical fusion is O(N ), while the current median-based disagreement procedure gives the complete assessment an O(N log N ) worst-case complexity. The fact that compatibility/disagreement assessment is faster than numerical fusion in the tested range does not contradict this asymptotic result; for the evaluated source counts, implementation constants dominate the theoretical growth-rate difference. 12
The additional validation experiments support the same interpretation. Under the deliberately harsh randomized stress distribution, the full trust policy becomes substantially more conservative than confidenceonly filtering and maintains markedly lower unsafe acceptance and accepted-output error. The sensitivity analysis confirms that stricter decision thresholds reduce unsafe acceptance at the cost of lower local coverage, while different weight priorities shift the balance among physical failure modes. Missing-telemetry tests show that outputs lacking required noise, drift, or freshness information are conservatively forwarded rather than locally accepted. Finally, all deterministic fog-compatibility tests correctly reject mismatched task, modality, decision-context, observation-window, and timestamp conditions. VI. CONCLUSION & FUTURE WORK
This paper introduced a trust-aware framework for managing PNN outputs after invocation in cloud-continuum systems. Instead of treating every returned result as reliable, the framework evaluates output quality at the edge, forwards uncertain cases to the fog, and applies compatibility checking, disagreement screening, and reliability-aware fusion when multiple PNN outputs are available. The results show a clear improvement in output reliability. Across 20 random seeds, unsafe acceptance decreased from about 61.5% for raw CP2 N2 output handling to about 5.0%, while accepted-output MAE decreased from approximately 0.504 to 0.242. The risk– coverage analysis also shows that this improvement is not only caused by lower local acceptance, since the full trust method performs better than confidence-only filtering at similar coverage levels. At the fog layer, the proposed fusion method provides a small but consistent improvement over strong uncertainty-aware baselines, while the main benefit comes from combining compatibility, disagreement, and reliability checks before fusion. The current prototype mainly relies on controlled and synthetic uncertainty scenarios. Future work will focus on trace-driven and real PNN experiments, where confidence, uncertainty, and drift can be estimated directly from substrate behavior. Further work will also study adaptive thresholds, disagreement sensitivity, and tighter integration with recalibration and digital-twin management. VII. DATA AVAILABILITY
https://github.com/MaliheHa93/Trust_PNNs References
[1] S. Fischer, N. Ay, O. Landsiedel, E. Mohammadi, S. Otte, B.-C. Renner, N. Rußwinkel, Beyond silicon: Materials, mechanisms, and methods for physVOLUME 4, 2016
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
ical neural computing, 2026. URL: https://arxiv. org/abs/2604.09833. doi:. arXiv:2604.09833. [2] G. Tanaka, T. Yamane, J. B. Héroux, R. Nakane, N. Kanazawa, S. Takeda, H. Numata, D. Nakano, A. Hirose, Recent advances in physical reservoir computing: A review, Neural Networks 115 (2019) 100–123. [3] D. Marković, A. Mizrahi, D. Querlioz, J. Grollier, Physics for neuromorphic computing, Nature Reviews Physics 2 (2020) 499–510. [4] J. E. Pedersen, S. Abreu, M. Jobst, G. Lenz, V. Fra, F. C. Bauer, D. R. Muir, et al., Neuromorphic intermediate representation: A unified instruction set for interoperable brain-inspired computing, Nature Communications 15 (2024). [5] J. Xue, L. Xie, F. Chen, L. Wu, Q. Tian, Y. Zhou, R. Ying, P. Liu, Edgemap: An optimized mapping toolchain for spiking neural network in edge computing, Sensors 23 (2023) 6548. [6] W. Poole, A. Pandey, A. Shur, Z. A. Tuza, R. M. Murray, Biocrnpyler: Compiling chemical reaction networks from biomolecular parts in diverse contexts, PLOS Computational Biology 18 (2022) e1009987. [7] B. J. Kagan, A. C. Kitchen, N. T. Tran, F. Habibollahi, M. Khajehnejad, B. J. Parker, A. Bhat, B. Rollo, A. Razi, K. J. Friston, In vitro neurons learn and exhibit sentience when embodied in a simulated game-world, Neuron 110 (2022) 3952– 3969.e8. [8] F. D. Jordan, M. Kutter, J.-M. Comby, F. Brozzi, E. Kurtys, Open and remotely accessible neuroplatform for research in wetware computing, Frontiers in Artificial Intelligence 7 (2024) 1376042. [9] D. Hogan, A. Doherty, B. K. Khoo, J. Zhou, R. Salib, J. Stewart, K. Lawson, A. Loeffler, B. J. Kagan, Cl api: Real-time closed-loop interactions with biological neural networks, 2026. URL: https: //docs.corticallabs.com/. doi:. [10] S. Fischer, M. Hariri, S. Otte, phys-mcp: A control plane for heterogeneous physical neural networks, arXiv preprint arXiv:2605.04256 (2026). [11] D. S. Parra, J. Spahr-Summers, Model Context Protocol Specification, Technical Report, Model Context Protocol, 2025. URL: https:// modelcontextprotocol.io/specification/2025-11-25. [12] M. Lagally, R. Matsukura, M. McCool, K. Toumura, Web of Things (WoT) Architecture 1.1, W3C Recommendation, World Wide Web Consortium, 2023. URL: https://www.w3.org/TR/wot-architecture11/, 5 December 2023. [13] E. Kamburjan, N. Bencomo, S. L. T. Tarifa, E. B. Johnsen, Declarative lifecycle management in digital twins, in: Proceedings of the ACM/IEEE 27th International Conference on Model Driven VOLUME 4, 2016
Engineering Languages and Systems, Association for Computing Machinery, 2024, pp. 353–363. doi:. [14] N. B. Agostini, C. G. Johnson, W. C. Cannon, A. Tumeo, Chemcomp: A compilation framework for computing with chemical reaction networks, in: Proceedings of the 30th Asia and South Pacific Design Automation Conference (ASPDAC ’25), 2025, pp. 872–878. doi:. [15] Y. Shi, K. Yang, T. Jiang, J. Zhang, K. B. Letaief, Communication-efficient edge ai: Algorithms and systems, IEEE Communications Surveys & Tutorials 22 (2020) 2167–2191. [16] D. Xu, T. Li, Y. Li, X. Su, S. Tarkoma, T. Jiang, J. Crowcroft, P. Hui, Edge intelligence: Empowering intelligence to the edge of network, Proceedings of the IEEE 109 (2021) 1778–1837. [17] M. G. S. Murshed, C. Murphy, D. Hou, N. Khan, G. Ananthanarayanan, F. Hussain, Machine learning at the network edge: A survey, ACM Computing Surveys 54 (2021) 1–37. [18] E. Baccour, N. Mhaisen, A. A. Abdellatif, A. Erbad, A. Mohamed, M. Hamdi, M. Guizani, Pervasive ai for iot applications: A survey on resourceefficient distributed artificial intelligence, IEEE Communications Surveys & Tutorials 24 (2022) 2366–2418. [19] L. Qendro, J. Chauhan, A. G. C. P. Ramos, C. Mascolo, The benefit of the doubt: Uncertainty aware sensing for edge computing platforms, in: Proceedings of the 2021 IEEE/ACM Symposium on Edge Computing, 2021. doi:. [20] M. Gruber, S. Eichstädt, W. Pilar von Pilchau, J. Hähner, V. Gowtham, A. Willner, N.-S. Koutrakis, J. Polte, M. Riedl, Uncertainty-aware sensor fusion in sensor networks, in: SMSI 2021 – Sensor and Measurement Science International, 2021, pp. 246–247. doi:. [21] M. Gruber, W. Pilar von Pilchau, V. Gowtham, T. Dorst, S. Eichstädt, J. Hähner, N.-S. Koutrakis, J. Polte, M. Riedl, A. Willner, Application of uncertainty-aware sensor fusion in physical sensor networks, in: 2022 IEEE International Instrumentation and Measurement Technology Conference, 2022, pp. 1–6. doi:. [22] A. P. Vedurmudi, et al., Automation in sensor network metrology: An overview of methods and their implementations, Measurement: Sensors (2025). [23] X. Wang, B. Wang, Y. Wu, Z. Ning, S. Guo, F. R. Yu, A survey on trustworthy edge intelligence: From security and reliability to transparency and sustainability, IEEE Communications Surveys & Tutorials (2024). [24] R. S. Hallyburton, et al., Security-aware sensor fusion with mate, arXiv preprint arXiv:2503.04954 (2025). [25] B. Tan, A. Matta, The digital twin synchronization 13
Hariri et al.: Trust-Aware Management of Physical Neural Network Outputs
problem: Framework, formulations, and analysis, IISE Transactions 56 (2024) 652–665. STEFAN FISCHER is a full professor of computer science at the University of Lübeck, Germany, and the Director of the Institute of Telematics. He received the Diploma degree in information systems and the doctoral degree in computer science from the University of Mannheim, Germany, in 1992 and 1996, respectively. After a postdoctoral year at the University of Montreal, Canada, he held positions at the International University in Germany and the Technical University of Braunschweig before joining the University of Lübeck in 2004. His research interests include distributed systems, sensor and ad hoc networks, the Internet of Things, smart cities, molecular communications, and physical neural computing. MALIHEH HARIRI is a Ph.D. student in computer science at the University of Lübeck, Germany, working on distributed intelligent systems, edge and fog computing, and emerging computing paradigms, especially in scheduling and resource management. Currently, her research focuses on the integration of Physical Neural Networks into cloud-continuum environments. She has also been working on scientific workflow scheduling, Service Function Chain placement, and resource optimization in fog and cloud computing systems. Her broader research interests include physical AI, edge intelligence, distributed AI infrastructures, and reliable decision-making under uncertainty.
14
VOLUME 4, 2016