IntCC: A Framework for Interactive Confidential Computing Qingzhe Bing∗
Kaiyuan Zhang∗
Yinqian Zhang✉
SUSTech Shenzhen, China [email protected]
SUSTech Shenzhen, China [email protected]
SUSTech Shenzhen, China [email protected]
arXiv:2609.35552v1 [cs.CR] 28 Sep 2026
Abstract
processing [8]. At its core, this paradigm relies on hardwarebased Trusted Execution Environments (TEEs) to establish secure execution contexts through strict memory isolation and transparent hardware-level encryption. Modern TEEs provide protection at various granularities, ranging from process-level enclaves (e.g., Intel SGX [28]) to full-system Confidential Virtual Machines (CVMs) (e.g., AMD SEV-SNP [2] and Intel TDX [29]). By shielding the execution state from the underlying untrusted host operating system or hypervisor, these technologies promise to protect sensitive assets from cloud providers, effectively enabling the secure deployment of cloud-native workloads [13, 40]. However, a fundamental paradox lies at the heart of confidential computing: the rigid, static security model of TEEs is inherently incompatible with the dynamic nature of modern interactive workloads. TEEs rely on remote attestation to guarantee the integrity of their initial memory state [7]. This mechanism binds the trust of the system to a static measurement of the code and data loaded at boot time. This model is fundamentally at odds with interactive workflows. In scenarios like Large Language Model (LLM) fine-tuning [35, 42, 58] or Exploratory Data Analysis (EDA), data processors require human-in-the-loop capabilities, including dynamic code injection, intermediate state inspection, and hyperparameter tuning. These operations are inherently dynamic and cannot be anticipated at launch time, placing them beyond the coverage of traditional remote attestation. This conflict is explicitly acknowledged by leading industry frameworks. For instance, the Confidential Containers (CoCo) architecture [13] warns that enabling interactive features like kubectl exec fundamentally disrupts the static chain of trust established by remote attestation; consequently, such tools are strictly disabled in production profiles [12]. Red Hat practitioners similarly highlight this dilemma, noting that while diagnosing production issues is crucial, enabling standard debugging “risks exposing sensitive data to the entities we aim to exclude from the trust model, undermining the core principles of confidential computing” [18]. Furthermore, the Confidential Computing Consortium warns that unrestricted debugging interfaces render it “almost impossible to ensure the protection of the confidentiality and integrity of the workload” [14]. Crucially, while existing industry guidelines primarily focus on mitigating threats from untrusted cloud providers,
Confidential computing leverages Trusted Execution Environments (TEEs) to ensure the confidentiality and integrity of data in use. However, TEEs rely on remote attestation to guarantee the integrity of their initial memory state. This model is fundamentally at odds with interactive development workflows. In scenarios like LLM fine-tuning and exploratory data analysis, data processors need human-in-the-loop capabilities, including dynamic code injection, intermediate state inspection, and hyperparameter tuning, all of which inherently violate the static, one-time integrity guarantees of traditional remote attestation. To reconcile this tension, we propose the interactive confidential computing paradigm, a system architecture enabling untrusted data processors to execute dynamic, nondeterministic operations within TEEs without compromising data confidentiality. Driven by the insight that inherently unmeasurable human interaction must be excluded from the Trusted Computing Base (TCB), we logically partition the TEE into an interactive controller and a verifiable runtime. To realize this paradigm, we present IntCC, a framework featuring three key mechanisms: (1) a proxy-based dispatch system to preserve the native development experience; (2) a fine-grained information flow control mechanism based on a security lattice to prevent data leakage; and (3) a privacypreserving verifiable execution mechanism to guarantee the runtime compliance of dynamic workflows. We implement IntCC on AMD SEV-SNP using Confidential Containers and evaluate it across diverse real-world workloads. Our experiments demonstrate that IntCC effectively balances security and interactivity, incurring a practical overhead of less than 5% for LLM fine-tuning and under 17% for data analysis relative to baseline execution. CCS Concepts: • Security and privacy → Domain-specific security and privacy architectures. Keywords: Confidential Computing; Trusted Execution Environments; LLM Fine-Tuning; Data Analysis
1
Introduction
The confidential computing paradigm is reshaping cloud security by ensuring that data remains protected during ∗ Both authors contributed equally to this research.
1
interactive workloads exacerbate the inherent mutual distrust between data owners and data processors. On one hand, data owners cannot safely provision interactive access: once post-launch operations are permitted, a data processor could pivot from a benign tenant executing a pre-measured workload into an active adversary injecting malicious logic to exfiltrate private assets. On the other hand, data processors are reluctant to trust data owners or external auditors. Because interactive workflows embed proprietary algorithms, hyperparameters, and custom logic, processors are fundamentally unwilling to expose their source code or plaintext execution traces for compliance auditing. This dichotomy leaves a critical research gap: how can we enable untrusted data processors to execute non-deterministic, human-in-the-loop workflows on sensitive data, while guaranteeing verifiable runtime compliance for the data owner without compromising the processor’s intellectual property? To address this problem, we present interactive confidential computing, a paradigm that enables untrusted entities to perform dynamic, non-deterministic operations within TEE-protected boundaries while still providing guarantees for data confidentiality and verifiable runtime compliance. The key insight is that because human interaction is inherently arbitrary and unmeasurable, they must be excluded from the Trusted Computing Base (TCB). The underlying architecture for interactivity must therefore be a logical partition: the unmeasurable interactive logic is stripped from the TCB and isolated into an Interactive Controller, while measurable sensitive operations are confined within a Verifiable Runtime. Crucially, the fundamental contract bridging these two partitions is a prescribed security policy. Under this paradigm, data processors can utilize familiar ecosystems (e.g., PyTorch [41]) seamlessly; however, all cross-boundary interactions and operations on sensitive assets are strictly mediated by the Verifiable Runtime to ensure absolute compliance with the agreed-upon security policy. To realize this partitioned architecture, we propose the IntCC framework to address the critical challenges of maintaining confidentiality and ensuring integrity under dynamic interaction. First, to guarantee performance while preserving the native development experience, we design a proxy-based dispatch system. This system maintains lightweight shadow objects in the controller that reference concrete data within the runtime, enabling data processors to operate on remote assets using standard APIs without incurring the prohibitive overhead of cross-domain serialization. Second, simply separating the domains is insufficient if the interactive controller can infer sensitive data. We solve this by designing a fine-grained information flow control mechanism [15, 45] based on a security lattice that enforces relaxed noninterference, ensuring that only data adhering to declassification policies can cross the security boundary. Third, to establish runtime compliance for dynamically determined execution paths without exposing the data processor’s proprietary code
to data owners or auditors, we introduce a privacy-preserving verifiable execution mechanism. This mechanism generates canonical audit evidence for dynamic operations, commits to the execution trace, and proves in zero-knowledge that the committed trace satisfies predefined policies. We instantiate IntCC on AMD SEV-SNP [2] based Kata Containers [54] and evaluate it across real-world interactive workloads, including LLM fine-tuning and data analysis. The results demonstrate that IntCC supports these workflows with practical overhead relative to the baseline, while effectively mitigating data leakage risks [19, 21, 59]. In summary, the contributions are as follows: • We introduce the interactive confidential computing paradigm to reconcile the fundamental tension between the static integrity guarantees of TEE remote attestation and the inherent dynamism of interactive development. • We architect the IntCC framework to operationalize this paradigm. IntCC integrates three core mechanisms: (1) a proxy-based dispatch system to preserve native usability and performance; (2) a fine-grained information flow control mechanism to strictly enforce data confidentiality; and (3) a privacy-preserving verifiable execution mechanism to guarantee runtime compliance. • We implement IntCC on AMD SEV-SNP and evaluate it against real-world workloads. Our results demonstrate that IntCC effectively thwarts data exfiltration while incurring highly practical overhead (e.g., < 5% for LLM fine-tuning and < 17% for data analysis).
2
Background
2.1
Confidential Computing
Confidential computing leverages TEEs to protect data in use. Modern CVMs, such as AMD SEV-SNP [2] and Intel TDX [29], incorporate the Guest OS and application stack into the TCB. This architectural design substantially reduces migration barriers for complex workloads like deep learning [1, 41]. Furthermore, to accommodate cloud-native deployments, projects like CoCo seamlessly encapsulate these unmodified workloads within CVMs [13, 54]. Remote attestation serves as the foundational trust primitive in confidential computing. It enables remote verifiers to validate that a CVM initial state, encompassing firmware, kernel, and loaded images, conforms to expected values via measurement reports [7]. However, this rigid security model exposes a fundamental coverage gap for interactive workloads: it cannot attest to non-deterministic, post-launch interactions introduced by human operators. Consequently, to preserve strict security assurances, existing frameworks enforce a black-box model that severely restricts external interactions [12]. While maximizing security, this inherently conflicts with the dynamic, iterative nature and observability requirements of modern cloud-native applications. 2
2.2
Application Scenarios for Interactive Confidential Computing
model fine-tuning and exploratory data analysis, workloads are inherently dynamic and iterative [11, 35]. Data processors require the capability to observe intermediate results in real time, dynamically adjust hyperparameters, or even inject temporary debugging code to diagnose anomalies. This creates a fundamental dilemma between interactivity and measurability. Interactivity inherently implies the continuous injection of non-deterministic, unmeasured entropy (e.g., human inputs, dynamic code) into the system. Measurability, conversely, demands a static, deterministic boundary to establish a cryptographic root of trust. The current landscape places users in a zero-sum situation: to obtain TEE measurable security, they must forgo interactivity; conversely, to enable interactivity, they are forced to breach the integrity boundary, rendering the environment unmeasurable. Therefore, the critical question arises: how can we reconcile these two seemingly opposing properties within a unified framework?
While the black-box model of TEEs is effective for static tasks, production environments frequently demand interactive workflows where computational logic cannot be fully determined prior to launch. We identify three representative scenarios driving this need: • LLM Fine-tuning: LLM fine-tuning [26, 35, 58] is inherently iterative and human-in-the-loop. Data processors typically cannot determine the optimal configuration in a single pass; they require runtime monitoring during training. For instance, they must observe the loss function curve to detect overfitting, dynamically adjust hyperparameters such as learning rate or batch size based on intermediate results, and inspect problematic samples when model performance is anomalous. However, supporting such human-in-the-loop interactivity in existing TEEs requires exposing I/O boundaries, creating severe risks of leaking private training data to reconstruction attacks [21, 49, 59].
3.2
• Exploratory Data Analysis: In sensitive domains like healthcare or finance, data processors rarely formulate complete analysis programs in advance [11, 32]. Instead, they rely on iterative workflows to dynamically compose code, tune aggregation pipelines, and inspect distribution summaries without leaking raw records. • Debugging and Diagnosis of Confidential Applications: Maintaining complex applications in CVMs encounters a strict observability wall. When production containers crash, traditional diagnostics (e.g., attaching GDB or inspecting core dumps) are explicitly prohibited [12, 14, 18]. Consequently, data processors urgently need secure mechanisms to interactively inspect runtime states, such as stack traces and variables, without compromising sensitive assets. While these domains differ, they share a fundamental security dilemma: the necessity for dynamic code injection, runtime state observability, and non-deterministic operations directly conflicts with the static measurement and strict data confidentiality guarantees of traditional TEEs.
System Model
3
Motivation
We consider a representative system architecture comprising four distinct entities: Data Owners: The entities with high-value sensitive data. They provision data access to data processors for specific computational tasks under the strict mandate that data confidentiality is preserved and misuse is strictly prevented. Data Processors: The entities responsible for authoring code, debugging programs, or executing analytical tasks. They demand interactive access to the computing environment (e.g., via SSH) to dynamically submit non-deterministic instructions. Crucially, they view their source code as proprietary intellectual property, refusing to disclose plaintext execution traces for external auditing. Auditors: The verifying entities responsible for ensuring that the data processor’s runtime behaviors strictly adhere to the owner’s prescribed security policies. They may act directly on behalf of the data owners or serve as independent third-party verifiers. Cloud Service Provider (CSP): The entity providing the underlying confidential computing environments.
3.1
The Dilemma between Interactivity and Measurability
3.3
Threat Model
Our study makes the following assumptions: CSP is not trusted by data owners and data processors: Consistent with traditional confidential computing models, we treat the CSP as an external adversary that controls the software stack outside the TEE, such as the hypervisor and host OS. This adversary may attempt to observe or tamper with protected execution through host-side attacks such as memory dumping, bus snooping, or privileged introspection, thereby targeting both the data owner’s sensitive assets and the data processor’s proprietary code or workflow artifacts.
TEEs rely on remote attestation [7, 13] to guarantee the integrity of their initial memory states. This mechanism binds system trust to static measurements of the code and data loaded at boot time. Consequently, this security model establishes a strict chain of trust, requiring verifiers to depend entirely on these pre-launch measurements to confirm that the environment has booted into a known-good state. However, a fundamental dichotomy exists between this rigid security model and the development lifecycle of complex applications. In scenarios such as machine learning 3
However, it is not assumed to break the TEE’s hardware isolation guarantees [2, 29]. Data processors are not trusted by data owners: Interactive confidential computing introduces a second adversarial principal: the data processors. Unlike traditional confidential computing, where the attested workload is typically fixed and the main adversary is the CSP, interactive workloads intentionally permit data processors to issue new post-launch actions through legitimate interfaces. This threat does not assume that the TEE is compromised. Rather, it reflects a limitation of static remote attestation: attestation authenticates the initial measured state, but cannot pre-approve the data processor’s future interactions. A malicious data processor may therefore abuse authorized interactivity to trigger policy-violating behavior, directly exfiltrate sensitive data, or infer protected information indirectly [19, 59]. Auditors are not trusted by data processors: Establishing runtime compliance necessitates an auditing mechanism, which introduces a fundamental trust dilemma. Data processors inherently distrust auditors—including data owners acting as verifiers. Data processors consider their interactive workflows (e.g., proprietary code structure, API usage patterns and workflow strategy) to be highly valuable intellectual property. If the auditing mechanism requires plaintext execution traces or source code disclosure, a malicious auditor could exploit this access to steal the data processor’s trade secrets. Consequently, the data processor views the auditor as a potential threat to its workflow privacy. Security Goals: We target the following aims: • Confidentiality: The system must ensure the confidentiality of the data owner’s assets against both the data processors and the CSP as a fundamental security invariant, which must hold even under interactive execution. Additionally, the confidentiality of the data processor’s intellectual property (e.g., source code or workflow traces) must be protected from the CSP and the data owners.
runtime compliance. Crucially, the verification process itself must not undermine any confidentiality goals.
3.4
Design Goals and Technical Challenges
In light of the aforementioned conflicts and threats, we propose the paradigm of interactive confidential computing. We define this paradigm as a system that permits untrusted subjects to execute dynamic, non-deterministic operations within the protected boundary of a TEE, while simultaneously guaranteeing data confidentiality and verifiable runtime compliance. The Necessity of Logical Partitioning. To resolve the dilemma described in Sec. 3.1, we argue that a logical partitioning of the TCB is necessary for the following reasons. First, human interaction is inherently unmeasurable. Unlike deterministic algorithms, human operators introduce external, unpredictable entropy into the system. If these behaviors occur directly within the TCB, the runtime memory state is continuously mutated by unmeasured instructions. Second, measurability is the prerequisite for trust. The security model of a TEE is predicated on the capability to measure the TCB. If the TCB includes unmeasurable interactive logic, the entire TCB becomes unmeasurable, invalidating the chain of trust established by remote attestation. Consequently, we posit that interactivity is incompatible with a monolithic TCB. Logical partitioning is not merely an architectural choice but a necessity. To maintain security, we must strip the unmeasurable interactive logic from the TCB and isolate it into an Interactive Controller, while confining measurable sensitive operations within a Verifiable Runtime. This separation constitutes the only viable approach to reconciling dynamic control with static integrity. To realize interactive confidential computing, the system must simultaneously satisfy the following three properties, which introduce non-trivial technical challenges: Property 1: Transparent Compatibility. As a primary usability requirement, the system must enable data processors to use existing technology stacks without being aware of the underlying partitioned architecture, while keeping performance overhead within an acceptable range. • Challenge 1: The Performance Challenge and Compatibility Issues of Partitioned Execution. The logical partitioning established above necessitates isolating interactive logic from data states, breaking the assumption of a single address space [1]. If massive data such as model weights were transferred via traditional RPC, the overhead of serialization and deserialization would lead to unacceptable interactive latency. Therefore, a critical engineering challenge lies in how to create an illusion for upper layer applications into perceiving data as local within an isolated environment, while simultaneously eliminating the performance bottlenecks associated with cross-domain data copying.
• Integrity: The system must guarantee the authenticity and policy compliance of dynamic workloads at runtime. While enforcing low-level control-flow integrity or arbitrary program functional correctness is intractable for nondeterministic, human-in-the-loop interactions, our integrity goal ensures runtime compliance. Specifically, it guarantees that the execution trace genuinely reflects the operations processed within the hardware-isolated environment, and that these dynamic behaviors strictly adhere to the data owner’s prescribed IFC policies. • Verifiability: The system must provide cryptographic mechanisms to prove that all aforementioned security invariants were strictly maintained throughout the interaction. Unlike static remote attestation, which authenticates only the initial state, the system must generate cryptographic evidence allowing auditors to formally verify the 4
Property 2: Interactivity. The system must empower data processors to orchestrate the computational workflow and inspect intermediate states (e.g., dynamically modifying code, viewing statistical results). However, this interactivity must not devolve into a conduit for data leakage. • Challenge 2: Reconciling Flexibility with Information Flow Security. In an interactive environment, distinguishing benign debugging interactions from malicious secret exfiltration is exceptionally challenging. Traditional access control lists only govern static file access. While information flow control can track data propagation paths [45], strict noninterference properties often render the system unusable. The challenge lies in constructing a fine-grained security lattice within the TEE to constrain unstructured interactive behaviors within security boundaries compliant with relaxed noninterference [33].
Framework Design
4.1
Overall Architecture
To resolve the tension between interactivity and measurability, IntCC materializes the logical partitioning strategy proposed in Sec. 3.1, separating an interactive control domain from a measured execution domain. As illustrated in Figure 1, the IntCC-Controller hosts human-driven, post-launch logic and the proxy-based dispatch system. Conversely, the IntCC-Runtime hosts the compute engine operating on sensitive assets, securely governed by a fine-grained IFC mechanism and a verifiable execution module. While both partitions are physically protected from the untrusted host by the TEE, they assume distinct security roles: the Controller is excluded from the TCB for integrity, whereas the Runtime serves as the static root of trust for secure computation, strict policy enforcement, and runtime evidence generation. We formally define the roles and security properties of these two partitions as follows: • IntCC-Controller: This partition functions as the Interactive Controller, exposing standard development interfaces for dynamic post-launch interaction, including Python scripts, Jupyter kernels, and workload-specific control logic. It is structurally designed to host human-facing logic that cannot be fully pre-measured at boot time. Accordingly, although the confidentiality of this partition remains protected by the TEE against the CSP, it is strictly treated as outside the TCB for verifiable runtime compliance. From a policy enforcement perspective, the Controller is untrusted and may issue arbitrary interactive requests through the explicitly prescribed interfaces.
Property 3: Verifiable Execution and Compliance. The system must generate cryptographic evidence of its dynamic runtime behavior. We refer to this property as verifiable runtime compliance: it empowers data processors to cryptographically verify that their interactive instructions were indeed executed within the attested TEE, while enabling data owners and auditors to verify that the resulting execution trace strictly adhered to the prescribed security policies. • Challenge 3: The Efficiency and Privacy Challenge of Runtime Verification. Prior approaches such as Operation Execution Integrity (OEI) [52] assume fixed operations whose legal control paths and critical data can be identified in advance; they are therefore ill-suited to interactive confidential computing, where execution paths are workloaddependent and not exhaustively enumerable. Furthermore, finer-grained attestation or instruction-level tracing would incur prohibitive overhead [4, 48]. Directly disclosing execution traces may reveal the data processor’s proprietary code structure or workflow strategy. Consequently, a key challenge is how to prove that dynamic operation sequences have been executed and adhered to prescribed policies, without sacrificing performance or data processor’s privacy.
3.5
4
• IntCC-Runtime: This partition serves as the Verifiable Runtime and constitutes the system’s measured root of trust for verifiable runtime compliance. It hosts the private data, execution backends, and the IFC monitor that mediates all dynamic interactions on sensitive assets. Initialized from an attested measured state, the Runtime is the foundational component whose identity is authenticated by remote attestation and whose audit evidence is later utilized to prove compliance with the prescribed IFC policies. This design ensures that all operations on sensitive data remain verifiable, fulfilling the integrity requirements. As shown in the data flow in Figure 1, the IntCC-Controller cannot access any sensitive data directly. Instead, it interacts with the IntCC-Runtime through a proxy-based dispatch system (Sec. 4.2). The Runtime enforces a fine-grained information flow control mechanism (Sec. 4.3) to prevent data leakage and emits privacy-preserving runtime evidence via a verifiable execution mechanism (Sec. 4.4) to show that the trace was generated by the attested runtime and strictly satisfies the designated IFC constraints. Notably, this partitioned architecture supports diverse implementation strategies. It can be realized using two distinct SGX enclaves, isolated containers within a single CVM (e.g.,
Our Solution
To tackle the aforementioned challenges, we propose the IntCC framework and design the following mechanisms. First, a proxy-based dispatch system preserves the native development experience under partitioned execution while keeping sensitive objects inside the Verifiable Runtime. Second, a lattice-based IFC mechanism tracks data sensitivity and permits cross-boundary outputs only through approved declassification functions. Third, a privacy-preserving verifiable execution mechanism binds dynamic runtime traces to the attested Runtime and proves policy compliance in zeroknowledge without exposing the data processor’s workflow. 5
IntCC-Controller
IntCC-Runtime gRPC Response
② Fine-Grained IFC gRPC Request (Commands + Handles)
User Interface ① Proxy-based Dispatch System User Python Script … model.train() for batch in dataloader: optimizer.zero_grad() …
Proxy(model) Proxy(data) Handle
gRPC Private Dataset Server
gRPC
Real Objects
Boundary Checker (Taint Sink)
Model Weights Handle
Proxy(model) Proxy(data) Handle
Compute Engine
gRPC Response (Handles/Public Value)
Handle
③ Verifiable Execution Result is Public (L) → Return Value by gRPC Result is Sensitive → Create New Handle & Return Handle by gRPC
Figure 1. High-level architecture of the framework. • Interception: When an operation is triggered, the shadow object intercepts the data processor’s call and marshals the method name alongside opaque handles—randomized identifiers that reveal zero information about the underlying data’s content or memory layout—into a request.
leveraging CoCo [13]), or a nested TEE configuration (e.g., executing the runtime within an Intel SGX [28] enclave situated inside an Intel TDX [29] trust domain). 4.2
Proxy-based Dispatch System
• Dispatch: The request is transmitted across the trust boundary to the IntCC-Runtime via a secure channel.
To guarantee performance and preserve the native development experience under partitioned execution, IntCC incorporates a proxy-based dispatch system. By decoupling control flow from data flow, this system enables processors to manipulate remote assets using local syntax without exposing raw data. The design relies on two core mechanisms: shadow object abstraction and an opaque reference protocol.
• Resolve and Execute: The Dispatcher resolves the handle IDs to concrete memory addresses. The requested operation is then executed on the real data. • Register and Propagate: The Runtime registers the execution result locally and propagates a unique Handle ID back to the Controller, where a new shadow object is instantiated to wrap it.
4.2.1 Shadow Object Abstraction. In traditional remote execution, data is often serialized and transmitted to the client for processing. IntCC inverts this model by introducing shadow objects—lightweight, stateless abstractions in the IntCC-Controller that act as semantic mirrors of the concrete entities residing within the IntCC-Runtime. A shadow object acts as a proxy mirroring the public interface (attributes and methods) of the concrete class, encapsulating no physical data. When a data processor interacts with it, the object does not compute locally. Instead, it serves as a trap, intercepting invocations to preserve the native development experience while strictly separating states.
• Registration and Propagation: Depending on the operation’s semantics, the Runtime either updates the existing data in-place or registers a newly generated asset in its object store. When a new asset is produced, its unique Handle ID is propagated back to the Controller to instantiate a new shadow object, seamlessly sustaining the execution chain. Inspired by prior works like TrainCheck [30] and DLBox [27], the proxy layer functions as both a state-isolation boundary and a dynamic policy enforcement point to regulate interactive operations. Specifically, the system inspects the data sensitivity level at this boundary: low-sensitivity data is permitted to be serialized and returned as concrete values, which enables data processors to seamlessly inspect low-sensitivity artifacts necessary for debugging and tuning.
4.2.2 Opaque Reference Protocol. To secure the boundary, we employ a handle-based protocol that replaces direct memory access with opaque references. The workflow, as depicted in Figure 2, proceeds in four logical phases: 6
IntCC-Controller (Unmeasurable Domain) Script y = function(x)
Shadow Object (Interception) Wraps: ℎ𝑥
Invoke
Propagate
New Shadow Object Wraps handle ℎ 𝑦
(Return Handle ℎ 𝑦 )
IntCC-Runtime (Measurable Domain)
Dispatch (Method + Handles)
Dispatcher
(NO RAW DATA)
Object Store Registrate
Execute
Execution (Python Library)
Resolve
Real Objects
Figure 2. Workflow of the proxy-based dispatch system. 4.3 Security Lattice Based Information Flow Control
implicit downgrading with explicit declassification, executed via a trusted, audited exception process [46]. Explicit Declassification. IntCC maintains a centrally managed registry that maps pre-audited, trusted sanitization functions to specific input/output label transition rules:
To reconcile interactive flexibility with data confidentiality, IntCC employs a fine-grained IFC mechanism [11, 15]. This model formally defines data sensitivity levels within the IntCC-Runtime and strictly constrains information flow and declassification pathways during computation.
𝑓𝑡𝑟𝑢𝑠𝑡𝑒𝑑 : 𝑙𝑖𝑛 → 𝑙𝑜𝑢𝑡
4.3.1 Security Lattice and Taint Propagation. We formalize the policy mechanism’s foundation as a security lattice ⟨ℒ, ⊑, ⊔, ⊓⟩. The label set ℒ comprises five hierarchical levels dictating privacy obligations: H (High-Sensitivity) for strictly isolated raw data; T (Transformation) for data requiring masking or de-identification; A (Aggregation) for intermediate results needing batch averaging to prevent reconstruction; N (Noise Injection) for aggregated data mandating differentially-private noise [17] against differencing attacks; and L (Low-Sensitivity) for public artifacts cleared for export. These logically form a sequential privacy-preserving pipeline governed by a strict linear hierarchy (total order), satisfying 𝐿 ⊑ 𝑁 ⊑ 𝐴 ⊑ 𝑇 ⊑ 𝐻 . To enforce this, the mechanism utilizes ℒ as dynamic taint markers, intercepting all IntCC-Runtime operations for automatic label propagation. We adopt the lattice Join (⊔) operation to handle information flow convergence, while the corresponding Meet (⊓) operation is reserved for policy composition. When an operation takes inputs with labels 𝑙 1, . . . , 𝑙𝑘 , the output label is deterministically computed as the least upper bound of the inputs: 𝑙𝑜𝑢𝑡 =
𝑘 Ä
(where 𝑙𝑜𝑢𝑡 < 𝑙𝑖𝑛 )
For instance, the registry may define: • 𝑓𝑑𝑝 : 𝑁 → 𝐿 (a differential privacy function that discharges the outstanding noise-injection obligation). • 𝑓𝑎𝑛𝑜𝑛𝑦𝑚𝑖𝑧𝑒 : 𝑇 → 𝐿 (a text anonymization function that discharges the outstanding transformation obligation). The mechanism permits the output to bypass the default Join (⊔) propagation rule and acquire a lower sensitivity label only when the Runtime invokes these whitelisted functions. Boundary Audits. The system’s security perimeter is enforced at the taint sink—specifically, the gRPC server response point where data returns to the IntCC-Controller. Crucially, to guarantee absolute confinement, IntCC strictly prohibits all other potential sinks at the container level (e.g., outbound network connections and arbitrary host file I/O are utterly disabled). Consequently, the gRPC interface serves as the singular, heavily audited exit point. The policy mechanism inspects all data attempting to cross this boundary: only data labeled 𝐿 is permitted to be serialized and returned as concrete values. For data at all other sensitivity levels, the system returns only opaque handles. 4.4
𝑙𝑖
𝑖=1
Privacy-Preserving Verifiable Execution Mechanism
To guarantee verifiable runtime compliance, IntCC extends static remote attestation with runtime evidence. While remote attestation authenticates the IntCC-Runtime’s launch state, the verifiable execution mechanism binds post-launch dynamic behavior to this attested runtime, proving policy compliance without exposing plaintext traces.
This design rigorously adheres to the noninterference principle [22], ensuring information only flows from lower to higher sensitivity levels, thereby preventing the inadvertent downgrading of highly sensitive assets during execution. 4.3.2 Explicit Declassification and Boundary Audits. While strict propagation prohibits unauthorized flow from high to low levels, practical usability dictates the need for controlled sensitivity reduction. We address this by replacing
4.4.1 TEE-Bound Trace Commitment. The Runtime records each dynamic interaction as a canonical trace entry containing the operation name, operation class, input and 7
5
output labels, whether an approved declassification function is invoked, whether the result is exported across the boundary, and whether the operation is allowed or denied. Let 𝜏 = (𝑒 1, . . . , 𝑒𝑛 ) denote the execution trace. The Runtime computes 𝐶 = SHA256(ser(𝜏)), and embeds 𝐶 into the hardware attestation report. This binds the trace commitment to a genuine measured Runtime instance, preventing an untrusted data processors or CSP from substituting a different trace, while keeping the plaintext trace private.
Implementation
We implemented the core framework of IntCC in Python, comprising approximately 3,200 Lines of Code (LoC), and deploy it on AMD SEV-SNP [2] using Confidential Containers [13]. We chose Python to natively support dynamic computational graphs, runtime reflection, and dynamic object instantiation, which are prevalent in modern AI and data analytics stacks like PyTorch [41] and Dask [44]. Furthermore, this general design minimizes system extension effort, requiring no more than 200 LoC to support new complex orchestration frameworks (e.g., LLaMA-Factory [58], DuckDB [43]). Implementation of Proxy-based Dispatch System. We realize the shadow object abstraction through dynamic interception of method invocation, attribute access, and object construction on the controller side. Each intercepted operation is marshaled into a Protobuf request containing the operation name, opaque runtime-object handles, and any public literal arguments. On the IntCC-Runtime, a dispatcher resolves handles through a thread-safe object registry, invokes the backend operation, and registers returned objects under fresh handles. To support dynamic Python libraries that rely on runtime inspection and modification, we further implement a remote reflection path that forwards metadata queries to the Runtime and synthesizes matching controllerside stubs on demand, thereby preserving native Python development without source changes. Implementation of IFC mechanism. We implement IFC as a dynamic enforcement layer in the Runtime dispatcher. Each runtime object is stored together with label metadata, and every invocation is checked against a policy map. Concretely, the Runtime deserializes the receiver and arguments, computes an effective input label by joining their tags, resolves the corresponding policy rule, and validates that the invocation is admissible under that rule. The Runtime then either denies the operation or executes it and derives the output label according to the resolved policy. In the current policy map, all non-declassification operators use the default join rule, while trusted operators implement explicit declassification transitions such as anonymization and DP release. We integrate IBM Diffprivlib [23] for differentially private numerical release and Microsoft Presidio with spaCy [24] for text anonymization. A serialization guard at the response boundary then enforces the final export rule: only objects labeled 𝐿 may leave the Runtime as concrete values; higherlabeled outputs remain inside the Runtime and are returned as opaque handles or rejected. Implementation of Verifiable Execution. We implement runtime evidence generation as a service inside the Runtime. Whenever the IFC monitor processes a dynamic operation, it appends the corresponding IFCTraceEntry to the in-memory trace. At the end of task execution, the Runtime
4.4.2 Zero-Knowledge Proof and Verification. Revealing 𝜏 to data owners or auditors would expose the data processor’s proprietary workflow, including API usage patterns, model-tuning strategies, and intermediate control decisions. IntCC therefore treats 𝜏 as a private witness and proves only that it is both committed by 𝐶 and policy-compliant. Concretely, the Runtime generates a zero-knowledge proof 𝜋𝑧𝑘 for the relation: ℛ(𝐶, 𝜏) = 1 ⇐⇒ 𝐶 = SHA256(ser(𝜏)) ∧ CheckIFC(𝜏) = 1, where CheckIFC verifies the IFC rules in Sec. 4.3: ordinary operations must follow join-based taint propagation, declassification may occur only through approved transitions, and only 𝐿-labeled values may be exported. The public statement contains 𝐶 and the policy result, while the sequence of operations remains hidden. Verification has two steps. First, auditors verify the hardware attestation report to check that 𝐶 was produced by an IntCC-Runtime with the expected static measurement. Second, they verify 𝜋𝑧𝑘 against 𝐶 and the agreed policy circuit. A successful proof establishes that the committed trace satisfies the IFC policy without exposing the trace itself. Conversely, data processors can hash their local plaintext trace and compare it with the attested commitment 𝐶 to confirm trace authenticity and detect CSP-side omission or substitution. Thus, the mechanism gives data owners policy-compliance evidence and data processors trace-integrity evidence, without forcing either party to reveal private assets. 4.4.3 Verification Scope. The guarantee provided by this mechanism is intentionally narrower than full-program correctness. It does not prove that an arbitrary interactive workload implemented every unstated data processor’s intent, nor does it certify each low-level computation step in isolation. Such step-by-step proof is unnecessary for the threat model because remote attestation already establishes that the execution trace is generated by the genuine attested Runtime and faithfully reflects its actual execution. The remaining task is therefore not to re-prove each low-level execution step, but to verify that the recorded execution complied with the designated policies. To this end, IntCC provides verifiable runtime compliance. Under this guarantee, data processors can verify that their interactions were processed within the attested Runtime, while data owners or auditors can verify that the recorded execution remained policy compliant. 8
• IntCC (Partitioned TEE): This setup executes the same workflows using the proposed partition architecture. The interactive logic runs in the Controller, while sensitive data and computation reside in the Runtime. All interactions are mediated via the proxy-based dispatch system. We define overhead as the relative increase in end-toend latency induced by the IntCC partitioned architecture. By benchmarking IntCC against a monolithic confidential baseline rather than a native, untrusted host process, we effectively factor out standard TEE virtualization costs. This methodology strictly isolates the performance overhead of interactivity: the latency introduced by gRPC communication, IFC enforcement, zero-knowledge proof generation, and the proxy-based dispatch system. To comprehensively evaluate the performance overhead introduced by IntCC, we structured the experiments across two primary categories: Fine-tuning and data analysis.
canonicalizes the trace, computes a SHA-256 trace commitment 𝐶, and embeds 𝐶 into the attestation report. For privacypreserving verification, we implement a zero-knowledge proof backend based on SP1 [51]: the prover takes the plaintext execution trace as private input, re-encodes each entry into the compact witness format expected by the circuit, and generates a proof that (1) the witness trace hashes to the attested commitment 𝐶, and (2) every trace entry satisfies the IFC rules; the verifier then checks the resulting proof against the public output and the verification key deterministically derived from the guest ELF. Finally, we instantiate IntCC across two representative domains. For interactive fine-tuning, IntCC supports native PyTorch scripts [41], PEFT-based adapter tuning [35], and LLaMA-Factory [58]. For analytical workloads, it supports Dask [44], DuckDB [43], and Pandas [36] over the TPC-H benchmark [55]. Together, these instantiations show that the implementation is not tied to any single stack, but to a reusable Runtime for heterogeneous interactive workloads. Crucially, framework extension overhead is minimal: supporting any of these complex ecosystems requires modifying fewer than 200 LoC within the core orchestration logic.
6
6.1.1 LLM Fine-tuning. Table 1 details the experiment configurations, listing the evaluation dimensions, datasets, and model architectures employed for each dimension. Unless otherwise specified, all tasks utilized Hugging Face PEFT [35] to replicate mainstream industrial usage patterns.
Evaluation
Table 1. Details of datasets and models used in evaluation.
In order to clearly demonstrate the practical impact of IntCC, we design an experiment to empirically evaluate the performance overhead and security enhancement of IntCC. We aim to answer the following research questions: • Performance Efficiency: Does the logical partitioning architecture introduce prohibitive latency compared to standard, monolithic confidential execution? (Sec. 6.1) • Security Guarantees: Can the system reliably prevent data exfiltration and enforce verifiable runtime compliance under adversarial conditions? (Sec. 6.2) Experimental Setup. All experiments were conducted on a server equipped with an AMD EPYC 9334 processor, employing AMD SEV-SNP [2] for hardware-level memory encryption. An NVIDIA H100 GPU was attached via PCIe passthrough. The software stack was hosted on Ubuntu 24.04.3 LTS, utilizing Confidential Containers to orchestrate the confidential execution sandbox [13]. 6.1
Dimension
Dataset
Model
Size
Dataset Characteristics
CNN/DailyMail XSum Alpaca
Pythia
6.9B
Model Scale
CNN/DailyMail
Pythia Family
160M, 1B, 2.8B, 6.9B
Diverse Architectures
CNN/DailyMail
Pythia Llama3 Qwen3
6.9B 8B 8B
We first evaluated the end-to-end performance under diverse dataset characteristics. We selected three widely-used benchmark datasets—CNN/DailyMail [47], XSum [38], and Alpaca [42]—to represent diverse sequence lengths and workload patterns (evaluating 1,000 instances each). Specifically, the processing time incurred marginal overheads of 4.50% for CNN/DailyMail, 4.55% for XSum, and 4.53% for Alpaca. To evaluate the impact of workload complexity, we performed a scaling analysis on 10,000 samples using the Pythia family [6], ranging from 160M to 6.9B. Figure 3(b) demonstrates a clear overhead amortization effect: as the model size increases, the relative performance overhead consistently decreases. This phenomenon occurs because the fixed system costs are effectively amortized over the dominating GPU computation time required by larger models. Specifically, the lightweight Pythia-160M incurred a 2.5% overhead, which dropped to 1.7% for Pythia-1B, and further to 0.9% for
Performance Evaluation
We focus on quantifying the specific overhead imposed by the partitioned architecture of IntCC. To do this rigorously, we define two distinct architectural configurations: • Baseline (Monolithic TEE): This configuration executes the target computational workflows within a single, monolithic Confidential Container. This represents the current state-of-the-art for static confidential workloads, where the training script, data, and computation engine all reside in the same trusted domain with local access. 9
Latency (s)
700 Baseline IntCC 600 500 400 300 200 100 0 CNN/DailyMail XSum
Alpaca
(a) Dataset Characteristics
3500 3000 2500 2000 1500 1000 500 0 160M
1B
2.8B
(b) Model Scale
4000
400
3000
300
2000
200
1000
100
0 Pythia-6.9B Llama-3-8B Qwen3-8B
6.9B
0
(c) Diverse Architectures
Dask
DuckDB
Pandas
(d) Data Analysis
Figure 3. End-to-end performance analysis. Pythia-2.8B. For the largest model, Pythia-6.9B, the overhead stabilized at a mere 0.9%. To verify that this performance consistency extends beyond a single architecture, we further evaluated IntCC across other prominent foundation models, including Llama-3-8B [3] and Qwen3-8B [53]. For each model, we fine-tuned 10,000 randomly selected samples. Figure 3(c) demonstrates that IntCC maintains high efficiency regardless of the specific foundation model. For Llama-3-8B, IntCC was only 0.81% slower than the native baseline, while the overhead for Qwen38B was strictly controlled at 0.75%. In summary, the evaluation across these three categories comprehensively validates the efficiency and robustness of IntCC in fine-tuning applications. Furthermore, the overhead amortization effect observed with increasing model scales confirms that fixed CPU-bound costs such as gRPC and security enforcement are effectively amortized in computeintensive fine-tuning scenarios. Ultimately, the consistent performance across complex, modern model families like Llama-3 [3] and Qwen3 [53] highlights the generality of IntCC, seamlessly supporting large-scale LLM fine-tuning tasks without requiring manual adaptation or incurring additional translation penalties.
applicability and broad generality of IntCC, proving its capacity to securely orchestrate a wide spectrum of real-world computing paradigms.
Table 2. Overhead breakdown across different workloads. Workload Total (s) Dask DuckDB Pandas Finetune
36.78 12.68 20.15 27.70
gRPC
IFC
ZK Proof
Proxy
0.16 (0.4%) 0.16 (1.3%) 0.20 (1.0%) 0.06 (0.2%)
0.00 (0.0%) 0.00 (0.0%) 0.00 (0.0%) 0.07 (0.2%)
8.57 (23.3%) 8.81 (69.5%) 8.72 (43.3%) 9.27 (33.5%)
28.06 (76.3%) 3.71 (29.3%) 11.23 (55.8%) 18.30 (66.1%)
6.1.3 Micro-benchmark. To further investigate the composition of system overhead, we conducted a comprehensive micro-benchmark across both fine-tuning and data analysis workloads, as detailed in Table 2. We selected the Pythia-6.9B task as the representative of LLM fine-tuning. The breakdown reveals that across all categories, gRPC network latency and IFC enforcement contribute negligibly to the total overhead. With the exception of DuckDB, the Proxy System Overhead, stemming from the proxy-based dispatch system, constitutes the largest percentage of the latency: specifically, it accounts for 76.3% in Dask, 66.1% in Pythia-6.9B fine-tuning, and 55.8% in Pandas. We attribute this phenomenon primarily to the proxybased dispatch system execution environment. Data-intensive operations are inherently sensitive to the runtime state; when executing within a gRPC worker thread rather than a clean local main thread, concurrent live RPC dispatching and poller activities significantly elongate the wall-clock coordination path of the underlying task schedulers. Conversely, DuckDB exhibits a unique distribution where zeroknowledge proving accounts for the majority of its overhead (69.5%). This is due to the overhead amortization effect: because DuckDB’s absolute execution penalty is exceptionally low, the cryptographic costs become proportionally dominant. Overall, these results demonstrate that although proxybased dispatching constitutes the primary source of latency, its absolute impact remains highly practical across architecturally diverse workloads.
6.1.2 Data Analysis. To evaluate the efficacy of IntCC on data analysis workloads, we benchmarked IntCC using a TPC-H dataset generated at scale factor 10 (SF=10) [55]. We compared three mainstream data analysis frameworks that have been successfully integrated into IntCC: Dask [44] DuckDB [43], and Pandas [36]. As illustrated in Figure 3(d), the frameworks exhibit distinct performance behaviors under IntCC architecture.Specifically, Pandas incurred the lowest end-to-end overhead of 6.32%, followed by Dask with a moderate overhead of 11.16%, while DuckDB exhibited the highest percentage overhead at 16.56%. In summary, maintaining an end-to-end overhead strictly below 17% represents a highly acceptable and competitive trade-off for practical deployments. More importantly, the demonstrated ability to seamlessly support architecturally diverse and complex frameworks highlights its strong practical 10
Table 3. Efficacy of the PII anonymization.
Security Evaluation
In this section, we evaluate the security effectiveness of IntCC in the context of LLM fine-tuning. The evaluation is twofold: we first validate the resilience of IntCC against adversarial interactions, and subsequently quantify the efficacy of its declassification mechanisms in realistic scenarios.
6.2.1 Defense against Adversarial Interactions. We evaluate the system’s resilience against three primary threat vectors: explicit data theft by data processors, implicit data reconstruction by data processors, and execution pipeline tampering by the untrusted cloud provider. Defense against Direct Data Theft. Malicious data processors may attempt to exfiltrate sensitive assets via I/O channels, including standard output streams, file system updates, or network transmission. To evaluate this, we simulated a malicious data processor and injected interactive commands attempting to invoke prohibited I/O operations on the sensitive data, such as printing plaintext tensors to the terminal, writing variables to .pth files, and sending outbound HTTP requests. IntCC successfully neutralized 100% of these attempts. The defense is attributed to the IFC enforcement, which correctly identified the target objects as High-sensitivity and intercepted the serialization requests at the system boundary. Consequently, the system returned only opaque handles to the controller, preventing any bitstream transmission to the external environment. Defense against Data Reconstruction Attacks. Attackers may attempt to reconstruct private training data by performing gradient inversion [21]. Specifically, the adversary captures the gradient ∇𝑊 generated during a training step and iteratively optimizes a dummy input until its resultant gradient matches the intercepted ∇𝑊 , thereby recovering the original input. We simulated this threat by crafting requests to export exact gradient tensors during the backward pass . IntCC successfully thwarted these attempts. The defense is attributed to the strict enforcement of DP policies on intermediate computation states. The exposed gradients injected by noise lack the precision required for the inversion optimization to converge, rendering reconstruction of the private training samples infeasible. Defense against Pipeline Tampering. A lazy or malicious cloud provider may attempt to compromise the integrity of the training pipeline, for example, by skipping computations to save GPU resources. We simulated this threat by configuring a malicious gRPC proxy to intentionally drop execution requests for the optimizer.step() call during training loop iterations. The evaluation shows that the omission of these specific steps was successfully detected upon audit. This defense capability is attributed to the verifiable execution mechanism. Any deviation from the canonical execution path generates a runtime hash that mismatches
Entity Type
Precision
Recall
F1-Score
EVENT FAC GPE PRODUCT TIME
0.8000 0.9643 0.9620 0.9371 1.0000
0.9836 0.9469 0.9701 0.9982 0.8579
0.8824 0.9555 0.9660 0.9667 0.9235
Overall
0.9465
0.9556
0.9510
17.5
1.0 MSE RMSRE
15.0
0.8
MSE
12.5 10.0
0.6
7.5
0.4
RMSRE
6.2
5.0 0.2
2.5 0.0
0.0 0.1
0.5
1.0
2.0
5.0
Privacy Budget (ε)
(a) Error Metrics vs. Privacy Budget 2500 True Distribution DP Distribution (ε=0.1) DP Distribution (ε=0.5) DP Distribution (ε=1.0)
Token Length
2000 1500 1000
90th pct. True: 1539 DP (ε=1.0): 1556
500 0 0
20
40
60
80
100
Percentile (%)
(b) Token Length Distribution
Figure 4. Evaluation of the privacy-utility trade-off.
the cryptographic commitment anchored in the attestation report, immediately flagging the pipeline tampering. 6.2.2 Efficacy of PII Sanitization Engine. To validate the utility of safe data previews, we evaluated the PII sanitization engine in identifying and masking sensitive entities. We utilized a subset of the OntoNotes 5.0 dataset [25] (5,000 samples). Table 3 presents the evaluation results. The engine showed high efficacy across critical entity categories, achieving an overall F1-Score of 0.9510. These results indicate that the system effectively minimizes the attack surface during data previews while still allowing data processors to comprehend structural features and debug their code. 6.2.3 Privacy-Utility Trade-off in Differential Privacy. We further evaluated the inherent trade-off between privacy guarantees and debugging utility when applying DP mechanisms. This evaluation assesses whether the sanitized 11
statistics provided by IntCC are sufficient for data processors to make correct decisions regarding hyperparameter tuning and training stability. Utility Quantification. As a baseline heuristic for statistical fidelity, we first observed the deviation of noisy gradients from ground-truth values during a Pythia-160M fine-tuning task (using a subset of 2,000 CNN/DailyMail [47] samples, trained over 10 epochs with a batch size of 16, and a chunk size of 100 for step aggregation). As illustrated in Figure 4 (a), the Mean Squared Error (MSE) and Root Mean Squared Relative Error (RMSRE) decline rapidly as the privacy budget (𝜖) increases. At a moderate budget of 𝜖 = 2.0, the MSE stabilizes at 0.7535. While MSE serves primarily as a macrolevel indicator, this bounded error margin suggests that the sanitized gradients successfully preserve the overall distributional trends of the original data. This provides data processors with the necessary directional signals to monitor for training anomalies (e.g., gradient vanishing or exploding). Case Study: Assisted Hyperparameter Configuration. To validate practical utility, we conducted a case study on configuring the max_length hyperparameter, a critical setting that balances computational efficiency against information loss due to truncation. We computed the quantile distribution of token lengths for the private dataset (a subset of 2,000 samples from the CNN/DailyMail dataset) via the DP mechanism. Figure 4 (b) compares the ground-truth distribution with the DP-protected output. The results show that the noisy 90th percentile value (1556) is statistically equivalent to the ground truth (1539). A data processor relying solely on the sanitized view would derive the same optimal configuration (e.g., setting max_length=2048) as one with direct data access. This confirms that IntCC successfully preserves the utility required for interactive development workflows while maintaining strict privacy guarantees.
7
declassification rules are statically registered. We argue that data processors should be empowered to upload custom declassification logic encapsulated in sandboxed formats like WebAssembly. The core challenge then shifts to automated verification. Future iterations of IntCC could adopt a Proof-Carrying Code (PCC) [39] model, where data processors submit code alongside a formal proof demonstrating adherence to security policies. The IntCC-Runtime would be responsible for validating the proof [10]. Limitations against Covert Channels. While IntCC effectively constrains explicit information flows and partial implicit flows, these guarantees do not extend to covert channels (e.g., timing, termination, scheduling, or resource contention) [45]. Achieving progress-sensitive or timingsensitive noninterference requires supplementary interventions, such as rigorous program analysis, deterministic padding, or secure multi-execution [16, 31, 37]. Consequently, IntCC provides robust semantic-level confinement, while comprehensive defense against covert channels remains an orthogonal challenge requiring complementary mechanisms.
8
Related Work
Confining Untrusted Software in TEEs. NestedSGX [56] and vSGX [57] aim to minimize the TCB via guest-level privilege separation, which necessitates extensive modifications to the hypervisor or guest OS. In contrast, IntCC functions as an application-level logical partition, rendering it orthogonal to, and deployable atop, these system-level designs. Another line of research enforces constraints at the language or framework level. For instance, PoBF [9] enforces security objectives through rigorous compile-time static analysis in Rust. Although this approach ensures safe execution paths for deterministic tasks, it relies on the assumption of a predictable computational graph, making it ill-suited for the exploratory, dynamic nature of data science (e.g., Python programs). Unlike these strict static constraints, IntCC targets the flexibility required by interactive analysis. Privacy-Preserving Model Tuning and Data Analytics. Privacy-Preserving fine-tuning systems represent one important line of work. Atlas [50] focuses on auditing the machinelearning lifecycle but does not address secure interactivity against untrusted data processors. DLBox [27] pioneers partition and proxy mechanism for sensitive training data, but its scope is restricted to model training workloads, lacking the versatility of a general-purpose, interactive confidential computing interface. Privacy-preserving data analytics systems like Laputa [32] secure Apache Spark through isolated execution and finegrained policy checking over physical plans. Similarly, PICACHV [11] formally verifies data-use policy enforcement over relational-algebra query plans, demonstrating this approach in Polars. While these systems provide workloadspecific guarantees for analytics engines, IntCC offers a
Discussion
We reflect on the limitations of IntCC and outline future directions. It is worth noting that the framework remains agnostic to specific implementations, enabling components to evolve with technological advancements. Overhead of Zero-Knowledge Proving. A practical limitation of current design is the cost of zero-knowledge proving: while it avoids exposing plaintext audit traces, proof generation is not lightweight enough for interactive confidential workloads. In particular, prover latency and resource consumption can become a bottleneck when audit frequency grows or when the system scales to larger distributed deployments. An important next step is to reduce proofgeneration overhead through more efficient SNARK/STARK backends [5, 20] or hardware acceleration [34]. Programmable Security Policies. The effectiveness of IntCC relies heavily on the expressiveness and correctness of the security policies defined by data owners. Currently, 12
workload-agnostic framework capable of supporting LLM fine-tuning, data analysis, and other interactive tasks.
9
[10] Hongbo Chen, Quan Zhou, Sen Yang, Sixuan Dang, Xing Han, Danfeng Zhang, Fan Zhang, and XiaoFeng Wang. 2025. Agora: Trust Less and Open More in Verification for Confidential Computing. Proc. ACM Program. Lang. 9, OOPSLA2, Article 321 (Oct. 2025), 28 pages. doi:10.1145/3763099 [11] Haobin Hiroki Chen, Hongbo Chen, Mingshen Sun, Chenghong Wang, and XiaoFeng Wang. 2025. PICACHV: Formally Verified Data Use Policy Enforcement for Secure Data Analytics. In 34th USENIX Security Symposium (USENIX Security 25). USENIX Association, Seattle, WA, 5053–5070. https://www.usenix.org/conference/usenixsecurity25/pre sentation/chen-haobin [12] Confidential Containers Community. 2025. Release Notes for v0.17.0. https://github.com/confidential-containers/confidential-containers /blob/main/releases/v0.17.0.md. Accessed: 2026-01-07. [13] Confidential Containers Community. 2026. Confidential Containers: Deploy Cloud Native Applications inside Confidential Enclaves. [EB/OL]. https://confidentialcontainers.org/ Accessed: 2026-05-01. [14] Confidential Computing Consortium. [n. d.]. Confidential Computing: Logging and Debugging – Confidential Computing Consortium. https: //confidentialcomputing.io/2023/06/14/confidential-computinglogging-and-debugging/ [15] Dorothy E. Denning. 1976. A lattice model of secure information flow. Commun. ACM 19, 5 (1976), 236–243. doi:10.1145/360051.360056 [16] Dominique Devriese and Frank Piessens. 2010. Noninterference through secure multi-execution. In 2010 IEEE Symposium on Security and Privacy. IEEE, 109–124. [17] Cynthia Dwork. 2006. Differential privacy. In 33rd International Colloquium on Automata, Languages and Programming (ICALP). Springer, 1–12. doi:10.1007/11787006_1 [18] Emanuele Giuseppe Esposito and Pradipta Banerjee. 2025. How to debug confidential containers securely. Red Hat Developer. https:// developers.redhat.com/articles/2025/07/10/how-debug-confidentialcontainers-securely Accessed: 2026-01-07. [19] Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. 2015. Model inversion attacks that exploit confidence information and basic countermeasures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security (CCS). 1322–1333. [20] Ariel Gabizon, Zachary J. Williamson, and Oana-Madalina Ciobotaru. 2019. PLONK: Permutations over Lagrange-bases for Oecumenical Noninteractive arguments of Knowledge. IACR Cryptol. ePrint Arch. 2019 (2019), 953. https://api.semanticscholar.org/CorpusID:201685538 [21] Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, and Michael Moeller. 2020. Inverting gradients–how easy is it to break privacy in federated learning? Advances in Neural Information Processing Systems (NeurIPS) 33 (2020), 16937–16947. [22] Joseph A Goguen and José Meseguer. 1982. Security policies and security models. In 1982 IEEE Symposium on Security and Privacy. IEEE, 11–20. [23] Naoise Holohan, Stefano Braghin, Pól Mac Aonghusa, and Killian Levacher. 2019. Diffprivlib: the IBM differential privacy library. ArXiv e-prints 1907.02444 [cs.CR] (July 2019). [24] Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength Natural Language Processing in Python. (2020). doi:10.5281/zenodo.1212303 [25] Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2006. OntoNotes: The 90% Solution. In Proceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers. Association for Computational Linguistics, New York City, USA, 57–60. https://aclanthology.org/N06-2015 [26] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR). https://openreview.net/forum?id= nZeVKeeFYf9
Conclusion
In this paper, we introduced the interactive confidential computing paradigm to reconcile static TEE integrity with dynamic workflows. To operationalize this paradigm, we designed IntCC. By logically decoupling control from execution, enforcing lattice-based information flow control, and guaranteeing verifiable runtime compliance, IntCC securely enables human-in-the-loop operations. Evaluations confirmed IntCC supports complex tasks like LLM finetuning and data analysis with practical overhead.
References [1] Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zhang. 2016. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, Savannah, GA, 265–283. https://ww w.usenix.org/system/files/conference/osdi16/osdi16-abadi.pdf [2] Advanced Micro Devices, Inc. 2026. AMD Secure Encrypted Virtualization (SEV). [EB/OL]. https://www.amd.com/en/developer/sev.html Accessed: 2026-05-01. [3] AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/metallama/llama3/blob/main/MODEL_CARD.md [4] Mahmoud Ammar, Ahmed Abdelraoof, and Silviu Vlasceanu. 2024. On Bridging the Gap between Control Flow Integrity and Attestation Schemes. In 33rd USENIX Security Symposium (USENIX Security 24). USENIX Association, Philadelphia, PA, 6633–6650. https://www.usen ix.org/conference/usenixsecurity24/presentation/ammar [5] Eli Ben-Sasson, Iddo Bentov, Yinon Horesh, and Michael Riabzev. 2018. Scalable, transparent, and post-quantum secure computational integrity. IACR Cryptol. ePrint Arch. 2018 (2018), 46. https: //api.semanticscholar.org/CorpusID:44557939 [6] Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 2397–2430. https: //proceedings.mlr.press/v202/biderman23a.html [7] Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan. 2023. Remote ATtestation procedureS (RATS) Architecture. RFC 9334. doi:10.17487/RFC9334 [8] Philip Bues. 2025. Unlocking the Future of Data Security: Confidential Computing as a Strategic Imperative. White Paper US53866125. IDC. Sponsored by Confidential Computing Consortium. [9] Hongbo Chen, Haobin Hiroki Chen, Mingshen Sun, Kang Li, Zhaofeng Chen, and XiaoFeng Wang. 2023. A Verified Confidential Computing as a Service Framework for Privacy Preservation. In 32nd USENIX Security Symposium (USENIX Security 23). USENIX Association, Anaheim, CA, 4733–4750. https://www.usenix.org/conference/usenixsecurity23/pre sentation/chen-hongbo 13
[27] Jaewon Hur, Juheon Yi, Cheolwoo Myung, Sangyun Kim, Youngki Lee, and Byoungyoung Lee. 2025. DLBox: New Model Training Framework for Protecting Training Data. In Network and Distributed System Security (NDSS) Symposium 2025. The Internet Society, San Diego, CA, USA. doi:10.14722/ndss.2025.243001 [28] Intel Corporation. 2026. Intel® Software Guard Extensions (Intel® SGX). [EB/OL]. https://www.intel.com/content/www/us/en/products /docs/accelerator-engines/software-guard-extensions.html Accessed: 2026-05-01. [29] Intel Corporation. 2026. Intel® Trust Domain Extensions (Intel® TDX). [EB/OL]. https://www.intel.com/content/www/us/en/developer/tool s/trust-domain-extensions/overview.html Accessed: 2026-05-01. [30] Yuxuan Jiang, Ziming Zhou, Boyu Xu, Beijie Liu, Runhui Xu, and Peng Huang. 2025. Training with confidence: catching silent errors in deep learning training with automated proactive checks. In Proceedings of the 19th USENIX Conference on Operating Systems Design and Implementation (Boston, MA, USA) (OSDI ’25). USENIX Association, USA, Article 18, 17 pages. [31] Vineeth Kashyap, Ben Wiedermann, and Ben Hardekopf. 2011. Timingand Termination-Sensitive Secure Information Flow: Exploring a New Approach. In 2011 IEEE Symposium on Security and Privacy. 413–428. doi:10.1109/SP.2011.19 [32] Byeongwook Kim, Jaewon Hur, Adil Ahmad, and Byoungyoung Lee. 2025. Secure Data Analytics in Apache Spark with Fine-grained Policy Enforcement and Isolated Execution. In Proceedings 2025 Network and Distributed System Security Symposium. Internet Society, San Diego, CA, USA. doi:10.14722/ndss.2025.242496 [33] Peng Li and Steve Zdancewic. 2005. Downgrading policies and relaxed noninterference. SIGPLAN Not. 40, 1 (Jan. 2005), 158–170. doi:10.114 5/1047659.1040319 [34] Tao Lu, Chengkun Wei, Ruijing Yu, Yi Chen, L. xilinx Wang, Chaochao Chen, Zeke Wang, and Wenzhi Chen. 2023. cuZK: Accelerating ZeroKnowledge Proof with A Faster Parallel Multi-Scalar Multiplication Algorithm on GPUs. IACR Cryptol. ePrint Arch. 2022 (2023), 1321. https://api.semanticscholar.org/CorpusID:252734156 [35] Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, Benjamin Bossan, and Marian Tietz. 2022. PEFT: State-ofthe-art Parameter-Efficient Fine-Tuning methods. https://github.com /huggingface/peft. [36] Wes McKinney. 2010. Data Structures for Statistical Computing in Python. SciPy 2010 (2010). doi:10.25080/Majora-92bf1922-00a [37] Scott Moore, Aslan Askarov, and Stephen Chong. 2012. Precise enforcement of progress-sensitive security. In Proceedings of the 2012 ACM Conference on Computer and Communications Security (Raleigh, North Carolina, USA) (CCS ’12). Association for Computing Machinery, New York, NY, USA, 881–893. doi:10.1145/2382196.2382289 [38] Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, Brussels, Belgium, 1797–1807. doi:10.18653/v1/D18-1206 [39] George C. Necula. 1997. Proof-carrying code. In Proceedings of the 24th ACM SIGPLAN-SIGACT symposium on Principles of programming languages (New York, NY, USA, 1997-01-01) (POPL ’97). Association for Computing Machinery, 106–119. doi:10.1145/263699.263712 [40] Olga Ohrimenko, Felix Schuster, Cédric Fournet, Aastha Mehta, Sebastian Nowozin, Kapil Vaswani, and Manuel Costa. 2016. Oblivious multi-party machine learning on trusted processors. In Proceedings of the 25th USENIX Conference on Security Symposium (Austin, TX, USA) (SEC’16). USENIX Association, USA, 619–636. [41] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein,
Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 32. Curran Associates, Inc., 8024–8035. http://papers.neurips.cc/paper/9015pytorch-an-imperative-style-high-performance-deep-learninglibrary.pdf [42] Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction Tuning with GPT-4. arXiv:2304.03277 [cs] doi:10.48550/arXiv.2304.03277 [43] Mark Raasveldt and Hannes Mühleisen. 2019. DuckDB: an Embeddable Analytical Database. In Proceedings of the 2019 International Conference on Management of Data (Amsterdam, Netherlands) (SIGMOD ’19). Association for Computing Machinery, New York, NY, USA, 1981–1984. doi:10.1145/3299869.3320212 [44] Matthew Rocklin. 2015. Dask: Parallel Computation with Blocked Algorithms and Task Scheduling. (2015). doi:10.25080/Majora7b98e3ed-013 [45] Andrei Sabelfeld and Andrew C. Myers. 2003. Language-based information-flow security. IEEE Journal on Selected Areas in Communications 21, 1 (2003), 5–19. doi:10.1109/JSAC.2002.806121 [46] A. Sabelfeld and D. Sands. 2005. Dimensions and principles of declassification. In 18th IEEE Computer Security Foundations Workshop (CSFW’05). 255–269. doi:10.1109/CSFW.2005.15 [47] Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vancouver, Canada, 1073–1083. doi:10.18653/v1/P17-1099 [48] Zhanyu Sha, Carlton Shepherd, Amir Rafi, and Konstantinos Markantonakis. 2024. Control-Flow Attestation: Concepts, Solutions, and Open Challenges. arXiv:2408.06304 [cs.CR] doi:10.48550/arXiv.2408.06 304 [49] Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov. 2017. Machine Learning Models that Remember Too Much. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (Dallas, Texas, USA) (CCS ’17). Association for Computing Machinery, New York, NY, USA, 587–601. doi:10.1145/3133956.3134077 [50] Marcin Spoczynski, Marcela S. Melara, and Sebastian Szyller. 2025. Atlas: A Framework for ML Lifecycle Provenance & Transparency. In 2025 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW). 448–461. doi:10.1109/EuroSPW67616.2025.00058 [51] Succinct Labs. 2024. SP1: A performant, 100% open-source, modular zkVM. https://github.com/succinctlabs/sp1. [52] Zhichuang Sun, Bo Feng, Long Lu, and Somesh Jha. 2020. OAT: Attesting Operation Integrity of Embedded Devices. In 2020 IEEE Symposium on Security and Privacy (SP). 1433–1449. doi:10.1109/SP40000.2020.000 42 [53] Qwen Team. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https://arxiv.org/abs/2505.09388 [54] The Kata Containers Community. 2026. The speed of containers, the security of VMs. [EB/OL]. https://katacontainers.io/ Accessed: 2026-06-01. [55] Transaction Processing Performance Council (TPC). [n. d.]. The TPC-H Benchmark. https://www.tpc.org/tpch. Accessed: 2026-05-06. [56] Wenhao Wang, Linke Song, Benshan Mei, Shuang Liu, Shijun Zhao, Shoumeng Yan, XiaoFeng Wang, Dan Meng, and Rui Hou. 2025. The Road to Trust: Building Enclaves within Confidential VMs. In Network and Distributed System Security (NDSS) Symposium. The Internet Society, San Diego, CA, USA. doi:10.14722/ndss.2025.240385 [57] Shixuan Zhao, Mengyuan Li, Yinqian Zhangyz, and Zhiqiang Lin. 2022. vSGX: Virtualizing SGX Enclaves on AMD SEV. In 2022 IEEE Symposium on Security and Privacy (SP) (2022-05). 321–336. doi:10.1 14
109/SP46214.2022.9833694 [58] Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (Bangkok, Thailand, 2024-08), Yixin Cao, Yang Feng, and Deyi
Xiong (Eds.). Association for Computational Linguistics, 400–410. doi:10.18653/v1/2024.acl-demos.38 [59] Ligeng Zhu, Zhijian Liu, and Song Han. 2019. Deep leakage from gradients. Curran Associates Inc., Red Hook, NY, USA.
15