arXiv:2609.11411v1 [cs.CR] 10 Sep 2026
“They don’t care about this”: A Systematic Study of TEE Build Reproducibility in the Wild Annika Wilde
Marco Gutfleisch
Felix Reichmann
Anirban Chakraborty
Ruhr University Bochum Bochum, Germany [email protected]
LMU Munich Munich, Germany [email protected]
Ruhr University Bochum Bochum, Germany [email protected]
Max Planck Institute for Security and Privacy Bochum, Germany [email protected]
Yuval Yarom
M. Angela Sasse
Ghassan Karame
Ruhr University Bochum Bochum, Germany [email protected]
Ruhr University Bochum Bochum, Germany [email protected]
Ruhr University Bochum Bochum, Germany [email protected]
Abstract
in the cloud. At its core, confidential computing relies on Trusted Execution Environments (TEEs), specialized hardware components within modern CPUs that create isolated regions of execution, often called enclaves or confidential virtual machines (CVMs). These environments safeguard both the data being processed and the integrity of the code performing the computation, offering strong assurances even when the surrounding system or cloud provider cannot be fully trusted. To provide these security guarantees, TEEs minimize the Trusted Computing Base (TCB)—the set of components that must be trusted to remain secure—to contain only the CPU, the TEE firmware, and the code running inside the enclave. Users can verify the integrity of this TCB and the code running inside the enclave through remote attestation (details in Section 2.1). As such, they are expected to entrust their sensitive data to an enclave if and only if its attested measurement matches a known, trusted reference. The TEE ecosystem has enabled a wide range of applications, including secure messaging platforms, such as Signal [49], privacypreserving databases [12], distributed machine learning [37], and blockchain systems [47, 58], that leverage the ability to delegate computation to remote infrastructure while maintaining strong guarantees of data confidentiality and integrity. For example, consider a privacy-preserving machine learning (PPML) service [48] that processes anonymized medical data using open-source models hosted on a remote server. Multiple users, such as hospitals, diagnostic centers, insurance companies, etc., can securely contribute their datasets to train a shared model for tasks like early disease detection or patient health assessment. The security assurances provided by TEEs ensure that data is protected, even from strong adversaries, such as compromised or corrupt cloud providers. Although remote attestation guarantees that the application running in the enclave is signed and authenticated by the developer, it does not provide any assurances about the trustworthiness of the application code itself. A key challenge in the adoption of TEEs is ensuring that the code executed inside the enclave is trustworthy and free from malicious backdoors or vulnerabilities that could leak sensitive data to the application developer. In essence, TEEs aim to shift trust from developers to code: rather than relying on developers’ integrity, users can inspect and verify the enclave code to confirm that it meets required security and privacy guarantees.
Trusted Execution Environments (TEEs) have become a cornerstone of modern cloud computing, providing strong confidentiality and integrity guarantees for both code and data. A critical component of this trust model is remote attestation, which enables external entities to verify the authenticity and integrity of code executing within a TEE through cryptographic measurements. However, the effectiveness of remote attestation fundamentally depends on the verifier’s ability to trace the reported measurement back to the original source code—a property that can only be guaranteed through reproducible builds. In this paper, we investigate the reproducibility of TEE builds through a technical analysis of 115 TEE deployments. Our analysis spans popular TEEs such as Intel SGX, Intel TDX, and AMD SEV, and reveals that a striking 91% of those deployments were not reproducible, with 80% failing to provide both source code and a reference build, the two essential prerequisites for reproducibility. To explore the root causes, we contacted the maintainers of 50 SGX projects and managed to recruit 12 developers from industry and academia for interviews. Only one of our participants reported that reproducibility is a priority during development, effectively confirming our technical findings. Beyond technical barriers (e.g., timestamps included in the binary) that can be readily addressed, we identify broader ecosystem-level challenges, such as the lack of control over the build environment in projects involving multiple stakeholders. We argue that achieving reproducibility in TEEs requires a holistic development approach that extends beyond individual developers and calls for stronger commitments—rather than treating TEEs as a “security badge”.
1
Introduction
The rapid adoption of cloud computing has increased the demand for distributed and decentralized applications, where infrastructure resources—including hardware, platforms, and software—are owned and operated by different entities. This paradigm shift in the IT landscape has also intensified concerns about the confidentiality and integrity of sensitive data and computations performed on remote systems. Confidential computing has emerged as a promising approach to address these challenges and ensure data protection 1
</>
2
remained unreproducible. We further demonstrate that these challenges are not specific to any particular TEE technology, persisting across SGX, TDX, and SEV-SNP alike; however, we observe that reproducibility rates vary depending on the primary build system used, with notable differences observed across systems such as prebuilt Docker images and dependency-pinning tools such as Nix [6] (cf. Section 3). These results are echoed in our interview study, where 10 of the 12 developers reported that reproducibility is neither prioritized nor considered during development (cf. Section 4). We analyze the reasons behind this lack of reproducibility and identify several issues that prevent users from reproducing the enclave binary, as well as challenges faced by developers. To our surprise, many of the technical obstacles, such as embedded timestamps or non-deterministic file system ordering, seem to be wellknown in the broader, traditional reproducible-builds community. Other reasons include an occasional inclusion of secret material within the enclaves, which cannot be reproduced by anyone outside the project. Beyond these technical factors, our qualitative study uncovers several ecosystem-level and procedural challenges. In particular, the traditional software development model process, where components are built in isolation, is incongruent with the requirements of TEE development. Since every component in the build pipeline can influence the enclave’s final measurement, developers must adopt a broader, system-level view of the build process. This mismatch undermines one of the core security mechanisms of TEEs: if an enclave’s measurement cannot be reproduced, remote attestation loses its true meaning (cf. Section 5). We conclude that ensuring build reproducibility in TEE applications requires more than simple technical solutions—it demands a rethinking of development practices. We therefore argue for a TEE-specific development process that explicitly prioritizes reproducibility and enables developers to regain control over their build environments to ultimately seize the value of cloud computing (cf. Section 6).
0110 1011
3 reproducbile builds
1 remote attestation
Figure 1: Interplay between remote attestation and reproducible builds. Remote attestation links a service to the specific deployed binary (Step 1 ), while source code inspection (Step 2 ) and reproducible builds (Step 3 ) establish trust in the source code and connect it to the attested binary. Returning to the example of the PPML service, before sharing private patient information, users could examine the open-source enclave code to ensure it does not intentionally or inadvertently expose sensitive data. Similarly, for other sensitive applications, such as cryptocurrencies, smart contracts, or electronic-voting systems, it is essential to validate the trustworthiness of the code before entrusting the service with sensitive information. Once users are satisfied with the source code, they must be able to confirm that the deployed binary was genuinely built from the vetted source code. This requires linking the enclave’s measurement to the source code, typically by recompiling the source locally and comparing the resulting binary to the attested one.1 For this comparison to succeed, the locally compiled binary must be bit-by-bit identical to the deployed version. Achieving such identity depends on reproducible builds—ensuring that identical inputs and environments always yield identical outputs. Figure 1 shows such a verification process, depicting the interplay between remote attestation and reproducible builds. The example above highlights that the reproducibility of the build process is a critical component of TEE security. However, it remains unclear whether state-of-the-art and widely deployed TEEs actually support reproducible builds. Furthermore, there is limited understanding of whether TEE developers themselves recognize the importance of reproducibility in ensuring software integrity and trust. In this paper, we present the first comprehensive study of reproducibility in TEE builds, combining two complementary approaches: 1 an empirical analysis of 115 TEE projects drawn from [19, 29, 36], covering widely used TEEs including Intel SGX, Intel TDX, and AMD SEV-SNP; and 2 a qualitative interview study with 12 SGX developers. Our findings paint a stark picture: reproducible builds remain the exception rather than the norm across today’s TEE ecosystem. Of the 115 projects examined, 94% lack the documentation or tooling needed to reproduce enclave measurements, and 77% fail to provide both source code and a reference build—two fundamental prerequisites for reproducibility. Even when developers supplied additional information directly, 91% of the projects
2 Background and Related Work 2.1 Trusted Execution Environments Trusted Execution Environments (TEEs) use hardware-based mechanisms to protect memory regions at runtime, enabling secure execution of sensitive code in isolated environments. Depending on their granularity, TEEs can isolate either a single process—known as process-based TEEs [2, 21, 28]—or an entire virtual machine (VM)— referred to as VM-based TEEs [1, 23]. Examples include Intel Software Guard Extensions (SGX) [21] and Arm TrustZone [2] for process-based TEEs, and Intel Trust Domain Extensions (TDX) [23] and AMD Secure Encrypted Virtualization (SEV) Secure Nested Paging (SNP) [1] for VM-based TEEs. In this work, we generally refer to a TEE instance initialized with user code as an enclave. The TEE trust model assumes that only the CPU, TEE firmware, and the enclave’s specific code form the Trusted Computing Base (TCB). In other words, all privileged system software, such as the operating system and hypervisor, is excluded from the TCB and treated as untrusted, since it could potentially be compromised. Each enclave is characterized by two cryptographic identifiers. The first, referred to as measurement, is a hash of the enclave’s initial code and data, uniquely defining its content. The second, called signer identity in Intel TEEs, is derived from the developer’s
1 Note that the process of auditing source code and compiling it to derive a trusted
reference measurement is inherently time-consuming and requires technical expertise that exceeds the capabilities of most end users. However, such verification need not be performed by every individual. A subset of technically skilled and security-conscious users can conduct these steps and publish verified reference measurements for the community. In the PPML example, this collaborative verification model allows ordinary users to rely on community-vetted attestations, establishing collective trust in the service’s confidentiality guarantees. 2
2.3
public key and uniquely identifies the entity that signed the enclave. To verify an enclave’s authenticity, TEEs such as Intel SGX, Intel TDX, and AMD SEV provide a remote attestation mechanism. During attestation, the TEE firmware generates a signed certificate of the enclave’s measurement using a hardware-specific key. A remote verifier can validate the enclave’s integrity by checking this signature through an attestation service (e.g., one operated by the hardware vendor) and comparing the reported measurement to a known trusted reference. In contrast to Intel SGX and AMD SEV-SNP, which report a single measurement in the attestation, Intel TDX exposes five measurement registers; four of them can be updated at runtime. Successful verification confirms that the enclave is executing the expected code within a genuine TEE, thereby upholding the TEE’s security guarantees.
2.2
Related Work
To the best of our knowledge, no prior work has systematically examined build reproducibility within TEEs. That said, several studies have explored reproducibility in traditional software contexts. Lamb and Zacchiroli [27] provide a comprehensive overview of sources of non-determinism in build processes and discuss the technical barriers to achieving reproducible builds. Similarly, de Carné de Carnavalet and Mannan [8] analyze 16 versions of the opensource encryption tool TrueCrypt and identify multiple causes of non-determinism that hinder verifiable builds. Fourné et al. [17] investigate the motivations and challenges faced by open-source developers pursuing reproducible builds through interviews with 24 contributors of the Reproducible-Builds. org project. Their findings highlight both technical difficulties, as noted by Lamb and Zacchiroli [27], and broader ecosystem challenges like limited awareness and poor documentation. In contrast, our study focuses on SGX developers to assess their awareness, understanding, and attitudes toward reproducible builds in TEEs. Moreover, we target a community that is inherently interested in reproducible builds, which is essential for the attestation process. In the context of TEEs, Hugenroth et al. [22] propose attestable builds, where the build process and input verification occur within a TEE, which then produces a cryptographic attestation of the resulting binary’s hash. Similarly, Delignat-Lavaud et al. [10] introduce a Code Transparency Service for tracking code provenance and holding developers accountable for released binaries. Both approaches assume developers’ interest in achieving reproducible builds. In contrast, our work examines the actual awareness, motivation, and challenges faced by SGX developers, offering insights and recommendations to foster reproducibility within TEE ecosystems.
Reproducible Builds
Reproducible (or deterministic) build is a software development practice that ensures a build process consistently produces identical output across different environments. In other words, when a build is reproducible, the same source code, compiler version, and build configuration always yield the same output, regardless of where or when the build is performed. Build reproducibility can be categorized into two types: bit-by-bit reproducibility and semantic reproducibility. Bit-by-bit reproducibility guarantees that two independently produced binaries are identical at the byte level, while semantic reproducibility is less strict, requiring only identical behavior, even if the binary representations differ. Reproducibility is critical for verifying the integrity and authenticity of software binaries. It allows developers and users to independently rebuild software and confirm that the resulting binary matches the expected output. This is particularly important for open-source projects, where transparency and verifiability are key—anyone can inspect the source code, reproduce the build, and confirm its trustworthiness. By supporting reproducible builds, developers strengthen transparency and trust in their software.
3
Reproducible TEE Builds in the Wild
We now evaluate the reproducibility of TEE builds. To this end, we conduct a systematic analysis of TEE-based open-source applications, focusing on their ability to deterministically reproduce enclave measurements.
Technical Challenges in Reproducing Builds. Traditional build processes depend on numerous environmental and non-deter-ministic factors, such as the operating system, file system, and even timestamps. These sources of variability introduce non-determinism that must be mitigated to achieve reproducibility. Lamb and Zacchiroli [27] identify several common sources: • Timestamps: Macros like __DATE__ often embed build timestamps into the binary, creating variability between two builds produced at different points in time. • Build paths: Macros like __FILE__ record absolute file paths in binaries, which vary between build environments unless a fixed directory structure is enforced. • File system ordering: File systems may order directory contents differently, causing inconsistent file processing during the build. • Archive metadata: Archive metadata, such as timestamps and ownership, can differ between builds. • Compiler randomness: Some compilers generate random identifiers or values during compilation, introducing non-determinism. • Uninitialized memory: In low-level languages (e.g., C/C++), uninitialized memory can cause inconsistent build behavior.
3.1
Methodology
Project selection. In our analysis, we selected candidate projects independently from two established sources to minimize selection bias: (i) the marketplaces from Microsoft Azure [36] and Google Cloud Platform (GCP) [19], and (ii) the GitHub repository Awesome SGX Open Source [29] (which, to the best of our knowledge, provides the most comprehensive list of publicly available TEE applications). In the former, we identified candidate projects by searching for the terms TEE, Trusted Execution Environment, SGX, TDX, and SEV-SNP, discarding products that appeared solely due to loose search policies rather than genuine support or implementation of these technologies. Our initial candidate set consisted of 232 applications from [29] and 35 products from the marketplaces— specifically, seven products from the GCP marketplace [19] and 33 products from the Azure marketplace [36], with five products available on both platforms. We filtered this candidate set by applying the following three criteria (C1–C3): C1: Suitable projects for our study must run within a TEE or contribute to the binary running in the TEE. 3
Table 1: Overview of dataset filtering criteria (C1–C3). Each row reports the number of projects that met each criterion individually. The last row shows how many met all of them. Stage Initial candidate set TEE-integrated (C1) Directly purchasable (C2) Active (C3) Final dataset (C1–C3)
Awesome SGX 232 203 NA 122 95
Finally, we compared the measurement of our compiled artifact with the corresponding reference value. When discrepancies occurred, we engaged with developers to identify possible causes of variation and obtain guidance on achieving reproducibility. For example, a measurement mismatch occurred when reproducing CCF’s SGX enclave binary despite successful compilation. We emailed the repository’s contact with details of the release, build procedure, and environment, requesting clarification on possible reproducibility issues. Similar messages were sent to other teams—e.g., to Phala [42] via email and Ternoa [58] through a Discord Ticket.
Azure & Google cloud marketplaces 35 27 29 35 21
3.2
Results
Table 2 provides an overview of the selected projects, while Table 3 summarizes our analysis of the 27 TEE applications that provide both source code and a reference measurement. We evaluated each application in two scenarios: an unguided scenario, where we attempted to build and reproduce the enclave without explicit developer assistance, and a guided scenario, in which we engaged with the developers to achieve a successful and/or reproducible enclave build if the unguided attempt failed. Each entry in Table 3 reports (for both scenarios) whether the application was successfully compiled, whether the enclave hash was reproduced, and whether the development team responded to our inquiries. We successfully compiled 25 projects independently; the remaining two could not be compiled due to dependencies being unavailable. Among the evaluated projects, ten provide sufficient information to reproduce the enclave measurement: seven rely solely on public documentation, while three require non-public information from the developers. In other words, 91.3% of applications in our dataset are not reproducible. Specifically, 77% fail to provide both source code and a reference measurement. More surprisingly, among those projects that do provide a reference measurement, only 37% are reproducible.
C2: Marketplace products must be directly purchasable on the platform (i.e., offers requiring prior vendor interaction are excluded). This ensures that relevant technical details are accessible without additional communication. C3: We are interested in active projects (which we define as having a commit on the default branch after January 2020, not being archived by the end of December 2025, and not being marked as unsuitable for production use). After applying these criteria (cf. Table 1), our final dataset comprises 115 TEE applications: 95 from [29] and 21 from the marketplaces, with one application present in both sets. Six applications appear multiple times among the 115 due to multi-TEE support: three support both SGX and SEV-SNP, two support both TDX and SEV-SNP, and one supports all three TEEs. In our analysis, we treat the same application on different TEEs as separate observations, since reproducibility can vary across deployments, even among projects from the same vendor, as shown by the cases of CCF and the Secret Network (cf. Section 3.3). We describe our analysis process below.
Academic vs. commercial projects. Among the analyzed applications, 27% are academic artifacts, all but one based on SGX (cf. Table 2). Reproducibility is 0% for academic projects and 12% for non-academic ones. Notably, only a single academic project provides a reference measurement, a prerequisite for reproducibility. This suggests that, as expected, reproducibility is more naturally incentivized in production settings, while academic artifacts are often prototypes. Given the dataset’s imbalance (only 8.7% reproducible) and skew toward SGX (82.6%), we restrict our analysis in what follows to the 84 non-academic projects, where reproducibility is more relevant.
Analysis workflow. To evaluate reproducibility, our analysis is necessarily limited to open-source applications, as access to the source code is required to rebuild the enclave. We therefore consider only applications that provide both source code (with build instructions) and a reference measurement, either as a concrete measurement value or as a release binary. Out of the 115 TEE applications in our dataset, 88 do not meet this criterion and are therefore unreproducible. Conversely, 27 applications fulfill this requirement (cf. Table 2). We assess the reproducibility for each of these projects as follows. We first attempted to compile the latest release of each selected project using the build instructions provided in its documentation. If compilation failed, we contacted the corresponding developers for assistance via the official listed communication channels (e.g., email, Discord, or GitHub issues), following this order of preference. For each successfully compiled application—either independently or with developer support—we analyzed the reproducibility of the resulting enclave build. Specifically, we extracted the enclave measurement from the compiled artifacts using the tools provided by the respective SDK. For the official Intel SGX SDK, for example, we employed the sgx_sign utility to retrieve enclave metadata, including the measurement. When an explicit reference measurement was unavailable, we derived it from the provided release binary or the official Docker image.
Is there a correlation between the TEE technology (i.e., SGX, TDX, SEV-SNP) and reproducibility? Among the filtered projects, Intel SGX leads with 65 commercial implementations, followed by AMD SEV-SNP with 14 and Intel TDX with three. This distribution reflects the relative maturity of each technology: Intel TDX and AMD SEV-SNP are newer additions to the landscape (TDX being the most recent) and are still gaining market traction, a trend that is equally reflected in academic projects as well. We observe that reproducibility varies notably across these platforms: none of the TDX (most recent TEE) projects are reproducible, only 9.2% of SGX projects (oldest and most widely studied TEE) are, whereas 28.6% of SEV-SNP (released somewhere between the two) projects are reproducible. To assess whether this variation reflects a 4
Table 2: Overview of the selected projects. TEE Technology Dataset Total projects Project type Commercial Academic Primary build system Unspecified Docker file Docker image Yocto Bazel Nix Reference availability Neither source code nor reference Source code only Reference only Source code + reference Reproducibility Unreproducible Reproducible guided scenario Reproducible unguided scenario
SGX
SEV-SNP
TDX
Unspecified
95
15
3
2
65 30
14 1
3 0
2 0
74 11 8 0 1 1
8 3 1 1 1 1
1 0 0 1 1 0
2 0 0 0 0 0
14 57 2 22
5 2 4 4
1 0 1 1
2 0 0 0
89 3 3
11 0 4
3 0 0
2 0 0
Table 3: Summary of our analysis of 27 TEE applications. The project source is denoted by 𝑎 ([29]) and 𝑚 ([19, 36]). We report the reproducibility of the enclave hash in an unguided scenario (independently) and a guided scenario (with developer assistance). In the guided case, “–” indicates success without assistance; we additionally report whether developers responded. ★ indicates inherently unreproducible projects. Unguided Guided Built Reproduced Replied Built Reproduced Intel SGX CCF [31] 𝑚 ✓ ✗ ✓ – ✓ 𝑎 Edgeless RT [53] ✓ ✗ ✗ – ✗ 𝑎 EGo [54] ✓ ✓ – – – 𝑎 enclave-vrf [30] ✓ ✗ ✗ – ✗ Google Asylo [18] 𝑎 ✓ ✗ ✗ – ✗ Gramine [43] 𝑎 ✓ ✗ ✓ – ✓ Marblerun [55] 𝑎 ✓ ✗ ✓ – ✓ 𝑎 MobileCoin [15] ✓ ✗ ✗ – ✗ 𝑚 MongoDB-SGX [12] ★ ✓ ✗ ✓ – ✗ 𝑎 Mystikos [32] ✓ ✗ ✗ – ✗ 𝑎 Nginx-SGX [13] ★ ✓ ✗ ✓ – ✗ 𝑎 Oasis Sapphire [14] ✓ ✗ ✗ – ✗ 𝑎 Occulum [56] ✓ ✗ ✗ – ✗ Phala [42] 𝑎 ✓ ✗ ✗ – ✗ Safeheron Arweave [45] 𝑎 ✗ ✗ ✓ ✗ ✗ Linux SGX [46] 𝑎 ✓ ✓ – – – SGXSSE [60] 𝑎 ✗ ✗ ✓ ✗ ✗ 𝑎𝑚 Secret Network [25] ★ ✓ ✗ ✓ – ✗ Signal CDSI [50] 𝑎 ✓ ✓ – – – 𝑎 SnowHaze zka-sgx [51] ✓ ✗ ✗ – ✗ 𝑎 Teaclave SDK [59] ✓ ✗ ✗ – ✗ Ternoa enclaves [57] 𝑎 ✓ ✗ ✓ – ✗ AMD SEV-SNP 𝑎 CCF [33] ✓ ✓ – – – Cosmian KMS [7] 𝑚 ✓ ✓ – – – Secret Network [26] 𝑎 ✓ ✓ – – – TamaGo [3] 𝑎 ✓ ✓ – – – Intel TDX Secret Network [26] 𝑎 ✓ ✗ ✗ – ✗ Project
genuine association, we test the null hypothesis that reproducibility rates are equal across TEE technologies. Fisher’s exact test with the Freeman–Halton extension yields a p-value of 0.099, indicating no statistically significant relationship between TEE technology and reproducibility. Does the choice of the build system correlate with reproducibility? We classify non-academic projects according to their primary build system, distinguishing four categories: dedicated reproducibility-oriented tools such as Bazel, Nix, or Yocto (seven projects); prebuilt Docker images (nine projects); Dockerfile-based builds (eleven projects); and unspecified systems (57 projects). The highest reproducibility rate, 42.9%, belongs to the first category, consistent with these tools’ native support for dependency pinning and hermetic build environments. Projects relying on prebuilt Docker images or Dockerfiles achieve lower yet comparable rates of 22.2% and 18.2%, respectively, while only 5.3% of projects with unspecified primary build systems are reproducible. Fisher’s exact test with the Freeman–Halton extension decisively rejects the null hypothesis of equal reproducibility rates across build systems (p = 0.012), confirming a statistically significant association between build system choice and reproducibility.
the identified reproducibility blockers in selected case studies below and summarize them in Table 4. Secret Network. Secret Network [25] is a blockchain platform that leverages TEEs for confidential smart contract execution, using Intel SGX to protect both execution integrity and contract state. The network promises that contracts are “private-by-default” and “can’t be viewed by others” [47]. The Secret Network node is offered as an Azure application for one-click deployment [38] and its source code is publicly available, allowing users to inspect it, e.g., to ensure the absence of vulnerabilities or backdoors that may allow the developers to leak secret information. However, the Secret Network source repository notes that production systems must use the enclave binary from the latest release, as builds are not reproducible since they are signed with an undisclosed key. We built version v1.14.0 in a Docker container and, as expected, failed to reproduce the measurement. The developers confirmed that builds are non-reproducible and that the production enclave is indeed signed. As the inclusion of the signature is a design feature, we conclude that the Secret Network is non-reproducible by design. In other words, the lack of reproducibility prevents users from independently verifying that the production enclave runs the expected code and is free of vulnerabilities or backdoors. This limitation has important implications: although the source code is publicly available, users cannot verify that the deployed
Main Takeaways: Our analysis reveals that the challenges of achieving reproducibility (approximately 91% of all selected TEE projects were not reproducible) are orthogonal to the underlying TEE technology (i.e., SGX, TDX, SEV-SNP). Build system choice, by contrast, emerges as a strong predictor; in fact, dependency-pinning tools such as Nix and Yocto attain the highest reproducibility rate in our dataset (42.9%).
3.3
Case Studies
During our analysis, we identified 20 unreproducible projects in the unguided scenario. We contacted the developers of the corresponding projects and recovered the root cause in nine cases. We discuss 5
enclave corresponds to the audited code. As a result, trust in the system ultimately depends on the developers, weakening the intended transparency guarantees. Despite this limitation, the platform is actively used in practice. While the SGX-based Secret Network node (also available on the Azure marketplace) lacks reproducibility, a newer version leveraging AMD SEV-SNP and Intel TDX now offers reproducibility support, including build and verification instructions [39], indicating a stronger emphasis on reproducibility in recent TEE deployments. Following these instructions, we successfully reproduced the SEV-SNP measurement for Secret VM v0.0.25. For TDX, reproducibility was only partially achieved. Recall that TDX exposes four additional measurement registers alongside the initial VM measurement at boot (cf. Section 2.1). While the initial VM measurement and one of the additional registers matched the values retrieved from the release, the remaining three did not. This discrepancy may be due to incomplete instructions for retrieving the measurements, which required modifications to the command invoking the suggested measurement tool. However, since the same command was used for both the released and locally built artifacts, this explanation is unlikely. At the time of writing, the issue remains unresolved, as we did not receive a response from the developers. Similar patterns appear in other SGX-based systems. For instance, based on communication with their developers, neither NginxSGX [13] nor MongoDB-SGX [12] is reproducible. Both are enclavized versions of NGINX and MongoDB built with Gramine [43], and both take deliberate shortcuts that favor ease of deployment over reproducibility. Rather than injecting certificates after successful remote attestation, they generate self-signed certificates dynamically at build time. Because these certificates differ from build to build yet are covered by the enclave measurement, the measurement itself cannot be reproduced.
For Ternoa, building enclave v0.4.5-mainnet failed due to unresolved transitive dependencies. Although we obtained a working build after manually fixing multiple issues, the resulting binary did not match the release. According to the developers, some dependencies required for full reproducibility remain unknown, making reproduction infeasible. As with CCF and Gramine, the issue was insufficient build documentation; however, the missing information could not be recovered in this case. MobileCoin and Mystikos. For eleven of the 17 unreproducible projects, the cause remains unknown: unlike the Secret Network and Gramine cases, we received no response from the developers. Because enclave measurements are one-way hashes over every bit of the binary, we can only speculate about the source of non-determinism. Some causes—timestamps, build paths, file systems—are easy to spot; others, such as a mismatched dependency version, are far harder to pin down. Given how individualized the remaining build processes are, we selected two projects for manual analysis in order to rule out candidate causes: MobileCoin [15], a privacy-preserving payment network, and Mystikos [32], an SGX runtime for Linux applications. Both builds turned out to be deterministic and independent of build time. The Mystikos build, however, depended on the build path. We recovered the original path from the published binary and rebuilt with it, yet the result still did not reproduce, so at least one further factor must influence the output. Varying the file system across ext4, xfs, and btrfs left MobileCoin’s measurement unchanged, but altered the Mystikos binary hash between ext4 and xfs/btrfs. Recompiling Mystikos with the corrected build path produced yet another hash. Together, these results illustrate how hard it is to identify the root cause of non-reproducibility without access to the original build environment.
CCF. The Confidential Consortium Framework (CCF) [44] is an open-source framework for high-performance decentralized applications, providing confidentiality, integrity, and availability via a TEE-based consensus protocol. Starting with version 6.0.0, CCF discontinued support for Intel SGX and now exclusively targets AMD SEV-SNP. However, the Azure service “Confidential ledger” [35] still utilizes the SGX-based CCF in the backend [31]. We successfully compiled CCF v5.0.6, but the enclave measurement did not match the release. Achieving reproducibility required the CI/CD build path (/__w/CCF), which was kindly provided by the developers. With the transition from Intel SGX to AMD SEV-SNP, Microsoft appears to have prioritized reproducibility, similar to the Secret Network. The updated documentation includes a dedicated section on reproducible builds [34], providing sufficient information to achieve reproducibility without developer assistance.
Main Takeaways: As shown earlier, our analysis of the case studies did not reveal any technical challenges that were inherently difficult to overcome. Most issues could be addressed by adding more detailed information to the documentation, suggesting that fundamental technical limitations are not the main obstacle. In the following section, we explore the underlying reasons for the lack of reproducible TEE builds in practice through an interview study.
4
Interview Study
Having established that reproducibility challenges extend across SGX, TDX, and SEV-SNP, we turn to understanding the deeper factors at play behind these challenges. To this end, we conducted an exploratory qualitative interview study with 12 maintainers of open-source SGX projects from our dataset. All participants had several years of experience developing TEE software, primarily for SGX, with some also having experience with other TEE platforms. We focused on Intel SGX given its dominant ecosystem and comparatively mature developer base, while also inviting participants to reflect on their experiences with other TEEs where relevant. Our aim was to investigate developer practices, surface challenges,
Gramine and Ternoa. Similar to CCF, reproducing Gramine [43] and Ternoa [58] was hindered by incomplete build environments. For Gramine v1.8, building on the documented platform (Ubuntu 22.04) produced enclave measurements that did not match the official release. Reproducibility was only achieved by replicating the full release environment, including Debian 11 (with backports), exact dependencies, compiler flags, and environment variables provided by the developers. 6
Table 4: Summary of identified reproducibility blockers, including excluded sources of non-determinism for two projects without a developer response.
and identify opportunities for improvement in the reproducibility of TEE-based build pipelines. In what follows, we present our methodology and the results of the interviews.
4.1
Project Reproducibility blockers Excluded root causes Reproducible in guided scenario CCF [31] Build path – Gramine [43] Undisclosed dependencies – Marblerun [55] Undisclosed build configuration – Unreproducible in guided scenario MongoDB-SGX [12] Self-signed certificate – Nginx-SGX [13] Self-signed certificate – Safeheron Arweave [45] Unavailable dependency – Secret Network [25] Embedded secret – SGXSSE [60] Unavailable dependency – Ternoa enclaves [57] Unknown dependencies – Unreproducible and no guided scenario (no developer response) MobileCoin [15] – Compiler randomness, build path, file system, timestamp Mystikos [32] Build path, file system Compiler randomness, timestamp
Methodology
To understand how our target group approaches reproducibility in the context of SGX, we adopt a qualitative, exploratory research design based on semi-structured interviews. This approach is appropriate given the relatively small size of the population and the suitability of interviews for gaining in-depth insights into individual practices, processes, structures, and perceptions. Based on the repositories identified in Section 3, we reached out to potential participants to discuss their development practices. To collect demographic information in a structured way, save time during the interviews, and simplify scheduling, we distributed a short prequestionnaire prior to the 25-minute interviews. Figure 2 provides an overview of our methodology and its main components.
participants’ aggregated demographic information to prevent reidentification. Table 5 summarizes these aggregated demographics, including an overview of the participants’ positions and roles.
4.1.1 Recruiting. We reached out by email to the contributors of all SGX GitHub projects identified in Section 3, where contact information was available. In our recruitment email (cf. Figure 3 in Appendix D), we transparently explained that we are researchers seeking to advance TEE security in the context of reproducibility. We explicitly noted that no prior experience with reproducibility was required to participate. Participants were offered a €50 gift card (5000+ digital reward options in 170+ countries) as compensation and could schedule an interview directly at the end of the prequestionnaire. If no response was received after the initial contact, we sent one polite reminder and then refrained from further followups. Three participants chose to waive their compensation in the pre-questionnaire.
4.1.4 Interview Structure & Procedure. The semi-structured interviews were conducted with two researchers present. One served as the interviewer, while the other ensured adherence to the interview guide and handled technical aspects, with roles alternating across sessions. The researcher in the assistant role introduced themselves at the beginning, but then turned off their video and remained silent to avoid distracting the participant. The interview guide is structured into the five main areas: 1 Onboarding: During onboarding, the researchers briefly introduced themselves, the institution, and the purpose of the study to build a rapport with the participants. We asked if they had any questions about the consent form, and asked them to agree again to the interview being recorded. 2 Understanding & Practices: We first asked what reproducibility means for the participants, capturing their individual perception of the concept and the perceived importance. If a participant’s interpretation diverged from the study definition, the interviewer reiterated our definition of reproducibility developed following the pilot interviews to ensure a consistent baseline. The participants then explained the extent to which reproducibility was relevant in their current project. 3 Challenges: We followed up on previously mentioned issues and explicitly invited participants to discuss any additional challenges or problems they had encountered in their own projects, as well as those related to other TEEs. 4 Improvements: We asked participants how the reproducibility challenges in SGX and other TEEs could be addressed, and discussed appropriate channels for disseminating reproducibility-related information and ways the community could strengthen reproducibility practices. 5 Offboarding: We thanked the participants for taking part, gave them the opportunity to ask final questions, and outlined next steps, including issuance of the voucher and the interview evaluation process.
4.1.2 Piloting. The interview guide (cf. Section C) was developed deductive-iteratively. Four researchers—two with technical expertise and two with a background in human-centered security—independently proposed questions based on the earlier identified reproducibility issues (cf. Section 3). The questions were then collaboratively selected, grouped, and refined. The guide was piloted twice internally and twice externally with participants (P1 and P2) from our professional network, who worked with SGX in the past. The most substantial change resulting from the pilot concerned the definition of reproducibility in the context of SGX, which we refined to ensure a consistent understanding among participants for the subsequent interview questions. The pilot participants are therefore not included in our analysis or results. 4.1.3 Participants. From 180 SGX GitHub projects, we identified contact information for 67 contributors across 55 projects. Among those, the contact information of five projects was outdated. Of the 62 email invitations sent, we received 13 responses, and eleven participants agreed to schedule an interview. One participant, due to language difficulties, asked to communicate in writing rather than verbally. We accepted this mode of participation, meaning 12 interviews form our final dataset. Our participants came from eleven different countries, with an average age of 33.08 years and almost 16 years of programming experience. In terms of their roles, half of the participants hold academic positions, and the others work in industry. In line with common practice, we only report
Analysis. All interviews were recorded and transcribed locally using Whisper [40], then manually checked for accuracy, and potentially identifying information was redacted where necessary. We conducted the analysis following the principles of thematic analysis [4]. After each interview, the two researchers who conducted the interview sessions exchanged first impressions and took initial 7
Recruitment
Survey Structure Landing Page & Consent Form
Onboarding 25-minute Interview
Reaching out to Project Owners (n = 50)
Pre-Questionnaire
Table 5: Demographics of the participants (N=12).
Interview Guide Structure
SGX Github Repositories (n = 180)
Gender Male 12 100% Age [years] Min. 26.00 Mean, Std. 33.08 ±4.44 Highest Education Level Bachelor’s degree 1 8% Doctoral degree (PhD) 6 46% Country of Residency Germany 2 15% China 1 8% Japan 1 8% Slovenia 1 8% Switzerland 1 8% UK 1 8% Coding Experience [years] Total Median 15.00 Total Mean, Std. 15.92 ±5.32 Total Min. 9.00 Total Max. 25.00 Weekly Dev. Time [hours] less than 5 hours 4 31% 20–30 hours 4 31% Employment Status Employed full-time 10 77% I am not gainfully employed. 1 8% Position/role Academic researcher/scientist Developer Engineer CEO/Manager/Team lead Educator Security specialist
Understanding & Practices Challenges Improvements
Demographics Offboarding Interview Scheduling Transcription & Analysis
Figure 2: Overview of our methodology.
notes. Subsequently, the research team held joint analysis sessions to deepen their understanding of the data, clarify open questions, and establish an initial version of the codebook. Following the joint session, one researcher coded the data and iteratively refined the codebook as new aspects emerged. After finalizing the codebook and the initial coding, a second researcher reviewed all codes and discrepancies were resolved through discussion. The resulting codebook was used as a shared analytical framework for the subsequent steps and is included in the appendix (cf. Section D). In several iterative discussions within the research team, we used the refined codes to synthesize higher-level themes and structure the presentation of results. These sessions also informed the development of overarching themes and insights discussed in Section 5.
4.2
Max. Median
40.00 33.00
Master’s degree
5
38%
Brazil India Singapore Sweden Thailand
1 1 1 1 1
8% 8% 8% 8% 8%
Prof. Median Prof. Mean, Std. Prof. Min. Prof. Max.
6.00 7.08 ±4.25 2.00 18.00
10–20 hours 30–40 hours
3 1
23% 8%
Self-employed
1
8%
P3, P5, P6, P7, P8, P11, P12 P4, P6, P12 P4, P14 P4 P3, P6, P11 P4, P9, P10, P13, P14
lack of a consistent understanding of what reproducibility actually entails. Participants also gave us multiple ideas and definitions themselves. For seven participants, “reproducible” meant publishing results or application output that could be replicated: “Once a research or a group of researchers publish their work, other researchers around the world are able to get similar results by running, by executing, by reproducing exactly what these researchers made available.” – [P8]. Three participants described reproducibility as creating the same binary from the source code, which was very close to our definition of reproducibility. Five associated the notion of reproducibility with applications “working as expected”. We further asked participants in what context they had learned about reproducibility. Although not all participants responded (9/12), four said they had gained their knowledge in an academic setting and three described it as a general understanding. Only one participant said they had learned about it in an educational setting or in the workplace.
Results
We now present the interview results, following the structure of our codebook (cf. Table 7 in Appendix D). We begin with participants’ perceptions of reproducibility, followed by the associated challenges and how reproducibility is prioritized within their projects. We then discuss use cases and enablers of reproducibility, and conclude with current practices and suggested improvements. 4.2.1 Perception of Reproducibility. Participants discussed two areas: the importance of reproducibility for TEEs and what reproducibility meant to them. Afterward, the researchers provided their definition of reproducibility to ensure a shared understanding of the concept. Relevance of reproducibility. Eight participants commented on the relevance of reproducibility. Five participants stated that it is essential for ensuring security, and three participants stated that a lack of reproducibility directly leads to a lack of security: “It’s not a feature. It’s an assumption. So if this assumption breaks, then our system breaks.” – [P7] However, two participants also considered reproducibility irrelevant for end-users of products, as it might be challenging for them to test reproducibility: “You want to enable people to do it themselves, but only a few people will do it themselves, right? Most people will want to trust some service or some statement or some public value. And doing this transparently is also difficult, right.” – [P9]
4.2.2 Challenges. Participants reported various challenges they encountered within their project, or which they perceived as challenging. While most participants referred to SGX, four of the 12 noted that the challenges are often not specific to a given TEE, but apply to confidential computing in general. An overview of the participants’ mentioned challenges is presented in Table 6 (here we also highlight which of these challenges we encountered—either fully or partially—during our own analysis in Section 3.3). Limited control over the environment. The challenge that was most frequently cited by participants (6/12) was limited control over the environment. This included the impact of dependencies
Definition of reproducibility. We identified three different patterns about reproducibility in the context of SGX, highlighting the 8
Table 6: Overview of challenges reported by the participants. The dots indicate whether we encountered the challenges in our technical analysis completely, partially, or not at all.
on reproducibility, the complexity and dynamics of the development environment, and the fact that the build environment often does not support reproducibility. P12 explained the complexity of considering all parts that might impact the hash of a binary: “In my experience a lot of moving parts, system or program which is complex enough has so many moving parts, it’s pretty hard to count them all in with at least regular tools.”. P4 also elaborated on the challenges of distributed build environments: “We’re aware that for bigger build systems, you might have a situation where you’re distributing the build across machines and then maybe the optimizations or the, the final binary that’s produced depend on how the distribution has been done.” – [P4] P6 further added, that in some project one does not have control over the build environment, as they might not be maintained by the development team. Dependencies might be integrated in the environment (e.g., in Ubuntu itself), they further elaborated on their concern, that “the developers have to be very precise about the dependency they’re using, make sure that they’re coming from the same source, compiling using the same tools, etc.” Five other participants also mentioned challenges related to software dependencies. P3, P4, and P6, for example, explained that major changes to packages can also affect the compiler tool set, and that everything must be managed accordingly. In the same context, P3 mentioned that they used multiple build environments: “Actually, for that, I think in one of my earlier projects, I just like have multiple gcc versions. So for different enclaves, I need to, like, switch to that version, and then switch back.”. It was also noted that dependencies can change rapidly over time and must be managed accordingly. Legacy software and packages that are no longer maintained also pose a challenge. Software packages may no longer be available, or patches may need to be applied manually, which can consequently affect the build process or the build environment(s). One participant also criticized the compilers’ lack of support for reproducible builds: “But it’s more like the tooling. The tooling is not aware of this. And I don’t know how aware the compiler writers are of the importance of maintaining reproducibility. They could make optimization decisions that would make the builds faster or more efficient in some way, but at the expense of reproducibility.” – [P4]
Challenge # P. Encountered Limited control of the environment 6 Lack of Awareness 4 High complexity 4 Debugging non-reproducibility is hard 2 Challenges are often not TEE specific 2 Non-deterministic signatures 1 Timestamps in Binaries 3 Path resolution 2 Hotfix release before fix 1
P10, for example, pointed out that SGX development is “notoriously complex” and that this “discourages developers from focusing on reproducibility in already highly pressured environments”. P14 also emphasized the complexity of achieving reproducibility by stating that they knew people at a large tech enterprise who believed reproducibility could not be achieved within the available budget. P10 specifically explained that reproducibility becomes even more challenging in the future: “With newer CVM technologies like TDX or SEV-SNP, the enlarged TCB makes reproducibility virtually impossible to guarantee. Attempting to ensure reproducibility across such massive stacks is clearly infeasible, further shifting attention away from reproducibility.”. Debugging non-reproducibility is hard. Three participants also reported that compilation alone, without accounting for reproducibility, posed a major challenge in general. Two participants reported that debugging was a significant challenge when reproducibility failed. Factors contributing to this included insufficient documentation of the components that influence the hash (P10) and difficulty in identifying sources of non-determinism (P14). Non-deterministic signatures & hotfix release before fix. One participant (P4) mentioned that generating and embedding the signature early in the build process makes it non-deterministic and that the rapid release of the source code after a hotfix—while necessary for reproducibility—also allows attackers to exploit vulnerabilities before end users have applied the patch.
Build-induced irregularities. Irregularities introduced during the build process were mentioned by five participants. Four participants mentioned that timestamps left in the build binary “are a really common issue”. Paths embedded during the compilation process also posed a challenge for two participants. Related to paths introducing irregularities, P4 stated: “Another thing is that anything that sets symbols in the binaries tends to capture the absolute path where you run the builds. For SGX, we tended to kind of sidestep that.”
4.2.3 Prioritization of Reproducibility. We found that only one participant (P14) perceived reproducibility as a priority in their project. The participant described it as “one of our major goals” with nonreproducibility being a “release blocker”. On the other hand, eight participants reported that reproducibility had little or no priority in their projects. One participant explicitly elaborated on the situation of projects from an academic context: “This intense academic competition meant that issues like reproducibility—although important—were often de-prioritized if they did not immediately lead to security breakage.” – [P10]. They continued elaborating on commercial products: “On the market side, products such as PowerDVD or blockchain-related solutions did appear. However, here too, due to commercial competition or the involvement of maintainers who may not be deeply familiar with TEEs, specialized issues like reproducibility were again often neglected.” – [P10]. P6 reflected on their own experience, where they prioritized other tasks over achieving reproducibility: “I think that there are things that are a bigger priority
Lack of awareness. Four participants either directly stated that they or others are not aware of the importance of reproducibility in the context of the security of SGX applications, or they had misconceptions about it. For example, one participant assumed that reproducibility depends primarily on their own code rather than on the “underlying dependencies”, or it was pointed out that researchers may think of artifact evaluation instead when it comes to reproducibility. High complexity. Four participants stated that the complexity of the concepts behind secure enclaves makes reproducibility difficult. 9
This included that participants used a “single machine” – [P4] or the same “environment for a couple of years, so it’s kind of stable” – [P3]. The use of static build paths, fixed timestamps, or fixed entropies for signatures was also mentioned by individual participants. Among the participants, two elaborated on the sources from which they learned how to make applications reproducible. While both participants stated that trial and error was part of the learning process, P4 also pointed out that “a lot of the techniques and the tooling come from Linux distro people” and that they are unaware of a single source of reference for reproducibility: “It was very much a sort of look, look for things, look for techniques. I don’t think we have a single, single source of reference.” – [P4].
right now. For example, when I was trying to, when I was developing my own SGX software, there was a big issue with attestation because it’s really difficult to use [...] I think that there are a list of priorities in SGX deployments and we are kind of working down the list. We still haven’t gone past some really important issues, like even to get things to compile, even to get the framework to stabilize and usable and user-friendly.” – [P6]. They also explained that companies “know this is a problem, but they don’t care about this” and that they “don’t want to invest manpower”. In contrast to the other participants, reproducibility was a priority for P14, but they also stated that customers might not value it: “For us, it’s a selling feature that our stuff is reproducible, and you can validate it. But yeah, the usually contact to a customer is like, okay, they think it is really important and then they learn about what it means to verify it. And then they are, maybe it’s not that important. It’s enough if you can like have a checkbox and yeah.” – [P14]
4.2.5 Suggested improvements. All but one participant provided suggestions on how to improve reproducibility. Their ideas included both technical improvements and general recommendations, as well as specific actions researchers could take and channels through which information about reproducibility should be disseminated.
4.2.4 Practices to achieve reproducibility. In the interviews, we also discussed our participants’ experiences who had addressed, or attempted to address, reproducibility. As outlined in Section 4.2.3, reproducibility was not a priority during development for most participants. However, two participants elaborated in more detail about the time they invested in achieving reproducibility. P4, for example, explained that reproducibility had not been considered from the beginning of the project and that they later assigned one full-time employee to address it: “I think [the employee] spent about five to six weeks, I would say, improving the reproducibility of the process, making it path independent, having a CI job that was doing that, producing a script so that end users could just run the script with a metadata file and reproduce things.” – [P4]. They continued explaining that “reproducibility is not necessarily something you have straight away”. But after they achieved reproducibility, they also made clear that they do expect to invest much effort in the future. P14 also stated that, depending on their situation, they spent a considerable amount of time on reproducibility, but in general it did not take much from the development budget: “Maybe 5% to 10% over the years. But that that greatly differs what we are working on, right? So we had times where we spent like full months on reproducibility.” – [P14]. P7 guessed, that for commercial projects the effort should remain below 30%. Eight participants further elaborated on tools they use or consider to use for achieving reproducibility. All of them mentioned Docker. While four were already using Docker, and four reported that Docker is a valuable approach for achieving reproducibility, it was also criticized. P14 noted that “The inputs are not pinned. So if you do an apt install or an update in the Docker container, you will just get anything”. Other tools named as valuable were Nix (2), Bazel (1), and the Yocto Project (1). Specifically, Nix and Bazel are described as useful for achieving reproducibility, because they require developers to be more strict about their input specifications: “Nix has quite a strong sandboxing. So we know that all the inputs that we require are pinned by hash because that’s enforced by Nix, actually. [...] We also use Bazel in some other projects, but the concept is the same, right? You need to pin all your inputs by hash.” – [P14]. Four participants also commented on the practices they used during the build process to achieve reproducibility. The focus was particularly on the need to pre-define sources for non-determinism.
General desirable improvements. Four participants demanded that “we should increase awareness.” – [P7]. P7 further elaborated stakeholders are unlikely to prioritize reproducibility without a triggering security incident: “So yeah, it’s not easy. So unless, so if they see some problem from reproducibility, so they will invest efforts on this. So yeah, we need to maybe try to provide some kind of like attack. It’s not real attack, but we try to present the problem to the industry to let them know that this is a big problem. [...] So the very big zero-day vulnerability, so the company will pay attention very quickly.” – [P7]. P10 also highlighted the importance of raising awareness: “Perhaps most importantly, raising awareness of reproducibility as a critical issue will itself have a major impact.” – [P10] Other general improvements included, for example, “extending version-managed builds like Yocto Project” – [P10], building reproducibility tools in GitHub, for example, “some kind of script that checks if it’s following all the reproducible guidelines” – [P8], or developing tools that simplify debugging (P14). Technical improvements. One technical improvement mentioned was that the entire compilation environment could be built and offered in Docker (P6, P7). Other suggestions were the use of SBOMs to track the tools used in the environment (P3, P9), introducing reproducibility features in package managers (P7), or a trusted authority issuing code-signing certificates for enclave code (P10). Support through research. Six participants expressed wishes for how research could support the dissemination of reproducible builds through appropriate projects. P5 and P8 asked for specific best practices for achieving reproducibility that could be defined by research. P4 suggested creating sample projects in which tools could be evaluated and reproducibility practices could be traced. P10 hoped that “having researchers directly join projects as maintainers could improve reproducibility quality”. P14 also suggested research teams could evaluate the reproducibility of different ecosystems. Information channels. Six participants elaborated on the channels they preferred for information on SGX reproducibility. Overall, however, there seemed to be no consensus on this point. While three participants mentioned the Intel website, other channels, such as 10
scientific conferences, attestation services, and GitHub repositories, were discussed only by individual participants.
5
undisclosed secrets and non-deterministic signatures was identified as a major obstacle, consistent with our empirical findings (cf. Section 3.3). Finally, dependency management emerged as a particularly challenging aspect in the interviews. This is reflected in our empirical results, which show that tools such as Nix [6], Bazel [41], and Yocto [16]—all three enforcing externally defined and versioned build inputs (e.g., dependencies)—substantially improve reproducibility (cf. Section 3.2).
Discussion
We now discuss and interpret the results from our analysis of TEE projects presented in Section 3 and the interview study described in Section 4. Table 7 in the appendix maps the themes discussed in this section to the corresponding codes of our codebook. Our analysis revealed a general lack of reproducibility among open-source TEE projects. Specifically, only ten of the 115 analyzed projects, i.e., 8.7%, were reproducible (cf. Section 3.2). Among the three projects that required additional developer input to achieve reproducibility, the lack of relevant documentation was the primary reason (cf. Section 3.3); once the missing information was obtained, achieving reproducibility was straightforward. Nonetheless, this raises the question of why reproducibility support remains so scarce, given that the ability to reproduce enclave binaries is an important component for remote attestation—and therefore essential for the secure deployment of TEEs.
5.1
Organizational Barriers: Problems Exceed the Developers’ Responsibility. A key reason identified in our interviews for the lack of practical reproducibility support—beyond known technical challenges— is developers’ limited control over the build environment (cf. Section 4.2.2). In many projects, builds are executed in CI/CD2 pipelines that are centrally managed by DevOps teams or shared across multiple teams, rather than controlled by individual developers. As a result, even straightforward technical changes can become significant organizational hurdles, hindering the adoption of deterministic build practices. Moreover, software development is typically collaborative, often spanning multiple teams or organizations. If any component in this chain fails to ensure reproducibility, the entire build process may become non-deterministic. One participant noted that a single non-reproducible dependency can compromise the entire application’s reproducibility (cf. Section 4.2.2). Addressing such issues typically requires coordination with external parties or manual tracing of non-determinism—both time-consuming efforts: in our empirical analysis, only 35% of contacted development teams responded, and in four cases, this information was still insufficient to reproduce the build (cf. Section 3.2). Similarly, participants emphasized that debugging non-determinism in large, dependency-rich code bases is complex and resource-intensive. These findings suggest that practically improving reproducibility in TEE-based applications requires more than technical knowledge among developers.
Trends Across TEE Technologies
Our paper examined three TEE technologies: Intel SGX, Intel TDX, and AMD SEV-SNP. Among these, SGX is the oldest and most extensively studied, TDX is the most recent, and SEV-SNP was released in between. All three ground their security guarantees in remote attestation, for which reproducible builds are a key prerequisite to enable independent verification of enclave measurements. Arm TrustZone, by contrast, does not natively support remote attestation, making build reproducibility less critical in that context—though it remains beneficial for mitigating supply-chain attacks. Our technical analysis of 115 TEE deployments finds no statistically significant association between the TEE technology and reproducibility (cf. Section 3.2). This finding is reinforced by our interview study, in which participants reported similar challenges across platforms, including non-deterministic build inputs, insufficient documentation, and complex dependency management (cf. Section 4.2.2). Taken together, these results suggest that reproducibility challenges are TEE-agnostic and that our findings (along with the broader implications discussed below) generalize across TEE architectures and likely extend to future TEE designs that similarly rely on remote attestation and enclave measurement verification.
5.2
5.3
Security Is A Badge
Recall that security is a central feature of TEE projects, with remote attestation forming a core component of this model. Eight interview participants consider reproducibility highly important—or even as a prerequisite—for ensuring security in SGX-based systems, confirming this perspective. Despite this, our empirical findings show that reproducibility is rarely achieved in practice. This aligns with feedback from all but two participants who noted that reproducibility is generally not treated as a development priority (cf. Section 4.2.3). These results indicate a disconnect between the perceived importance of reproducibility and its prioritization in practice, driven by decision-making rather than developer awareness. Over a project’s lifecycle, evolving organizational structures, increasing dependencies, and personnel turnover make retrofitting reproducibility increasingly difficult. One participant reported that achieving reproducibility required five to six weeks of dedicated developer effort once it was prioritized—an investment many organizations avoid without clear incentives. Prior work similarly shows the difficulty of dependency discovery, even with specialized
Reproducibility Challenges
Developers’ Challenges: Known Issues. Our interview study revealed several technical issues that complicate realizing reproducibility in practice (cf. Section 4.2.2), many of which are known from prior work on traditional reproducible builds [27]. Five participants, for instance, mentioned timestamps included in the binary— introduced via macros (e.g., C’s __DATE__) or build tools—as a common source of non-determinism. Similarly, divergent file system ordering across hosts can alter processing order, and absolute file paths embedded in binaries (e.g., via the __FILE__ pre-processor macro) vary with the build directory, eventually leading to nonidentical artifacts. We also observed the latter in our analysis of an SGX project (cf. Section 3.3). Additionally, randomness from
2 Continuous Integration and Continuous Deployment
11
tools [5, 52]. This highlights the need for stronger organizational and economic incentives to address reproducibility early. Given that remote attestation has limited value without reproducible builds, a key question arises: how do customers and stakeholders perceive the value of attestation, and is the mere label “uses TEE” considered sufficient? Without a technical understanding of attestation, customers may not demand reproducibility, thereby weakening incentives for companies to invest in it. Similar patterns appear in other security domains, where the cost of implementing robust protections often outweighs the perceived financial risk of incidents [20]. Consistent with this, one interview participant noted that reproducibility would only become a priority if customers explicitly demanded end-to-end attestation support (cf. Section 4.2.3).
5.4
importance of reproducibility for trusted execution and account for developers’ diverse backgrounds. TEE manufacturers such as Intel and AMD could incorporate this guidance into their developer documentation and whitepapers. (2) Adopt practices for reducing non-determinism. Although the common sources of non-determinism are well known [10], many projects still rely on them (cf. Section 3.3), and two interviewees reported substantial team effort spent tracking such issues down. We recommend established best practices, such as avoiding the __DATE__ and __FILE__ macros and using compiler options like GCC’s -ffile-prefix-map for locationindependent builds. Addressing these issues early can substantially improve reproducibility and avoid costly revisions later. (3) Document the build process and dependencies. Our analysis shows that many open-source TEE projects lack sufficient documentation to reproduce enclave measurements, even when the required information is available to the developers. We recommend documenting the build process and its dependencies systematically, including exact versions and cryptographic hashes, and using standardized formats for build metadata. Debian’s .buildinfo files [9], for example, record build environments, toolchains, and parameters, enabling verification across systems.
Study Limitations
Although we contacted developers for all 50 out of 180 SGX projects for which we could locate contact information, only 12 agreed to be interviewed, and we were unable to reach additional developers. Reasons for this limited response can range from a lack of time and legal constraints to a lack of prioritization of reproducibility. Nonetheless, we note that 71% of the development teams contacted for our technical study were unresponsive (cf. Table 3), suggesting that support for aiding reproducibility could be improved. While the response rate limited the sample size of our interview study, it covers diverse participants’ backgrounds and TEE expertise, capturing a range of perspectives, and thus offers exploratory evidence about the perception of reproducibility among TEE developers and the barriers they face. Note that nearly half of the participants were from academia, where reproducibility is often not a primary objective and enclave implementations are typically prototypes that may not support remote attestation. Nonetheless, our technical analysis—including production-grade industry projects—shows that poor reproducibility is equally prevalent in industrial-scale TEE deployments. Finally, our interview-based findings reflect the usual constraints of qualitative self-reported data, including recall bias and social desirability bias, particularly among maintainers who may be reluctant to portray their projects or organizational practices negatively.
6
6.2
Beyond individual practices, our analysis highlights organizational barriers to reproducibility. We therefore propose the following recommendations for project managers and decision-making bodies: (1) Prioritize reproducibility during development. Several interviewees reported that reproducibility is not an organizational priority, and developers in large organizations often lack the authority to modify standardized CI/CD workflows. We therefore recommend treating reproducibility as a release requirement, supported by dedicated training, coding guidelines, and process audits. One participant’s organization already enforces reproducibility as a prerequisite for deployment acceptance, a practice we argue should extend to all TEE-based applications. (2) Enforce code provenance and transparency. Users currently must either trust developers to publish binaries built from the correct source or verify the correspondence between source and binary manually, which is error-prone. We recommend that organizations adopt code provenance services to improve transparency and accountability. Delignat-Lavaud et al. [10], for example, propose a code transparency service that tracks TEE code provenance through a public, append-only ledger. Integrated into CI pipelines, such services can verify provenance automatically before release. (3) Standardize build environments and development workflows. Half of the developers in our study reported challenges from inconsistent build environments, particularly in large, multi-team projects with external dependencies. Tools such as Docker [11], Nix [6], and Bazel [41] isolate the build process and ensure consistent toolchains. Docker encapsulates dependencies and configurations in portable container images, although versions must be pinned explicitly. Bazel sandboxes build actions and tracks input hashes, and its hermeticity mode
Recommendations
Based on our technical analysis and expert interviews, we recommend that TEE stakeholders raise awareness of reproducibility, provide practical guidance, and establish it as a foundational concern in TEE development.
6.1
Organizational Managers and Stakeholders
Application Developers
To improve reproducibility without disrupting established development workflows, we propose the following recommendations for application developers: (1) Increase awareness and understanding of reproducibility. Our interviews reveal limited consensus on reproducibility concepts, with 33% of developers unaware of its importance for TEEs and many reporting diverging interpretations shaped by academic or general software perspectives. We therefore recommend standardized educational materials that explain the 12
pins tools and dependencies. Nix is designed for reproducibility from the outset, enforcing exact dependency sources and hashes.
6.3
innovation program (REWIRE, Grant Agreement No. 101070627), and an ARC Discovery Project number DP210102670. Views and opinions expressed are, however, those of the authors only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them. Moreover, this paper was edited for grammar using Grammarly, DeepL, and ChatGPT.
End Users and Customers
Our findings suggest that the main obstacle to reproducibility is not technical difficulty but a lack of organizational incentive. End users and customers can therefore drive adoption by requiring reproducibility explicitly in procurement specifications, requesting verifiable build information, and favoring products built with reproducible development practices. Such market pressure can elevate reproducibility from an optional feature to a core component of trust and security in TEE-based software.
6.4
References [1] AMD. 2020. Strengthening VM isolation with integrity protection and more. White Paper, January (2020). https://www.amd.com/content/dam/amd/en/ documents/epyc-business-docs/white-papers/SEV-SNP-strengthening-vmisolation-with-integrity-protection-and-more.pdf Accessed: 2025-11-08. [2] ARM Architecure. 2009. ARM Security Technology Building a Secure System using TrustZone Technology. ARM Limited (2009). https://documentationservice.arm.com/static/5f212796500e883ab8e74531 Accessed: 2025-11-08. [3] USB armory. 2026. The Go Programming Language: Repository on GitHub. https://github.com/usbarmory/tamago-go. Tag: tamago-go1.26.2. [4] Victoria Clarke and Virginia Braun. 2014. Thematic analysis. In Encyclopedia of critical psychology. Springer, 1947–1952. [5] Serena Cofano, Giacomo Benedetti, and Matteo Dell’Amico. 2024. SBOM Generation Tools in the Python Ecosystem: an In-Detail Analysis. In TrustCom. 427–434. doi:10.1109/TrustCom63139.2024.00077 [6] NixOS contributors. 2026. Nix & NixOS | Declarative builds and deployments. https://nixos.org/. [7] cosmian. 2026. Cosmian KMS: Repository on GitHub. https://github.com/ Cosmian/kms. Tag: 5.17.0. [8] Xavier de Carné de Carnavalet and Mohammad Mannan. 2014. Challenges and implications of verifiable builds for security-critical open-source software. In ACSAC. 16–25. doi:10.1145/2664243.2664288 [9] debian. 2017. Reproducible Builds BuildinfoInfrastructure. https://wiki.debian. org/ReproducibleBuilds/BuildinfoInfrastructure. Accessed: 2026-04-22. [10] Antoine Delignat-Lavaud, Cédric Fournet, Kapil Vaswani, Sylvan Clebsch, Maik Riechert, Manuel Costa, and Mark Russinovich. 2024. Why Should I Trust Your Code? Commun. ACM 67, 1 (2024), 68–76. doi:10.1145/3624578 [11] Docker Inc. 2026. Docker: Accelerated Container Application Development. https://www.docker.com/. [12] enclaive. 2022. MongoDG-SGX: Repository on GitHub. https: //github.com/enclaive/enclaive-docker-mongodb-sgx. Commit: 464f22f619fd729cc2669315ec872aabc2042305. [13] enclaive. 2022. Nginx-SGX: Repository on GitHub. https: //github.com/enclaive/enclaive-docker-nginx-sgx. Commit: 1372c154096f9b03a40918dd75f5b8e89a690484. [14] Oasis Protocol Foundation. 2025. Sapphire ParaTime: Repository on GitHub. https://github.com/oasisprotocol/sapphire-paratime. Tag: v1.2.0. [15] Sentz Foundation. 2024. MobileCoin: Repository on GitHub. https://github.com/ mobilecoinfoundation/mobilecoin. Tag: v6.1.1. [16] The Linux Foundations. 2026. Yocto Projects. https://www.yoctoproject.org/. [17] Marcel Fourné, Dominik Wermke, William Enck, Sascha Fahl, and Yasemin Acar. 2023. It’s like flossing your teeth: On the Importance and Challenges of Reproducible Builds for Software Supply Chain Security. In IEEE SP. 1527–1544. doi:10.1109/SP46215.2023.10179320 [18] Google. 2022. Asylo: Repository on GitHub. https://github.com/google/asylo. Commit: e73365548454950929d339d37d810a789e5c383b. [19] Google Cloud. 2026. Google Cloud Marketplace. https://console.cloud.google. com/marketplace Accessed: 2026-03-24. [20] Lawrence A. Gordon and Martin P. Loeb. 2002. The economics of information security investment. ACM Trans. Inf. Syst. Secur. 5, 4 (Nov. 2002), 438–457. doi:10. 1145/581271.581274 [21] Matthew Hoekstra, Reshma Lal, Pradeep Pappachan, Vinay Phegade, and Juan del Cuvillo. 2013. Using innovative instructions to create trustworthy software solutions. In HASP. 11. doi:10.1145/2487726.2488370 [22] Daniel Hugenroth, Mario Lins, René Mayrhofer, and Alastair R. Beresford. 2025. Attestable builds: compiling verifiable binaries on untrusted systems using trusted execution environments. CoRR abs/2505.02521 (2025). arXiv:2505.02521 doi:10. 48550/ARXIV.2505.02521 [23] Intel. 2023. Intel Trsut Domain Extensions. (2023). https://cdrdv2.intel.com/v1/ dl/getContent/690419 Accessed: 2025-11-08. [24] Internet Engineering Task Force (IETF). [n. d.]. Supply Chain Integrity, Transparency, and Trust (SCITT). https://datatracker.ietf.org/group/scitt/about/. Accessed: 2025-11-13. [25] Secret Labs. 2024. Secret Network: Repository on GitHub. https://github.com/ scrtlabs/SecretNetwork. Tag: v1.14.0.
Special Interest Groups and Researchers
Finally, special interest groups such as the IETF Supply Chain Integrity, Transparency, and Trust (SCITT) Working Group [24] can advance reproducibility in the TEE ecosystem by defining standards, tools, and best practices, and by raising awareness in developer and policy communities. Researchers can complement these efforts by identifying reproducibility barriers systematically, developing improved verification mechanisms, and building open-source tools and reference implementations that make reproducibility practical and accessible across the ecosystem.
7
Conclusion
In this paper, we investigated a fundamental aspect of the security guarantees offered by TEEs: the reproducibility of TEE builds. We empirically analyzed 115 TEE deployments and conducted an interview study with 12 SGX developers from both industry and academia. Our analysis revealed that 91% of the examined projects were not reproducible, indicating that reproducible builds remain the exception rather than the norm in today’s TEE ecosystem. Our findings indicate that reproducibility challenges are largely TEE-agnostic. Instead, they stem from a combination of technical and organizational factors, including non-deterministic build processes, inadequate documentation, limited control over build environments, and low prioritization. Interestingly, developers often perceive TEE reproducibility as inherently difficult. Yet, both prior research and our findings indicate that many sources of nondeterminism in build processes are well-understood and can be addressed using straightforward technical solutions. This points to organizational and awareness-related factors as the primary obstacles to achieving reproducibility. Building on these insights, we proposed targeted recommendations for TEE ecosystem stakeholders—developers, project managers, and standardization bodies—to raise awareness, establish reproducibility as a core quality and security requirement from project inception, and offer practical guidance for achieving it without disrupting existing workflows.
Acknowledgments This work is partly funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy - EXC 2092 CASA - 390781972, DFG project number 560392681, the European Union’s Horizon 2020 research and 13
[59] The Apache Software Foundation. 2023. Apache Teaclave SGX SDK: Repository on GitHub. https://github.com/apache/teaclave-sgx-sdk. Commit: 1b1d03376056321441ef99716aa0888bd5ef19f7. [60] Viet Vo, Shangqi Lai, Xingliang Yuan. 2025. SGXSSE: Repository on GitHub. https://github.com/MonashCybersecurityLab/SGXSSE. Tag: 1.25.2. [61] Annika Wilde, Marco Gutfleisch, Felix Reichmann, Anirban Chakraborty, Yuval Yarom, M. Angela Sasse, and Ghassan Karame. 2026. Artifact Testing Reproducibility of TEE Projects: Repository on GitHub. https://github.com/RUBInfSec/reproducibility_of_tee_projects
[26] Secret Labs. 2026. Secret VM Build System: Repository on GitHub. https://github. com/scrtlabs/secret-vm-build. Tag: v0.0.25. [27] Chris Lamb and Stefano Zacchiroli. 2022. Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Softw. 39, 2 (2022), 62–70. doi:10.1109/ MS.2021.3073045 [28] Dayeol Lee, David Kohlbrenner, Shweta Shinde, Krste Asanović, and Dawn Song. 2020. Keystone: an open framework for architecting trusted execution environments. In EuroSys. 38:1–38:16. doi:10.1145/3342195.3387532 [29] Maxul Lee. 2024. Awesome SGX Open Source Projects: Repository on GitHub. https://github.com/Maxul/Awesome-SGX-Open-Source/ Commit: 008a9e657a53b0aa5acbcbfe01b27197172f29b6. [30] Microsoft. 2022. enclave-vrf: Repository on GitHub. https://github.com/microsoft/ CCF. Tag: v0.1.0. [31] Microsoft. 2024. CCF: Repository on GitHub. https://github.com/microsoft/CCF. Tag: ccf-5.0.6. [32] Microsoft. 2024. Mystikos: Repository on GitHub. https://github.com/microsoft/ mystikos. Tag: v0.13.0. [33] Microsoft. 2024. Reproducible Build - CCF documentation. https://ccf.dev/release/ 6.x/audit/reproducible_build.html. Accessed: 2026-02-26. [34] Microsoft. 2026. CCF: Repository on GitHub. https://github.com/microsoft/CCF. Tag: ccf-6.0.23. [35] Microsoft. 2026. Confidential Ledger. https://marketplace.microsoft.com/enus/product/Microsoft.ConfidentialLedger?tab=Overview. Accessed: 2026-04-27. [36] Microsoft. 2026. Microsoft Azure Marketplace. https://azuremarketplace. microsoft.com Accessed: 2026-03-24. [37] Mithril Security. 2023. BlindAI: Repository on GitHub. https://github.com/mithrilsecurity/blindai. Tag: v0.6.3. [38] Enigma MPC. 2026. Secret Network Node. https://marketplace.microsoft.com/enus/product/enigmampc1596722883851.secret-node-01?tab=Overview. Accessed: 2026-04-27. SecretVM Full Verification. https://docs.scrt. [39] Secret Network. 2024. network/secret-network-documentation/secretvm-confidential-virtualmachines/verifying-a-secretvm/full-verification. Accessed: 2026-04-14. [40] OpenAI. 2025. Whisper. https://github.com/openai/whisper. Accessed: 2025-1016, commit: 31243bad24cc746f07d4c8bfdd2d974872cb1803. [41] Bazel organization. 2026. Bazel. https://bazel.build/. [42] Phala Network. 2024. Phala Blockchain: Repository on GitHub. https://github.com/Phala-Network/phala-blockchain. Commit: 536f4aacd99815ebf64708eb40890319adb6fb05. [43] The Gramine Project. 2024. Gramine Library OS with Intel SGX Support: Repository on GitHub. https://github.com/gramineproject/gramine. Tag: v1.8. [44] Mark Russinovich, Edward Ashton, Christine Avanessians, Miguel Castro, Amaury Chamayou, Sylvan Clebsch, Manuel Costa, Cédric Fournet, Matthew Kerner, Sid Krishna, et al. 2019. CCF: A framework for building confidential verifiable replicated services. Microsoft, Redmond, WA, USA, Tech. Rep. MSR-TR-2019-16 (2019). [45] Safeheron. 2023. Safeheron’s TEE Based RSA Key Sharding Service: Repository on GitHub. https://github.com/Safeheron/sgx-arweave-cpp. Commit: 61cfa6857b5ccbc6536f50d3b9b94bed63b1876b. [46] Safeheron. 2026. Intel(R) Software Guard Extensions for Linux* OS: Repository on GitHub. https://github.com/intel/confidential-computing.sgx. Tag: sgx_2.28. [47] Secret Network. [n. d.]. About Secret Network. https://scrt.network/about. Accessed: 2025-11-12. [48] SecretFlow. 2025. SecretFlow: Repository on GitHub. https://github.com/ secretflow/secretflow. Commit: 11425fa5282cf0898393cf881df09d210ef384e2. [49] Signal. 2026. Signal. https://signal.org/en/. [50] Signal Messenger, LLC. 2020. CDSI (Contact Discovery Service on Icelake): Repository on GitHub. https://github.com/signalapp/ContactDiscoveryServiceIcelake. Commit: 052069563303c6b2fa8dc3069cef450b678e2ea3. [51] SnowHaze. 2020. SnowHaze Zero-Knowledge Verification: Repository on GitHub. https://github.com/snowhaze/zka-sgx. Commit: a4b7f4ac497b03697c0558bb0d39c97ae443cf6a. [52] Trevor Wayne Stalnaker, Nathan Wintersgill, Oscar Chaparro, Massimiliano Di Penta, Daniel M. German, and Denys Poshyvanyk. 2024. BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems. In ICSE. 1–13. doi:10.1145/3597503.3623347 [53] Edgeless Systems. 2025. Edgeless RT: Repository on GitHub. https://github.com/ edgelesssys/edgelessrt. Commit: e1f37c0d0473e405105bc085cb12728b40930623. [54] Edgeless Systems. 2025. EGo: Repository on GitHub. https://github.com/ edgelesssys/ego. Tag: v1.9.0. [55] Edgeless Systems. 2025. MarbleRun: Repository on GitHub. https://github.com/ edgelesssys/marblerun. Tag: v1.9.0. [56] Occlum team. 2024. Occlum: Repository on GitHub. https://github.com/occlum/ occlum. Commit: 6eaad6994148bc50ef403755f58189ce0270a829. [57] Ternoa. 2023. Ternoa TEE server for Secret-NFT: Repository on GitHub. https: //github.com/capsule-corp-ternoa/ternoa-enclaves. Tag: v0.4.5-mainnet. [58] Ternoa. 2025. Ternoa 2.0. https://www.ternoa.network/. Accessed: 2025-11-12.
A
Open Science
In compliance with the open-science policy, we make all research artifacts required to evaluate the contributions of this paper publicly available at [61]. Artifacts. The following artifacts are provided: (i) a table containing the detailed results of our analysis of TEE projects from [19, 29, 36]; (ii) detailed instructions on our reproducibility tests for the 27 applications listed in Table 3; and (iii) scripts used to process the data and generate the tables in the paper.
B
Ethical Considerations
At the time of our study, our institution has no institutional review board (IRB) or ethics review board (ERB). However, we complied with strict institutional and (inter-)national data protection and privacy regulations. The research team included two experts from the Human-Centered Security department, who supervised the planning and ensured compliance with ethical standards for studies that involve humans. We used a consent form, which our data protection officer had previously verified. The consent form covered all information that would usually be required for IRB/ERB approval. Participants were informed that participation was voluntary and that they could withdraw from the study at any time without any consequences. Our contact information was provided in case participants had any questions regarding data protection or the survey. After identifying relevant projects, we used email addresses from GitHub to contact the experts and ask for participation in our study. We paid participants €50 for the 25-minute interview, which corresponds to €120 per hour. Before recording, all participants were reminded of the privacy policy and asked if they had any questions. All interviews were transcribed locally and then anonymized to ensure that no sensitive information was published.
C
Interview Guide
1. Onboarding • Introduction of interviewers and the institution • Motivation: Exploring the reproducibility of open-source SGX applications • Introduction of the interviewee • Clarification of any questions regarding the consent form and data protection • Beginning audio recording (with consent) 2. Narrative Prompt We’ll begin with your general thoughts on reproducibility, especially in the context of your most recent SGX project. (1) Introduction Reproducibility 14
• What does reproducibility mean to you when working with SGX? • How did you first learn about reproducibility? (2) Clarifying Terminology • For this study, reproducibility refers to the ability to reproduce the enclave hash, linking source code to its binary output. (3) Reproducibility in Your Current Project • Did you consider reproducibility in development? • How important is reproducibility in this project? Why? – Is it important to your users, community, or other stakeholders? • What steps have you taken to support reproducibility? • Have you tested for reproducibility? If so, how? • How much effort are you willing to invest (or have already invested) in ensuring reproducibility?
• Why do you think reproducibility is not widely addressed? – Who could make reproducibility a priority? • Have you worked with other TEEs (e.g., TrustZone, TDX, SEV)? – Have you encountered and addressed reproducibility challenges in those platforms? – What are the key differences between SGX and other TEEs regarding reproducibility? 4. Improving Reproducibility In this final section, we’d like to explore ideas and solutions to support more reproducible enclave builds.
3. Challenges in Reproducible Enclave Builds Let’s now discuss obstacles you may have encountered in building reproducible TEE applications. (1) Project-Specific Challenges • What challenges have you faced in achieving reproducibility in your current project? (2) Wider Context • We observed that reproducibility is often problematic across many open-source SGX projects. (3) General and Cross-TEE Challenges • Have you experienced reproducibility issues in other projects? Which?
• How would you address the challenges you mentioned? • One idea we’re exploring is using Docker containers to simplify reproducibility testing. Would this be feasible for your project? Do you think it would help others? • Through which channels would you share reproducibilityrelated information (e.g., Intel’s website, GitHub, documentation)? • How can the research community better support developers in achieving reproducibility? 5. Offboarding • Do you have any final questions or topics you’d like to discuss? • Concluding the session and stopping the recording
D
Emails & Codebook
Interview recruiting email Subject: Chat on reproducibility (SGX) - Academics need help Dear Recepient, We are researchers from Institution, Country, currently conducting a study on the reproducibility of open-source projects that use Intel SGX. While exploring various projects on GitHub, we found that the majority of SGX-based projects were not easily reproducible—a challenge we are aiming to understand more deeply. To that end, we are reaching out to project maintainers, developers, and contributors like yourself to hear your perspective on reproducibility—its relevance, challenges, and how it fits into your development process. Even if you are not familiar with reproducibility, or if it was not a specific focus for your projects, your input would still be extremely valuable to our study. Would you be available for a brief interview (∼25 minutes) at your convenience? Our goal is to support the open-source community in building more secure, reliable, and trustworthy software. In appreciation, we will provide you with a 50€ (∼$58) gift card from [xxxxx] after the interview, which you can redeem for vouchers covering a broad selection of services. If you are interested, please fill out a short (∼5 minutes) pre-survey and book a slot for the interview here: link-to-pre-survey. If you have any questions, feel free to reply to this email. Best regards, Author names
Figure 3: Email to SGX project developers and maintainers for interview regarding reproducibility challenges.
15
Table 7: Codebook. (*) denotes a container for sub-codes and therefore is not used during coding. The themes map to the discussion in Section 5.2 and Section 5.3: Organizational Barriers (OB), Known Issues (KI), and Security is a Badge (SB). Code TEE Technologies (*)
Description Example Quote Theme Subcategories are used for cross-code correlations to support sensemaking and — the structuring of results. Intel SGX Coded if SGX or related technologies are explicitly mentioned. “But yeah, SGX is part of that, but also AMD’s TEEs as well.” Intel TDX Coded if Intel TDX or related tech. are explicitly mentioned. “But when you go to SEV-SNP or on TDX, it’s, it’s more complex [...]” AMD SEV Coded if AMD SEV or related tech. are explicitly mentioned. “I have a reasonably good estimate actually for SNP because this was, this is work that was mostly done by an intern recently. [...]” Arm TrustZone Coded if Arm TrustZone or related tech. are explicitly mentioned. “I mean, so the TrustZone, ArmCCA has attestation. [...]” Arm CCA Coded if Arm CCA or related tech. are explicitly mentioned. “ [...] it cannot do attestation. So CC can, ArmCCA. [...]” Perceptions of Reproducibility (*) General attitudes, beliefs, and understandings of what reproducibility means in — the participants’ context. General Relevance of Reproducibility This code captures participants’ opinions, including statements that directly assess “And so I think it’s really important that you have the ability to do that.” SB or describe the relevance of reproducibility, as well as those that imply its perceived importance indirectly. Definition of the Concept Statements in which participants describe how they define or understand repro- “The ability to create the same results from the same source codes [...]” KI ducibility. Prioritization of Reproducibility (*) Descriptions of how reproducibility is actually handled in day-to-day work, — projects, or organizations. Reproducibility is a priority Situations where reproducibility is treated as an explicit goal or priority, including “It’s like one of our major goals.” SB dedicated efforts or resources. Reproducibility not prioritized Situations where reproducibility is downplayed, neglected, or consciously traded “They know this is a problem, but they don’t care about this. They don’t SB want to invest manpower.” off against other concerns. Relevance of Reproducibility (*) Explanations of why reproducibility matters (or does not) in specific contexts, — including stakeholders and use cases. Parties Requesting Reproducibility Mentions of actors who demand or expect reproducibility “Mostly customers.” SB Parties Who Can Make Statements about people or roles who have the power to prioritize or enforce “In the end, it’s the customers, right.” SB Reproducibility a Priority reproducibility. Use Cases Where Concrete situations or scenarios in which reproducibility is particularly critical “during remote attestation and all that [...]” SB Reproducibility is Important (e.g., incident response, certification, debugging). Challenges (*) Descriptions of obstacles and problems that hinder reproducibility. — Challenges in Achieving Specific factors that make it difficult to obtain reproducible builds. — Reproducibility (*) Timestamps Problems caused by time-dependent data (e.g., embedded timestamps) that lead “You also provide a build timestamp, which does change.” KI to non-identical outputs. (Absolute) paths Issues due to absolute or environment-specific file paths affecting build or runtime “Paths is a really common issue. For example [...]” KI outcomes. Signatures are Challenges where cryptographic signatures or similar artifacts vary between “[...] you have signatures that you have embedded in the binary that can KI non-deterministic builds and break reproducibility. also affect reproducibility.” Lack of awareness Descriptions of situations where stakeholders are unaware of reproducibility “So this is the first time I’m actually aware of this being an issue in SB issues or their implications. production.” High complexity Difficulties stemming from complex systems, toolchains, or configurations that “SGX development is notoriously complex.” KI are hard to control or replicate. Hotfix release before fix Problems caused by rapid hotfixes or ad-hoc releases that bypass reproducible “[...] if you find a security issue and you immediately patch it in your OB TEE’s and you immediately release the source code for reproducibility, processes or documentation. you are potentially making things easier for attackers [...]” Debugging NonStatements emphasizing that diagnosing the causes of non-reproducible behavior “the hardest thing I think for reproducibility is the debugability of it” KI is time-consuming or technically challenging. Reproducibility is hard Limited Control of Difficulties arising from partial control over hardware, OS, toolchain, or external “in my experience a lot of moving parts, system or program which is OB services that affect reproducibility. complex enough has so many moving parts, it’s pretty hard to count them the Environment all in [...]” Challenges are often Observations that many obstacles stem from general software-engineering issues “They’re not really specific to SGX or to confidential computing in OB not TEE specific rather than TEEs themselves. general.” Transferrability to Concerns about whether reproducibility-related solutions or problems transfer “I think basically it’s the same problem” OB across different TEE platforms. other TEEs Dealing with challenges (*) Ways in which participants respond to or try to manage the above challenges. — Practices to Achieve Reproducibility Concrete methods, workflows, or tools used to improve or enforce reproducibility “build the enclave inside the same Docker container environment on both KI in practice. the Attester and Relying Party sides” Suggestions for Improvement Proposed ideas or wishes for improving reproducibility. — of Reproducibility (*) Technical Improvements Suggestions for technical changes (e.g., tooling, build systems, TEE features) that “pinning the Docker, like pinning this strict version, I mean, that’s what KI would facilitate reproducibility. SBOM is for” General Desirable Broader improvement suggestions, e.g., better processes, documentation, or orga- “we should increase the awareness. [...]” OB, SB Improvements nizational practices around reproducibility. What Research Can Do Expectations or wishes regarding contributions from academic or industrial re- “Having researchers directly join projects as maintainers could improve OB, SB search to support reproducibility. reproducibility quality [...]” Channels for Spreading Mentions of how knowledge about reproducibility should be disseminated (e.g., “For more general concerns, I often use technical blogging platforms [...] SB or my company’s engineering blog.” Reproducibility-related documentation, talks, communities, standards). Information
16