Conceptio › Archive › arXiv CS
arXiv CSopen access

Secrets That Survive Everything: Runtime Credential Exposure in Production Web Applications

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Secrets That Survive Everything: Runtime Credential Exposure in Production Web Applications

arXiv:2609.23042v1 [cs.CR] 19 Sep 2026

Hemanth Gorijala∗

Abstract Pre-deployment secret scanning operates only on source code, never on what a production application serves. We document two exploitation chains in which Azure AD client credentials and APIM subscription keys from production JavaScript bundles enabled account takeover and mass data exposure. An authorized engagement covered approximately 2,000 enterprise web assets in one organization; 113 (5.65%) served live credentials. To quantify the shiftright gap, we built an independent Ground Truth (GT-194) of 194 secret-grade credentials through Claude Opus 4.7 extraction and manual analyst review, with the 247 LLM-extracted candidates independently validated by GPT-5.5 (Brennan-Prediger κ = 0.676). The principal finding is structural: 13.9% of GT-194 (27 of 194) is surfaced only by manual analysis and recovered by none of the nine evaluated production scanners, a tool-agnostic blind spot the ground-truth model also misses. CryptoJS encrypted configuration separately defeats every static scanner: the credential exists only after decryption with a co-located key, reached only by runtime-aware detection. Combined coverage plateaus at 86.1%. Among the nine scanners, the best static scanner recovers 36.6% and the best runtime-aware scanner 77.8% (F1 = 0.818, McNemar p < 0.001); the ground-truth model is reported separately as a reference comparator, not an evaluated detector. On 63 of 86 secret-exposed applications (73.3%), the full Azure AD token-mint chain is co-located in one bundle, reachable from browser code. We characterize five paths by which credentials reach production undetected and present a layered runtime detection methodology and remediation framework. Recall is scoped to a single-organization Azure-heavy corpus (Section 10).

Index Terms— runtime secret detection, credential exposure, JavaScript security, webpack, shiftleft security, shift-right gap, API key exposure, web application security, penetration testing, Azure AD, APIM, CryptoJS, LLM-assisted ground truth, McNemar test, UpSet plot

1

Introduction

Despite mature shift-left tooling for secret detection [8, 16], live credentials continue to appear in the production-served content of enterprise web applications. Static secret scanners (GitLeaks, TruffleHog, GitHub Advanced Security) and SAST tools (Semgrep, SonarQube, Checkmarx) operate on source code repositories and build configurations and do not examine what a deployed application serves to its users. Credentials that reach production through build-time substitution, CI/CD pipeline variable injection, runtime-fetched configuration, third-party script inclusion, or ∗ Independent Security Researcher, ORCID 0009-0006-9810-4001. E-mail: [email protected]. Published in IEEE Access, 2026. DOI: 10.1109/ACCESS.2026.3734984. Open access under CC BY 4.0. An earlier version was posted as a preprint at Zenodo (DOI: 10.5281/zenodo.19464446); this manuscript is substantially extended, adding the GT-194 benchmark and the nine-scanner evaluation.

1

deliberate scanner suppression therefore enter served content without crossing any layer those tools cover. This paper characterizes the resulting exploitation surface and quantifies the corresponding shift-right tooling gap. Prior population-scale work established that browser-delivered credential exposure exists; the contribution here is the complementary per-tool recovery benchmark and runtime-detection analysis for enterprise JavaScript bundles, not the existence of the phenomenon. Stated precisely, the specific problem addressed is: credentials that reach the production-served layer are not examined by any repository-stage detector, and the extent to which detectors operating on served content recover them has not been measured. We therefore ask three questions on a real enterprise corpus: (1) by what structural paths do live credentials reach served content without crossing a repository-stage scanner; (2) how much of a served-content credential ground truth do current static, template-driven, and runtime-aware detectors actually recover; and (3) how weaponizable are the credentials that do reach served content, measured by end-to-end exploitation chains. An authorized security engagement covered approximately 2,000 enterprise web application assets across one organization. Of these, 113 (5.65%, 95% Clopper-Pearson exact CI: 4.7%–6.8%) contained at least one live credential in served content (JavaScript bundles, HTML view-source, or JSON and XML API responses) that had bypassed every pre-deployment secret scanning control in place. Of those 113 vulnerable applications, 63 (55.8%, 95% Clopper-Pearson exact CI: 46.1%– 65.1%) contained a complete Azure AD credential set (client ID, client secret, tenant ID, and resource URI co-located in the same JS bundle) sufficient to execute the full exploitation chain documented in this paper. Not every credential exposure produces an exploitable end-to-end chain. A complete chain requires both the full credential set and over-permissive service principal scopes, as documented in Sections 3.4 and 4.3. Applications were selected based on engagement scope, not screened for likelihood of exposure. Confidence intervals are Clopper-Pearson exact intervals at the 95% level. To measure the shift-right tooling gap on the same corpus, we constructed an independent Ground Truth (GT-194) of 194 unique secret-grade credentials extracted from production JavaScript bundles across 113 enterprise applications. The benchmark is anchored to extraction by Claude Opus 4.7 (Anthropic) with no access to any of the evaluated tools, and its 247 LLMextracted candidates were independently validated by an LLM from a different vendor (GPT-5.5, OpenAI; 207 concordant, Brennan-Prediger κ = 0.676), with the 27 manual-only additions verified by analyst review. After human review, manual analyst additions, and the field-deployment register of 8 May 2026, 194 unique secret-grade credentials were locked as the benchmark across eight credential types (Azure APIM subscription key, Azure AD client secret, CryptoJS-AES blob, plaintext user credential, JWT, CyberArk AIM, Google API key, and Other API key/token including App Insights iKey). Nine production scanners were evaluated against GT-194 at their tightened configurations, with Claude Opus 4.7 reported separately as an LLM-assisted groundtruth reference comparator (it co-constructed the benchmark and therefore cannot be scored as an independent detector). The evaluation uses pairwise McNemar significance testing with HolmBonferroni correction, an UpSet-style detection-set overlap analysis, and a separately tracked set of 249 chain-completion identifiers that are public-by-design and therefore not counted in the recall denominator. Section 8.5 reports the full results. The remainder of this paper is organized as follows. Sections 3 and 4 document two end-to-end exploitation chains drawn from the corpus. Section 5 characterizes the five structural paths by which credentials reach production undetected. Sections 6 and 7 analyze why both shift-left and shift-right tooling categories systematically miss these paths. Section 8 presents a layered runtime detection methodology and the GT-194 evaluation. Sections 9 through 12 cover recommendations, limitations, future work, and responsible disclosure. 2

1.1

Scope and Authorization

Both exploitation chains described in this paper were identified during authorized security assessments. Organization names and application identifiers have been anonymized throughout.

1.2

Contributions

This paper makes the following contributions: 1. Two documented exploitation chains, end-to-end, reproducible attack paths in which Azure Active Directory client credentials and API subscription keys found in production JavaScript bundles were combined with over-permissive service principal scopes to achieve full account takeover and mass user data exposure in real enterprise applications that had passed all deployed shift-left security controls. 2. A structural taxonomy of five paths to production (build-time environment injection, CI/CD pipeline variable substitution, runtime configuration fetching, third-party script inclusion, and scanner suppression at the organizational level), each of which produces live credentials in deployed artifacts without those credentials appearing in the application owner’s repository at any point. 3. A characterization of the tooling gap. A systematic analysis of why both shift-left scanning tools (GitLeaks, TruffleHog, SAST) and shift-right tools (DAST, WAF, RASP) fail to detect this class of vulnerability, grounded in the specific technical properties of minified production JavaScript. 4. A detection methodology for the runtime layer. A layered approach combining anchored vendor token patterns, Shannon entropy analysis, and key-value context scanning, designed for the constraints of production JavaScript: minification, noise suppression, and no outbound verification calls. 5. A prioritized remediation framework, covering immediate credential rotation, architectural remediation via the Backend for Frontend pattern, service principal scope correction, and ongoing monitoring controls for Azure AD and Azure API Management environments, with generalizations to AWS and GCP stacks. 6. A cross-vendor LLM-validated benchmark (GT-194) and a quantified toolagnostic blind spot. Nine production scanners are evaluated against a locked benchmark assembled from independent Claude Opus 4.7 extraction and manual analyst review, with its LLM-extracted candidates independently validated by GPT-5.5 (Brennan-Prediger κ = 0.676, 207 of 247 concordant); Claude Opus is reported separately as an LLM-assisted groundtruth reference comparator rather than as an evaluated detector, since it co-constructed the benchmark. The evaluation uses pairwise McNemar significance testing with Holm-Bonferroni correction and UpSet-style detection-set overlap analysis. The principal finding is structural: 13.9% of GT-194 (27 credentials) is recovered by none of the nine production scanners and is surfaced only by manual analysis. Separately, the CryptoJS encrypted-configuration class defeats every static scanner (the credential exists only after decryption with a co-located key) but is recovered by runtime-aware detection. Methodological mitigations of the developerauthor conflict of interest are detailed in §8.5; recall figures are scoped to this corpus rather than cross-industry estimates (Section 10).

3

1.3

Related Work

Secret leakage in version control. The leakage of credentials through public repositories is well-documented. Meli et al. [8] conducted a large-scale measurement study of public GitHub repositories, identifying thousands of exposed API keys, passwords, and tokens across millions of commits, and established that secret leakage in source code is a systemic problem providing the empirical foundation for tools such as GitLeaks and TruffleHog. Demir et al. [15] extended this analysis to live web content, crawling 10 million web pages and identifying 1,748 distinct credentials from 14 cloud and SaaS providers embedded in JavaScript, HTML, and JSON resources, finding that 84% of credentials appeared in JavaScript files and persisted for an average of twelve months. The present paper addresses a complementary structural problem: credentials that reach production without ever appearing in any repository, through build-time injection, pipeline variable substitution, and runtime configuration delivery. Industry-scale measurement of credentials in served JavaScript. Intruder’s December 2025 measurement [28] applied a Nuclei-based JavaScript spider to ≈5 million applications and identified ˜42,000 exposed tokens across 334 secret types, providing the first industry-scale lowerbound on credentials in served JavaScript. The present paper is methodologically complementary: a smaller per-application benchmark with cross-vendor LLM-validated ground truth, designed to quantify the structural detection gaps a regex-driven sweep cannot reach. Secret scanning tool evaluation. Basak et al. [16] conducted a comparative evaluation of nine secret detection tools (including GitHub Secret Scanner, Gitleaks, SpectralOps, and TruffleHog) against a benchmark dataset of real-world credential leaks, providing precision and recall measurements and identifying the primary sources of false positives and false negatives across tool categories. Their work establishes the limits of existing scanner coverage. This paper extends that analysis to a scanning layer those tools were not designed to address: the runtime-served application surface. Credential leakage in mobile and mini-app ecosystems. Shi et al. [17] identified systematic credential exposure in mini-applications hosted within super-app platforms such as WeChat, finding 15 categories of vulnerable services in which credentials embedded in mini-app bundles enabled account hijacking and phishing. While the delivery mechanism differs from web SPA bundles, the structural cause is identical: build-time credential injection into client-delivered artifacts. The exploitation patterns and the tooling gap are the same. JavaScript bundle security. Rack and Staicu [18] conducted an empirical study of JavaScript bundling practices across large-scale web deployments, analyzing bundle composition, dependency inclusion, and the security implications of bundler behavior. Their analysis of AST-level bundle reversibility is directly relevant to the detection methodology described in Section 8: understanding how bundled JavaScript can be reverse-engineered to recover credential context informs both attacker technique and defender detection strategy. Lauinger et al. [19] documented the widespread inclusion of outdated and vulnerable JavaScript libraries in web applications, establishing the broader pattern of client-side JavaScript as an understudied attack surface. Static analysis tool limitations. Brito et al. [20] evaluated JavaScript static analysis tools against a dataset of 957 real-world vulnerabilities in npm packages, finding significant gaps in detection coverage across tool categories. Their findings on the limitations of AST-based static analysis for JavaScript are consistent with the shift-left scanning gaps documented in Section 6: static analyzers operating on pre-build source code cannot detect credentials that are introduced at build time or delivered at runtime. Client-side JavaScript as an attack surface. Prior work on client-side JavaScript security has focused primarily on XSS vulnerabilities [21], third-party script inclusion risks, and subresource 4

integrity enforcement. DAST tooling has evolved to cover injection vulnerabilities and authentication flaws in running applications. Neither category was designed to identify credentials embedded in served JavaScript bundles or injected into HTML at runtime. This paper maps that gap and presents a detection methodology suited to the runtime attack surface, addressing the structural limitations identified across the related work reviewed above.

1.4

Terminology

• Shift-left: Security controls applied before or at the point of code commit, secret scanning, SAST, pre-commit hooks • Shift-right: Security controls applied after deployment, DAST, runtime monitoring, WAF • SPA: Single-Page Application. A web application where the entire frontend is delivered as a JavaScript bundle loaded once by the browser • APIM: Azure API Management. A gateway service that proxies and manages access to backend APIs • BFF: Backend for Frontend. A server-side proxy layer owned by the frontend team that holds service credentials and proxies requests to backend APIs

1.5

Threat Model

Attacker profile and capabilities. This paper models a network-capable attacker with no prior authentication to the target application, the same capabilities available to any web visitor, bug bounty researcher, or penetration tester. The attacker can load the application in a browser, inspect served HTTP responses including JavaScript bundles and JSON API replies, and issue authenticated API calls using any credentials recovered from those responses. The attacker does not need network interception capabilities, special tooling, or insider access. The attack surface is entirely passive: credentials are read from content the application actively delivers to every visitor. What is in scope. Credentials embedded in client-delivered artifacts (JavaScript bundles, HTML source, statically served JSON configuration files, and JSON or XML API responses) that are accessible without prior authentication. Service credentials (Azure AD client secrets, APIM subscription keys) that are structurally reachable by this attacker model and that grant API access beyond what the current user session is authorized for. The exploitation chains in Sections 3 and 4 operate entirely within this model. What is out of scope. Server-side credential stores (Azure Key Vault, environment variables not reflected in served content, database connection strings not returned in API responses). Attacks requiring network interception, active injection, or compromise of infrastructure components. Social engineering, phishing, and supply chain attacks. The structural gap documented in this paper is not a network-layer vulnerability. It is a consequence of credentials materializing in client-accessible artifacts through the deployment paths described in Section 5.

2

Background, Azure AD, APIM, and Why These Credentials Matter

2.1

The Azure AD + APIM Architecture

Azure API Management acts as a reverse proxy gateway sitting in front of backend APIs. It handles authentication and authorization by validating tokens and subscription keys, rate limiting, request 5

and response transformation, and analytics. The typical flow in an Azure-native SPA application is:

Figure 1: Azure AD + APIM request flow. The browser SPA authenticates via Azure AD using client credentials and calls backend APIs through the APIM gateway using both a Bearer token and a subscription key. APIM enforces access through two separate credential types: Azure AD tokens (JWTs issued by the identity provider) and subscription keys (gateway-specific pass keys scoped to API products) [7]. Both are required to reach protected endpoints.

2.2

The Four Credential Values

In the applications examined, four values were present in the client-side JavaScript bundle: Table 1: The four Azure AD and APIM credential values present in the client-side bundle, their purpose, and whether each legitimately belongs in the browser. Credential

Purpose

Should it be in the browser?

AppID (Client ID) AppKey (Client Secret) Resource SubscriptionKey

OAuth2 client identifier registered in Azure AD. Tells the identity provider which application is authenticating Password paired with AppID. Used in confidential client flows (server-to-server) API resource URI or scope used during token acquisition APIM gateway pass key (Ocp-Apim-Subscription-Key header) required to reach APIM-protected APIs

Yes, public by OAuth2 spec for public clients Never Harmless alone, combined with AppKey enables full token generation Never

The AppID is legitimately public in OAuth2 public client flows. The AppKey is not. It is the client secret for a confidential client flow, designed for server-to-server authentication where the application code is not visible to end users [1, 4]. Client secrets are architecturally valid for server-side web applications. The problem is not that the secret exists. It is that it appeared in browser-accessible JavaScript, where it is visible to every visitor who opens developer tools. A browser-based SPA is a public client and should not be issued a client secret at all; where it must authenticate users, the correct pattern is the OAuth 2.0 Authorization Code flow with Proof Key for Code Exchange (PKCE) [29], which requires no client secret (see Section 9). OAuth2 grant types operate independently on a single application registration. A client secret present on the registration activates the client credentials grant unconditionally, the authorization server applies no constraint between grant types, so the intended authorization flow does not restrict which grants the server will accept [1].

6

2.3

Why All Four Values Together Is Critical

The AppKey alone is insufficient: it requires the AppID and resource URI to generate a token. The SubscriptionKey alone is also insufficient, since API endpoints require a valid Bearer token. Individually, each value has limited reach. Together, they represent a complete authentication package: the ability to authenticate as the application itself and call any API the application is authorized to access. The tenant ID (required to construct the token endpoint URL) was visible in each application’s login redirect URL, hardcoded alongside the other values in the same bundle.

3

Exploitation Chain 1, Azure AD + APIM

3.1

Discovery

The client-side JavaScript bundle of a production application (Azure AD authenticated, sitting behind an API Management gateway, serving thousands of users) contained all four credential values described in Section 2.2. The bundle was served over HTTPS to every visitor of the application with no authentication required to retrieve it. The tenant ID was present in the application’s login redirect URL, which was also hardcoded in the bundle.

3.2

Token Generation

The Azure AD token endpoint was called using the extracted credentials: POST https://login.microsoftonline.com/{tenant_id}/oauth2/token grant_type=client_credentials client_id={AppID} client_secret={AppKey} resource={Resource}

The endpoint returned a fully authenticated Bearer access token. The grant type used (client_credentials) is a confidential client flow designed for server-to-server communication where the client secret is stored securely on the server [1]. (The v1 endpoint format is shown, the v2 endpoint uses /oauth2/v2.0/token with a scope parameter in place of resource. The difference is that v1 identifies the target API by a single audience URI (resource), whereas v2 follows the OAuth 2.0 convention of naming granular per-permission scopes (for example api://<app-id>/.default); the underlying client-credentials grant is identical. Both endpoints remain in active use across enterprise Azure environments.) The application was using it in a public client context where the secret was visible to every user.

3.3

API Endpoint Reconstruction

Rather than scanning the bundle for additional secrets, it was read as a map of how the application was built. Buried in the minified code were API endpoint definitions, GET and POST endpoints with complete schema structures showing exactly which parameters each endpoint expected. The frontend had documented its own backend. Using the Bearer token and APIM subscription key together:

7

Authorization: Bearer {token} Ocp-Apim-Subscription-Key: {SubscriptionKey}

The reconstructed endpoints were called. User profile data returned immediately. A password reset endpoint was identified using the same schema reconstruction approach, same token, same subscription key, same reconstructed schema. The finding was reported at this point and the assessment did not proceed further.

3.4

The Scope Condition

The exposed credentials alone constitute a significant finding regardless of what they unlock. A client secret in a public JavaScript file means any visitor can authenticate as the application, the blast radius depends entirely on what permissions that application has been granted. In this case, the amplifying factor was that the service principal had been granted over-permissive API scopes. Application permissions (acting as the application itself, not as a specific user) had been granted where delegated permissions (acting as the signed-in user) would have been appropriate and sufficient [3, 4]. The result: the application’s client identity had been given the ability to perform user-level operations across every account in the system. Without this misconfiguration, the Bearer token would have had access only to application-level resources. With it, the token was effectively an administrative credential for every user account.

3.5

Subscription Key Blast Radius

One APIM subscription key typically maps to a product containing multiple APIs: SubscriptionKey -> APIM Product "Internal APIs" |-- /users/* |-- /payments/* |-- /admin/* ‘-- /reports/*

One exposed subscription key grants access to every API in that product. The scope of exposure is determined by how the APIM product is configured. A detail that is not visible from the client side and must be assessed in the APIM portal. Result: Full account takeover from four values in a JavaScript file served to every visitor of the application. This was not a sophisticated attack. It required no exploit framework, no vulnerability scanner, and no special tooling. It required reading the JavaScript file the application was already serving to everyone and understanding what the credentials unlocked.

4

Exploitation Chain 2, Client-Side Encryption Bypass

4.1

Discovery

A different application, a different codebase, and a different obfuscation approach, but the same Azure AD credential pattern waiting at the end. The credentials reached production via Path 1 (Section 5.1): Azure AD client credentials were stored in an environment file and baked into the JavaScript bundle at build time. Rather than

8

leaving them as plaintext, the development team had encrypted the configuration object using CryptoJS before bundling. A pattern intended to obscure the credentials from casual inspection. The developer intent is that the configuration is protected because it is encrypted. The problem: the decryption key was hardcoded in the same JavaScript file, three lines away from the encrypted string [2, 12].

4.2

Decryption

The hardcoded key was used to decrypt the environment string using the same CryptoJS call visible in the source: CryptoJS.AES.decrypt(encryptedConfig, hardcodedKey).toString(CryptoJS.enc.Utf8)

The output was a complete configuration object. Every Azure credential, every service key, every internal endpoint the application needed to function. The encryption had provided exactly zero protection. The key to unlock everything was sitting next to the lock. The decrypted object contained the same pattern: AppID, AppKey, Resource, Azure AD client credentials embedded in what the development team believed was a secured configuration.

4.3

Exploitation

As with Chain 1, the service principal had been granted over-permissive scopes, Application permissions allowing an application-level token to perform user-level data access operations. From the decrypted credentials: • A Bearer token was generated using the extracted Azure AD client credentials • API endpoints were located hardcoded elsewhere in the bundle • Endpoint schemas were reconstructed from the minified bundle structure • Authenticated GET endpoints were called with the Bearer token • Full user profile data (names, email addresses, account identifiers) was returned for any user in the application Result: Personal data for thousands of users was accessible to anyone who opened the JavaScript file and understood what the encrypted string was hiding.

4.4

Why CryptoJS Obfuscation Is Particularly Dangerous

The CryptoJS pattern creates a false sense of security that is difficult to identify in a standard code review. Developers implement it believing the configuration is protected. Security reviewers see encryption and move on. Static secret scanners see an encrypted string and find nothing to flag. The construct is meaningless only when you read the entire file rather than scan it for plaintext secrets, which is exactly what static scanners do not do. This pattern is not isolated to a single application or development team. Across the 113application corpus assessed in this study, the CryptoJS obfuscation construct (encrypted configuration object with a co-located decryption key) recurred across multiple unrelated codebases, with 16 applications (14.2% of vulnerable apps) containing CryptoJS-encrypted configurations and 37 distinct CryptoJS-AES encrypted-configuration blobs in the GT-194 benchmark. The recurrence across unrelated codebases suggests the pattern originates from a shared internal library, a shared development template, or a common architectural recommendation propagated across teams. A single flawed pattern adopted at the architecture or template level can introduce the same cryptographic misuse across an entire application portfolio simultaneously. 9

5

How Secrets Reach Production

Credentials reach production undetected through structural paths in modern build and deployment pipelines, regardless of organizational security maturity. Shift-left tools do not fail in these scenarios. The path from source code to production contains branches that no pre-deployment scanner covers. Five such paths are characterized in Sections 5.1 through 5.5.

5.1

Path 1, Build-Time Environment Injection

React applications using REACT_APP_* environment variables, and Angular applications using environment.prod.ts, pass credentials to the build tool at compile time. webpack’s DefinePlugin and the Angular CLI substitute these values directly into the output bundle [6]. The credential never exists in the repository. It is read from the build environment at compile time and written into the output artifact. Every secret scanner that ran before the build completed saw nothing, because there was nothing to see. The Angular CLI’s production build pattern deserves particular attention. The 113 credentialbearing applications break down by framework as follows: Table 2: Front-end framework distribution across the assessed application corpus. Framework

Apps

Share

Angular (CLI) React REST API / JSON config Legacy / jQuery ASP.NET WebForms React (CRA)

75 25 5 4 2 2

66% 22% 4% 4% 2% 2%

Angular’s dominance (66%) is not coincidental. The Angular CLI treats environment.prod.ts as the documented and recommended approach for environment-specific configuration, credential injection is the framework’s intended pattern. webpack’s DefinePlugin structurally ensures that any value placed in the environment file is compiled directly into the output bundle. The prevalence of credential exposure across this portfolio is a direct consequence of framework architecture, not isolated developer error. Where a framework’s recommended pattern produces credentials in clientdelivered artifacts by design, the risk is systemic rather than individual.

5.2

Path 2, CI/CD Pipeline Variable Substitution

This is a frequently observed path. A placeholder value lives in the repository while the real credential is stored as a pipeline variable in Azure DevOps, GitHub Actions, or a similar CI/CD platform. The placeholder is what every repository scanner, pre-commit hook, and SAST tool sees. The credential materializes only in the build artifact, after every scanner has already completed. It never exists in git at any point.

10

Figure 2: Two structural paths by which credentials reach production without repository exposure. Left, build-time environment injection: the credential is injected by webpack or the Angular CLI at compile time and never exists in git. Right, CI/CD pipeline variable substitution: a placeholder lives in source while the real credential is written into the artifact only after every scanner has completed.

5.3

Path 3, Runtime Configuration Injection

Some applications fetch their configuration after the browser loads the initial bundle: • A /assets/config.json file served statically at runtime alongside the application • A window.__APP_CONFIG__ object injected into index.html by the server at request time • SSR state blobs (the __NEXT_DATA__ object injected by Next.js or window.__INITIAL_STATE__ injected by Nuxt) which regularly contain tokens, API endpoints, and internal service configuration These configurations are loaded after deployment, sometimes from CDN nodes, sometimes from origin servers. They do not exist as files in the repository. They are assembled and served at runtime, invisible to any scanner that ran before deployment. In the applications assessed for this study, two targets were REST API configuration endpoints (not JavaScript bundles) returning credential-bearing JSON responses directly to the browser. In both cases the endpoint was called by the SPA on load to retrieve runtime configuration, and the response contained Azure AD client credentials in plaintext JSON fields. These endpoints were publicly accessible without authentication. This confirms that the runtime attack surface extends beyond JavaScript files to include configuration APIs that serve secrets to the browser on demand, and that no JavaScript-only scanner would detect credentials delivered through this path.

5.4

This Pattern Extends Beyond Azure

The exploitation chains documented in this paper involve Azure AD and Azure API Management because that is the stack present in the applications assessed. The same structural paths appear across other providers. The remediation guidance in Section 9 is Azure-specific and should be adapted to the relevant platform. The same paths appear across every stack: • AWS: Cognito user pool client IDs and client secrets embedded in React bundles via REACT_APP_* variables. API Gateway API keys substituted into JavaScript at build time. S3 presigned URL generation credentials hardcoded in SPA configuration. 11

• GCP: Firebase API keys and project configuration objects (apiKey, authDomain, projectId) served in firebase-init.js to every visitor. These are legitimately public for Firebase Authentication but are frequently accompanied by service account credentials that are not. • Generic SaaS: Twilio account SIDs and auth tokens, SendGrid API keys, Stripe secret keys, and Mailgun API keys substituted into build artifacts from CI/CD pipeline variables. The credential format changes. The path to production does not. The exploitation technique is identical across all of these: read what the application serves, identify the credential, authenticate with it. The shift-left tooling gap is identical: none of these credentials touched a repository at any point in their path to the browser.

5.5

Credentials Persist Across All Environments

Because credential injection occurs at build time, every environment that runs a build receives the same secrets baked into its artifact. Development, SIT, UAT, pre-production, and production environments all receive credentials through the same pipeline, and all produce the same exposure. Across the applications assessed in this study, the same credential set was confirmed present in production, UAT, SIT, and pre-production environments of the same application in multiple cases. Lower environments typically have weaker access controls, broader team access, and no security monitoring, making them at least as attractive a target as production for an attacker who has identified the pattern. This has a direct consequence for remediation: rotating credentials in production without simultaneously fixing the build pipeline leaves every lower environment still exposed. An attacker who has already extracted credentials from a UAT environment retains access regardless of what happens in production. Remediation is only complete when the credential is removed from the build pipeline and rotated across all environments simultaneously.

5.6

Path 4, Third-Party Script Inclusion (Supply-Chain Credentials)

Modern web applications routinely embed third-party scripts. Analytics platforms, tag managers, customer-support widgets, error monitoring SDKs, advertising libraries, and CDN-hosted dependencies all execute in the application’s origin and run with full access to its DOM, cookies, and storage. These scripts also carry their own credentials, including API keys, project identifiers, instrumentation tokens, and customer-tenant identifiers that are public-by-design from the third party’s perspective but materialize in the served bundle of every site that integrates them. Two structural problems follow. First, the credentials embedded by third parties are outside the application owner’s repository, so no scanner the owner runs can flag them. They enter the browser through a script tag that the owner did not author. Second, the security posture of those third-party endpoints, including rate limiting, scope enforcement, and audit logging, is determined by the third party rather than the application owner. Demir et al. [15] reported that 16% of verified credential exposures originated from third-party inclusions, with the highest third-party rates observed for OpenAI (24%), Twilio (24%), Mailchimp (23%), Alibaba (23%), Stripe (22%), and SendGrid (22%) credentials served via embedded SDKs that the integrating application did not control. In the corpus assessed for this study, the Google API keys (6 instances in GT-194) and a subset of the Other API key/token class are consistent with this path, integrated through third-party SDKs whose credential material enters the served bundle outside any first-party scanning workflow. The exposure pattern is structurally indistinguishable from build-time or runtime injection. A credential reaches a deployed artifact, is served to every visitor, and has no automated detection on either side of the integration. 12

5.7

Path 5, Scanner Suppression at the Organizational Level

A fifth contributing factor operates at the organizational level and compounds the four structural paths above. Many SPA architectures structurally require credentials in the browser, including Google Maps API keys, Firebase configuration, Stripe publishable keys, and Twilio client tokens. When shift-left scanners flag these, the operational path of least resistance is to suppress the alert through ignore-listing files, dismissing alerts as false positives, adding inline suppression comments, or allowlisting specific patterns in the SAST ruleset. The suppression is deliberate. The developer knows the credential is real but believes it is an accepted architectural risk or lacks the authority or time to implement the server-side proxy pattern (Section 9.1) that would eliminate the in-browser requirement entirely. The credential is live, required, and invisible to the security tooling stack. The build process did not bypass the scanner. The scanner was deliberately told to look away. The outcome is identical to the structural paths. A live credential reaches a production artifact with no automated monitoring.

6

Why Shift-Left Scanners Miss This

The five paths in Section 5 produce credentials in deployed artifacts through different mechanisms but converge at the same boundary. Shift-left scanners run before or at the point of commit, before the build executes, before the pipeline substitutes variables, before any third-party script is fetched, and before the artifact is assembled. The credential does not exist when the scanner runs. By the time it exists in the served artifact, every shift-left scanner has already completed. The table below maps this timing mismatch precisely. Table 3: Where each credential is present across the delivery pipeline, and which scanning layer, shift-left or shift-right, covers each stage. Pipeline Stage

Credential present?

Shift-left scans here?

Shift-right covers here?

Source code / git repository CI/CD pipeline variable store Build artifact (dist/) Third-party SDK delivered at runtime Deployed to CDN / Blob Storage Runtime-fetched config (/assets/config.json)

No, placeholder or absent No, stored server-side Yes, injected at compile time Yes, embedded by integration Yes, served to every visitor Yes, fetched by browser

Yes No No No No No

No No No No No No

The gap is not merely a configuration failure or a tool weakness. It is a structural property of when each scanning layer operates relative to when the credential materializes. More aggressive rule tuning may narrow the static-tool gap on credentials that are present in the artifact (§10), but no amount of pre-deployment rule tuning can reach a credential that only materializes at runtime; closing that portion of the gap requires the scanning layer itself to change. The shift-left tooling market is mature. GitLeaks, TruffleHog, detect-secrets, and GitHub Advanced Security are widely deployed [8]. SAST tools (Semgrep, SonarQube, Checkmarx) are standard in enterprise CI/CD pipelines. Each tool does exactly what it is designed to do. None of them are designed to scan build artifacts, third-party SDK output, or live deployed applications.

6.1

The Minification Problem

Pattern-based secret scanners rely on variable names and known key formats to identify secrets. Minification destroys variable names: 13

Figure 3: Where shift-left scanning stops versus where secrets materialize. Shift-left tooling covers only the source code layer. Secrets introduced through build-time injection, pipeline variable substitution, runtime configuration fetching, or third-party script inclusion exist only in deployed artifacts and served content. These are layers no pre-deployment scanner reaches.

14

Table 4: Representative shift-left scanners, what each scans, and why each misses credentials that materialize after the source-code stage. Scanner

What It Scans

Why It Misses

GitLeaks SAST (Checkmarx, SonarQube, Veracode) Semgrep / CodeQL GitHub Advanced Security

Git commits, staged files, history Source .ts/.js pre-build AST of source code Repository content

Pipeline variables are substituted after git. Placeholder is in git, real value is not. Never sees dist/ Sees process.env.SUBSCRIPTION_KEY, no secret present. Does not execute the build Pre-build source only, same gap as SAST ${{ secrets.KEY }} injected into build output never touches the repository

// Source, scanner identifies this: const subscriptionKey = "abc123xyz789..." // Minified, scanner may miss this: const a={b:"abc123xyz789..."}

Entropy-based detection addresses part of this problem (high-entropy strings are flagged regardless of variable name) but entropy alone generates significant noise against the minified JavaScript of a modern SPA, which contains many high-entropy strings that are not credentials (hashed asset filenames, base64-encoded resources, obfuscated library code).

6.2

Source Maps, The Silent Bypass

Angular and webpack generate .map files alongside minified bundles by default: dist/main.abc123.js dist/main.abc123.js.map

<- minified <- full original source, variable names, logic intact

If source maps are deployed to production (a common default) any credential “hidden” in minified JavaScript is fully readable in browser developer tools via the Sources panel. SAST scans the .ts source. The .map file exposes everything at runtime regardless of what the scanner found. Disabling source maps in production builds (sourceMap: false in angular.json) removes this bypass entirely.

7

Why Shift-Right Scanners Also Miss This

Dynamic Application Security Testing tools receive JavaScript files as part of application spidering, but they treat those files as attack delivery vehicles, not as targets. DAST tools spider the application looking for injection points, authentication flaws, and misconfigurations. They do not parse main.js for Ocp-Apim-Subscription-Key patterns. The shift-right tooling landscape by category: Tools that partially address this class of finding: Table 5: Representative shift-right tools that partially address deployed-secret detection, their approach, and their limitations. Tool

Approach

Limitation

Nuclei TruffleHog (filesystem/web mode) Detectify / Probely GitGuardian perimeter

Community templates matching secret patterns in HTTP responses Entropy and pattern detection on files DAST with secret detection, downloads JS assets Monitors public-facing endpoints

Regex-based, misses minified variable names, template coverage varies Web mode experimental. Reliable on artifact filesystem, not live application scanning Does not execute JavaScript, misses runtime-fetched configurations Focused on public GitHub exposure, not internal deployed applications

None of these tools provide comprehensive coverage of the full runtime attack surface of a modern SPA: split webpack bundle chunks loaded dynamically, SSR state objects injected into 15

Figure 4: Approximate tooling coverage by runtime security category. Mature tooling exists for injection vulnerabilities and network anomalies. Deployed secrets in static assets and served content represent a largely unsolved coverage gap. HTML, JSON and XML API responses returning credentials in fields never meant to be clientfacing, and outbound request headers carrying subscription keys and bearer tokens on every API call. The result: once a secret evades shift-left controls and reaches production, it is effectively invisible to automated tooling. The only things finding runtime secrets in production today are manual penetration testers, bug bounty researchers, and attackers. Two of those three report what they find.

8

What Runtime Layer Detection Requires

Closing the shift-right gap requires a different scanning model than the tools currently deployed. This section describes what a tool operating at the runtime layer must address, the attack surface it must cover, the detection approach required for minified production JavaScript, and the operational constraints that determine whether findings are actionable in practice.

8.1

Attack Surface Coverage

A runtime-layer scanner must cover what the application actually serves, not what exists in its source repository. The full attack surface of a modern web application includes: • JavaScript bundles, including split webpack chunks that modern SPAs load dynamically as users navigate, not just the initial bundle • HTML source, SSR state blobs (__NEXT_DATA__, window.__INITIAL_STATE__) injected by serverside rendering frameworks, which regularly contain tokens and internal configuration [3] • JSON API responses, credentials returned in fields never intended to be client-facing, which static scanners never see • XML responses, enterprise services carrying connection strings and service credentials in structured response bodies 16

• Outbound request headers, subscription keys and bearer tokens sent on every API call in plaintext For bulk assessment, following <script src> references and tracing chunk URLs is required to scan an application’s full JavaScript surface rather than only the files present in the initial page load.

8.2

Detection Approach

Production JavaScript presents a different detection problem than source code. A layered approach is required: Anchored vendor token patterns, Known credential formats from cloud providers (Azure, AWS, GCP), SaaS platforms (Twilio, SendGrid, Stripe), and authentication services have identifiable structure. Anchored regex patterns matched against specific token formats produce fewer false positives than generic high-entropy matching and are resistant to minification because they target the value format, not the variable name. Shannon entropy analysis (For credentials without known formats) internal API keys, session tokens, symmetric encryption keys, Shannon entropy identifies high-randomness strings that are statistically unlikely to be application logic [11]. Entropy analysis catches credentials that minification has stripped of their variable name context. However, entropy analysis alone against minified JavaScript generates significant noise: hashed asset filenames, base64-encoded resources, and obfuscated library code all produce high-entropy strings that are not credentials. Key-value context scanning, Examining the structural context surrounding a high-entropy value reduces noise by requiring that the value appear in a credential-relevant context. A JSON key named subscriptionKey, apiKey, or clientSecret paired with a high-entropy string is significantly more likely to be a credential than an isolated high-entropy string. This context requirement is the primary mechanism for reducing false positives in minified JavaScript where variable names are not available. PII detection with validation, Pattern matching alone is insufficient for PII. Credit card numbers require Luhn algorithm validation [14]. Social Security Numbers require format and checksum verification. Without validation, the false positive rate for these patterns in arbitrary JavaScript is too high to be actionable.

8.3

Operational Constraints

Two operational constraints determine whether runtime scanning findings are actionable in practice: No outbound verification calls. Verifying whether an exposed API key is valid by calling the issuing service generates log entries at the target provider, can trigger security alerts, and reveals the assessment to defenders. Runtime scanning must run locally with no outbound verification calls to third-party APIs or services. Noise suppression is a first-class requirement. A scanner that generates hundreds of false positives per scan is a scanner that gets disabled. In the practitioner workflows where runtime scanning is most needed (penetration testing engagements, bug bounty assessments, security reviews) findings must be trustworthy enough to act on immediately. Every false positive erodes that trust and increases the likelihood that a real credential is dismissed as noise. Noise suppression is not a secondary quality concern. It is what determines whether the tool gets used. The detection methodology described in this section informed the design of SecretSifter, referenced in Section 9.3. Algorithm 1: Runtime Credential Detection

17

Algorithm 1: Runtime Credential Detection Input: URL , target application URL Output: F , set of (credential_type, value, context) findings 1. 2. 3.

F <- {} R <- HTTP_GET(URL, follow_redirects=true) C <- extract_content_units(R) // C = {JS bundles, HTML blobs, JSON responses, headers} for each chunk_url in extract_script_srcs(R.html) do C <- C U {HTTP_GET(chunk_url)} end for

4. 5. 6. 7. 8. for each content_unit u in C do 9. // Layer 1: Anchored vendor token patterns 10. for each pattern p in VENDOR_PATTERNS do 11. matches <- regex_findall(p.regex, u) 12. for each m in matches do 13. F <- F U {(p.credential_type, m, surrounding_context(m, u))} 14. end for 15. end for 16. 17. // Layer 2: Shannon entropy + key-value context 18. tokens <- tokenize_kv_pairs(u) 19. for each (key, value) in tokens do 20. if shannon_entropy(value) >= ENTROPY_THRESHOLD then 21. if key in CREDENTIAL_KEY_NAMES then 22. F <- F U {(infer_type(key), value, key)} 23. end if 24. end if 25. end for 26. 27. // Layer 3: PII with validation 28. for each pattern p in PII_PATTERNS do 29. matches <- regex_findall(p.regex, u) 30. for each m in matches do 31. if p.validator(m) = true then 32. F <- F U {(p.credential_type, m, surrounding_context(m, u))} 33. end if 34. end for 35. end for 36. end for 37. 38. return deduplicate(F)

18

VENDOR PATTERNS includes anchored regex for Azure AD (client id, client secret), APIM subscription keys, AWS access keys, Google API keys, and 40+ additional provider formats. ENTROPY THRESHOLD = 3.5 bits/char. CREDENTIAL KEY NAMES includes apiKey, subscriptionKey, clientSecret, appKey, encryptionKey, and synonyms. No outbound verification calls are made at any step.

8.4

Preliminary Case Study: Detection Outcomes (N=2)

This section reports detection outcomes from the two confirmed-positive applications in the authorized assessments described in Section 1. This is a preliminary case study with N=2 confirmedpositive applications, not a controlled benchmark evaluation. No ground-truth labeled corpus of production JavaScript exists for this class of finding. Constructing one from unauthorized applications is not ethically feasible. What follows is a factual account of what the runtime detection methodology identified in both confirmed-positive cases, and why no other automated control was positioned to find the same. Credentials in the detection gap. In both confirmed-positive applications, the exposed credentials had reached production through the structural paths described in Section 5, build-time environment injection and CI/CD pipeline variable substitution. By the time the credentials were present in the served application, they had passed every automated control in the pre-deployment pipeline. This is not a statement about whether those controls ran or how well they were configured. It is a structural property of the paths: credentials introduced after the build scanner runs, or substituted by the pipeline after the repository scanner completes, do not exist in any layer those tools examine. The shift-left scanning layer had nothing to find because the credentials were not there when it looked. No shift-right control covered this layer. Once deployed, both applications were running in production environments with standard enterprise security controls, including web application firewalls and network monitoring. None of those controls are designed to examine the content of served JavaScript files for embedded credentials. DAST tools spider applications for injection vulnerabilities, WAF and network monitoring tools inspect inbound traffic for attack patterns. Neither category asks whether the JavaScript being served to users contains an Azure AD client secret. The credentials were live, publicly served, and invisible to the entire deployed security stack, not because any tool failed, but because no tool in any category was scanning that layer. Runtime detection results. The layered detection approach described in Sections 8.1–8.2 (anchored vendor token pattern matching combined with key-value context scanning) identified the complete credential sets in both confirmed-positive applications. In both cases, the credentials were located in served JavaScript bundles. In Chain 1 (Section 3), the AppID, AppKey, Resource, and SubscriptionKey were matched directly by anchored Azure AD and APIM credential patterns. In Chain 2 (Section 4), the encrypted configuration was located by key-value context scanning identifying a high-entropy value paired with a credential-relevant key name. Decryption was performed manually using the co-located key, after which the same credential patterns applied. Both findings were confirmed by successful exploitation before disclosure. Credential set completeness. Of the 113 applications with confirmed credential exposure, 63 (55.8%) contained a complete Azure AD credential set (client id, client secret, tenant id, and resource URI co-located in the same JS bundle) sufficient to execute the exploitation chain documented in Section 3 given over-permissive service principal scopes. The remaining 50 applications contained partial credentials. Across all 113 vulnerable applications the most common credential types observed were Azure App Insights instrumentation keys (62 apps, 54.9%), Azure APIM subscription keys (46 apps, 40.7%), JWT tokens issued to the browser (25 apps, 22.1%), CryptoJS19

encrypted configuration objects (16 apps, 14.2%), and CyberArk AIM tokens (11 apps, 9.7%). Applications frequently contained multiple credential types simultaneously. Partial credentials represent significant exposure but do not independently enable the full account takeover chain. The high proportion of complete Azure AD sets (56% of vulnerable applications) indicates that the credential injection pattern typically carries the full credential bundle rather than isolated values. Per-app prevalence and per-credential GT-194 counts can diverge in either direction (for example, 62 apps share 15 unique App Insights workspace iKeys, while the 16 CryptoJS-using apps contain 37 distinct encrypted blobs in GT-194). Scope and limitations. The sample of 113 applications is drawn from a single authorized engagement scope of approximately 2,000 enterprise web application assets, not a randomly selected population across multiple organizations. Detection was performed by a skilled practitioner applying the methodology described in this section, not by a fully automated tool running without guidance. A formal evaluation (with a labeled benchmark corpus, automated tool execution, and precision and recall measurements at scale) remains an area for future empirical work. The outcomes reported here establish that the methodology successfully identifies credentials in the layer that existing tooling cannot reach, and that this layer contains real, exploitable credentials in production systems that have passed all deployed automated controls.

8.5

Tool Comparison Study, GT-194 Cross-Vendor LLM-Validated Benchmark

Section 8.4 established that the runtime detection methodology successfully identifies credentials in two confirmed-positive applications. This section quantifies how the nine production scanners (established static tools at default and tightened-custom configurations, and the runtime-aware extension SecretSifter) perform against an independent Ground Truth, GT-194, on the same corpus, with the LLM-assisted ground-truth reference comparator reported separately. The evaluation is deliberately structured to remove the developer-author conflict of interest from the headline numbers: the Ground Truth is constructed by an LLM with no access to any evaluated detector, the LLM-extracted candidates were independently classified by an LLM from a different vendor, and the field-deployment register was locked on 8 May 2026. 8.5.1

Benchmark Construction

What “GT-194” means. GT-194 is the Ground Truth set of 194 unique secret-grade credentials assembled from production JavaScript bundles across 113 enterprise applications. “Ground Truth” here means the labelled, manually validated set of credentials that scanners are measured against, the canonical denominator for recall calculations throughout this section. GT-194 includes only credential instances that are actually secret-bearing, partitioned into eight credential types: Azure APIM subscription key (n=54), Azure AD client secret incl. btoa-encoded OAuth (n=50), CryptoJS-AES Salted blob (n=37), Other API key/token incl. App Insights iKey (n=29), JWT (n=11), Google API key (n=6), CyberArk AIM (n=4), and plaintext user credential (n=3). The eight type counts sum to 194. Corpus. The 113-application benchmark is the full set of credential-bearing applications identified during the engagement (drawn from approximately 2,000 enterprise web application assets in scope). It includes all applications providing downloadable JavaScript bundles suitable for both static-scanner evaluation and LLM-assisted ground-truth extraction, and captures the full credential-bearing production deployment set including builds containing CryptoJS-AES blobs, Other API keys, and plaintext user credentials. Independent Ground Truth extraction. Claude Opus 4.7 (Anthropic) [22] was prompted

20

to extract every value from the bundles that it considered a real high-impact secret. The model was given only the bundle content and a strict-secret rubric, with no access to SecretSifter rules, output, or any other tool’s findings. After human review, the locked GT-194 was assembled as the union of this independent extraction and a manual analyst pass that added secret-grade credentials the model did not surface. Because the ground truth therefore contains credentials no single detector produced on its own, no scanner reaches full recall against it: even the extracting reference (Claude Opus) recovers only 166 of 194 (85.6%, Table 8), and the 27 credentials in the tool-agnostic blind spot (§8.5.11) are those surfaced solely by manual analysis, detected by none of the nine evaluated production scanners and missed by the GT-construction reference comparator. Each of these 27 manual-only credentials was verified as a secret-grade credential by analyst review against the same strict-secret rubric; because they were not produced by the LLM extractor, the automated crossvendor (GPT-5.5) validation reported below applies to the LLM-extracted candidate set, and the 27 rest on manual verification rather than on the automated validator. The blind-spot arithmetic reconciles as follows: the reference comparator misses 28 of 194 (194 minus 166); of those, one is recovered by SecretSifter (the single credential SecretSifter catches that the reference misses), leaving 27 that no production scanner recovers, which is the tool-agnostic blind spot. Finally, GT194 covers first-party application JavaScript bundles only; it does not evaluate scanner performance against Path 4 (third-party script inclusion, §5) credentials, whose remediation ownership and detection surface differ. Cross-vendor independent validation. Each Opus-extracted candidate was independently classified by GPT-5.5 (OpenAI) [23] via the OpenAI Codex CLI. GPT-5.5 received the candidate value, key name, and source-file context, and was asked to apply the same strict-secret rubric independently. Across the 247 LLM-extracted candidates, GPT-5.5 returned 207 concordant SECRET classifications, giving Brennan-Prediger [24] κ = 0.676, substantial on the Landis-Koch interpretive scale [26]. The 27 manual-only additions to GT-194 were not produced by the LLM extractor and are excluded from the κ denominator; they rest on analyst review against the same strict-secret rubric. Because the adjudicated candidate set is single-class in the SECRET label (every candidate is an Opus-labelled SECRET), Cohen’s κ is undefined here (the Feinstein-Cicchetti paradox [25]); Brennan-Prediger κ is the bias-corrected metric appropriate for single-class prevalence. Why public-by-design Azure identifiers are tracked separately (“below the line”). The corpus also contains 249 chain-completion identifiers, 153 Azure AD App IDs / client ids / tenant ids and 96 Azure AD resource URIs (resource, Resource, *_Resource, resourceId). These are public-by-design Azure identifiers, not credentials, and are therefore excluded from the GT194 recall denominator. Two reasons drive the structural choice. First, including them would artificially inflate scanner scores, since these are not credentials a peer reviewer would accept as “leaks.” Second, excluding them entirely would hide the operational reality that all four components (client_secret, client_id / App ID, tenant_id, and resource URI) must be combined to mint an Azure AD access token. Listing them separately preserves both honesty about the recall metric and visibility into the exploitation chain. The chain-completion identifiers are reported in Table 10 below the strict GT-194 total (“below the line”) and as a separate corpus-property finding in Table 11. Per-credential counting protocol. All per-tool hit counts in this section are computed as per-credential matches against GT-194. Each row in the Ground Truth corresponds to a distinct unique credential value. The unique-credential count is the canonical recall metric throughout. For SecretSifter and the reference comparator Claude Opus (the two highest-recall configurations), summing per-type detection counts yields totals higher than the unique-credential total, 189 versus 151 for SecretSifter, and 175 versus 166 for Claude Opus. The eight lower-recall scanners have no such discrepancy. The reason: a small number of credential values in the corpus are tagged with 21

multiple type labels (e.g., a value found as both a client_secret in one HTTP request and an Other API key/token elsewhere). For tools with broad cross-type coverage (SS, Claude), this overlap shows up as inflated per-type sums; for low-recall tools, it does not. The unique-credential count remains the canonical recall metric and is what appears in headline tables and figures. 8.5.2

Tool Corpus and Evaluation Protocol

Nine production scanners were evaluated against GT-194: SecretSifter (curated runtime extension), TruffleHog, JSluice, SecretFinder, Titus, JSMiner, Nuclei, Sensitive Discoverer, and Cariddi (static / template-driven scanners). Claude Opus 4.7 is reported alongside them as the LLM-assisted ground-truth reference comparator, included for reference because it serves as the GT extractor and, having co-constructed the benchmark, cannot be scored as an independent detector. Each tool was run at its tightened configuration. For the rule-extensible tools (TruffleHog, SecretFinder, Titus, JSluice), tightened configuration adds the same 6-rule overlay covering the most prevalent credential shapes in the corpus: Azure AD client id (UUID v4), Azure AD client secret (40-character base64 with embedded tildes), APIM subscription key (32-hex), JWT (eyJ-prefix three-segment), App Insights instrumentation key, and CryptoJS U2FsdGVkX1-prefix encrypted blobs. The full regex set is reproduced in Appendix A. Nuclei was run with the secret-scanning template categories exposures/tokens, exposures/apis, and exposures/keys. JSMiner has hard-coded patterns and does not expose a user-rule API, it was run at its built-in configuration. SecretSifter and Cariddi were run at their built-in configurations. Per-tool finding outputs were matched to GT-194 via exactvalue substring lookup with no other post-processing. Cariddi achieves 2/194 = 1.0% recall on GT-194, a non-zero data point that contributes signal to the bottom of the recall ranking. Cariddi’s primary mode is web-crawling and endpoint enumeration; secret detection is a secondary capability with limited overlap with this corpus’s credential mix. Author-developer disclosure. The SecretSifter edition evaluated here (the Burp Suite extension, version 1.0.1) is open-source software developed by the first author (Gorijala) and published on the PortSwigger BApp Store and GitHub. The evaluation methodology is structured to make the comparison independent of this fact: GT-194 was constructed by an LLM (Claude Opus 4.7) with no access to SecretSifter’s rules, output, or source code; its LLM-extracted candidates were independently validated by an LLM from a different vendor (GPT-5.5, OpenAI); the same 6-rule overlay was applied identically to all rule-extensible tools with no per-tool tuning (Appendix A); and the rule overlay was selected from the credential shapes most prevalent in GT-194 without consulting any tool’s output. The headline numbers reflect performance on a single-organization corpus and should be interpreted as benchmark performance on this corpus rather than as crossorganization population estimates (Section 10). Tool versions and source. All scanners were evaluated at the versions listed in Table 6. Each version is the latest release available at the GT-194 lock date (8 May 2026); commit hashes are recorded for tools without semantic versioning. The evaluation environment was macOS 25.1 (Darwin), Burp Suite Professional 2025.10 hosting the Burp-extension scanners (SecretSifter, JSMiner, Sensitive Discoverer), and Python 3.12 for the JSMiner re-implementation. 8.5.3

Configuration, Sensitivity, and the Fairness of the Comparison

This evaluation is deliberately not a pure default-configuration fairness contest between interchangeable products. Its purpose is to measure how much of a real runtime-exposure corpus each detector recovers, under the strongest configuration a practitioner could reasonably deploy. Be-

22

Table 6: Scanner versions and sources. Each row records the exact version evaluated against GT194 and the public source where the same version can be obtained. Claude Opus 4.7 and GPT-5.5 are listed at the bottom of the table because they serve as the LLM-assisted ground-truth reference and cross-vendor validator respectively, not as evaluated detection scanners. Scanner

Version evaluated

Source

SecretSifter TruffleHog JSluice SecretFinder Titus JSMiner Nuclei Sensitive Discoverer Cariddi Claude Opus 4.7 (GT-construction reference) GPT-5.5 (cross-vendor validator)

1.0.1 3.83.7 0.0.7 commit a07d215 (Apr 2026) 1.4.0 Trustwave SpiderLabs distribution v1.0; Python re-implementation included with this paper 3.3.10 with template repository commit 2026-04-30 9.4 (BApp Store, May 2026) 1.4.1 claude-opus-4-7 (Anthropic API, Mar 2026 release) gpt-5.5 (OpenAI API via OpenAI Codex CLI v0.6.x)

https://portswigger.net/bappstore (BApp Store) and https://github.com/secretsifter/burp-secret-scanner https://github.com/trufflesecurity/trufflehog https://github.com/BishopFox/jsluice https://github.com/m4ll0k/SecretFinder https://github.com/Falconcyber-research/Titus https://github.com/SpiderLabs/JS-Miner https://github.com/projectdiscovery/nuclei https://portswigger.net/bappstore and https://github.com/CYS4srl/SensitiveDiscoverer https://github.com/edoardottt/cariddi https://docs.anthropic.com/en/docs/about-claude/models https://platform.openai.com/docs/models

cause SecretSifter is authored by the first author, the configuration protocol is stated explicitly so the comparison can be judged on its face. Table 7 records the exact operating mode each detector was evaluated under. Table 7: Evaluated configuration per detector. The condition under which each detector was run against GT-194. The rule-extensible static tools all received the identical six-rule overlay (Appendix A); SecretSifter received no GT-194-derived rules. Detector

Category

Evaluated configuration

SecretSifter TruffleHog SecretFinder JSluice Titus JSMiner Nuclei Sensitive Discoverer Cariddi

runtime-aware static, rule-extensible static, rule-extensible static, rule-extensible static, rule-extensible static, fixed patterns template-driven Burp extension web-crawler / scanner

Shipped built-in rules (v1.0.1); no GT-194-derived rules Default detectors + identical 6-rule overlay Default patterns + identical 6-rule overlay Default extraction + identical 6-rule overlay Default rules + identical 6-rule overlay Built-in patterns (no user-rule API) exposures/{tokens,apis,keys} templates Built-in ruleset Built-in secret detection

Claude Opus 4.7 (ref)

LLM reference

Strict-secret rubric (GT constructor)

Two facts make the asymmetry cut against SecretSifter, not for it. First, the six-rule overlay was added to help the rule-extensible static tools: it targets the credential shapes most prevalent in GT-194, and on a strict hard-credential subset (n = 129, used only to isolate the overlay’s effect) it lifted their recall by 3.9× to 7.1× over their out-of-the-box configuration (for example TruffleHog from 7.0% to 49.6%; the full default-versus-overlay figures are in the supplementary sensitivity analysis). Every rule-extensible competitor is therefore reported at a strengthened configuration, not a handicapped one, and its recall is an upper estimate of its deployable performance. Because the overlay raises the static tools rather than lowering them, a fully matched, no-overlay protocol, the like-for-like condition a reader may prefer, widens the runtime-versus-static gap rather than narrowing it: on the same subset SecretSifter recovers 75.2% at its shipped configuration against 7.8% for the best static tool at default. Second, SecretSifter was evaluated at its frozen, publicly shipped v1.0.1 rule set, published before GT-194 was constructed and derived without reference to it; no GT-194-specific rule was added to SecretSifter. The one detector that could in principle have been tuned to the benchmark was the only one held to its frozen, version-locked shipped configuration. We nonetheless do not claim a like-for-like product ranking. The honest reading of Table 8 and Table 9 is a coverage statement scoped to this corpus: on a single-organization, Azure-heavy 23

runtime-exposure corpus, a runtime-aware curated scanner recovers substantially more served credentials than static and template scanners do even after those tools are given a corpus-informed overlay. We do not extend this to a claim of universal scanner superiority across corpora or credential mixes (Section 10). We read this honestly in both directions: targeted static rules can substantially narrow, and with a sufficiently large curated ruleset could further close, the staticscanner gap on credentials that are present in the artifact. What such tuning cannot reach is the structural remainder, the credentials that only materialize at runtime and the CryptoJS encryptedconfiguration class whose plaintext exists only after decryption with a co-located key; together with SecretSifter’s shipped out-of-the-box coverage, these remain the durable differentiators of the runtime-aware approach on this corpus. 8.5.4

Headline Recall Comparison

Figure 5: Per-tool recall on GT-194. The structural finding is the gap between repository-era static scanning and runtime-aware detection: the best static scanner (TruffleHog) plateaus at 71/194 = 36.6%, while runtime-aware detection recovers far more (SecretSifter 151/194 = 77.8%). Claude Opus (166/194 = 85.6%) is the LLM-assisted ground-truth reference comparator that helped construct GT-194; it is shown separately as a reference, not as a deployable evaluated detector, and its recall is a near-definitional property of how the benchmark was built (see Table 6 and §8.5). SecretSifter is developed by the first author; the conflict-of-interest mitigations are detailed in §8.5. The shift-right tooling gap is now quantified end-to-end on GT-194. Among the nine evaluated production scanners, SecretSifter (77.8%) leads the best static tool (TruffleHog at 36.6%) by +41.2 percentage points, equivalent to 2.13× the static-scanner recall. Relative to the LLM-assisted ground-truth reference (Claude Opus, 85.6%, reported separately and not an evaluated detector), SecretSifter closes all but 7.7 percentage points of the distance (a gap of 15 of the 194 credentials, 166 24

Table 8: Per-tool overall recall on GT-194. Counts are per-credential matches against the locked GT-194 secret-grade Ground Truth. The nine production scanners are ranked; Claude Opus 4.7 is listed below the rule as the LLM-assisted ground-truth reference comparator (marked “ref”), not ranked among the evaluated scanners, because it co-constructed GT-194 and therefore cannot miss anything except the manual analyst additions. Its 85.6% is a reference figure, close to a definitional property of how the benchmark was built rather than an independent detection measurement. Rank

Scanner

Role

Detected

GT total

Recall

1 2 3 4 5 6 7 8 9

SecretSifter TruffleHog SecretFinder JSluice Titus JSMiner Nuclei Sensitive Discoverer Cariddi

curated runtime scanner static scanner static scanner static scanner static scanner Burp extension template-driven scanner Burp extension web-crawler / scanner

151 71 63 62 40 12 7 7 2

194 194 194 194 194 194 194 194 194

77.8% 36.6% 32.5% 32.0% 20.6% 6.2% 3.6% 3.6% 1.0%

ref

Claude Opus 4.7

GT-construction reference (not an evaluated detector)

166

194

85.6%

minus 151); because that reference co-constructed the benchmark its 85.6% is a near-definitional upper bound rather than an achievable detection target, so the 7.7-point figure is an indicative bound on remaining headroom, not a like-for-like tool gap. 8.5.5

Precision and False-Positive Rate

Recall is half the operational story. The other half is precision, how many findings each scanner emits beyond GT-194, and what proportion of those emissions are credentials versus noise. A scanner that finds every credential but emits ten thousand false positives per application is unusable in a security-engineering workflow. This subsection reports precision, recall, and F1 score for the nine production scanners against GT-194, with the reference comparator listed separately. Three structural observations: 1. SecretSifter has the highest F1 of any evaluated production scanner at 0.818; the GT-construction reference records 0.851. Among the nine evaluated production scanners, SecretSifter leads on F1 = 0.818 (precision 86.3%, recall 77.8%). The LLM-assisted ground-truth reference (Claude Opus 4.7) records F1 = 0.851 (precision 84.7%, recall 85.6%); because it coconstructed GT-194 this is a reference figure rather than an evaluated-detector result, and the 0.033 difference is an indicative bound on remaining headroom for curated detection rather than a like-for-like gap. SecretSifter’s lead over the best static scanner (TruffleHog, F1 = 0.520) is +0.298. The lead is structural rather than precision-driven: TruffleHog has nominally similar precision (89.9% vs SecretSifter 86.3%) but recovers only 36.6% of GT-194, while SecretSifter’s higher recall (77.8%) under comparable noise translates into the higher harmonic mean. 2. Regex architecture without context awareness catastrophically inflates falsepositive rate on minified JavaScript. SecretFinder produces 1,402 false positives on GT194. The structural reason: generic UUID, 32-hex, and base64-shaped patterns match webpack chunk hashes, build fingerprints, asset content-hashes, internal element identifiers (questionPanelId, FRAUDNET FNCLS, proposalId), and Angular framework constants (Inject, providedIn, NullInjectorError) that are ubiquitous in minified production JavaScript. The same regexes that produce signal in repository code produce overwhelming noise on built artifacts. TruffleHog’s lower

25

Figure 6: Precision vs. recall scatter on GT-194. Diagonal contours show iso-F1 lines. Static scanners cluster at low recall; runtime-aware detection (SecretSifter, marked with ⋆) moves up and to the right by combining high precision with broad recall. Claude Opus is the LLM-assisted ground-truth reference comparator and is plotted separately as a reference rather than an evaluated deployable detector (§8.5).

26

Table 9: Precision, recall, and F1 on GT-194. TP = true positives (GT-194 credentials correctly detected); FP = false positives; FN = false negatives (GT-194 credentials missed). Precision = TP / (TP + FP). Recall = TP / (TP + FN). F1 = 2·P·R / (P + R). False positives were counted as follows: every string a scanner emitted that was not a GT-194 credential was adjudicated by manual review against the same strict-secret rubric used to build GT-194, and those judged non-credential (for example webpack chunk hashes, build fingerprints, and framework constants) were counted as false positives for that scanner. FP counts are thus measured over each scanner’s full emission set on the benchmark corpus, not only over GT-194 rows. The nine production scanners are listed first; Claude Opus 4.7 appears below the rule marked “(ref)” as the LLM-assisted ground-truth reference comparator, not an evaluated detector, because it co-constructed GT-194. Scanner

TP

FP

FN

Precision

Recall

F1

SecretSifter TruffleHog SecretFinder JSluice Titus JSMiner Nuclei Sensitive Discoverer Cariddi

151 71 63 62 40 12 7 7 2

24 8 1,402 89 11 6 0 8 0

43 123 131 132 154 182 187 187 192

86.3% 89.9% 4.3% 41.1% 78.4% 66.7% 100.0% 46.7% 100.0%

77.8% 36.6% 32.5% 32.0% 20.6% 6.2% 3.6% 3.6% 1.0%

0.818 0.520 0.076 0.359 0.327 0.113 0.070 0.067 0.020

Claude Opus 4.7 (ref)

166

30

28

84.7%

85.6%

0.851

FP count (8) reflects its built-in entropy and verification heuristics; with custom rules disabled and verification off, TruffleHog would emit a similar volume to SecretFinder. 3. Parser-based scanners achieve precision through context but lose recall on minified output. Titus reaches 78.4% precision but only 20.6% recall. The architectural pattern is consistent: AST-aware parsers reject string literals that do not appear in credential-typed contexts, which is precisely why they emit few findings overall, but it is also why they miss credentials whose context is destroyed by minification (variable names mangled, key-value pairs collapsed into runtime objects, base64-encoded UUIDs that defeat the cleartext UUID-v4 regex). JSluice reaches 41.1% precision and 32.0% recall, higher noise than Titus because the key-context extraction returns every UUID-shaped value bound to any key, including form-field IDs and internal element IDs that are not credentials. SecretSifter’s curated runtime architecture (anchored vendor token patterns, key-value context scanning, entropy filtering, and CDN/key-name blocklists, §8.2) sits in the precision–recall quadrant that no other measured non-LLM scanner reaches on this corpus. The runtime layer differs from static-tool capability in operational regime, addressing both halves of the precision–recall trade-off simultaneously and narrowing the distance to the GT-construction reference to 0.033 in F1. 8.5.6

Per-Credential-Class Breakdown and Chain-Context Identifiers

GT-194 comprises eight strict credential classes. Below the strict total are 249 chain-completion identifiers (Azure AD App IDs, client ids, tenant ids, and resource URIs) that are public-by-design and therefore tracked separately rather than included in the recall denominator (§8.5.1). ⋆ = SecretSifter is the only detector for that type.

27

Figure 7: Per-credential-type recall heatmap on GT-194. Top panel: strict GT-194 secret-grade credentials. Bottom panel (separated by a red bold header): chain-context, public-by-design Azure identifiers (NOT in the GT-194 strict total). Cells show per-type recall percentage. Table 10: Per-credential-type recall on GT-194 with chain-context identifiers below the line. Cells show detected count / type denominator. Below-the-line entries are excluded from the strict GT194 recall denominator because they are not credentials, but are listed because they are required to convert leaked client secrets into valid access tokens. Column abbreviations: Claude = Claude Opus 4.7, SS = SecretSifter, TH = TruffleHog, SF = SecretFinder, JS = JSluice, Tit = Titus, JSM = JSMiner, Nuc = Nuclei, SD = Sensitive Discoverer, Car = Cariddi. Section

Type

strict GT strict GT strict GT strict GT strict GT strict GT strict GT strict GT STRICT GT-194 TOTAL chain context chain context CHAIN-CONTEXT SUBTOTAL

Azure APIM subscription key Azure AD client secret (incl. btoa OAuth ×2) CryptoJS-AES blob User credential (plaintext + btoa) JWT CyberArk AIM Google API key Other API key / token (App Insights iKey ×15, Braintree ×1) (unique credentials) Azure AD App ID / client id / tenant id Azure AD resource / * Resource (separate, public-by-design)

n

Claude

SS

TH

SF

JS

Tit

JSM

Nuc

SD

Car

54 50 37 3 11 4 6 29 194 153 96 249

50 46 32 3 10 3 6 25 166 135 85 220

54 ⋆ 48 37 ⋆ 3 11 4 5 27 151 153 ⋆ 96 ⋆ 249 ⋆

0 41 0 0 11 0 5 14 71 0 0 0

0 35 0 0 11 0 5 12 63 0 0 0

0 33 0 0 11 0 5 13 62 0 0 0

0 22 0 0 4 0 5 9 40 0 0 0

0 4 0 0 1 0 1 6 12 0 0 0

0 0 0 0 1 0 1 5 7 0 0 0

4 0 0 0 0 0 1 2 7 0 0 0

0 0 0 0 0 0 0 2 2 0 0 0

For the eight lower-recall scanners, the per-type column sums match the unique-credential totals in the highlighted Strict GT-194 total row exactly. For SecretSifter and Claude Opus, the per-type sums (189 and 175 respectively) exceed the unique totals (151 and 166) because some credentials in the corpus are tagged with multiple type labels (§8.5.1). The unique credential count remains the canonical recall metric. The chain-context subtotal, 249 / 249 = 100% for SecretSifter, 0 for every other static scanner, quantifies SecretSifter’s unique contribution to chain-completion. These identifiers are not creden-

28

tials and therefore not part of the recall comparison, but operationally they are required to convert a leaked Azure AD client_secret into a working access token. Without surfacing them, a leaked secret cannot be reduced to an exploitable chain.

Figure 8: Per-tool detection breakdown by credential type on GT-194. Each stacked bar shows what each tool detects, decomposed by credential type. Bar heights are per-type detection events summed; the label above each bar shows the unique-credential total and recall percentage.

8.5.7

Chain-Completion in the Production Corpus

In Azure AD’s OAuth client-credentials flow, four components, client_secret, client_id (App ID), tenant_id, and resource URI, must be combined to mint an access token. We measured how often these components are co-located in the same client-side JavaScript bundles. This is a corpus property, not a tool-comparison metric: the question is “in how many production applications is the full token-mint chain leaked client-side?” rather than “which scanner finds the most chains?” Table 11: Chain-completion in the production corpus. Counts and percentages relative to the 86 applications where an Azure AD secret is exposed in client-side JavaScript. Finding Apps where Azure AD secret is exposed in client-side JS Apps where full token-mint chain (secret + App ID + tenant + resource) is co-located in same JS bundles Apps where chain components are partially present (tenant or resource missing)

Count

% of secret-exposed apps

86 63 23

100% 73.3% 26.7%

This finding characterizes the corpus, not any specific scanner. These are the same 63 applications reported in Section 8.4; there they are expressed as a share of all 113 credential-bearing applications (55.8%), and here as a share of the 86 applications that leak an Azure AD secret specifically (73.3%). It establishes that for the majority of credential-bearing applications in this engagement, the entire authentication chain, not just the secret, is reachable from passive browser inspection. An attacker who recovers the client secret from a JavaScript bundle does not need to perform additional reconnaissance to identify the App ID, tenant id, or resource URI; all four are 29

Figure 9: Chain-completion in the production corpus. Of 86 production applications where an Azure AD secret (client_secret or APIM subscription key) is exposed in client-side JavaScript, 63 (73.3%) have all four chain components co-located in the same JS bundles, so the full token-mint chain is reachable from browser-visible code with no additional reconnaissance. The remaining 23 applications (26.7%) have a partial chain (some chain components missing from the JS).

30

already co-located in the same delivered artifacts. The exploitation chains documented in Sections 3 and 4 are therefore not edge cases but representative instances of a structural pattern present in roughly three-quarters of secret-exposed production deployments. Note on tool-comparison framing. Chain reachability is reported as a corpus property rather than as a per-tool metric. “73.3% of secret-exposed apps have a full chain leaked” is a fact about the production deployment landscape, independent of which scanner one uses. The per-tool comparison story is captured in Tables 8, 10, 9, and 13; Table 11 reports the corpus property. 8.5.8

Credential Categories No Static Scanner Detected

Three credential categories in GT-194 received zero detections from every static scanner evaluated, and a fourth was detected only by SecretSifter: • Azure APIM subscription key (54 instances). SecretSifter is the only scanner that detects this class with full coverage on this corpus, surfacing 54 of 54. Sensitive Discoverer surfaces 4 of 54 incidentally, its apikey:"..." rule captures the value when it is bound to a key whose name contains the literal apikey substring (the corpus has 4 such cases under names like apimSubKey); every other static scanner returned 0. The APIM key format (32-character hex paired with the Ocp-Apim-Subscription-Key header name in HTTP requests) is not represented in the standard rule libraries shipped with TruffleHog, JSluice, SecretFinder, Titus, JSMiner, Nuclei, or Cariddi. • CryptoJS-AES Salted blob (37 instances). SecretSifter is again the only evaluated production scanner that detects this class, surfacing all 37. The CryptoJS construct is invisible to pattern-based detection by design: the credential exists only after decryption, and the decryption key is co-located with the ciphertext (CWE-321 [12]). Encrypting a credential with a co-located key is in fact worse than shipping it in cleartext: it adds no protection against an attacker, who simply reads the adjacent key and decrypts, while it defeats the reviewer’s eye and every static scanner, which see only a high-entropy blob. The net effect is to convert an obvious plaintext exposure into a hidden one, lowering the odds of detection without raising the cost of exploitation. Detecting this class requires identifying the construct (the U2FsdGVkX1-prefix base64 envelope, paired with a co-located decryption key) rather than a credential signature. The GT-construction reference (Claude Opus) reaches 32 of 37 in this category, high but not complete, because the model treats some heavily-truncated blobs as obfuscated configuration rather than encrypted credentials. • User credential / plaintext password (3 instances). Only Claude Opus and SecretSifter recover this class (3 of 3 each). All other scanners return 0. Plaintext password values do not match any standard credential-shape regex; they are identified only by context (key name + value adjacency) or by semantic understanding of the surrounding code. • CyberArk AIM coordinates (4 instances). SecretSifter recovers all 4; only Claude Opus comes close (3 of 4). All other static scanners return 0. Together, these four credential classes total 98 of GT-194 (50.5%). For all four classes the structural reason for the static-tool gap is the same reason no shift-left scanner detects them in production code: they were never encountered during the pattern-library calibration phase of any of the evaluated static tools, because they appear primarily in production-deployed Angular bundles, the surface those tools were not designed to scan. Beyond the four SecretSifter-leading categories, the chain-context identifiers tracked separately in Section 8.5.6 (153 Azure AD App IDs / client ids / tenant ids and 96 Azure AD resource URIs,

31

249 in total) are also detected only by SecretSifter (249 of 249) and Claude Opus (220 of 249); every other static scanner returns 0 on chain-context identifiers as well. 8.5.9

Validator Confidence Calibration and Sensitivity

GPT-5.5’s confidence field is a self-report, so two acceptance thresholds are reported. The unit of agreement is the LLM-extracted candidate and the decision is binary (SECRET vs. not-SECRET); every candidate is by construction an Opus-labelled SECRET candidate, so the Opus margin is single-class and the 2×2 table has one degenerate margin. GPT-5.5 adjudicated the 247 LLMextracted candidates, concurring on 207 and dissenting on 40 at any confidence, and concurring on 195 at high confidence only. The 27 manual-only additions to GT-194 were not produced by the LLM extractor and are excluded from this denominator; they were verified by analyst review against the same strict-secret rubric. These 247 records are per-occurrence candidate adjudications; after de-duplication, the Opus-positive candidates contribute to the 166 unique Claude Opus credentials reported in GT-194 (§8.5.1). Agreement is therefore reported per rated candidate, while scanner recall is reported per unique credential. Because expected agreement by chance for a two-category decision is 1/k = 0.5, the bias-corrected Brennan-Prediger coefficient is κBP = (po −0.5)/(1−0.5) = 2po − 1. The any-confidence figure is po = 207/247 = 0.8381 (83.8%), giving κBP = 2(0.8381) − 1 = 0.676 (substantial, Landis-Koch). The high-confidence-only variant is reported as a conservative sensitivity bound: po = 195/247 = 78.9%, giving κBP = 0.579 (moderate). Cohen’s κ is not reported because the single-class Opus margin makes it undefined here (the Feinstein-Cicchetti paradox). Table 12 gives the contingency underlying κBP = 0.676. Table 12: Cross-vendor agreement contingency on the LLM-extracted candidates (any-confidence). Rater A is Claude Opus 4.7 (extraction/classification); Rater B is GPT-5.5 (independent validation). The unit is one LLM-extracted candidate; the decision is binary SECRET vs. not-SECRET. Every candidate is an Opus-labelled SECRET candidate by construction, so the Opus-negative row is empty (single-class prevalence). The 27 manual-only additions to GT-194 are excluded from this denominator. Observed agreement po = 207/247 = 0.8381; κBP = 2po − 1 = 0.676.

Opus SECRET Opus not-SECRET Total

GPT-5.5 SECRET

GPT-5.5 not-SECRET

Total

207 0 207

40 0 40

247 0 247

A blind-extraction sensitivity analysis was also performed, in which GPT-5.5 was given the raw bundles (no Opus rationale, no field hints) and asked to extract independently. That analysis confirmed that the SecretSifter recall figure is not artifactually inflated by pattern leakage from SecretSifter into the Opus extraction prompt. 8.5.10

Statistical Significance of Pairwise Differences

McNemar’s test (with Edwards’ continuity correction) was computed pairwise across the nine production scanners and the reference comparator (45 comparisons in total) on the per-credential GT-194 detection matrix to determine which recall differences are statistically distinguishable from sampling variation. The test conditions on the disagreement cells (items one tool catches and the other misses, and vice versa), discarding the both-hit and both-miss cells which are uninformative for pairwise comparison. 32

Table 13: Selected pairwise McNemar tests on GT-194. Counts are unique credentials (per §8.5.1’s canonical recall metric); each cell of the 2×2 contingency table treats one row of GT-194 as one observation. A only = unique GT-194 credentials the first scanner catches and the second misses; B only = unique credentials the second catches and the first misses. Chi-square computed with Edwards continuity correction (df = 1). The full matrix of all 45 pairwise comparisons is provided in the supplementary materials. Tool A

Tool B

SecretSifter SecretSifter SecretSifter SecretSifter SecretSifter SecretSifter SecretSifter SecretSifter Claude Opus Claude Opus TruffleHog JSluice JSMiner

TruffleHog SecretFinder JSluice Titus JSMiner Nuclei Sensitive Discoverer Cariddi SecretSifter TruffleHog SecretFinder SecretFinder Nuclei

A only

B only

Chi-square

p-value

Sig

80 88 89 111 139 144 144 149 16 96 8 2 5

0 0 0 0 0 0 0 0 1 1 0 1 0

78.0 86.0 87.0 109.0 137.0 142.0 142.0 147.0 11.5 91.1 6.13 0.0 3.20

< 0.001 < 0.001 < 0.001 < 0.001 < 0.001 < 0.001 < 0.001 < 0.001 < 0.001 < 0.001 0.013 1.000 0.074

✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ × ×

Three findings emerge. First, SecretSifter’s lead is robustly significant against every other tool. All eight SS-vs-static comparisons reach p < 0.001, the 77.8% versus ≤36.6% headline gap cannot be attributed to sampling variation. For reference, the SecretSifter-vs-reference-comparator (Claude Opus) difference also reaches p < 0.001; because the comparator co-constructed the ground truth, this indicates only that SecretSifter has not closed the near-definitional reference gap, not that a deployable detector outperforms it. Second, the static-tool field is highly stratified. TruffleHog (36.6%) versus SecretFinder (32.5%) is significant at p = 0.013. SecretFinder (32.5%) versus JSluice (32.0%) is not statistically distinguishable (p = 1.000). The static-tool ceiling on GT-194 sits at approximately 32–37%, with TruffleHog significantly above SecretFinder/JSluice but those two tied within sampling error. Third, the bottom of the ranking is also stratified. JSMiner (6.2%) versus Nuclei (3.6%) is not statistically distinguishable (p = 0.074), but both are significantly below the top static cluster. Nuclei versus Cariddi (1.0%) is also not distinguishable (p = 0.074); the data does not support reading those two as significantly different from each other on this corpus. Sensitive Discoverer (3.6%) ties Nuclei on aggregate recall but operates over a different subset of GT-194 (4 plaintext APIM keys vs Nuclei’s 1 Google API + 1 JWT + 5 Other API tokens), McNemar SD vs Nuclei is not significant (chi-square ≈ 0.07, p ≈ 0.79). All hit counts in this table are computed against the unique-credential GT-194 detection matrix (one observation per unique credential), so the per-row arithmetic |A only ∪ B only ∪ overlap| reconciles directly with the unique-credential totals in Table 8 (Claude 166, SS 151). The per-type detection-event totals (Claude 175, SS 189) reported in §8.5.1 are not used here, because McNemar requires a single-class observation unit and the per-type tagging of multi-type credentials would otherwise inflate the disagreement cells.

33

 Multiple-comparison correction. Because all 10 2 = 45 pairwise tests are drawn from the same detection matrix, the Holm-Bonferroni step-down procedure was applied to the full 45-test family to control the family-wise error rate at α = 0.05. Thirty-seven of the 45 comparisons remain significant after correction. Every conclusion this paper rests on survives: all eight SecretSifter-vsproduction-scanner comparisons stay significant (adjusted p < 0.001), and every static-vs-runtime gap remains p < 0.001. The SecretSifter-vs-reference-comparator difference (Claude Opus) is also significant after correction (adjusted p = 6.85 × 10−3 ), but is reported for reference only because the comparator co-constructed the ground truth and is not an evaluated detector. Of the eight comparisons that are not significant after correction, six were already non-significant before it, all among the bottom-tier static and template tools (for example SecretFinder vs. JSluice, JSMiner vs. Nuclei, Nuclei vs. Cariddi); only two flip from significant to non-significant as a result of the correction, and both are weak-static-vs-weak-static pairs (TruffleHog vs. SecretFinder, adjusted p = 0.093; TruffleHog vs. JSluice, adjusted p = 0.061) that are not load-bearing for any claim made here. The correction therefore strengthens rather than weakens the significance argument. The full corrected table of all 45 adjusted p-values is provided in the supplementary materials. 8.5.11

Tool-Coverage Overlap and the Tool-Agnostic Blind Spot

Beyond pairwise comparisons, the question of which credentials each tool catches and how detection sets overlap is shown in Figure 10 below as an UpSet-style plot [27]. UpSet generalizes the Venn diagram to arbitrary numbers of sets and is the established replacement when more than four sets are compared. The most prominent patterns in GT-194 are: (P1) 79 credentials detected by Claude Opus + SecretSifter only, (P2) 28 credentials detected by 6 tools (no JSMiner, Nuclei, Sensitive Discoverer, Cariddi), (P3) 27 credentials detected by no tool, the tool-agnostic blind spot, (P4) 22 credentials detected by 5 tools, and (P5) 16 credentials detected only by Claude Opus. The pattern P3 is the operationally critical observation: 13.9% of GT-194 is not findable by any combination of the nine production scanners evaluated (and, as noted in §8.5.1, surfaced only by manual analysis). A complementary view is the cumulative recall when scanners are run in combination. Figure 11 below shows what fraction of GT-194 is recovered when scanners are added one at a time, with each next scanner chosen to maximize the new (previously-uncovered) credentials caught. Three structural observations: 1. The tool-agnostic blind spot is real, sized, and stable across measurement angles. 27 of GT-194 (13.9%) credentials are missed by every evaluated production scanner and were surfaced only by manual analysis. This is visible from the UpSet plot (Figure 10, the all-empty column) and from the coverage-saturation curve (Figure 11, the dashed line plateau at 86.1%). Both numbers derive from the same underlying detection matrix and are guaranteed consistent. The blind spot lies outside the CryptoJS and APIM classes that SecretSifter fully covers (37/37 and 54/54, Table 10); it consists of credentials surfaced only by manual analysis, whose shape and context defeat every evaluated production scanner and were also missed by the reference comparator (Claude Opus). 2. SecretSifter’s lead comes from disjoint findings, not from a superset relationship to static tools. Pattern P1 (79 credentials caught by Claude + SS only) and pattern P5 (16 caught by Claude only) together account for 95 of the 167 detectable credentials, meaning 95 of 167 credentials are caught only by the reference comparator (Claude Opus) and SecretSifter, not by any static tool. Among the static tools, intersection patterns dominate over exclusive contributions: TruffleHog and SecretFinder rarely catch a credential that SecretSifter misses. This addresses the structural concern that a curated tool’s headline recall might reflect overlap with 34

Figure 10: Detection-set overlap on GT-194 (UpSet-style). Each column is a unique detection signature, a unique combination of tools that jointly detect a credential. The bar at the top shows how many of the 194 GT credentials match that signature. The dots at the bottom show which tools participate in the signature (filled coloured dot = tool detects credentials in this pattern; gray dot = it does not). The blind-spot column (highlighted red) has zero filled dots because no evaluated production scanner detects those 27 credentials; they were surfaced only by manual analysis: this is the 13.9% tool-agnostic blind spot.

35

Figure 11: Coverage saturation. Tools are added one at a time, picking the next tool that catches the most credentials the previous tools missed. Even after running every scanner together, the curve plateaus at 86.1% (167 of 194 detected) - leaving 27 of 194 GT credentials (13.9%) that no evaluated production scanner detects (surfaced only by manual analysis). This tool-agnostic blind spot is the core motivation for runtime credential exposure research beyond static scanning.

36

established detectors rather than additive coverage. On GT-194, SecretSifter’s contribution is primarily additive, not redundant. 3. Beyond a runtime-aware scanner, additional detectors deliver diminishing returns. Figure 11 shows cumulative recall as detectors are added in recall order. Once the groundtruth reference comparator (Claude Opus) and SecretSifter are combined, the union reaches 86.1%, and no additional static or template scanner (tools 3 through 9) contributes any new detection: every credential they catch is already covered. Because the reference comparator co-constructed GT-194 and is not a deployable detector, the operationally meaningful coverage among deployable tools is SecretSifter’s 77.8%. The practical implication for 2026 is that the choice among deployable tools is between a runtime-aware curated scanner (77.8%) and static scanning, which plateaus near 32–37%: the static tools are complementary to each other only on the easy targets and fail together on the structural blind spots (Azure AD client secret, APIM keys, CryptoJS blobs). 8.5.12

Comparison with Concurrent Population-Scale Measurement

Demir et al. [15] reported the first population-scale dynamic measurement of credential exposure on the rendered web, covering 10M HTTP Archive landing pages, 1,748 verified credentials, and 14 vendor-API-verifiable service classes. Their work and ours measure different properties of the same exposure surface. Demir et al. establish population-scale prevalence on credential classes whose vendors publish verification endpoints. This work establishes per-tool detection performance at enterprise depth on eight credential classes that lack public verification endpoints, including the CryptoJS-encrypted-configuration construct that defeats every static scanner and is recovered only by runtime-aware detection. Demir et al. uniquely contribute population prevalence at 10Mpage scale, longitudinal persistence (mean twelve months exposure), vendor-API verification across fourteen service classes, and remediation-outcome measurement across 2,435 notified entities. This work uniquely contributes a per-tool benchmark across nine production scanners (with an LLMassisted ground-truth reference comparator) on a locked ground truth, cross-vendor LLM-validated candidate adjudication (κ = 0.676 over the 247 LLM-extracted candidates), pairwise McNemar significance testing (all static-vs-runtime gaps p < 0.001), chain-completion analysis showing 73.3% of affected applications co-locate the full Azure AD token-mint chain client-side, and structuralclass detection of the CryptoJS-AES construct. Figure 12 visualizes the complementary detection scope at the credential-class level. Both studies reach the same systemic conclusion. Credential exposure on the rendered web is widespread, persistent, and structurally undetected by repositoryonly scanning.

9

Recommendations

9.1

Immediate, Days

Rotate all exposed credentials immediately, across all environments. AppKeys, SubscriptionKeys, and any other credentials found in client-side code should be treated as compromised and rotated before any other remediation step [9]. Rotation does not confirm they were abused, but it closes the window. Critically, rotation must be applied simultaneously across production, UAT, SIT, and pre-production environments. Because credentials are injected at build time, every environment that ran a build received the same secret. Remediating production while leaving UAT exposed means an attacker who extracted credentials from a lower environment retains access.

37

Figure 12: Credential-class detection scope by methodology. Left column: Demir et al. [15] verifiedcredential counts per category, normalized to 0-100. Right column: number of credentials of each class in the locked GT-194 benchmark (n in benchmark). Empty cells indicate absence from the corresponding study. The diagonal pattern illustrates the complementarity: vendor-API verification reaches the population-scale services in the upper rows; multi-tool detection plus an LLM-assisted ground truth reaches the structural and enterprise classes in the lower rows.

38

Remediation is only complete when the credential is removed from the build pipeline and rotated everywhere at once. Audit APIM subscription key product scope. Each exposed subscription key maps to a product in APIM containing one or more APIs. Access the APIM portal and identify exactly which APIs each key grants access to. This determines the blast radius and informs incident response scope. Review Azure AD app registration permissions. Check whether the application’s app registration uses Application permissions (acting as the application itself, no user context) or Delegated permissions (acting as the signed-in user). Application permissions combined with a client secret in browser code grant the ability to access all users’ data, not just the current user’s. Remove Application permissions that are not strictly required. Enable APIM anomaly alerting. Check APIM Analytics for anomalous usage of the exposed subscription key, off-hours access, high-volume calls, unexpected geographic sources. Check Azure AD sign-in logs for client_credentials grants originating from browser IP ranges. This flow is designed for server-to-server communication and should not be initiated from end-user browsers regardless of application architecture. Disable source maps in production builds. Set sourceMap: false in angular.json under the production configuration. Do not deploy .map files alongside production bundles.

9.2

Short-Term, Weeks

Use the Authorization Code flow with PKCE for user authentication. The presence of a client secret in the browser is a symptom of a deeper misconfiguration: a browser-based single-page application is a public client and should never be issued a client secret. For the user-authentication path, the application should be registered as a public client and use the OAuth 2.0 Authorization Code flow with Proof Key for Code Exchange (PKCE) [29], which authenticates the user and obtains tokens without any client secret. PKCE binds the authorization request to the token exchange through a dynamically generated code verifier, eliminating the need for the static AppKey that this study found leaking in production bundles. PKCE addresses the credential problem, but it does not, on its own, resolve the token-exposure problem: with PKCE alone, the resulting access and refresh tokens still reside in the browser, where they remain reachable by cross-site scripting. This is why current IETF guidance for browser-based applications favors the Backend for Frontend pattern described below [30]. The two are complementary. Authorization Code with PKCE is the correct public-client pattern for authenticating the user with no client secret, while the BFF pattern keeps service credentials and long-lived tokens server-side. An architecture needing only user-delegated access can rely on PKCE; one that must also call downstream services with service credentials, as in the case documented here, requires the BFF. Implement the Backend for Frontend (BFF) pattern. The correct long-term architecture removes all service credentials from the browser entirely [10, 31]: The BFF stores AppKey and SubscriptionKey in Azure Key Vault or App Service configuration, never in code or environment files that reach the build artifact. It receives the user’s Azure AD JWT, validates it, and makes downstream calls using its own server-side credentials. The browser never holds anything other than the user’s own access token. Implementation options by effort level: Correct the service principal scope. Replace Application permissions with Delegated permissions where possible. An application acting on behalf of a user should see only what that

39

Figure 13: Current architecture (left) versus the Backend for Frontend pattern (right). In the current architecture, AppKey and SubscriptionKey are embedded in the browser JavaScript bundle. The BFF pattern moves all service credentials to a server-side proxy that holds them in Azure Key Vault, so the browser holds only the user’s own JWT. Table 14: Backend-for-frontend implementation options for removing credentials from the browser, by effort level. Option

Effort

Notes

Azure Function proxy Node.js / Express proxy APIM validate-jwt policy

Low Low-Medium Low

Lightweight, scales automatically, fits alongside existing Azure infrastructure Familiar to Angular teams, can share TypeScript types APIM validates user JWT and rejects requests without one, SubscriptionKey removed from browser entirely

user is authorized to see, not every user’s data. This is the configuration that converted a credential exposure into an account takeover chain in both cases documented above. Restrict MSAL.js token caching to in-memory storage. The MSAL.js default caches tokens in localStorage, which is accessible to any JavaScript executing on the page, including XSS payloads [5]. Configuring cacheLocation: "memory" limits token extraction risk at the cost of tokens being lost on page refresh. Audit CORS policy on APIM. Development CORS policies allowing wildcard origins (*) with credentials enabled are commonly left open in production. With exposed credentials and permissive CORS, a malicious third-party site can make authenticated API calls using a visiting user’s session. Lock CORS to specific, explicitly allowed origins. For non-Azure stacks: The same architectural principle applies across providers. AWS applications should move Cognito client secrets and API Gateway keys to Lambda or ECS task roles using IAM instance profiles, never into the React bundle. GCP applications should replace Firebase service account keys in the browser with Firebase App Check and server-side token exchange. SaaS credentials (Twilio, Stripe, SendGrid) should be proxied through a server-side endpoint in all cases, no SaaS secret key belongs in client-side JavaScript regardless of platform.

9.3

Ongoing, Process

Add artifact scanning to the CI/CD pipeline before deployment. This closes part of the build-time injection gap by scanning the build artifact before it reaches production. It catches pipeline-substituted secrets that materialized in dist/ but does not cover runtime-fetched configurations. The scanning step should be inserted after the build stage and before the deployment stage, scanning the compiled output directory rather than source files:

40

Figure 14: Recommended CI/CD pipeline with artifact scan stage inserted between Build and Deploy. The artifact scan targets the compiled output directory (e.g. dist/), not source files. Secrets introduced by pipeline variable substitution are caught before reaching production. The artifact scan step targets the compiled output directory (e.g. dist/), not source files. These are different scan targets. On secrets found, the pipeline should fail and block deployment. Any secret scanning tool with filesystem scanning capability is applicable. Monitor for abuse of already-exposed credentials. Rotation confirms exposure is closed. It does not confirm the credentials were not already used. After rotation: • Filter APIM Analytics by the rotated subscription key for historical usage anomalies • Search Azure AD sign-in logs for client_credentials grants originating from browser IP ranges. This flow is designed for server-to-server communication and should not be initiated from enduser browsers • Alert on any usage of a rotated key post-rotation. This is an active attacker with a cached credential • Azure Sentinel can correlate APIM logs with Azure AD sign-in anomalies for sustained monitoring Eliminate server-side secrets with Azure Managed Identity. For the BFF or backend layer, Managed Identity removes the need to store or rotate AppKey entirely [13]: Azure App Service (BFF) - Managed Identity enabled - No credentials in code, config, or pipeline - Azure handles authentication transparently - BFF calls APIM / Key Vault / Storage with zero stored secrets

Combined with Key Vault references in App Service configuration for any remaining secrets, this eliminates the credential rotation problem for the server-side layer. Add periodic runtime scanning of live deployed applications. Artifact scanning before deployment does not cover runtime-fetched configurations, lazy-loaded chunks served dynamically from CDN, or credentials embedded in third-party scripts that never touch the build pipeline. Periodic scanning of the live application closes this remaining gap. Tools designed for this layer passively monitor live HTTP traffic (JavaScript bundles, HTML source, JSON and XML responses, request headers) and flag credentials as they flow, without outbound verification calls. SecretSifter (the open-source Burp Suite extension evaluated in §8.5) is one such implementation; comparable runtime-aware tooling is appropriate provided it operates on served content rather than source repositories.

41

10

Limitations

The exploitation chains documented in this paper are specific to Azure Active Directory and Azure API Management. The structural paths by which secrets reach production (build-time injection, CI/CD variable substitution, runtime configuration fetching, and scanner suppression) are platformagnostic and appear across AWS, GCP, and SaaS stacks, as noted in Section 5.3.1. However, the specific credential types, token endpoint formats, and remediation guidance in Sections 2 and 9 are Azure-specific and should be adapted to the relevant platform. The empirical figures cited in Section 1 are drawn from a single authorized security engagement covering approximately 2,000 enterprise web application assets within one organization, of which 113 contained credentials in served content. The sample is not randomly selected across organizations. Clopper-Pearson exact 95% confidence intervals are reported alongside each figure: 5.65% (CI: 4.7%–6.8%) for at least one live credential across the 2,000-asset scope, and 55.8% (CI: 46.1%– 65.1%) for a complete Azure AD credential set among the 113 credential-bearing applications. Because the sample is drawn from a single organization’s portfolio (predominantly Angular SPAs with a shared build pipeline) these figures should not be interpreted as a cross-industry populationlevel prevalence estimate. A multi-organization prospectively designed study would be required to produce generalizable estimates. Independent industry-scale measurement provides external corroboration of the underlying exposure pattern: Intruder’s December 2025 sweep [28] reported approximately 42,000 exposed tokens across roughly 5 million applications using regex-driven static detection, a finding consistent with the prevalence direction reported here while differing in scale and detection methodology. Section 8.5 presents a comparison evaluation of nine production scanners against the locked Ground Truth GT-194, 194 unique secret-grade credentials extracted across 113 enterprise applications by Claude Opus 4.7 (reported separately as the LLM-assisted ground-truth reference comparator, not as an evaluated detector), its LLM-extracted candidates cross-validated by GPT-5.5 (Brennan-Prediger κ = 0.676 over the 247 LLM-extracted candidates), and locked at the fielddeployment register on 8 May 2026. The evaluation has five documented limitations. (1) Single-organization corpus and author-developer role. The application pool is drawn from one engagement within one organization (predominantly Angular SPAs with a shared build pipeline). The first author (Gorijala) is also the developer of SecretSifter, one of the evaluated scanners. The headline recall numbers in Table 8 should therefore be read as scanner performance on this specific Azure-heavy corpus, not as cross-organizational population estimates. The structural mitigations applied (independent LLM Ground Truth extraction, cross-vendor classification, exactvalue substring matching with no per-tool tuning of the comparison protocol, the same 6-rule overlay applied identically to all rule-extensible tools) prevent the comparison from being engineered toward a specific tool, but they do not eliminate the corpus-specificity of the absolute numbers. A multiorganization replication on an independently constructed held-out benchmark would be required to claim cross-industry generalization, and is committed to as a direct follow-on study. (2) Sample frame. The 113-application benchmark subset is drawn from the same engagement that produced the Ground Truth, not from an independently sampled population. Recall and per-class results should be interpreted as measurements of scanner performance against GT-194, not as cross-organizational performance estimates. The ground-truth construction methodology (independent LLM extraction with cross-vendor validation) is the contribution that generalizes. The recall figures are anchored to this corpus. (3) Tightened-configuration scope. The 6-rule overlay applied to the rule-extensible tools (TruffleHog, SecretFinder, Titus, JSluice) was sized to cover the most prevalent credential shapes in this corpus and applied identically across all four tools (Appendix A). A more aggressive tuning 42

frontier, full vendor-rule libraries, customer-specific patterns, or per-tool optimization, may further close the static-tool gap. The point of the comparison is not that 6 rules are enough; it is that even a corpus-targeted overlay applied uniformly leaves a residual gap of 41 percentage points to a runtime-aware scanner, on this corpus. (4) Ground-truth reference limitations. The LLM-assisted ground-truth reference is itself a model with its own biases. Claude Opus did not surface 5 of 37 CryptoJS-AES blobs and 18 of 153 Azure AD App IDs in the chain-context tier, demonstrating that the reference model has structural blind spots that curated detection rules complement. The CryptoJS blobs the reference comparator missed were recovered by SecretSifter and are therefore not part of the tool-agnostic 27; that blind spot instead comprises credentials surfaced only by manual analysis, so the gap is a property of the underlying credential corpus rather than an LLM-vs-static comparison artefact. (5) Per-credential vs per-application generalization. GT-194 is a per-credential ground truth; the chain-completion finding in §8.5.7 is a per-application corpus property derived from 86 secret-exposed applications. The 73.3% (63/86) figure is descriptive of this engagement, not inferential of cross-industry chain-completion rates. A multi-organization replication is required for inferential claims about full-chain reconstruction rates. Relatedly, credentials are clustered within applications (multiple GT-194 credentials often originate from the same application bundle), so the per-credential observations that the McNemar and Holm-Bonferroni analyses (§8.5.10) treat as units are not fully independent deployment events. The significance results should therefore be read as pairwise detection-difference tests over the credential set, not as inferences over independent applications; a cluster-aware or per-application analysis is left to multi-organization replication.

11

Future Work

Multi-organization prevalence study. The 5.65% credential exposure rate reported in this paper is drawn from a single organization’s application portfolio. A prospectively designed study across multiple organizations (stratified by industry, stack, and deployment maturity) is required to produce a generalizable prevalence estimate. A multi-organization replication study, comparable in scope to the 2,000-application engagement reported here, is planned as a direct follow-on. Cross-organization precision benchmark. §8.5.5 reports direct precision, recall, and F1 for the nine production scanners against the GT-194 corpus (with Claude Opus reported separately as the ground-truth reference comparator). Replicating the same precision measurement on a heldout, multi-organization production corpus, with full TP / FP classification of every emitted finding by an independent validator, would test whether the precision ordering reported here transfers across stacks. The qualitative ordering (parser-based scanners produce few false positives, regexbased scanners produce many, SecretSifter’s noise-suppression layer materially differentiates it from comparable static tools) is robust under any reasonable re-measurement, but the exact precision and F1 numbers per scanner would tighten. Multi-organization held-out benchmark. GT-194 is constructed from a single engagement. Replicating the same construction methodology on a held-out benchmark drawn from multiple unrelated organizations would test whether the recall numbers reported here transfer across stacks, build pipelines, and credential conventions. The construction methodology (Claude Opus extraction + GPT-5.5 cross-vendor validation) is fully described in §8.5.1 and is reproducible by any practitioner with API access to both vendors. Runtime-fetched configuration coverage. GT-194 evaluates scanners against downloadable JavaScript bundles only, the surface for which static-file scanners can be applied uniformly. 43

Modern SPAs load credentials via runtime-fetched configuration endpoints (/config.json, /api/ settings) and HTML SSR state blobs that are never present in the initial bundle. Extending the benchmark methodology to runtime-fetched and SSR-injected credentials is the natural next step and would close the gap between the bundle-scanning protocol described here and the productiontraffic surface that real attackers actually inspect. Automated exploitation chain validation. The current study manually confirmed exploitability for both documented chains. Automated scope-aware validation, checking whether exposed Azure AD credentials have client_credentials grant enabled, or whether exposed APIM keys map to APIs with sensitive data, would allow prioritization of findings at scale without manual verification per application. Longitudinal remediation tracking. Both applications documented in this paper remediated within the assessment window. Whether organizations with incidentally discovered runtime credential exposure (found via bug bounty or passive scanning) remediate at comparable rates is unknown. A longitudinal study tracking time-to-remediation across disclosure channels would inform remediation guidance.

12

Responsible Disclosure

Both exploitation chains documented in this paper were identified during authorized security assessments conducted within the scope of formal engagements. In both cases: • Findings were reported immediately to the affected organization upon identification • Proof-of-concept evidence was limited to what was necessary to confirm exploitability, no bulk data extraction occurred • Remediation was implemented by the affected organization • Findings were verified through retest to confirm remediation before this publication No user data was retained beyond the assessment period. No credentials were used outside the scope of confirming exploitability. Organization names and application identifiers have been anonymized throughout this paper.

13

Conclusion

The shift-left investment the security industry has made is real, justified, and valuable. Catching secrets before they reach production is always preferable to finding them after. GitLeaks, TruffleHog, GitHub Advanced Security, and the SAST tools deployed in enterprise CI/CD pipelines do exactly what they are designed to do. But secrets are still reaching production, not because these tools fail, but because the path from development to deployment has branches that no pre-deployment scanner covers. Build-time environment injection, CI/CD pipeline variable substitution, and runtime configuration fetching each produce secrets in production that never existed in any repository at any point. The tools designed to prevent credential exposure scan a layer that those secrets never passed through. The exploitation chains documented above are not theoretical edge cases. They are patterns found repeatedly in real production applications serving real users, applications that had passed 44

every security review and had shift-left tooling in place. None of those tools provided automated visibility into what each application served, or automated detection when credentials appeared in delivered JavaScript. The shift-right tooling gap is now quantified, and includes a tool-agnostic blind spot. On the locked Ground Truth GT-194, 13.9% of credentials are missed by every one of the nine production scanners evaluated and are surfaced only by manual analysis. The CryptoJS encryptedconfiguration class is invisible to every static scanner by design and is recovered only by runtimeaware and LLM-based detection. The combined coverage of the nine production scanners together with the ground-truth reference comparator plateaus at 86.1%, leaving that 13.9% as a structural property of the credential surface itself rather than a tooling deficit that more rules can close. Beyond detection, the chain-completion finding (73.3% of affected applications co-locate the full Azure AD token-mint chain client-side) shows that the credentials reachable in served JavaScript are operationally weaponizable from browser-visible code alone. Industry-scale measurement corroborates this gap: Intruder’s December 2025 sweep, an industry vendor report [28], identified approximately 42,000 exposed tokens across roughly 5 million applications using regex-driven static detection alone, a lower-bound observation that the pattern documented here is widespread at industrial scale. The shift-left layer is covered. The runtime layer is not. That asymmetry, and the additional 13.9% blind spot that no evaluated production scanner closes on this corpus, is what is being exploited. Closing it requires the same community effort that built shift-left tooling into what it is today, practitioners documenting the gap, security teams demanding runtime scanning in their programs, and tool authors building for the layer that has been ignored. Bug bounty researchers finding runtime credentials should escalate beyond the credential itself and document the full exploitation chain. Penetration testers should treat runtime JavaScript as a first-class target surface, not an afterthought. Security teams should ask their tooling vendors a direct question: does your scanner cover what the application serves, or only what developers commit? The patterns in this paper will not stop appearing until the runtime layer gets the same tooling investment the repository layer already has.

References [1] OAuth 2.0 Authorization Framework (RFC 6749), Hardt, D. (Ed.), Internet Engineering Task Force, 2012. Sections 2.1 (Client Types) and 4.4 (Client Credentials Grant). https://www.rfc-editor.org/rfc/rfc6749 [2] OWASP Top 10:2021 (A02: Cryptographic Failures) OWASP Foundation, 2021. http s://owasp.org/Top10/A02_2021-Cryptographic_Failures/ [3] OWASP Top 10:2021 (A05: Security Misconfiguration) OWASP Foundation, 2021. https://owasp.org/Top10/A05_2021-Security_Misconfiguration/ [4] Microsoft Identity Platform (Client Credentials Flow) Microsoft Documentation. ht tps://learn.microsoft.com/en-us/entra/identity-platform/v2-oauth2-client-cre ds-grant-flow [5] Microsoft MSAL.js Token Caching, Microsoft Authentication Library for JavaScript, Token cache documentation. https://learn.microsoft.com/en-us/azure/active-directo ry/develop/msal-js-token-cache 45

[6] webpack DefinePlugin, webpack Documentation. https://webpack.js.org/plugins/def ine-plugin/ [7] Azure API Management Subscription Keys, Microsoft Documentation. https://lear n.microsoft.com/en-us/azure/api-management/api-management-subscriptions [8] How Bad Can It Git: Characterizing Secret Leakage in Public GitHub Repositories, Meli, M., McNiece, M.R., Reaves, B., Proceedings of the 2019 Network and Distributed System Security Symposium (NDSS). https://www.ndss-symposium.org/ndss-paper/how -bad-can-it-git-characterizing-secret-leakage-in-public-github-repositories/ [9] NIST SP 800-53 Rev. 5 (Security and Privacy Controls: IA-5 Authenticator Management) National Institute of Standards and Technology, 2020. https://csrc.nist.gov/ publications/detail/sp/800-53/rev-5/final [10] Backends For Frontends, Newman, S., samnewman.io, November 2015. https://samnew man.io/patterns/architectural/bff/ [11] Recommendation for the Entropy Sources Used for Random Bit Generation, Turan, M.S., Barker, E., Kelsey, J., McKay, K., Baish, M., Boyle, M., NIST Special Publication 80090B, National Institute of Standards and Technology, January 2018. https://doi.org/10.6 028/NIST.SP.800-90B [12] CWE-321: Use of Hard-coded Cryptographic Key, MITRE Common Weakness Enumeration, 2023. https://cwe.mitre.org/data/definitions/321.html [13] Azure Managed Identity Overview, Microsoft Documentation. https://learn.micros oft.com/en-us/entra/identity/managed-identities-azure-resources/overview [14] Luhn Algorithm (ISO/IEC 7812-1:2017, Identification cards) Identification of issuers. [15] Keys on Doormats: Exposed API Credentials on the Web, Demir, N., Vekaria, Y., Smaragdakis, G., and Durumeric, Z., arXiv preprint arXiv:2603.12498, March 2026. https: //arxiv.org/abs/2603.12498 [16] A Comparative Study of Software Secrets Reporting by Secret Detection Tools, Basak, S.K., Cox, J., Reaves, B., and Williams, L.A., arXiv preprint arXiv:2307.00714, July 2023. https://arxiv.org/abs/2307.00714 [17] The Skeleton Keys: A Large Scale Analysis of Credential Leakage in Mini-Apps, Shi, Y., Yang, Z., Zhong, K., Yang, G., Yang, Y., Zhang, X., and Yang, M., Proceedings of the 2025 Network and Distributed System Security Symposium (NDSS). https://www.ndss -symposium.org/wp-content/uploads/2025-273-paper.pdf [18] Jack-in-the-box: An Empirical Study of JavaScript Bundling on the Web and its Security Implications, Rack, J., and Staicu, C.-A., Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security (CCS), November 2023. https://pu blications.cispa.saarland/4036/1/Bundlers_Study_Submission-camera-ready.pdf [19] Thou Shalt Not Depend on Me: Analysing the Use of Outdated JavaScript Libraries on the Web, Lauinger, T. et al., Proceedings of the 2017 Network and Distributed System Security Symposium (NDSS). https://www.ndss-symposium.org/wp-content/upl oads/2017/09/ndss2017_02B-1_Lauinger_paper.pdf 46

[20] Study of JavaScript Static Analysis Tools for Vulnerability Detection in Node.js Packages, Brito, T., Ferreira, M., Monteiro, M., Lopes, P., Barros, M., Fragoso Santos, J., and Santos, N., arXiv preprint arXiv:2301.05097, January 2023. https://arxiv.org/pdf/23 01.05097 [21] On Detecting and Measuring Exploitable JavaScript Functions in Real-World Applications, Kluban, M., Mannan, M., and Youssef, A., ACM Transactions on Privacy and Security (TOPS), 27(1): 1–37, 2023. https://users.encs.concordia.ca/ ~ mmannan/publications/JS-vulnerability-TOPS.pdf [22] Claude Opus 4.7 System Card, Anthropic, April 2026. https://www.anthropic.com/cl aude-opus-4-7-system-card [23] GPT-5.5 System Card, OpenAI, 2026. https://openai.com/index/gpt-5-5-system-c ard/ [24] Coefficient kappa: Some uses, misuses, and alternatives, Brennan, R.L., Prediger, D.J., Educational and Psychological Measurement, 41(3): 687–699, 1981. [25] High Agreement but Low Kappa: I. The Problems of Two Paradoxes, Feinstein, A.R., Cicchetti, D.V., Journal of Clinical Epidemiology, 43(6): 543–549, 1990. [26] The Measurement of Observer Agreement for Categorical Data, Landis, J.R., Koch, G.G., Biometrics, 33(1): 159–174, 1977. [27] UpSet: Visualization of Intersecting Sets, Lex, A., Gehlenborg, N., Strobelt, H., Vuillemot, R., Pfister, H., IEEE Transactions on Visualization and Computer Graphics (Proc. InfoVis), 20(12): 1983–1992, 2014. https://doi.org/10.1109/TVCG.2014.2346248 [28] Intruder Uncovers New Secrets Detection Techniques, Finds Thousands of Exposed Tokens Unaddressed by Traditional Methods, Intruder Security Ltd., Business Wire, December 2025. https://www.businesswire.com/news/home/20251211585215/en/I ntruder-Uncovers-New-Secrets-Detection-Techniques-Finds-Thousands-of-Exposed -Tokens-Unaddressed-by-Traditional-Methods [29] Proof Key for Code Exchange by OAuth Public Clients (RFC 7636), Sakimura, N. (Ed.), Bradley, J., Agarwal, N., Internet Engineering Task Force, September 2015. https: //www.rfc-editor.org/rfc/rfc7636 [30] OAuth 2.0 for Browser-Based Applications, Parecki, A., De Ryck, P., Waite, D., IETF OAuth Working Group Internet-Draft draft-ietf-oauth-browser-based-apps-26 (Best Current Practice), December 2025. https://datatracker.ietf.org/doc/draft-ietf-oauth-brows er-based-apps/ [31] Microservices Patterns: With Examples in Java, Richardson, C., Manning Publications, 2018. ISBN 978-1617294549.

A

Tightened-Configuration Rule Set

The four rule-extensible scanners evaluated in §8.5 (TruffleHog, SecretFinder, Titus, JSluice) were each run with the same 6-rule overlay. The rules cover the credential shapes most prevalent in 47

GT-194. The full rule set is reproduced below verbatim. The same overlay was applied identically to all four tools; no per-tool tuning was performed. - id: custom.azure.client_id name: Azure AD Client ID (UUID v4) pattern: ’\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b’ - id: custom.azure.ad_secret name: Azure AD Client Secret (tilde-format, 34-40 char) pattern: ’\b[A-Za-z0-9_.\-]{1,16}~[A-Za-z0-9_.\-~]{16,40}\b’ - id: custom.azure.apim_key name: Azure APIM Subscription Key (32-hex) pattern: ’\b[0-9a-f]{32}\b’ - id: custom.jwt.token name: JSON Web Token (eyJ-prefix three-segment) pattern: ’\beyJ[A-Za-z0-9_\-]{10,}\.[A-Za-z0-9._\-]{10,}\.[A-Za-z0-9._\-]+’ - id: custom.azure.app_insights_ikey name: App Insights Instrumentation Key pattern: ’InstrumentationKey=[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}’ - id: custom.cryptojs.salted name: CryptoJS U2FsdGVkX1-prefix encrypted blob pattern: ’U2FsdGVkX1[A-Za-z0-9+/=]{16,}’

These six rules were selected from the credential shapes most prevalent in GT-194, derived without consulting any tool’s output. They are not a comprehensive coverage set; they are a minimal corpus-targeted overlay applied identically to all four rule-extensible tools (TruffleHog, SecretFinder, Titus, JSluice) so that the comparison reflects scanner-architecture differences rather than rule-library differences. JSluice was run with the same six patterns translated to the JSluice user-pattern JSON schema; the built-in JSluice secret rules silently produce zero output on minified webpack bundles in this corpus, and the user-pattern overlay is what allows JSluice to surface the key-value extractions that drive its reported recall. JSMiner exposes no user-rule API in the Burp extension distribution, so it was reimplemented in Python as a faithful clone of the published Trustwave SpiderLabs detection logic (21 keyword anchors plus Shannon-entropy scoring with the same Firm/Tentative classification thresholds) and run on the GT-194 corpus. Nuclei was run in the URL-mode JavaScript-spider configuration matching the Intruder December 2025 industryscale measurement [28]: bundles served via a local Python HTTP server on port 9101, scanned with the secret-scanning template categories http/exposures/tokens, http/exposures/apis, and http/ exposures/keys plus the secret/token/key/exposure template tags. The file-mode Nuclei configuration produces zero findings on raw .js files because Nuclei’s secret-scanning template library is HTTPresponse-oriented rather than file-content-oriented, a methodology mismatch documented here for reproducibility. Cariddi was run at its built-in default secret-detection configuration without any rule overlay (its primary mode is web-crawling and endpoint enumeration, with secondary secret detection). SecretSifter was run at its built-in configuration with the field-deployment scan output exported via the SecretSifter REST API.

Author contributions (CRediT): Gorijala, conceptualization, methodology, software (SecretSifter), investigation, validation, formal analysis, data curation, writing of original draft, writing of review and editing, project administration. 48

Declaration of generative AI use: This work used AI tools in two ways. As research instruments, Claude Opus 4.7 (Anthropic) built the GT-194 ground-truth benchmark and GPT-5.5 (OpenAI) independently validated its LLM-extracted candidates, with the manual-only additions verified by analyst review; Section 8.5 documents this and cites both systems as [22] and [23]. Generative AI also assisted with drafting and language editing of the manuscript. The author checked every statistic, citation, and technical claim, revised the text, and is fully responsible for its accuracy, originality, and integrity. No AI system is an author. No AI tool produced research data, results, or analysis beyond the benchmark construction described in Section 8.5. Conflict of interest disclosure: The author (Gorijala) is the developer of SecretSifter, which is referenced in Section 9.3 as a tool for runtime scanning and evaluated as one of the nine production scanners in Section 8.5. The author’s role as developer is mitigated structurally in the §8.5 evaluation: the Ground Truth (GT-194) was constructed by independent Claude Opus 4.7 (Anthropic) extraction and manual analyst review, and independently validated by GPT-5.5 (OpenAI), so no scanner (including SecretSifter) defines its own evaluation set; the union also contains credentials SecretSifter did not detect, making the benchmark strictly harder for it rather than easier. The SecretSifter edition evaluated here (the Burp Suite extension) is open source. The exploitation chains, structural gap analysis, and prevalence figures in this paper are independent of that tool and predate its development.

49

Record · ID 1028624 · SHA-256 d3ae5d7b5fcaf379
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.