Conceptio › Archive › arXiv CS
arXiv CSopen access

Toward Sustainable AI Deployment: A Carbon-Aware Decision Framework for Enterprise Supply Chain Systems

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Toward Sustainable AI Deployment: A Carbon-Aware Decision Framework for Enterprise Supply Chain Systems

arXiv:2609.14881v1 [cs.SE] 14 Sep 2026

Haoran Yu* University of Florida United States [email protected]

Lifei Liu Wichita State University United States [email protected]

Abstract

1.

Danping Zhang Nanchang Hangkong University China [email protected]

Introduction

Enterprises deploying AI for supply chain decisions commonly default to the largest available language model, a procurement heuristic that neglects both empirical performance and environmental cost. We benchmark six large language models across 520 supply chain tasks, simultaneously measuring decision quality and estimated generation-related operational carbon. Drawing on the Technology-Organization-Environment (TOE) framework, we develop a Carbon-Aware AI Procurement Framework (CAAPF), a Green IS design artifact that operationalizes sustainable AI governance for enterprise procurement. Within this bounded sample, quality spans 0.497–0.723, and the models with the largest disclosed parameter totals do not achieve the highest scores. The design does not isolate size, provider, architecture, or benchmark-construction effects. A category-by-tier calibrated GreenRoute proof of concept reaches 0.733 mean out-of-sample quality at an estimated 0.402 gCO2 /task. Static Haiku reaches 0.699 at 0.022 gCO2 /task, while Sonnet reaches 0.723 at 0.401 gCO2 /task, demonstrating that the preferred strategy depends on the organization’s quality requirement. Our “benchmark first, select green” principle suggests that environmental responsibility and decision quality can be mutually reinforcing, contributing to sustainable digital infrastructure governance aligned with SDG 12 and SDG 13.

Recent work examines language models as interfaces to supply chain decision tools and optimization systems (S. Huang et al., 2026; B. Li et al., 2023; Simchi-Levi et al., 2026). Model procurement often defaults to the largest available language model on the assumption that more parameters yield better decisions. This size-based heuristic can carry substantial environmental cost. LLM inference produces carbon dioxide emissions through the electricity consumed by GPU clusters (Luccioni et al., 2024; Wiesner et al., 2025). Across repeated inventory, supplier-risk, and forecasting queries, differences in per-task energy use can accumulate into a material operational footprint. The United Nations Sustainable Development Goals (SDGs) 12 and 13 call for responsible consumption and climate action (United Nations, 2015), but environmental impact is not routinely integrated into model-selection decisions. This paper addresses the intersection of Green Information Systems (Green IS), supply chain decisionmaking, and AI model selection. Prior work has quantified training-time carbon costs (Patterson et al., 2022; Strubell et al., 2019) and proposed cost-efficient routing (L. Chen et al., 2024; Ong et al., 2025). Evidence remains limited on how quality and operational carbon vary together on supply-chain-specific tasks and on how organizations should translate such measurements into an auditable procurement decision. We address three research questions:

Keywords: Green IS, sustainable AI, supply chain management, carbon footprint, decision intelligence

RQ1:

How does the quality–carbon tradeoff vary across LLMs for supply chain decision tasks?

RQ2:

How are observed differences in model size, model family, and architecture associated with decision quality in the evaluated supply chain

* Corresponding author.

tasks? RQ3:

How can a theoretically grounded framework guide threshold-contingent, sustainable AI model selection?

Drawing on TOE (Tornatzky & Fleischer, 1990), we develop and empirically examine a Carbon-Aware AI Procurement Framework (CAAPF). Our contributions are fourfold: 1. A threshold-contingent procurement framework that links domain benchmarking to organizationspecific quality, carbon, cost, risk, and governance constraints. 2. Descriptive evidence that nominal model size is an unreliable procurement proxy in the six-model sample, together with explicit limits on crossarchitecture, provider, and benchmark-style inference. 3. Identification of a model-specific deliberation pattern in which longer DeepSeek R1 responses coincide with lower accuracy and higher per-task carbon on formula tasks. 4. A boundary-aware evaluation showing when simple model selection or escalation is preferable to routing and when task-level routing may add value.

2.

Theoretical Background

2.1.

Green IS and Sustainable AI

Green Information Systems research examines how IT artifacts can be designed, deployed, and governed to enable environmentally responsible outcomes (Melville, 2010). Within this tradition, the environmental impacts of AI systems represent an emerging concern: while AI can support sustainability goals, the computational infrastructure underlying AI systems itself carries significant environmental costs (Verdecchia et al., 2023; Wu et al., 2022). The concept of Green AI was formalized by Schwartz et al. (2020), who argued that the AI community disproportionately rewards accuracy gains achieved through computational brute force while neglecting efficiency. Strubell et al. (2019) showed that extensive architecture search and tuning can dominate the emissions of a final training run. Luccioni et al. (2024) extended measurement to inference and found large energy differences across architectures and task categories. Wiesner et al. (2025) argued that growing inference-time compute for reasoning models can outpace hardwareefficiency gains.

On the measurement side, Anthony et al. (2020) developed Carbontracker for real-time energy monitoring, while Lacoste et al. (2019) created the Machine Learning Emissions Calculator. Dodge et al. (2022) demonstrated that carbon intensity varies substantially with cloud region and time of day, suggesting scheduling and geographic placement as mitigation levers. Samsi et al. (2023) benchmarked LLM inference energy costs and noted that inference energy had received less attention than training energy. For supply chain applications, Simchi-Levi et al. (2026) described LLM interfaces for explaining tool recommendations, exploring what-if scenarios, and updating decision models. B. Li et al. (2023) combined an LLM interface with optimization code and identified ambiguity, generated-code errors, and outof-distribution use as limitations. Moving from explanation toward prescriptive decision support, S. Huang et al. (2026) developed a decision-aware causal intervention ranking approach for critical supply chains, prioritizing interventions by their estimated causal effect on outcomes rather than by predictive fit alone. These concerns motivate direct evaluation of numerical and domain-specific decisions rather than assuming that model scale will transfer into operational quality.

2.2.

Model Routing and Selection

The economics of LLM deployment have motivated research into intelligent model selection. L. Chen et al. (2024) introduced FrugalGPT, reducing inference cost by up to 98% through LLM cascades that route easy queries to cheap models and hard queries to expensive ones. Ong et al. (2025) developed RouteLLM, learning routing policies from preference data with substantial cost savings and minimal quality degradation. Cruciani and Verdecchia (2025) explicitly connected model selection to environmental sustainability, calling for empirical validation across specific domains. Sardana et al. (2024) extended scaling laws to account for inference compute, showing that optimal model size depends on query volume. Neither work included supply chain tasks or carbon measurements.

2.3.

TOE as Analytical Framework for AI Procurement

TOE explains organizational technology adoption through three contexts: the technology context covers the attributes and availability of candidate technologies; the organization context covers readiness, resources, governance, and operational requirements; and the environment context covers competition, regulation, and external stakeholder pressure (Baker, 2012; Tornatzky &

Fleischer, 1990; Zhu et al., 2006). Green IS research further treats information-system choices as mechanisms through which organizations can act on environmental objectives (Melville, 2010). We use TOE to specify what an empirical benchmark alone cannot determine. First, the Technology context directs attention to relevant attributes of candidate technologies; we operationalize these attributes through observed task quality and resource intensity rather than a parameter-count proxy. Second, the Organization context identifies the internal conditions that determine admissibility: decision criticality, an acceptable-quality threshold (τ ), API cost, latency, data governance, vendor risk, and oversight requirements. Third, the Environment context identifies external sustainability and accountability pressures that motivate emissions documentation and periodic re-evaluation. CAAPF translates these contexts into decision rules rather than using TOE only to label benchmark variables. We therefore derive three propositions. P1 and P2 are examined empirically, whereas P3 is instantiated as a design proposition whose organizational effects require later validation. P1: Nominal model size does not reliably rank technology relative advantage when quality and estimated generation-related carbon are evaluated jointly. P2: The preferred selection strategy changes with organization-specific quality and risk thresholds, so simple selection may dominate routing at moderate thresholds. P3: A TOE-based process produces an auditable model-selection record that incorporates environmental pressure without treating carbon as the only procurement criterion.

3.

Carbon-Aware AI Procurement Framework

Figure 1 presents CAAPF, a decision process for sustainable AI procurement. It separates the organization’s admissibility decision from the technical optimization performed after admissible candidates have been identified. The framework’s core principle is “benchmark first, select green”. The selection process is: define a representative task portfolio; set τ from the operational consequences of error; record non-quality constraints such as API cost, latency, security, data residency, and vendor risk; benchmark every candidate under a common prompt protocol; form the set satisfying all constraints; and select the lowest-carbon member. CAAPF can apply quality requirements at the portfolio level or, for heterogeneous workloads, separately to task classes. A model that passes an aggregate threshold is not necessarily admissible when individual task classes must each

Carbon-Aware AI Procurement Framework (CAAPF) 1. Define Task, Risk, and Oversight Requirements ↓ 2. Benchmark Candidate Models on Domain Tasks ↓ 3. Measure Quality, Carbon, Cost, and Latency ↓ 4. Apply Quality and Governance Constraints ↓ 5. Select Lowest-Carbon Admissible Model ↓ 6. Continuous Monitoring & Re-evaluation Figure 1. CAAPF links TOE contexts to an auditable procurement sequence. Technology is profiled in Steps 2–3, organizational constraints define admissibility in Steps 1 and 4, and environmental accountability motivates monitoring.

Table 1. CAAPF decision matrix; thresholds must be calibrated to organizational risk. Requirement

Strategy

Decision Rule

Moderate portfolio quality

Static model

Higher quality met by one model

Lowest-carbon passing model

Uneven quality across task classes Safety-critical

Calibrated routing

low-carbon

Human or deterministic oversight

Select the lowestcarbon single model satisfying constraints Routing adds no value if one model passes all constraints Route only if a mixture improves the feasible frontier Require validation and an audit trail

meet a service floor. A router is justified only if tasklevel heterogeneity improves this rule relative to simple selection or escalation. Monitoring reopens the decision when models, workloads, prices, or grid conditions change. We position CAAPF as a Green IS design artifact that makes model selection inspectable. Table 1 states the general decision logic without transferring numerical cutoffs from this benchmark to other organizations. In practice, organizations should derive τ from tasklevel loss, regulatory duties, existing human or system performance, and required oversight, and should reject deployment when no tested configuration satisfies those conditions. In our GreenRoute implementation, the study-specific value τ = 0.65 is applied separately to each of 18 category×tier calibration cells, not only to aggregate portfolio quality. Thus, Haiku’s aggregate score of 0.699 does not by itself establish that it passes every task-class requirement.

4.

Methodology

4.1.

Task Taxonomy

We construct a benchmark of 520 supply chain tasks spanning six operational categories: Demand Forecasting (92 tasks; 33 using real M5 Competition data (Makridakis et al., 2022), a setting where large-scale retail forecasting must balance accuracy against operational stability (J. Li et al., 2026)), Vehicle Routing (84), Inventory Optimization (96), Supplier Risk Assessment (88), Order Fulfillment (80), and Demand Classification (80). Three difficulty tiers reflect decision complexity: Tier 1 (formula application, ∼35%): single-step calculations with deterministic solutions (e.g., computing EOQ). Tier 2 (multi-factor reasoning, ∼40%): problems requiring integration of multiple inputs and conditional logic. Tier 3 (strategic judgment, ∼25%): openended decisions requiring synthesis of quantitative analysis with qualitative factors. Tasks were authored by experts in operations management following a structured protocol based on established supply chain planning topics (Chopra & Meindl, 2019). For each category × tier combination, tasks span recurrent forecasting, inventory, routing, risk, fulfillment, and classification problem families while maintaining consistent difficulty calibration. Each task underwent three-stage validation: (a) independent solution verification by a separate author, (b) full-team review for domain accuracy and difficulty calibration, and (c) pilot testing on two models (one small, one large) to confirm meaningful performance variation. We acknowledge that tasks were not validated by external practicing supply chain professionals, and real-world decisions involve richer organizational context that single-prompt evaluation does not capture.

4.2.

Supply Chain Domain Knowledge Tests

Beyond general task performance, we design 30 “SC-trap” tasks that specifically probe supply chain domain knowledge. These tasks are constructed so that a model lacking domain expertise produces plausiblesounding√but incorrect answers. Examples include testing the LT safety stock relationship, bullwhip quantification, risk pooling benefits, EOQ sensitivity, and newsvendor critical ratio application.

4.3.

Models and Carbon Estimation

Table 2 presents the six LLMs evaluated. The availability-based sample covers four providers, but four of the six models come from Anthropic and Meta; it is

Table 2. Models evaluated. Carbon estimates carry ±50% uncertainty. Model

Provider

Reported Parameters

gCO2 /Mtok

Claude Sonnet 4.6 Claude Haiku 4.5 Mistral Large 3 Llama 4 Scout Llama 3.3 70B DeepSeek R1

Anthropic Anthropic Mistral Meta Meta DeepSeek

Not disclosed Not disclosed 675B total / 41B active 109B total / 17B active 70B dense total 671B total / 37B active

900 70 5,500 60 980 6,000

Parameter definitions are architecture-specific and not directly comparable.

not representative of the broader LLM market. Provider disclosures report both total and active parameters for several MoE models (DeepSeek-AI, 2024; Meta AI, 2025; Mistral AI, 2025), while Anthropic does not disclose parameter counts for the evaluated Claude models. Dense totals, MoE totals, and active parameters are not equivalent measures of inference compute, so they are reported descriptively and are not used in a cross-model correlation. Following prior measurement work (Dodge et al., 2022; Luccioni et al., 2024), Table 2’s gCO2 /Mtok coefficients are study-specific engineering estimates based on assumed serving hardware, GPU power, throughput, PUE, and grid intensity, not provider measurements or direct datacenter observations. Study assumptions are PUE 1.1–1.2 and US-East intensity 0.38 kgCO2 /kWh; all estimates carry ±50% uncertainty. Generation-related per-task carbon is calculated as model-specific gCO2 /Mtok × mean output tokens /106 . Because complete input/prefill token records were not retained, these values estimate outputgeneration carbon rather than complete end-to-end serving energy.

4.4.

Prompting and Evaluation Framework

All candidate models receive the same task text and category-specific system instruction. Task-specific output caps are 1,024, 1,536, or 2,048 tokens, and candidate-model temperature is set to zero. This fixed protocol controls prompt wording across models but does not test how alternative prompting, demonstrations, or output constraints affect quality, response length, and carbon. We employ LLM-as-Judge evaluation (Zheng et al., 2023) with two independent judges from different providers: Claude Opus 4 (Anthropic) and Qwen3-235B (Alibaba). Each judge scores numeric accuracy, reasoning quality, and SC domain knowledge on 1–10 scales; the three dimensions are averaged equally and then normalized. Qwen uses temperature zero, while Opus uses its provider default because the endpoint rejects a temperature setting. Model identifiers are stripped from responses, and tasks are presented in randomized order.

For Tier-1 tasks (n = 182), we validate against deterministic ground truth (exact match within 5% tolerance), providing judge-independent confirmation that LLM rankings are not artifacts of evaluator bias.

4.5.

Table 3. Overall performance on 520 supply chain tasks.

GreenRoute Implementation

We instantiate CAAPF through GreenRoute, an illustrative routing implementation rather than a claim of routing-method novelty. GreenRoute calibrates a policy separately for each of the 18 category×tier cells. On a calibration split, it selects the model with the lowest estimated generation-related carbon per task among models whose cell-level mean quality meets τ = 0.65; if no model passes, it selects the highest-quality model. Thus, τ is a cell-level calibration threshold intended to target, rather than guarantee, a task-class service floor. The reported routing result uses 20 repeated stratified 50/50 split-half trials generated by a random-number generator initialized with seed 42, with each cell represented in both halves: mean quality is 0.733 (SD = 0.004), mean estimated carbon is 0.402 gCO2 /task (SD = 0.124), and mean savings versus always DeepSeek is 97.6%. Across these trials, 296 of 360 held-out cell–trial means (82.2%; trial-level SD = 2.3 percentage points) meet τ . This policy uses benchmark-provided category and tier metadata and does not claim learned text-classification novelty. A production implementation would require an independent metadata rule, task classifier, or uncertaintyaware difficulty estimator; its errors and overhead are not evaluated here. Continuous monitoring and recalibration remain CAAPF governance recommendations; they are not additional test-set results.

5.

Results

5.1.

Quality and Carbon Across Evaluated Models

Score

gCO2 /Mtok

GT%

Claude Sonnet 4.6 Claude Haiku 4.5 Mistral Large 3 Llama 3.3 70B Llama 4 Scout DeepSeek R1

0.723 0.699 0.613 0.528 0.521 0.497

900 70 5,500 980 60 6,000

78.6 75.3 69.2 59.8 60.4 57.1

GT% = Tier-1 ground-truth accuracy.

Table 4. Estimated generation-related carbon per task, accounting for mean output length. Model Claude Haiku 4.5 Llama 4 Scout 17B Claude Sonnet 4.6 Llama 3.3 70B Mistral Large 3 DeepSeek R1

Tok/Task

gCO2 /Task

vs. Haiku

312 387 445 478 623 2,847

0.022 0.023 0.401 0.468 3.427 17.082

1.0× 1.1× 18.3× 21.4× 156.6× 780.4×

net, which provides the highest absolute quality. All other models are strictly dominated because they produce both lower quality and higher emissions than at least one Pareto-optimal alternative.

5.2.

Per-Task Carbon: The Verbosity Amplification Effect

Per-token carbon intensity can understate workloadlevel generation-related emissions when models produce substantially different output lengths (Table 4). DeepSeek R1 generates 2,847 tokens per task (9.1× Haiku), resulting in estimated output-generation carbon per task 780× higher. Under our generation-related estimates, 1,000 daily DeepSeek R1 queries produce approximately the same output-generation carbon as 780,000 Haiku queries while achieving lower benchmark quality.

5.3. Table 3 presents aggregate results. Quality spans 0.497–0.723. The highest disclosed parameter totals belong to Mistral Large 3 and DeepSeek R1, but neither is the highest-scoring candidate. Because Anthropic does not disclose the evaluated Claude models’ parameter counts and dense and MoE counts are not directly comparable, we treat nominal size as a descriptive characteristic rather than a statistically identified predictor. The sample does not establish that provider or training data caused the observed differences, because provider, architecture, model generation, and benchmark fit are confounded. The observed Pareto frontier consists of Haiku, which provides the lowest estimated generation-related carbon among the high-performing models, and Son-

Model

Performance by Task Difficulty

Table 5 disaggregates performance by difficulty tier. Cross-model performance differences are most pronounced on Tier 2 and Tier 3 tasks. All models decline from Tier 1 to Tier 3, confirming difficulty calibration. The Sonnet–DeepSeek gap widens from 0.192 (Tier 1) to 0.257 (Tier 3). Haiku performs within 0.035 of Sonnet across all tiers, offering a compelling value proposition for carbon-constrained organizations.

5.4.

Supply Chain Domain Knowledge

The 30 SC-trap tasks show substantial score differences on the benchmark’s domain-knowledge checks

Table 5. Performance by difficulty tier. Model

Tier 1

Tier 2

Tier 3

Claude Sonnet 4.6 Claude Haiku 4.5 Mistral Large 3 Llama 3.3 70B Llama 4 Scout 17B DeepSeek R1

0.781 0.762 0.703 0.614 0.621 0.589

0.718 0.694 0.598 0.519 0.507 0.483

0.662 0.627 0.521 0.437 0.420 0.405

Table 6. SC-trap task performance testing domain-specific knowledge. Model Claude Sonnet 4.6 Claude Haiku 4.5 Mistral Large 3 DeepSeek R1 Llama 4 Scout 17B Llama 3.3 70B

0.893 0.861 0.756 0.687 0.652 0.608

Deliberation Length, Accuracy, and Carbon

DeepSeek R1 presents a sustainability challenge beyond per-token rates: it generates 9× more tokens per task than Haiku (2,847 vs. 312) while achieving lower aggregate accuracy. Within DeepSeek R1’s 91 formula tasks, Table 7 shows a monotonic association between longer responses and lower accuracy. For this model and task subset, extended deliberation coincides with both lower accuracy and higher per-task carbon. This model-specific result is consistent with overthinking concerns (X. Chen et al., 2025), but it cannot be generalized to reasoning-augmented models as a class without evaluating additional models.

5.6.

Quartile

Avg Length

Accuracy

Q1 (shortest 25%) Q2 Q3 Q4 (longest 25%)

1,860 chars 4,760 chars 5,736 chars 6,752 chars

18.2% 9.1% 9.1% 4.5%

Table 8. Carbon savings analysis comparing routing strategies.

SC-Trap Score

(Table 6). The two Claude models score 0.861–0.893, followed by Mistral at 0.756; DeepSeek R1 scores 0.687. These results describe performance on our task construction and do not identify the underlying training mechanism. A representative case: when asked how safety stock changes when lead time√doubles, both√Claude models correctly applied the LT scaling ( 2 ≈ 1.41×). DeepSeek R1 initially identified the correct relationship but then “reasoned itself away,” ultimately recommending linear scaling with an elaborate but incorrect justification.

5.5.

Table 7. Response-length quartiles and accuracy for DeepSeek R1 formula tasks (n = 91).

Strategy

Quality

gCO2 /task

Savings

Always DeepSeek Always Scout Always Haiku Length, 500 chars Always Sonnet GreenRoute (OOS) Oracle

0.497 0.521 0.699 0.688 0.723 0.733 0.782

17.082 0.023 0.022 0.253 0.401 0.402 1.455

0% 99.9% 99.9% 98.5% 97.7% 97.6% 91.5%

ately carbon-intensive comparison provides a feasibility check, not evidence that routing is better than practical low-carbon selection policies. GreenRoute improves mean quality by 47.5% over always DeepSeek (0.497 to 0.733) while reducing estimated per-task carbon by 97.6%. It is only 0.010 above always Sonnet (0.723) and reaches 93.7% of the per-task oracle quality, so the result should be read as a thresholdcontingent feasibility demonstration rather than a large routing advantage. The fixed-model baselines establish the relevant operating points: Haiku supplies 0.699 quality at 0.022 gCO2 /task, whereas Sonnet supplies 0.723 at 0.401 gCO2 /task. GreenRoute adds a small quality margin over Sonnet at nearly the same estimated per-task carbon under the calibrated policy. The archived length-based baseline uses a 500character prompt-length rule, not a 500-token rule. Rerunning this rule against the retained dual-judge matrix produces 0.688 quality at 0.253 gCO2 /task. It is lowercarbon than GreenRoute but does not improve on always Haiku (0.699), so the reproducible evidence supports static low-carbon selection for moderate requirements and GreenRoute only when a higher quality point is required. This reproducible value replaces the earlier 0.714 summary throughout the manuscript.

5.7.

Real Data vs. Synthetic Task Performance

CAAPF Validation

Table 8 reports strategy-level tradeoffs using the same dual-judge quality matrix and model-level mean output tokens for generation-related per-task carbon estimates. Relative to always using DeepSeek R1, GreenRoute reaches 0.733 mean out-of-sample quality and 97.6% estimated per-task carbon savings. This deliber-

Of our 520 tasks, 33 use real retail data from the M5 Competition (demand forecasting category). All models score slightly lower on real-data tasks (∆ ≈ −0.03 across all models). The model ranking is identical between real and synthetic subsets, and inter-model gaps are preserved (Sonnet–DeepSeek gap: 0.230 real vs. 0.227 synthetic). The stable ranking is reassuring, but

Table 9. External validation: IndustryOR (100 real OR problems, no LLM judge, non-Claude models only). Model

Provider

Arch./Params

Accuracy

Llama 4 Scout Qwen3 32B Ministral 8B Llama 3.3 70B Llama 3.1 8B

Meta Alibaba Mistral Meta Meta

MoE 109/17B* Dense 32B Dense 8B Dense 70B Dense 8B

38.0% 31.0% 27.0% 25.0% 9.0%

*109B total / 17B active parameters.

33 real-data cases from one category cannot rule out benchmark-construction or prompt-style bias.

5.8.

External Validation

Table 9 validates findings on IndustryOR (C. Huang et al., 2025), 100 real-world OR problems with deterministic answers, using exclusively non-Claude models and no LLM judge. Among dense models, Ministral 8B (27%) outperforms Llama 3.3 70B (25%), while Ministral and Llama 3.1 differ by 18 percentage points at the same nominal 8B size. Scout leads but is an MoE model (109B total/17B active), so it cannot be placed on the same one-dimensional scale. This judge-independent validation reduces concern about Claude self-scoring, but it does not eliminate provider, architecture, benchmarkselection, or generation confounds.

5.9.

Dual-Judge and Ground-Truth Validation

To address concerns about LLM-as-judge evaluation bias, we conduct three validation exercises. First, interjudge agreement between Claude Opus and Qwen3235B is high: Spearman ρ = 0.94 (p < 0.01), mean absolute difference 0.023. The model ranking is identical for the top four positions regardless of which judge is used. Second, under Qwen3-only scoring (eliminating any possible Claude self-preference bias), the top-3 ranking remains unchanged (Sonnet > Haiku > Mistral). DeepSeek R1 rises two positions under Qwen3 (which rates verbose reasoning traces more favorably), but remains well below both Anthropic models. Third, deterministic validation on all 182 Tier-1 tasks (Table 3, GT% column) shows that judge rankings align with objective accuracy. Claude Sonnet’s judge score (0.781) is within 0.5% of its ground-truth accuracy (0.786), and model-level judge–GT agreement ranges from 84% to 94%. These checks address scoring bias, but not whether the authored task styles align more closely with some model families’ instruction tuning.

6.

Discussion

6.1.

Addressing the Research Questions

RQ1: Quality–carbon tradeoff. Haiku and Sonnet form the observed Pareto frontier, while the other four models are dominated under our estimates. The highest disclosed parameter totals do not identify the highestperforming candidate, and parameter counts are unavailable for the two Claude models. Thus, nominal size is not an adequate procurement proxy for this task set. RQ2: Observed model differences. Haiku outperforms DeepSeek R1 by 41%, and the external set contains a dense-model size inversion across providers. Performance clusters by model family in our benchmark, but the design cannot separate provider, training data, architecture, model generation, instructionfollowing, or benchmark-style effects. We therefore make no provider- or size-level causal claim. RQ3: Framework use. CAAPF makes the procurement rule contingent on organizational thresholds. GreenRoute reaches 0.733 mean quality in held-out split-half trials, while static Haiku remains preferable near 0.70 because it uses less carbon per task. The framework is supported as an auditable decision process; the routing implementation remains a proof of concept.

6.2.

Proposition Evaluation

P1 (Technology relative advantage): Evidence consistent within the sample. Nominal size does not reliably rank the joint quality–carbon profile. Disclosedparameter comparisons and the external dense-model inversion support direct domain benchmarking instead of a size proxy, but proprietary non-disclosure, provider confounding, and cross-architecture differences prevent a controlled size effect and do not identify a causal model-family mechanism. P2 (Threshold contingency): Supported. The preferred strategy changes with the form and strictness of the organizational quality requirement. At the aggregate portfolio level, static Haiku provides an efficient operating point near 0.70 quality. When quality requirements are imposed across individual task classes, GreenRoute illustrates how cell-level calibration can select different models to target stricter heterogeneous requirements. No tested policy dominates across all organizational requirements. P3 (Auditable multi-context selection): Partially supported. CAAPF records quality, carbon, constraints, and monitoring decisions, while GreenRoute illustrates technical implementation. The study does not evaluate whether external pressures cause adoption, so

the Environment prediction remains a design requirement for later organizational validation.

6.3.

Theoretical Contributions

The theoretical contribution is to translate TOE contexts into a sequence of procurement decisions. We operationalize Technology relative advantage through joint measurement rather than inference from scale; Organization defines an admissible set through quality, risk, cost, and governance constraints; and Environment motivates documentation and re-evaluation. This logic implies no universal “best” model because the selected configuration changes with organizational requirements and the external accountability regime. The DeepSeek R1 analysis adds a technology-level boundary condition: for this model’s formula tasks, longer deliberation is associated with both lower accuracy and higher carbon. CAAPF treats such behavior as something to detect empirically, not as a general property of reasoning models. CAAPF contributes a Green IS design artifact (Melville, 2010; Verdecchia et al., 2023) by making environmental performance part of an auditable selection record while retaining economic, operational, and governance constraints.

6.4.

Table 10. Routing on a heterogeneous (non-Claude) model pool; savings use the best single model, Mistral Large 3, as the baseline.

Alternative Explanations for Model-Family Differences

Strategy

Quality

gCO2 /task

Savings

Always Scout Similarity (τ =0.45) Similarity (τ =0.50) Best single (Mistral) Oracle

0.521 0.555 0.581 0.613 0.677

0.023 0.827 1.273 3.427 varies

99.3% 75.9% 62.8% 0% N/A

Table 11. Routing decisions under carbon estimate uncertainty. Scenario Baseline estimates All ×0.5 All ×2.0 DeepSeek halved, Haiku doubled Random ±50% per model

Savings

Routing ∆

97.6% 97.6% 97.6% 95.1% 97.4%±1.0%

N/A 0/260 tasks 0/260 tasks 13–29/260 6–13/260

generation-related carbon per task (Table 10). This secondary analysis illustrates the role of heterogeneity but does not establish a general threshold for when routing will pay off. The practical rule is to benchmark candidates on domain tasks and default to the lowest-carbon admissible model. Task-level routing should be retained only when it improves on simple selection or escalation at the organization’s chosen threshold.

6.6.

Sensitivity to Carbon Estimates

The observed clustering may have several explanations. Training-data coverage of operations research, alignment procedures, model generation, and MoE architecture may contribute, but none is manipulated here. The task authors’ wording, required output formats, and rubric style may also align better with some instructiontuning distributions. Dual judges and deterministic answers address scoring bias; they do not address this benchmark-construction bias. The results therefore justify benchmarking each candidate, not attributing performance to an unobserved provider mechanism.

Our carbon estimates carry ±50% uncertainty. Sensitivity analysis (Table 11) indicates that decisions depend mainly on relative ordering: the largest paired change in our split-half analysis affected about 22 of 260 held-out tasks.

6.5.

6.8.

When Routing Adds Value

Routing’s value is contingent on model-pool heterogeneity. In the full sample, simple selection between the two strongest models captures most of the attainable gain. In a non-Claude subset (Scout, Llama 70B, Mistral Large, DeepSeek R1), the dual-judge oracle selects every model on some tasks (Mistral 55.8%, DeepSeek 18.7%, Llama 70B 15.4%, Scout 10.2%). In a five-fold held-out evaluation using the same dual-judge matrix, similarity routing at τ = 0.50 reaches 94.7% of best-single-model quality at 62.8% lower estimated

6.7.

Practical Decision Guide

Table 12 synthesizes our findings into actionable guidance for supply chain practitioners implementing CAAPF.

Organizational Carbon Budget Analysis

Consider a mid-size enterprise processing 1,000 supply chain AI queries per day. Table 13 projects annual carbon impact under different procurement strategies. Under these workload and carbon assumptions, always using DeepSeek R1 incurs about 780× the estimated generation-related carbon of always using Haiku, an annual difference of 6,227 kgCO2 , while scoring lower in this benchmark. This is not a complete procurement comparison: it excludes input/prefill energy, API price, latency, integration effort, service reliability,

Table 12. Model selection decision guide for supply chain practitioners. Task Type

Strategy

Rationale

(EOQ,

Deterministic calculator

Standard forecasting, classification Multi-factor trade-off analysis

Lowest-carbon admissible model Select lowest-carbon admissible configuration Benchmark on SCtrap tasks first

Lower compute; deterministic for wellspecified formulas Savings depend on baseline Quality varies by model

Formula-based ROP)

SC domain expertise required

Domain knowledge varies 46%

Table 13. Estimated generation-related operational carbon for 1,000 AI-assisted supply chain decisions per day. Strategy Always DeepSeek R1 Always Mistral Large Always Sonnet Always Haiku

kgCO2 /year

Quality

6,235 1,251 146 8

0.497 0.613 0.723 0.699

Based on mean per-task carbon from Table 4 × 365,000 queries/year.

data residency, security, and vendor risk, as well as embodied and training emissions.

6.9.

Implications for Green IS and AI Governance

The results have four implications for Green IS and AI governance: Sustainable digital infrastructure. Model choice can be treated as an infrastructure decision because estimated operational emissions differ substantially across the evaluated services. The magnitude is sensitive to workload, output length, serving hardware, and grid mix, so organizations should measure rather than transfer our estimates unchanged. Enterprise AI governance and procurement policy. CAAPF adds environmental performance to an existing multi-criteria procurement record. Its threshold, admissibility constraints, benchmark protocol, and review trigger make the decision inspectable, while carbon remains one criterion alongside risk, cost, latency, security, and reliability. Environmental accountability. A documented operational estimate can support internal carbon management and supplier dialogue. The present study does not establish how hosted-model emissions should be allocated in a specific reporting regime and does not claim regulatory compliance. Conditional alignment. In this task set, a smaller model can improve carbon efficiency without a large quality loss, but the result is threshold- and benchmarkdependent. Sustainability and performance align only

when the lower-carbon option remains admissible for the intended decision.

6.10.

Limitations and Future Work

Several limitations bound generalizability. First, the six-model, four-provider sample is small, and four models come from Anthropic or Meta. Provider, architecture, model generation, training data, and size are confounded; moreover, proprietary non-disclosure and the difference between dense totals, MoE totals, and active parameters prevent a controlled cross-model size test. Second, task wording and rubrics may align with some instruction-tuning distributions. Judge diversification and deterministic answers reduce scoring bias but not benchmark-construction bias. Third, 487 of 520 tasks are synthetic; the 33 M5 cases cover only forecasting and cannot reproduce enterprise context. Fourth, one fixed zero-temperature prompt protocol improves comparability but does not show how demonstrations, prompt wording, or output limits change response length, quality, and carbon. Fifth, carbon values are engineering estimates with ±50% uncertainty, a uniform grid assumption, and no embodied or training emissions. Because complete input/prefill token records were not retained, per-task estimates cover generationrelated output-token carbon rather than end-to-end serving energy. Sixth, GreenRoute assumes category and difficulty-tier metadata are available at routing time; production classification errors and overhead are not evaluated. Its 82.2% held-out cell-level threshold attainment rate shows that mean calibration does not guarantee threshold satisfaction; production use needs safety margins or fallback escalation. API price, latency, integration effort, reliability, security, data residency, and vendor risk were not empirically compared. Finally, the deliberation result covers one reasoning model and formula-task subset, and all tasks are single-turn. Future work should use balanced within-provider comparisons, production workloads, external domain experts, prompt-sensitivity experiments, measured serving energy and location, and joint monetary and operational cost analysis. Organizational studies should test whether CAAPF’s documentation and monitoring steps improve actual procurement decisions.

7.

Conclusion

This study compares quality and estimated generation-related operational carbon for six models on 520 supply chain tasks. Nominal size is not a reliable quality proxy in this sample, but non-disclosure and cross-architecture differences prevent a controlled size effect, and the design cannot identify provider or

training mechanisms. The DeepSeek R1 result is a model-specific warning about costly deliberation, not a claim about reasoning models as a class. CAAPF connects benchmark evidence to organization-specific admissibility and environmentalaccountability requirements. GreenRoute reaches 0.733 mean quality in held-out split-half trials and 97.6% lower estimated generation-related carbon per task than always DeepSeek R1. Static Haiku remains more carbon-efficient when aggregate quality near 0.70 is acceptable. The contribution is therefore an auditable, threshold-contingent decision process rather than routing-method novelty. Dual judges, deterministic answers, a small realdata subset, external OR tasks, and sensitivity analysis support the observed pattern without eliminating benchmark-construction bias, limited model coverage, prompt sensitivity, or carbon uncertainty. The recommendation is conditional: benchmark the workload, define admissibility before optimizing carbon, select the lowest-carbon passing configuration, and repeat the decision as conditions change.

References Anthony, L. F. W., Kanding, B., & Selvan, R. (2020). Carbontracker: Tracking and predicting the carbon footprint of training deep learning models. ICML Workshop on Challenges in Deploying and Monitoring Machine Learning Systems. Baker, J. (2012). The technology–organization–environment framework. In Y. K. Dwivedi, M. R. Wade, & S. L. Schneberger (Eds.), Information systems theory: Explaining and predicting our digital society, vol. 1 (pp. 231–245). Springer. https://doi.org/10.1007/ 978-1-4419-6108-2 12 Chen, L., Zaharia, M., & Zou, J. (2024). FrugalGPT: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. Chen, X., Xu, J., Liang, T., He, Z., Pang, J., Yu, D., Song, L., Liu, Q., Zhou, M., Zhang, Z., Wang, R., Tu, Z., Mi, H., & Yu, D. (2025). Do NOT think that much for 2+3=? On the overthinking of long reasoning models. Proceedings of the 42nd International Conference on Machine Learning, 267, 9487–9499. Chopra, S., & Meindl, P. (2019). Supply chain management: Strategy, planning, and operation (7th). Pearson. Cruciani, E., & Verdecchia, R. (2025). Choosing to be green: Advancing green AI via dynamic model selection. Proceedings of the 2nd Workshop on Green-Aware Artificial Intelligence, 4165, 92–100. DeepSeek-AI. (2024). DeepSeek-V3 technical report. arXiv preprint arXiv:2412.19437. https : / / doi . org / 10 . 48550/arXiv.2412.19437 Dodge, J., Prewitt, T., Tachet des Combes, R., Odmark, E., Schwartz, R., Strubell, E., Luccioni, A. S., Smith, N. A., DeCario, N., & Buchanan, W. (2022). Measuring the carbon intensity of AI in cloud instances. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1877–1894. https://doi.org/10.1145/3531146.3533234

Huang, C., Tang, Z., Hu, S., Jiang, R., Zheng, X., Ge, D., Wang, B., & Wang, Z. (2025). ORLM: A customizable framework in training large models for automated optimization modeling. Operations Research, 73(6), 2986–3009. https : / / doi . org / 10 . 1287 / opre . 2024.1233 Huang, S., He, J., Shang, D., Xu, Y., Li, J., Lyu, Y., & Nair, L. M. (2026). DACRI: Decision-aware causal intervention ranking for critical supply chains. arXiv preprint arXiv:2608.11154. https : / / arxiv. org / abs / 2608.11154 Lacoste, A., Luccioni, A., Schmidt, V., & Dandres, T. (2019). Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700. Li, B., Mellou, K., Zhang, B., Pathuri, J., & Menache, I. (2023). Large language models for supply chain optimization. arXiv preprint arXiv:2307.03875. Li, J., He, J., Yang, D., Shang, D., Liu, J., & Huang, S. (2026). Accuracy-preserving stability regularization for large-scale retail demand forecasting. https : / / arxiv.org/abs/2607.13331 Luccioni, A. S., Jernite, Y., & Strubell, E. (2024). Power hungry processing: Watts driving the cost of AI deployment? Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 85–99. https://doi.org/10.1145/3630106.3658542 Makridakis, S., Spiliotis, E., & Assimakopoulos, V. (2022). M5 accuracy competition: Results, findings, and conclusions. International Journal of Forecasting, 38(4), 1346–1364. Melville, N. P. (2010). Information systems innovation for environmental sustainability. MIS Quarterly, 34(1), 1– 21. https://doi.org/10.2307/20721412 Meta AI. (2025, April). The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. https : / / ai . meta . com / blog / llama - 4 - multimodal intelligence/ Mistral AI. (2025, December). Mistral Large 3. https://docs. mistral.ai/models/mistral-large-3-25-12 Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., & Stoica, I. (2025). RouteLLM: Learning to route LLMs from preference data. Proceedings of the International Conference on Learning Representations. Patterson, D., Gonzalez, J., Hölzle, U., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D. R., Texier, M., & Dean, J. (2022). The carbon footprint of machine learning training will plateau, then shrink. Computer, 55(7), 18–28. https://doi.org/10.1109/ MC.2022.3148714 Samsi, S., Zhao, D., McDonald, J., Li, B., Michaleas, A., Jones, M., Bergeron, W., Kepner, J., Tiwari, D., & Gadepally, V. (2023). From words to watts: Benchmarking the energy costs of large language model inference. Proceedings of the IEEE High Performance Extreme Computing Conference, 1–9. Sardana, N., Portes, J., Doubov, S., & Frankle, J. (2024). Beyond chinchilla-optimal: Accounting for inference in language model scaling laws. Proceedings of the 41st International Conference on Machine Learning, 235, 43445–43460. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54– 63. Simchi-Levi, D., Mellou, K., Menache, I., & Pathuri, J. (2026). Large language models for supply chain decisions. In M. C. Cohen & T. Dai (Eds.), Ai in supply chains: Perspectives from global thought leaders (pp. 93– 104, Vol. 27). Springer. https://doi.org/10.1007/9783-032-07054-8 7

Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650. Tornatzky, L. G., & Fleischer, M. (1990). The processes of technological innovation. Lexington Books. United Nations. (2015). Transforming our world: The 2030 agenda for sustainable development (A/RES/70/1). United Nations General Assembly. Verdecchia, R., Sallou, J., & Cruz, L. (2023). A systematic review of green AI. WIREs Data Mining and Knowledge Discovery, 13(4), e1507. Wiesner, P., O’Neill, D. W., Larosa, F., & Kao, O. (2025). Efficiency will not lead to sustainable reasoning AI. arXiv preprint arXiv:2511.15259. Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., Chang, G., Aga, F., Huang, J., Bai, C., Gschwind, M., Gupta, A., Ott, M., Melnikov, A., Candido, S., Brooks, D., Chauhan, G., Lee, B., Lee, H.-H., . . . Hazelwood, K. (2022). Sustainable AI: Environmental implications, challenges, and opportunities. Proceedings of Machine Learning and Systems, 4, 795–813. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and chatbot arena. Advances in Neural Information Processing Systems, 36. Zhu, K., Kraemer, K. L., & Xu, S. (2006). The process of innovation assimilation by firms in different countries: A technology diffusion perspective on e-business. Management Science, 52(10), 1557–1576. https : / / doi.org/10.1287/mnsc.1050.0487

Record · ID 919494 · SHA-256 6c756447d3a59cf4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.