Conceptio › Archive › arXiv CS
arXiv CSopen access

Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection Rishi Bharadwaj, Yadati Narahari

arXiv:2609.20404v1 [cs.LG] 17 Sep 2026

Department of Computer Science and Automation, Indian Institute of Science (IISc)

Abstract Agricultural soils are a major untapped carbon sink. Carbon farming is emerging as a promising practice for tapping this potential. Smallholder farmers, who dominate agriculture across South Asia and sub-Saharan Africa, are key to scaling climate mitigation via carbon farming. It is ironic that real-world carbon programs largely fail to reach them. We study this important gap through the lens of contract design. An aggregator offers a single pooled contract to a heterogeneous population of smallholder farmers who have private adoption costs (adverse selection) and exert unobserved effort (moral hazard), with agronomic outcomes evolving over multiple seasons. We formulate this evolving contracting problem as a POMDP and use reinforcement learning to learn a dynamic profit-maximising contract. We analyse the performance of the aggregator under various conditions. We find that a profit-maximising aggregator does not merely inherit the exclusion of smallholders, it amplifies it. On large farms the aggregator realises 87.7% of achievable adoption, against only 8.2% on smallholdings. Per-hectare Measurement, Reporting and Verification (MRV) costs fall as farm size rises, and the aggregator’s pooling contract compounds this gradient rather than offsetting it. A counterfactual that makes MRV costs purely area-proportional eliminates this disparity. Our results and simulation can guide contract and policy design that opens carbon income to smallholders while enabling agricultural soils to contribute to climate mitigation at scale.

Code — [https://github.com/Rishi-Bharadwaj/carbonfarming-contract-design]

Introduction Carbon farming offers a rare opportunity to address two global challenges simultaneously, removing atmospheric carbon while directing climate finance to vulnerable farming communities. The technical potential of soil carbon sequestration in agroecosystems is 4.4–11.4 Gt CO2 e/yr (Lal 2011), but realised sequestration under existing carbon programs falls far short, in part because those programs do not reach smallholders. Smallholders under 2 hectares account for approximately 500 million of the world’s 582 million farms (Lowder et al. 2025) and dominate agricultural landscapes across South Asia and sub-Saharan Africa. They sustain crop diversity that large monocultural operations erode (Ricciardi et al. 2018). Smallholders are therefore where the

largest share of unrealised sequestration potential and the greatest climate vulnerability coincide. A market design that reaches them can bring a structurally excluded population into carbon markets, unlocking soil carbon mitigation at scale. Yet existing carbon-farming programs are systematically biased toward larger landholders (Barbato and Strong 2023; Cariappa and Krishna 2025; Johansson, Andersson, and Fischer 2025). Designing contracts that reach smallholders is difficult for several reasons. Adoption costs are privately known (adverse selection), effort is only partially observable (moral hazard), agronomic outcomes evolve over multiple seasons, and practical constraints require a single pooled contract to serve heterogeneous farmers. Fixed Measurement, Reporting and Verification (MRV) costs further disadvantage small farms by increasing per-hectare costs. The combination of private information, hidden effort, heterogeneous farmers, and multi-season dynamics creates a contract-design problem that is difficult to solve analytically and calls for methods that can learn adaptive policies under partial observability. Yet AI research has largely focused on firm-level carbon markets, leaving AI-driven contract design for carbon farming largely unexplored (Welsh, Grover, and Jaimungal 2026; Wang et al. 2024; Priyanka et al. 2025). We model carbon farming as a principal-agent Stackelberg game in which an aggregator commits to a single pooled contract and farmers best respond myopically. We derive the equilibrium for the single-season game analytically. The joint presence of moral hazard and adverse selection makes optimal contract computation APX-hard in discrete settings (Guruganesh, Schneider, and Wang 2021). While our formulation is continuous, it further adds multi-season dynamics over a heterogeneous partially observable population. We therefore formulate a POMDP for which reinforcement learning is a standard approach. We model commercially rational aggregators because this reflects prevailing market practice, and examine how these incentives shape market participation. Beyond learning contracts, our framework provides insight into who participates in carbon markets and why. Fixed MRV costs create a mechanical participation threshold that disadvantages smaller farms. We ask whether a profit-maximising aggregator merely inherits this threshold or amplifies it. To answer this, we benchmark against a first-best baseline computed under the same MRV cost structure and define the Re-

alised Adoption Share (RAS) as the proportion of first-best adoption achieved by the aggregator. RAS reaches 87.7% for large farms but only 8.2% for smallholdings, showing exclusion well beyond what cost alone explains. We attribute this amplification to the interaction between profit maximisation and the pooling constraint. A counterfactual that equalises MRV distribution removes this disparity while increasing both farmer participation and enrolled area. Finally, we show that moral hazard can act as a cost-limiting mechanism rather than a pure efficiency loss, so introducing it alongside adverse selection need not reduce contract performance. Our framework also serves as a decision-support tool, letting stakeholders evaluate policy interventions before implementation. It can be recalibrated to other regions and market conditions. Contributions. Learning approaches to contract design treat moral hazard and adverse selection in isolation and assume stationary settings. Analytical work on agricultural contracts handles both frictions but under stationary settings with a single dimension of heterogeneity. We aim to fill this gap and aid smallholder participation. Our paper develops a conceptual framework for examining how commercially rational contract design shapes participation and distributional outcomes. • We formulate commercial carbon-farming contract design as a multi-season principal-agent Stackelberg game embedded in a POMDP. We learn a dynamic pooled contract under joint moral hazard and adverse selection across a population heterogeneous in adoption cost, farm size and initial soil carbon. • We train a deep reinforcement-learning aggregator in this environment and measure it against a first-best welfare benchmark. This lets us ask not only which contracts the aggregator learns, but which farmers those contracts end up serving and why. • We introduce Realised Adoption Share (RAS), which separates participation losses caused by the MRV cost structure from those caused by profit maximisation under pooling. Using RAS, we show that commercially rational aggregators amplify rather than merely inherit smallholder exclusion. We identify socialisation of MRV costs as a policy intervention that equalises participation. • We release a configurable, open-source simulator that separates agronomic, population and market parameters from the environment implementation, so recalibration to a new region requires configuration changes alone. It provides a reusable testbed for evaluating carbon policy interventions before field deployment.

Related Work Automated contract design. Learning principal-agent contracts can be seen as a subfield of automated mechanism design (Conitzer and Sandholm 2003), an area that has seen substantial recent progress (Dütting, Roughgarden, and Talgam-Cohen 2019; Scheid et al. 2025; Wu et al. 2024; Gerstgrasser and Parkes 2023). For a comprehensive survey, see Dütting, Feldman, and Talgam-Cohen (2024).

B P f ci ki yi m u0

Set of possible carbon sequestering practices Carbon price ($/tCO2 e) Farm size (hectares) Carbon sequestered per practice i (tCO2 e/ha) Adoption cost for practice i ($/ha) Profit due to yield change per practice i ($/ha) Total MRV cost per farm ($) Farmer’s outside option (reservation utility) Table 1: Key notation.

Computing optimal contracts under joint frictions is difficult in general. Guruganesh, Schneider, and Wang (2021) prove that the principal-agent problem with moral hazard and adverse selection is APX-hard in discrete combinatorial settings, for both the profit-maximizing single contract and the profit-maximizing menu of contracts. Existing learning approaches each address only part of the problem. Ivanov et al. (2024); Wu et al. (2024) propose RL solutions for learning contracts, but neither considers adverse selection, while Gerstgrasser and Parkes (2023) explore learning Stackelberg equilibria in multi-agent settings, without either friction. Conversely, Bollini et al. (2026) consider learning in Bayesian Stackelberg games with adverse selection but without moral hazard, and Scheid et al. (2025) propose algorithmic solutions for a bandit contracting setting that likewise includes only adverse selection. Carbon contract design. Carbon sequestration has likewise motivated work on contract design. MacKenzie, Ohndorf, and Palmer (2012) study contracts with moral hazard and apply them to the problem of ensuring permanence in carbon sequestering activities. Priyanka et al. (2025) provide a comprehensive survey of carbon farming that explicitly identifies contract design with AI as a promising future direction. Raina, Zavalloni, and Viaggi (2024) examine how different incentive mechanisms for carbon farming are designed and implemented. Wang et al. (2024) use hierarchical MARL to model a government allocating credits to self-interested trading enterprises in a cap-and-trade system, while Welsh, Grover, and Jaimungal (2026) characterize Nash equilibria for greenhouse gas offset credit markets using Nash-DQN. Most closely related work. Dai Li, Immorlica, and Lucier (2022) and Hart and Latacz-Lohmann (2005) both study agricultural principal-agent contracts under moral hazard and adverse selection. However, they model simpler, stationary formulations with one level of heterogeneity, allowing analytical treatments. Our setting introduces heterogeneity across multiple farmer attributes, namely adoption costs, farm sizes and initial Soil Organic Carbon (SOC) concentration. More fundamentally, it adds non-stationary multi-season dynamics, which motivate our learning approach.

Model We begin with a single aggregator and a single farmer. The farmer is risk-neutral and protected by limited liability so that all payments are non-negative. Key parameters are summarised in Table 1.

MRV cost per farm is comprised of both a fixed component (m0 ) and a variable area-dependent component (m1 ) (Ducos, Dupraz, and Bonnieux 2009), with the δ term capturing the effect of scaling efficiency (Bellassen et al. 2015). m = m0 + m1 · f δ

(1)

We instantiate the environment with agronomic parameters for Indian paddy smallholders, following the agricultural literature (Supplementary Document (SD) E). Definition 1 (Action-Based Contract). An action-based contract is a tuple (a, α) where a = (a1 , . . . , a|B| ) ∈ [0, amax ]|B| is the vector of per-practice action payments and amax is the maximum per-action payment allowed. α ∈ [0, 1] is the farmer’s share of MRV costs.

Single-Season Game Structure The game proceeds as follows: (1) the aggregator offers a contract (a, α); (2) the farmer observes the contract and chooses a subset of practices S ⊆ B to adopt as well as whether or not to accept the contract; (3) payoffs are realised. Farmer utility. Given contract (a, α), the farmer’s utility from adopting practice set S is: X Uf (S) = (yi − ki + ai ) f − α · m (2) i∈S

Practice i is adopted iff it contributes non-negatively: S = {i ∈ B : yi − ki + ai ≥ 0}

(3)

The farmer accepts the contract iff Uf (S) ≥ u0 , breaking ties in favor of the principal per standard convention. Here S0 = {i ∈ B : yi − ki ≥ P0} is the farmer’s outside option practice set and u0 = f i∈S0 (yi − ki ). Farmers rejecting the contract may still adopt practices in S0 independently. Aggregator utility. Conditional on contract acceptance, the aggregator’s utility is X X UA (S) = P · f ci − f ai − (1 − α) · m (4) i∈S

i∈S

Otherwise, the aggregator’s utility is 0. Henceforth, all aggregator objectives are implicitly restricted to accepting farmers.

Equilibrium Analysis Proposition 1. At any solution to the aggregator’s problem in which the adoption constraint is slack, the Individual Rationality constraint binds, thus equilibrium farmer utility satisfies Uf∗ = u0 . (Standard result, proof available in SD-A.) Substituting the binding IR constraint into the aggregator’s objective yields the equilibrium aggregator utility: X X X UA∗ = P ·f ci +f (yi −ki )−f (yi −ki )−m (5) i∈S

i∈S

i∈S0

Extensions Moral Hazard The aggregator observes adopted practices but not the effort devoted to each. The farmer privately chooses ei ∈ [e, 1] for each adopted practice i. We set e = 0.2, reflecting that some effort is unavoidable under MRV. We perform a sensitivity analysis on e in SD-G, to show our conclusions are not an artifact of this choice. Effort scales cost, yield and carbon outcomes: realised adoption cost is ei · ki , realised yield ei · yi , and realised sequestration ei · ci . Under a pure actionbased contract (R = 0), the farmer has no incentive to exert effort beyond the minimum, so ei = e for all practices that are not individually profitable. To restore effort incentives, the aggregator augments the action-based contract with a result based payment R, yielding the hybrid contract (a, α, R). The farmer’s utility is: X Uf (S, e) = (ei · yi − ei · ki + ai ) f − α · m i∈S

(6)

! +

X

ei · c i

fR

i∈S

Since the farmer’s per-practice objective is linear in ei , the best response is a threshold policy:  1 if yi + ci R ≥ ki ei = (7) e otherwise The aggregator’s utility under the hybrid contract is: X X UA (S, e) = f (P −R) ei ·ci −f ai −(1−α)·m (8) i∈S

i∈S

Adverse Selection Adoption costs vary across farmers due to differences in soil quality, topography, access to capital, and local labour conditions (Eagle, Uludere Aragon, and Gordon 2022; Antle et al. 2003). We model contracting using a single pooled contract offered to the entire farmer population, consistent with current commercial practice. Prior work (Dai Li, Immorlica, and Lucier 2022; Hart and Latacz-Lohmann 2005) similarly studies pooling contracts rather than type-separating menus. Following Hart and Latacz-Lohmann (2005), we model the effective adoption cost for farmer j and practice i as: eff kji = θj · ki ,

θj ∼ Uniform[0, 1]

(9)

The scaling factor θj is private information, thus the aggregator cannot condition payments on the farmer’s realised type.

Multi-Season Dynamics A distinguishing feature of our framework is the explicit modelling of how current contracting decisions shape future states, captured by the following.

Yield ramp-up. Following the modelling of Dabbert and Madden (1986) and the results of Datta et al. (2025), we model yield as ramping up linearly to its equilibrium from season t = 0 to season TL = 3.   if yi ≤ 0 yi   tji yji = (10) tji  ,1 if yi > 0 yi · min TL where tji is the number of seasons farmer j has performed practice i. Negative yield effects are realised immediately, while positive effects emerge over time. Adoption cost decay. Motivated by evidence that imperfect knowledge of new-technology management constitutes a significant adoption barrier that diminishes rapidly with accumulated experience (Foster and Rosenzweig 1995), we model adoption costs as decaying linearly toward their typeeff scaled equilibrium values kji over TL = 3 seasons:   tji t eff eff kjiji = kji + 12 |kji ,0 (11) | · max 1 − TL Soil organic carbon dynamics. We adopt the C-saturation model of Stewart et al. (2007) with decomposition rate h = 0.02 yr−1 from Coleman and Jenkinson (1996):     Cjt Cjt+1 = max 0, Cjt + I˜jt · 1 − − h · Cjt (12) Cm where Cm is the soil-type-specific saturation capacity and the saturation deficit captures diminishing returns as soils approach capacity. Realized carbon input is the sum of effortdependent practice contributions and a stochastic shock. As effort is unobservable and the shock breaks the deterministic link between effort and outcome, effort is not directly contractible, creating moral hazard. Farmers make decisions based on the expected carbon sequestered. X I˜jt = etji · ci + εtj , εtj ∼ N (0, σ 2 ) (13)

Action space. At each season t, the aggregator offers a pooled contract (at , αt , Rt ). Each farmer j responds with a participation decision, adoption decision and effort profile. n  t t Sjt = i ∈ B : etji yjiji − kjiji  o Ct  + ci 1 − Cmj Rt + ati ≥ 0 (15) The farm-level MRV cost and decomposition terms are absent as neither depends on i and so neither affects the marginal comparison. A farmer accepts the contract only if the expected utility from participation equals or exceeds the baseline, i.e Uftj (Sjt , etj ) ≥ ut0j . Otherwise the farmer retains the outside option by adopting only practices that are inherently profitable without contract incentives. Farmer utility. Farmer j’s utility follows Eq. 6, with season-dependent yield, cost and carbon terms defined by the multi-season dynamics of Eqs. 10–12, and farmer j’s farm size fj :  X t t Uftj (Sjt , etj ) = etji yjiji − etji kjiji + ati fj i∈Sjt

− αt mj + (Cjt+1 − Cjt ) · fj Rt (16) Both parties are paid on the actual net carbon sequestered. Aggregator reward. Unlike the single-season objective (Eq. 8), the per-season reward here is computed on net carbon after saturation and decomposition, and summed across the farmer population. The aggregator’s per-season reward is: " N X X rt = fj (P − Rt )(Cjt+1 − Cjt ) − fj ati i∈Sjt

j=1

# − (1 − αt ) mj

(17)

i∈Sjt

Methodology

POMDP Formulation We extend the single-season Stackelberg game to a multiseason POMDP in which the aggregator is an RL agent offering a single pooled contract to N heterogeneous farmers over a horizon of T seasons. Farmers best-respond myopically each season. Observation space. Under pooled contracting, the aggregator observes the farmer population through aggregate statistics of the population. These aggregate observations condition the contract offered in the following season and define the observation  at state t: t  t (14) o = τ t , f¯, SOC , c̃¯t , d¯t , ρ̄t t

where τ t is the season index, f¯ the mean farm size, SOC the mean soil organic carbon, c̃¯t the mean carbon outcome realised in season t, d¯t ∈ [0, 1]|B| the vector of per-practice adoption rates, and ρ̄t ∈ [0, 1] the participation rate. All components are normalised to approximately [0, 1] for training stability. Farmer types θj remain unobserved, consistent with the adverse selection setting.

All main results are aggregated across 25 random seeds. Our objective is to learn the profit-maximising contract for a given farmer population rather than to generalise across populations, consistent with our design in which the aggregator is retrained per region. Each seed therefore fixes a population instance, and the aggregator is trained and evaluated on that instance. All confidence intervals are 95% percentile bootstrap intervals resampled per seed. All p-values use the Wilcoxon signed-rank test, paired over the 25 seeds for between-condition comparisons, and as a one-sample test against zero for the correlation coefficients. We train the aggregator with TQC (Truncated Quantile Critics; Kuznetsov et al. (2020)). In a preliminary comparison over 5 seeds per ablation, TQC matched or outperformed SAC (Haarnoja et al. 2018) across all conditions, from a statistical tie in the MH-only case to a ∼7% profit-capture-ratio gain (SD-C). On-policy PPO (Schulman et al. 2017) failed to converge to a competitive policy in our setting even after relaxing the training curriculum, so we focus on off-policy methods. Exact training details are available in SD-D.

Welfare Baseline

Ablation

Mean

Std

95% CI

To assess the learned contracts, we normalize the aggregator’s realized profit against a first-best welfare baseline: the maximum total welfare achievable under direct carbonmarket participation with myopically best-responding farmers. Comparison against the first-best is standard in contract theory (Laffont and Martimort 2002) and algorithmic contract design (Dütting, Feldman, and Talgam-Cohen 2024). We construct it by setting the carbon payment rate R equal to the market price P and the MRV cost-share α = 1, so the aggregator earns zero margin and farmers face the full social value of their actions, equivalent to transacting directly in the carbon market with no intermediary. Farmers best-respond as elsewhere, accepting if participating profit exceeds or equals their reservation utility u0 . Welfare is the sum of accepted farmers’ surplus above this outside option, counting only value created by the contract. Since this ceiling is invariant to how surplus is split between aggregator and farmers, an aggregator whose profit approaches it is capturing nearly all system surplus, while a large gap reflects either value destroyed by frictions or retained by farmers. We report the profit-capture ratio, RL reward over first-best welfare, computed per ablation. We also give the full decomposition into farmer surplus, aggregator surplus, and deadweight loss in SD-J.

MH+AS MH only AS only Neither

0.599 0.699 0.566 0.809

0.155 0.147 0.109 0.115

[0.538, 0.658] [0.640, 0.755] [0.520, 0.607] [0.761, 0.851]

Ablations Condition

Effort ei

Cost type θj

Farm size fj

MH+AS MH only AS only Neither

[0.2, 1.0] [0.2, 1.0] 1.0 1.0

∼ U [0, 1] 0.5 ∼ U [0, 1] 0.5

5.5 5.5 5.5 5.5

MH+AS MH only AS only Neither

[0.2, 1.0] [0.2, 1.0] 1.0 1.0

∼ U [0, 1] 0.5 ∼ U [0, 1] 0.5

∼ U [1, 10] ∼ U [1, 10] ∼ U [1, 10] ∼ U [1, 10]

Table 2: Ablations. We run 8 ablations (Table 2) by independently toggling adverse selection and moral hazard, each crossed with two farm-size regimes. When moral hazard is absent, effort is fixed at ei = 1.0. When it is present, farmers privately choose ei ∈ [e, 1.0] per practice, realised at e or 1, which the aggregator cannot observe. When adverse selection is absent, the cost multiplier is fixed at the population mean θj = 0.5. When it is present, θj is drawn independently per farmer and never revealed. Crossing these four combinations with fixed (5.5 ha) and heterogeneous (fj ∼ U [1, 10] ha) farm sizes isolates the effect of the information frictions from that of farm-size heterogeneity. Initial SOC is drawn uniformly (Cj0 ∼ U [30, 80] tCO2 e/ha) in all ablations.

Results Fixed Farm Size We first fix farm size to isolate the effect of the frictions.

Table 3: Profit-capture ratio, fixed farm size.

The profit-capture ratio is calculated as the cumulative aggregator reward over 5 seasons divided by the cumulative welfare baseline. Table 3 shows that the Neither condition outperforms all others. More surprisingly, MH+AS outperforms AS only, in aggregate though the difference is not statistically significant across seeds (p = 0.080). The effect arises from moral hazard, with practices unprofitable at full effort becoming profitable at low effort, raising adoption and ultimately profitability. Because effort scales realised loss but not the action payment, moral hazard lets the farmer limit downside on an otherwise unprofitable practice rather than bear it in full. Aggregated across the population, this effect raises adoption under MH+AS relative to AS only, outweighing the additional friction introduced by moral hazard. We provide a complete worked example of this effect in SD-H. Table 4 shows how adoption rate varies across conditions, averaged over farmers, seasons and seeds. More information regarding the contract payment structure itself is available in SD-K. Ablation

Mean

Std

95% CI

MH+AS MH only AS only Neither

0.542 0.592 0.428 0.691

0.204 0.179 0.093 0.154

[0.464, 0.623] [0.521, 0.661] [0.392, 0.464] [0.630, 0.751]

Table 4: Adoption rate, fixed farm size.

Variable Farm Size We next allow farm size to vary across farmers, drawn from a uniform distribution. Real smallholdings are skewed towards smaller areas and better described by a log-normal distribution, but the uniform draw provides coverage across the size range, which are better suited to characterizing adoption by farm size. Friction comparisons. Profit-capture ratios (Table 7) follow a similar overall ordering, but the gap between the best and worst performers narrows, and the best performer does substantially worse than under fixed farm size. We find two interesting results. First, MH+AS again outperforms AS only, this time significantly across seeds (p < 10−4 ). This is not driven by adoption rate differences (Table 8), but by a second effect of moral hazard. The aggregator learns to offer a lower average action payment, as it no longer has to compensate for the full cost of effort for otherwise unprofitable actions. We calculated the average payment across adopted practices,

Metric First-Best Adoption RL Adoption RL Adoption / FB Adoption Expected RL Profit/Farm Expected per-farm RL Profit/ha

Small [1-4 Ha]

Medium [4-7 Ha]

Large [7-10 Ha]

19.4% [16.6%, 22.3%] 1.7% [0.8%, 2.8%] 8.2% [4.3%, 12.7%] $7 [$3, $11] $2 [$1, $3]

68.7% [65.1%, 72.5%] 16.5% [13.2%, 20.4%] 23.8% [19.5%, 28.6%] $96 [$82, $110] $16 [$14, $19]

88.1% [86.4%, 89.7%] 77.3% [71.9%, 82.7%] 87.7% [81.7%, 93.4%] $404 [$375, $434] $47 [$44, $50]

Table 5: Adoption and profit by farm size tercile: (MH+AS), m0 = 300, m1 = 80, δ = 0.7 (95% CI). Metric First-Best Adoption RL Adoption RL Adoption / FB Adoption Expected RL Profit/Farm Expected per-farm RL Profit/ha

Small [1-4 Ha]

Medium [4-7 Ha]

Large [7-10 Ha]

72.1% [69.0%, 75.4%] 56.5% [47.2%, 65.9%] 78.8% [65.8%, 92.1%] $65 [$54, $77] $26 [$22, $30]

72.1% [68.2%, 76.0%] 57.4% [48.8%, 66.0%] 78.3% [68.3%, 87.9%] $157 [$137, $177] $28 [$25, $32]

71.3% [68.9%, 73.9%] 56.1% [47.9%, 64.2%] 77.6% [67.3%, 87.6%] $233 [$203, $264] $27 [$24, $31]

Table 6: Adoption and profit by farm size tercile: (MH+AS), m0 = 0, m1 = 101.3, δ = 1.0 (95% CI). Ablation

Mean

Std

95% CI

MH+AS MH only AS only Neither

0.620 0.612 0.568 0.680

0.109 0.096 0.104 0.100

[0.577, 0.663] [0.574, 0.650] [0.528, 0.610] [0.641, 0.720]

Table 7: Profit-capture ratio, variable farm size. Ablation

Mean

Std

95% CI

MH+AS MH only AS only Neither

0.320 0.315 0.337 0.411

0.075 0.046 0.076 0.095

[0.292, 0.350] [0.298, 0.333] [0.310, 0.369] [0.376, 0.449]

Table 8: Adoption rate, variable farm size. and it was $1.35/ha less (p < 10−4 ) in the MH+AS case compared with the AS only case. Second, MH+AS achieves comparable performance to MH-only (p = 0.916), suggesting the adverse selection mechanism does not meaningfully reduce performance when moral hazard is present. We leave full characterization of this interaction to future work. Amplified smallholder exclusion. Beyond the contracts themselves, our framework speaks to who participates and why. Table 5 reports adoption by farm-size tercile under RL and under first-best, together with the RAS. Two gradients emerge, with the first mechanical. Fixed per-farmer MRV costs do not scale with area, so the same cost is amortised over more hectares on larger farms, making larger farms more likely to clear the threshold. The first-best already embeds this gradient because it uses the same MRV cost structure. RAS captures the residual after this control. The aggregator realises 87.7% of achievable adoption on large farms but only 8.2% on smallholdings. Because the benchmark already

incorporates fixed costs, this residual exceeds the gradient attributable to cost structure alone. We attribute the residual to the interaction between profit maximisation and the pooling constraint. Under a single perhectare schedule, any increase in payment generous enough to bring marginal smallholders above their threshold must also be paid on every hectare already enrolled. That inframarginal cost scales with the area of the large farms in the pool and outweighs the thin margin recovered from newly included smallholders, so the profit-maximising contract stays below the level that would admit them. MRV counterfactual. To test whether this exclusion is rooted in the cost structure rather than a fundamental barrier, we consider a counterfactual in which MRV costs are made purely area-proportional. The fixed per-farmer component is removed, the scaling exponent is set to one, and the variable rate is recalibrated so that total MRV liability over the farmer population is unchanged from baseline (SD-I). The regime therefore redistributes MRV cost across farm sizes rather than lowering it. This requires large farmers in the pool. The same effect could also be reached through alternate strategies, such as government or NGO socialisation, or by lowering fixed costs through cheaper monitoring methods such as remote sensing. Evaluating these alternatives is exactly the kind of policy question the framework is built for. Table 6 shows that the size gradient vanishes. First-best and RL adoption are both flat across terciles, as is RAS, whose intervals fully overlap. Expected profit per hectare1 is now flat across farm sizes, removing the incentive to favour larger holdings that the fixed-cost component otherwise creates. The equalisation is two-sided, as smallholder adoption rises sharply while large-farm adoption falls, the latter reflecting the higher share of MRV costs large farms now bear. The change does not reduce overall participation or enrolled 1

Let sj denote the aggregator’s profit from farm j and fj its area. Expectations are taken over all farms in the tercile. E[sj /fj ] is the expected per-farm profit per hectare.

area, as both increase under the RL contract, so smallholder gains outweigh large-farm losses on both counts (SD-I). Aggregator profit falls by roughly 10%, which is the expected consequence of removing a cost asymmetry exploited by the profit-maximising contract. The restructuring is therefore not one an aggregator would adopt unilaterally, which is what makes it a policy instrument rather than a business recommendation. Farmer-level correlates. Table 9 corroborates this pattern at the farmer level. Across all information conditions, farm size has the strongest association with participation, with correlations that are statistically significant. Correlations with initial SOC and adoption cost are weaker, despite generally being significant. The substantive point is not that larger farms participate more, a generic consequence of area-scaled contracts, but that the effect is large enough to mask the informational frictions, which surface only once farm size is controlled for. Condition

Mean r

95% CI

p-value

Farm Size MH+AS MH only AS only Neither

0.693 0.748 0.642 0.787

[0.658, 0.726] [0.723, 0.771] [0.622, 0.661] [0.771, 0.803]

< 10−4 < 10−4 < 10−4 < 10−4

Initial SOC MH+AS MH only AS only Neither

-0.159 -0.209 -0.114 -0.155

[-0.191, -0.125] [-0.252, -0.166] [-0.151, -0.073] [-0.199, -0.108]

< 10−4 < 10−4 < 10−4 < 10−4

Adoption Cost MH+AS -0.114 MH only — AS only 0.033 Neither —

[-0.155, -0.075] — [-0.001, 0.071] —

< 10−4 — 0.1336 —

Table 9: Point-biserial correlation between farmer characteristics and adoption, variable farm size.

Limitations Additionality. We do not impose an additionality constraint in our model. However, as Proposition 2 (SDB) demonstrates, this yields a conservative lower bound on smallholder exclusion because enforcing additionality strictly raises the minimum viable farm size. We also abstract away from carbon permanence, which real-world registries enforce over multi-decade horizons rather than our five-season simulation. Farmer behaviour. Farmers best-respond each season, as this is the standard assumption in learning-based Stackelberg settings (Haghtalab et al. 2022), and the short effective planning horizons documented among smallholders (Duflo, Kremer, and Robinson 2011) make the asymmetry with our forward-looking aggregator realistic rather than a modelling

shortcut. Farmers are also risk-neutral and face no credit, tenure, or participation costs beyond MRV. These omitted frictions bind harder on smallholders, so our exclusion estimates are a lower bound. Non-learning baselines. We do not benchmark against non-learning optimization algorithms, as we make no claims regarding the absolute efficiency of the learner. Such a comparison would ask whether the size gradient comes from the environment or from the learner, and our cost counterfactual answers that directly. Calibration scope. Finally, our parameters characterise a single regional instantiation with farm sizes drawn uniformly and MRV charged per farm. The magnitudes should be read as illustrative of that setting, and thorough empirical validation is left to future work. The simulator is configurable, so a user can recalibrate to their own region and test with no changes to the model.

Conclusion We have studied how commercially rational carbon-farming contracts shape participation under adverse selection, moral hazard, and multi-season agronomic dynamics. Our results show that commercially rational pooled contracts systematically exclude smallholders beyond what the underlying MRV cost structure alone explains, concentrating participation among larger farms. This exclusion is not inevitable. Redistributing MRV costs in proportion to cultivated area, while preserving the overall expected MRV liability of the system, removes the participation gradient. Greater smallholder inclusion directly improves both equity and the share of soil carbon potential that markets can reach. Beyond climate finance, this formulation contributes directly to the growing subfield of automated contract design. We model the problem as a POMDP in which a learning principal offers a single contract to a heterogeneous, partially observed population under non-stationary dynamics. We treat moral hazard and adverse selection jointly and over a sequential horizon whereas prior learning approaches address the frictions in isolation and in stationary settings. While learned contracts are traditionally evaluated by principal utility, welfare, or regret, we demonstrate the necessity of quantifying their distributional consequences in public-good domains. The proposed Realised Adoption Share (RAS) distinguishes exclusion imposed by baseline cost structures from that amplified by profit maximisation under pooling. This dynamic extends beyond agriculture. Any pooled contract with fixed per-participant costs, such as microcredit or insurance, excludes those below a specific value cutoff. Our results demonstrate that profit maximisation pushes this cutoff well above the welfare-maximising baseline, a margin RAS is built to measure. More broadly, our work illustrates how AI can support the design and evaluation of economic mechanisms by revealing their societal consequences before deployment. We release our open-source environment to serve both as an experimental testbed for AI researchers and as a decisionsupport tool for policymakers evaluating market interventions.

References Antle, J.; Capalbo, S.; Mooney, S.; Elliott, E.; and Paustian, K. 2003. Spatial heterogeneity, contract design, and the efficiency of carbon sequestration policies for agriculture. Journal of Environmental Economics and Management, 46(2): 231–250. Barbato, C. T.; and Strong, A. L. 2023. Farmer Perspectives on Carbon Markets Incentivizing Agricultural Soil Carbon Sequestration. npj Climate Action, 2: 26. Bellassen, V.; Stephan, N.; Afriat, M.; Alberola, E.; Barker, A.; Chang, J.-P.; Chiquet, C.; Cochran, I.; Deheza, M.; Dimopoulos, C.; Foucherot, C.; Jacquier, G.; Morel, R.; Robinson, R.; and Shishlov, I. 2015. Monitoring, Reporting and Verifying Emissions in the Climate Economy. Nature Climate Change, 5(4): 319–328. Bollini, M.; Bacchiocchi, F.; Coutts, S.; Castiglioni, M.; and Marchesi, A. 2026. Learning in Bayesian Stackelberg Games With Unknown Follower’s Types. In Forty-third International Conference on Machine Learning. Cariappa, A. A. G.; and Krishna, V. V. 2025. Carbon farming in India: are the existing projects inclusive, additional, and permanent? Climate Policy, 25(5): 756–771. Coleman, K.; and Jenkinson, D. S. 1996. RothC-26.3 - A Model for the turnover of carbon in soil. In Powlson, D. S.; Smith, P.; and Smith, J. U., eds., Evaluation of Soil Organic Matter Models, 237–246. Berlin, Heidelberg: Springer Berlin Heidelberg. ISBN 978-3-642-61094-3. Conitzer, V.; and Sandholm, T. 2003. Automated mechanism design: complexity results stemming from the singleagent setting. In Proceedings of the 5th International Conference on Electronic Commerce, ICEC ’03, 17–24. New York, NY, USA: Association for Computing Machinery. ISBN 1581137885. Dabbert, S.; and Madden, P. 1986. The transition to organic agriculture: A multi-year simulation model of a Pennsylvania farm. American Journal of Alternative Agriculture, 1(3): 99– 107. Dahlgreen, J.; and Parr, A. 2024. Exploring the Impact of Alternate Wetting and Drying and the System of Rice Intensification on Greenhouse Gas Emissions: A Review of Rice Cultivation Practices. Agronomy, 14(2). Dai Li, W.; Immorlica, N.; and Lucier, B. 2022. Contract Design for Afforestation Programs. In Feldman, M.; Fu, H.; and Talgam-Cohen, I., eds., Web and Internet Economics, volume 13112, 113–130. Cham: Springer International Publishing. ISBN 978-3-030-94675-3 978-3-030-94676-0. Das, S. R.; Chatterjee, D.; Saha, S.; Sarkar, D.; Alam, R.; Dey, S.; Ghosh, S.; Nayak, B. K.; Smith, P.; and Pathak, H. 2025. Enhancing carbon sequestration potential of lowland rice agroecosystems for environmentally clean production system: A review. Climate Smart Agriculture, 2(2): 100054. Datta, A.; Wilke, B.; Charles, C.; Hasenick, M.; Ulbrich, T.; Singh, M.; Sears, M.; and Robertson, G. P. 2025. Crop performance and profitability for the initial transition years of a regenerative cropping system in the Upper Midwest United States. Journal of Environmental Quality, 54(6): 1572–1585.

Ducos, G.; Dupraz, P.; and Bonnieux, F. 2009. Agrienvironment contract adoption under fixed and variable compliance costs. Journal of Environmental Planning and Management, 52(5): 669–687. Duflo, E.; Kremer, M.; and Robinson, J. 2011. Nudging Farmers to Use Fertilizer: Theory and Experimental Evidence from Kenya. American Economic Review, 101(6): 2350–90. Dütting, P.; Roughgarden, T.; and Talgam-Cohen, I. 2019. Simple versus Optimal Contracts. In Proceedings of the 2019 ACM Conference on Economics and Computation, EC ’19, 369–387. New York, NY, USA: Association for Computing Machinery. ISBN 9781450367929. Dütting, P.; Feldman, M.; and Talgam-Cohen, I. 2024. Algorithmic Contract Theory: A Survey. Foundations and Trends® in Theoretical Computer Science, 16: 211–411. Eagle, A.; Uludere Aragon, N.; and Gordon, D. 2022. The Realizable Magnitude of Carbon Sequestration in Global Cropland Soils: Socioeconomic Factors. Technical report, Environmental Defense Fund, New York, New York. Foster, A. D.; and Rosenzweig, M. R. 1995. Learning by Doing and Learning from Others: Human Capital and Technical Change in Agriculture. Journal of Political Economy, 103(6): 1176–1209. Gerstgrasser, M.; and Parkes, D. C. 2023. Oracles & Followers: Stackelberg Equilibria in Deep Multi-Agent Reinforcement Learning. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 11213–11236. PMLR. Guruganesh, G.; Schneider, J.; and Wang, J. R. 2021. Contracts under Moral Hazard and Adverse Selection. In Proceedings of the 22nd ACM Conference on Economics and Computation, EC ’21, 563–582. New York, NY, USA: Association for Computing Machinery. ISBN 9781450385541. Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Dy, J.; and Krause, A., eds., Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, 1861–1870. PMLR. Haghtalab, N.; Lykouris, T.; Nietert, S.; and Wei, A. 2022. Learning in Stackelberg Games with Non-myopic Agents. In Proceedings of the 23rd ACM Conference on Economics and Computation, EC ’22, 917–918. New York, NY, USA: Association for Computing Machinery. ISBN 9781450391504. Hart, R.; and Latacz-Lohmann, U. 2005. Combating moral hazard in agri-environmental schemes: a multiple-agent approach. European Review of Agricultural Economics, 32(1): 75–91. Ivanov, D.; Dütting, P.; Talgam-Cohen, I.; Wang, T.; and Parkes, D. C. 2024. Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts. arXiv:2407.18074. Johansson, E.; Andersson, E.; and Fischer, K. 2025. The race for carbon in farmland: global mapping of the emerging

voluntary market for soil carbon credits. International Journal of Sustainable Development & World Ecology, 32(7): 810–826. Kuznetsov, A.; Shvechikov, P.; Grishin, A.; and Vetrov, D. 2020. Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics. In III, H. D.; and Singh, A., eds., Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, 5556–5566. PMLR. Laffont, J.-J.; and Martimort, D. 2002. The Theory of Incentives: The Principal-Agent Model. Princeton University Press. ISBN 9780691091846. Lal, R. 2011. Sequestering carbon in soils of agroecosystems. Food Policy, 36: S33–S39. Lal, R. 2015. Soil Carbon Sequestration in Agroecosystems of India. Journal of the Indian Society of Soil Science, 63(2). Lowder, S.; Arslan, A.; Cabrera Cevallos, C. E.; O’Neill, M.; and De La O Campos, A. P. 2025. A global update on the number of farms, farm size and farmland distribution – Background paper for The State of Food and Agriculture 2025. FAO Agricultural Development Economics Working Paper 25-14, FAO, Rome. MacKenzie, I. A.; Ohndorf, M.; and Palmer, C. 2012. Enforcement-proof contracts with moral hazard in precaution: ensuring ’permanence’ in carbon sequestration. Oxford Economic Papers, 64(2): 350–374. Priyanka, V.; Charan, G.; Suresh, R. P.; Sunkara, T.; Patil, M.; Sagar, K.; Trivedi, A.; Soumya, K.; Paul, S.; Hadimani, P.; Babu, G.; Trivedi, R.; and Narahari, Y. 2025. Carbon Farming: An Expository, Inter-Disciplinary Survey. Journal of the Indian Institute of Science, 105(2-3): 337–399. Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; and Dormann, N. 2021. Stable-Baselines3: Reliable Reinforcement Learning Implementations. Journal of Machine Learning Research, 22(268): 1–8. Raina, N.; Zavalloni, M.; and Viaggi, D. 2024. Incentive mechanisms of carbon farming contracts: A systematic mapping study. Journal of Environmental Management, 352: 120126. Ricciardi, V.; Ramankutty, N.; Mehrabi, Z.; Jarvis, L.; and Chookolingo, B. 2018. How much of the world’s food do smallholders produce? Global Food Security, 17: 64–72. Scheid, A.; Boursier, E.; Durmus, A.; Moulines, E.; and Jordan, M. I. 2025. Online Decision-Making in Tree-Like Multi-Agent Games with Transfers. arXiv:2501.19388. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. arXiv:1707.06347. Stewart, C.; Paustian, K.; Conant, R.; Plante, A.; and Six, J. 2007. Soil carbon saturation: Concept, evidence and evaluation. Biogeochemistry, 86: 19–31. Tyagi, A.; and Haritash, A. K. 2025. Climate-smart agriculture, enhanced agroproduction, and carbon sequestration potential of agroecosystems in India: a meta-analysis. Journal of Environmental Studies and Sciences, 15: 167–185.

Wang, H.; Li, W.; Zha, H.; and Wang, B. 2024. Carbon Market Simulation with Adaptive Mechanism Design. In Larson, K., ed., Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, 8824–8828. International Joint Conferences on Artificial Intelligence Organization. Demo Track. Welsh, L.; Grover, U.; and Jaimungal, S. 2026. Multi-Agent Reinforcement Learning for Greenhouse Gas Offset Credit Markets. arXiv:2504.11258. World Bank. 2021. Soil Organic Carbon MRV Sourcebook for Agricultural Landscapes. Washington, DC: World Bank. Wu, J.; Chen, S.; Wang, M.; Wang, H.; and Xu, H. 2024. Contractual Reinforcement Learning: Pulling Arms with Invisible Hands. arXiv:2407.01458.

A

IR Constraint

Proposition 1. At any solution to the aggregator’s problem in which the adoption constraint is slack, the Individual Rationality constraint binds: equilibrium farmer utility satisfies Uf∗ = u0 . Proof. Suppose Uf > u0 at an optimum (a∗ , α∗ ) with a participating farmer. If a∗ = 0 then S = S0 , so Uf = u0 − α∗ m ≤ u0 by Eq. 2, a contradiction. Hence a∗i > 0 for some i ∈ S. Let η = mini∈S (yi − ki + a∗i ) > 0 denote the slack in the adoption constraint and reduce a∗i by ε < min{η, a∗i , (Uf − u0 )/f }. The first bound preserves S, the second keeps payments non-negative, and the third preserves Uf ≥ u0 . Since UA is strictly decreasing in ai , this strictly increases the aggregator’s payoff, contradicting optimality.

B

Equilibrium under Additionality

Under an additionality constraint, the aggregator earns carbon credits only for practices not adopted in the farmer’s outside option (S \ S0 ), and only offers action payments for the same. The binding IR constraint reduces to: X X f ai = α · m − f (yi − ki ) (18) i∈S\S0

i∈S\S0

and the aggregator’s equilibrium utility is: X X UA∗ = P · f ci + f (yi − ki ) − m i∈S\S0

(19)

i∈S\S0

The problem reduces entirely to contracting over additional practices; the outside option determines S0 but drops out of the surplus calculation. Comparing this to the unconstrained equilibrium utility in Eq. 5 gives X UA∗,no-add − UA∗,add = P · f ci ≥ 0 (20) i∈S0

where non-negativity P follows from ci ≥ 0. The inequality is strict whenever i∈S0 ci > 0, that is, whenever the outside option already contains at least one practice with positive

sequestration. Since every practice in our setting satisfies ci > 0, this reduces to S0 ̸= ∅. Additionality therefore reduces aggregator surplus by the carbon value of outside-option practices while leaving the per-farmer cost m(f ) unchanged. Since viability trades perhectare surplus against a cost with a fixed component, the smallest farm size at which the aggregator breaks even is weakly larger under additionality. Proposition 2. Let f ∗ denote the smallest farm size at which UA∗ ≥ 0, under m(f ) = m0 + m1 f δ with δ P < 1. Then ∗ ∗ fadd ≥ fno-add , with strict inequality whenever i∈S0 ci > 0. P Proof. Immediate from Eq. 20. The wedge P i∈S0 ci lowers per-hectare surplus while leaving m(f ) unchanged, so ∗,add ∗,no-add U (f ) for every f , and strictly so when PA (f ) ≤ UA i∈S0 ci > 0. Every farm size viable under additionality is therefore viable without it. An additionality requirement thus raises the minimum viable farm size and strengthens the exclusion gradient reported in the main text. Our adoption results should be read as favourable to smallholder participation relative to a setting with additionality enforced.

C

Algorithm Comparison

Table 10: Cumulative profit-capture ratio: TQC vs SAC (mean across 5 seeds), with the relative improvement of TQC over SAC. Ablation

TQC

SAC

Improvement (%)

MH+AS MH only AS only Neither

0.590 0.679 0.555 0.807

0.561 0.682 0.519 0.783

+5.1 -0.4 +6.9 +3.0

TQC and SAC were trained with the final hyperparameters as reported in Table 12.

D

Model Training

Training TQC was done via Stable-Baselines3 (Raffin et al. 2021). Training proceeds through a 4-phase curriculum (Table 11) that progressively lengthens the season horizon. The entropy coefficient, set to "auto", is reset in each phase to encourage exploration. All experiments are repeated independently over 25 random seeds. In each run, the model is trained and evaluated using the same random seed, and reported results are aggregated across the 25 runs. During training, a callback evaluates the current policy every 500 gradient steps by running a deterministic episode on a fixed evaluation seed and records the cumulative reward. The checkpoint is saved whenever this reward exceeds the previous best, so all reported results use the bestperforming checkpoint rather than the final iterate, which avoids performance degradation from late-training instability. Selecting a checkpoint by its score on a single noise

realization and then reporting that same score would bias the reported value upward, since the selected checkpoint is the one whose score benefited most from that particular draw. We therefore separate selection from reporting. Every reported number comes from re-evaluating the selected checkpoint on an independent held-out noise draw, holding the farmer population fixed, so the reported values are free of selection bias. Phase

Timesteps

Seasons

Rejection penalty

1 2 3 4

25,000 25,000 25,000 100,000

1 3 5 5

0.3 0.1 0.1 0.0

Table 11: Four-phase training curriculum: horizon and rejection penalty schedule.

Hyperparameter

Value

Algorithm TQC (Truncated Quantile Critics) Network (actor, critic) (256, 256, 256) Learning rate 3 × 10−4 Batch size 1024 Buffer size 500,000 Discount factor γ 1.0 (undiscounted) Entropy coefficient auto (self-tuning, reset per phase) Learning Starts 5000 Parallel environments 24 Gradient steps 2 Training timesteps 175k Table 12: TQC hyperparameters. During training, the aggregator’s reward included an additional rejection penalty term, subtracted per farmer who did not adopt, following the schedule in Table 11. This penalty is reduced to zero by the final training phase, so all reported results reflect the aggregator’s true objective (Eq.17). If any additional details regarding the training setup are required, please refer to the Supplementary code.

E

Agronomic Parameters

The agronomic parameters in Table 14 are based on the agricultural literature (Dahlgreen and Parr 2024; Das et al. 2025; Tyagi and Haritash 2025; World Bank 2021) and have been validated by domain experts. They are intended as a plausible instantiation for one regional setting, i.e, Indian paddy farmers, rather than a definitive calibration. SOC parameters are based on Lal (2015). Here we define yi = 1000 · zi · w

(21)

F

Table 13: Key environment parameters Parameter

Value Symbol

Carbon market price $60/tCO2 e P Result-based payment [0, P ] R Action payment max $100/ha amax Initial SOC range ∼ U [30, 80] tCO2 e/ha Cj0 Maximum SOC 200 tCO2 e/ha Cm Carbon observation noise 0.5 tCO2 e σ Crop selling price $0.27/kg w MRV fixed cost $300 m0 MRV variable cost $80 m1 MRV scaling exponent 0.7 δ Number of farmers 100 N

Adoption Rate Change

Figure 1 illustrates how adoption rates evolve over the seasons, averaged across seeds and farmers. The gradual decline in adoption is driven by soil carbon saturation. As soil carbon accumulates over successive seasons, the marginal amount of additional carbon that can be sequestered decreases, reducing the expected carbon revenue from adopting the practices. Consequently, profit margins decline over time, making it increasingly difficult for the aggregator to learn contracts that are simultaneously attractive to farmers and profitable for the aggregator.

Figure 1: Adoption rates across seasons for the different conditions, variable farm size.

G Table 14: Carbon sequestration practices with associated parameters. AWD: alternate wetting and drying; INM: integrated nutrient management; DRV: deep-root variety. Practice Crop Residue Retention Reduced/Minimum Tillage Mid-Season Drainage AWD SRI Water Management Optimised N Application INM Manure Application Biochar Application Raised Bed Cultivation Early-Maturing Variety High-Biomass/DRV

ci ki zi (tCO2 e/ha) ($/ha) (t/ha) 1.10 0.32 0.40 0.55 0.60 0.35 0.85 0.50 2.20 1.50 0.30 0.70

−140 +0.15 −80 −0.10 −80 −0.10 −80 +0.05 0 +0.30 +120 +0.10 0 +0.10 +125 +0.12 +400 +0.30 +600 +0.45 −80 0.00 −80 +0.10

Sensitivity to the Cost-Limiter Effort Level e

The cost-limiter interpretation of moral hazard depends on the effort level at which the aggregator’s per-hectare losses are evaluated. To check that our conclusions are not an artifact of a particular choice, we re-ran the MH+AS and AS-only conditions under heterogeneous farm size for e ∈ {0.1, 0.2, 0.3}. Because this is a supplementary check, it uses a reduced set of ten seeds; the MH+AS and AS-only values and p-value therefore differ slightly from the 25-seed figures reported in the main text, and the intervals are correspondingly wider. Profit-capture is essentially flat in e (Table 15), and all three rows of MH+AS exceed the AS-only value of 0.546, though the confidence intervals overlap at this sample size. All three paired Wilcoxon tests against AS are significant. The price channel shows the effect most clearly. Paired by seed, the mean action payment rate for adopted practices is lower under MH+AS than under AS only at every effort level (Table 16). The direction is consistent across the range and the magnitude varies. We do not attempt to attribute that variation to a specific mechanism on ten seeds. The sweep establishes what it was designed to, the cost-limiter effect. Moral hazard operating through price rather than quantity under heterogeneous farm size is not an artifact of the particular e adopted in the main text.

H

Cost Limiter Mechanism

Consider this example for one isolated practice. Reduced tillage (base cost −$80/ha, yield effect −100 kg/ha, carbon

Experiment

Mean

Std

95% CI

p-value (vs AS only)

MH+AS (e=0.1) MH+AS (e=0.2) MH+AS (e=0.3) AS only (no MH)

0.603 0.603 0.596 0.546

0.137 0.126 0.130 0.121

[0.523, 0.691] [0.527, 0.682] [0.520, 0.680] [0.474, 0.622]

0.037 0.002 0.002 —

Table 15: Profit-capture ratio across cost-limiter effort levels e. Experiment

AP Rate

AS only

Diff

95% CI

p-value

MH+AS (e=0.1) MH+AS (e=0.2) MH+AS (e=0.3)

$6.88 $7.41 $8.07

$8.92 $8.92 $8.92

$-2.04 $-1.51 $-0.85

[-3.23, -0.86] [-2.50, -0.64] [-1.64, +0.12]

0.0195 0.0098 0.1309

Table 16: Mean action payment rate ($/ha) for adopted practices: each cost-limiter effort level e against AS only. 0.32 tCO2 e/ha) offered to a farmer with cost type θj = 0.45 and farm size 7 ha, in the first season (t = 0) with soil at saturation deficit 1 − Cj /Cm = 0.80, under a contract with R = $10/tCO2 e and action payment $3/ha: eff eff cost = kji + 21 |kji | = −36 + 18 = −$18/ha

(22)

so that total MRV liability over the farmer population is unchanged from the baseline schedule. Writing the baseline schedule as m(f ) = m0 + m1 f δ with m0 = 300, m1 = 80, δ = 0.7, and the counterfactual schedule as m′ (f ) = c f , the neutrality condition over N farms with areas fj is N X

(realised cost via Eq.11 at t = 0; the type-scaled saving eff kji = −80 × 0.45 = −36 is only half-realised early)

j=1

yield_revenue = −100 × $0.27 = −$27/ha

(23)

(negative yield, felt immediately with no ramp, Eq.10) carbon_payment = 0.32 × 0.80 × $10 = $2.56/ha (24) (practice contribution scaled by the saturation deficit) net = −27 + 2.56 − (−18) = −$6.44/ha

(25)

The net loss is therefore $6.44/ha, making the practice independently unprofitable. With MH (e = 0.2) : (0.2 × (−6.44) + 3.00) × 7 = $11.98

(26)

(scaled practice loss is only $1.29/ha, which the $3/ha action payment more than covers, yielding +$1.71/ha) Without MH (e = 1.0) : (1.0 × (−6.44) + 3.00) × 7 = −$24.08

(27)

Without MH, the full $6.44/ha loss exceeds the $3 AP, so the farmer does not adopt.

I

MRV Counterfactual

In the counterfactual regime we remove the fixed per-farmer MRV component (m0 = 0) and set the scaling exponent to δ = 1, so that MRV cost is strictly proportional to cultivated area and cost per hectare is constant across farm sizes. To ensure that the counterfactual isolates the effect of cost structure rather than cost level, we recalibrate the variable rate m1

c fj =

N X

 m0 + m1 fjδ ,

(28)

j=1

which solves to P   m0 N + m1 j fjδ m0 + m 1 E f δ P . = c = E[f ] j fj

(29)

Expressing c through the moments of the farm-size distribution rather than a particular sample makes it a property of the population, so a single constant can be held fixed across all seeds. For f ∼ Uniform[1, 10] we have E[f ] = 5.5 and Z   1 10 0.7 101.7 − 1 ≈ 3.210. (30) E f 0.7 = f df = 9 1 1.7 × 9 Substituting, 300 + 80(3.210) 556.8 = ≈ 101.3 (rounding up), 5.5 5.5 (31) that is, average baseline cost per farm divided by average area per farm. We therefore set m1 = 101.3 in the counterfactual. Two features of the counterfactual deserve discussion. First, removing m0 and setting δ = 1 are not independent changes but jointly define a single object. Both terms must be neutralised for it to be size-invariant. Setting m0 = 0 alone leaves a schedule still declining in f , and the scale economy the counterfactual is meant to measure would survive in the control condition. Second, as seen in tables 17 and 18, realised MRV expenditure rises, but through enrolment rather than price. The results reported in these two tables are cumulative across the 5 seasons. Total expected liability is equalised by construction, and cost per enrolled farmer falls even as the number c=

of enrolled farmers grows: total expenditure is an equilibrium outcome of the aggregator’s contracting problem, not a primitive of the schedule, and its increase is the mechanical consequence of covering farms the baseline excluded. Reported cost per hectare is affected the same way: under m′ it equals c whatever the composition of the enrolled pool, whereas under m it is an area-weighted average of a perhectare schedule that declines in f , so concentrating enrollment among larger farms lowers it mechanically. The gap between the two regimes therefore reflects who enrols rather than what MRV costs.

J

Efficiency Breakdowns

For completeness we report how welfare divides between the two parties under each ablation, for the fixed farm-size setting (Table 19) and the heterogeneous setting (Table 20). Each entry is expressed as a fraction of first-best welfare, so the three columns of a row sum to one. Farmer surplus is the aggregate net payoff to adopting farmers. Aggregator profit is as defined in the main text. Deadweight loss (DWL) is the residual: the share of first-best welfare the learned contract fails to realise, arising from factors such as the pooling constraint and the information frictions.

K

Contract Payments

Figure 2 shows mean carbon and action payments as fractions of revenue. Action payments are favored over result based payments across all ablations. This may be due to allowing greater freedom to the aggregator to incentivise specific actions, however exact characterisation of this behaviour is left to future work.

Figure 2: Contract payment structure, variable farm size.

Metric Farmers enrolled Land enrolled (ha) Total MRV spend MRV / farmer MRV / ha Agg. total profit

First-Best

RL

295 [285, 305] 2010 [1943, 2084] $177984 [$172091, $184220] $604 [$600, $607] $89 [$88, $89] –

160 [146, 175] 1302 [1189, 1420] $103426 [$94201, $112996] $646 [$642, $650] $79 [$79, $80] $84865 [$78967, $90747]

Table 17: Cumulative MRV expenditure and enrollment (MH+AS), m0 = 300, m1 = 80, δ = 0.7 (95% CI).

Metric Farmers enrolled Land enrolled (ha) Total MRV spend MRV / farmer MRV / ha Agg. total profit

First-Best

RL

359 [350, 369] 1979 [1911, 2054] $200498 [$193326, $208020] $558 [$546, $570] $101 [$101, $101] –

283 [242, 326] 1568 [1337, 1803] $158866 [$135017, $183366] $560 [$550, $570] $101 [$101, $101] $76292 [$68227, $84139]

Table 18: Cumulative MRV expenditure and enrollment (MH+AS), m0 = 0, m1 = 101.3, δ = 1.0 (95% CI).

Ablation

Farmer Surplus

Aggregator Profit

DWL

MH+AS MH only AS only Neither

0.154 [0.138, 0.169] 0.103 [0.091, 0.114] 0.175 [0.161, 0.188] 0.121 [0.108, 0.133]

0.599 [0.537, 0.656] 0.699 [0.639, 0.756] 0.566 [0.521, 0.607] 0.809 [0.761, 0.852]

0.247 [0.178, 0.320] 0.198 [0.135, 0.267] 0.259 [0.208, 0.313] 0.071 [0.020, 0.123]

Table 19: Profit-capture decomposition, fixed farm size (95% CI).

Ablation

Farmer Surplus

Aggregator Profit

DWL

MH+AS MH only AS only Neither

0.156 [0.142, 0.171] 0.131 [0.119, 0.142] 0.173 [0.160, 0.187] 0.164 [0.151, 0.177]

0.620 [0.577, 0.662] 0.612 [0.574, 0.651] 0.568 [0.529, 0.609] 0.680 [0.641, 0.719]

0.224 [0.170, 0.277] 0.257 [0.212, 0.303] 0.258 [0.207, 0.308] 0.156 [0.105, 0.205]

Table 20: Profit-capture decomposition, variable farm size (95% CI).

Record · ID 978448 · SHA-256 17028bc75bc3e539
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.