> For the complete documentation index, see [llms.txt](https://gotts.gitbook.io/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://gotts.gitbook.io/docs/gotts-vaults/vault/17-learning-economy.md).

# Learning and Strategy Economy

> **Part of**: [Vault PRD](/docs/gotts-vaults/vault.md) | **Last Updated**: 2026-02-20
>
> Specification for the agent learning loop and strategy marketplace: how agents earn returns, build insights from episodes, package insights as strategies, sell to other agents via x402, and prove performance claims with ZKML.

***

## §1 — Overview & Core Flywheel

The learning economy closes the loop between on-chain execution and continuous improvement. Every vault interaction generates data; that data feeds a self-improvement cycle that produces tradeable strategies; strategies earn revenue that compounds back into vault stakes.

**The four-stage loop:**

```
1. EARN      Agent manages ERC-4626 vault → generates on-chain returns
                           │
                           ▼
2. LEARN     Every vault operation captures an Episode (LanceDB)
             Reflexion inner loop: verbal reflection per operation
             ExpeL outer loop: cross-episode Insight distillation (every 50 episodes)
             Insights parameterize Strategy templates
                           │
                           ▼
3. PROVE     EZKL TypeScript library generates zk-SNARK proof
             Model hash committed on-chain
             Proof verifies strategy model produced claimed outputs
             EAS attestation for performance claims
                           │
                           ▼
4. SELL      StrategyMarketplaceRegistry lists strategies on-chain
             Buyers pay x402 micropayments (HTTP 402)
             Revenue compounds vault stakes
             Bazaar discovery indexes providers
```

> **MVP Scope Constraints (v1)**: The following simplifications apply for the initial release. Each is grounded in production evidence — see §8–§11 and §15 for rationale:
>
> * **Strategy verification**: opML + EAS model commitment + deterministic replay (not ZKML for routine ops; no specialized hardware required)
> * **Model complexity**: sub-100K parameter MLPs/classifiers only (provable in seconds with current tooling)
> * **Memory retention**: P\&L-impact scoring, not Ebbinghaus time-decay
> * **Regime sharing**: probability vector gossip (4 floats), not federated gradient sharing
> * **Pricing**: tiered access (fresh/stale/free) + optional Dutch auction, not continuous exponential
> * **Disputes**: automated → optimistic → expert jury (Kleros-style), not token-weighted UMA voting
> * **Chain**: Base only for v1

**Why this matters**: The research stack supporting this design is deep. Xu & Brini (AAAI 2025) show PPO agents outperform heuristic LP strategies in 7 of 11 out-of-sample windows. FQL (MDPI Electronics 2025) achieves 12.72% annualized return at Sharpe 1.12. The 4-state HMM (Koki et al.) achieves 60%+ directional accuracy on ETH. These results are only available to agents that persist and learn across episodes — not to stateless agents that reset on every vault operation.

**Relationship to DeFi Brain (shared/memory-architecture.md)**: This spec extends the DeFi Brain memory architecture with an economic layer. The DeFi Brain covers the mechanics of episodic/semantic/procedural memory. This spec covers what agents *do* with accumulated knowledge: package it as strategies, prove it, sell it.

***

## §2 — Episode Capture System

Every vault operation captures a `VaultEpisode` record that feeds the learning loop. Episodes are stored in LanceDB (vector table) for semantic retrieval — enabling "find similar market conditions to this current state."

```typescript
// packages/vault/sdk/src/learning/episode.ts
import type { Address } from "viem";

export interface PositionSnapshot {
  poolId: `0x${string}`;
  token0: Address;
  token1: Address;
  tickLower: number;
  tickUpper: number;
  liquidity: bigint;
  feesUncollected0: bigint;
  feesUncollected1: bigint;
  unrealizedPnlBps: number;
}

export interface VaultEpisode {
  id: string; // UUID v4
  timestamp: number; // Unix seconds
  agentId: bigint; // ERC-8004 token ID
  vaultAddress: Address;
  operation:
    | "deposit"
    | "rebalance"
    | "collect_fees"
    | "withdraw"
    | "emergency_exit";

  preState: {
    navPerShare: bigint;
    totalAssets: bigint;
    positions: PositionSnapshot[];
    regimeState:
      | "bull_low_vol"
      | "bull_high_vol"
      | "bear_low_vol"
      | "bear_high_vol";
    vpinScore: number; // 0.0–1.0 (Volume-synchronized PIN)
    gasGwei: number;
    blockNumber: bigint;
  };

  postState: {
    navPerShare: bigint;
    totalAssets: bigint;
    positions: PositionSnapshot[];
    gasUsed: bigint;
  };

  predictedOutcome: Record<string, unknown>;
  actualOutcome: Record<string, unknown>;
  reflection?: string; // LLM verbal analysis (Reflexion pattern)
  confidence: number; // 0.0–1.0 on pattern validity
  plImpactBps: number; // Actual realized P&L impact (positive or negative)
  regimeTag: string; // HMM-detected regime at operation time (reused for recall)
  importanceScore: number; // Computed: |plImpactBps| × recency_weight (0.0–1.0)
}
```

**Storage**: LanceDB vector table. Each episode is embedded (nomic-embed-text-v1.5) on the concatenation of `operation + regimeState + reflection`. This enables semantic queries like "find past rebalance episodes during bear\_high\_vol with VPIN > 0.7."

```typescript
// packages/vault/sdk/src/learning/episode-store.ts
import * as lancedb from "@lancedb/lancedb";
import { EmbeddingFunction } from "@lancedb/lancedb/embedding";

export class EpisodeStore {
  private table: lancedb.Table;

  async add(episode: VaultEpisode): Promise<void>;

  async searchSimilar(
    state: { regimeState: string; vpinScore: number; operation: string },
    k: number,
  ): Promise<VaultEpisode[]>;

  async getByOperation(
    op: VaultEpisode["operation"],
    limit: number,
  ): Promise<VaultEpisode[]>;

  async getRecent(agentId: bigint, limit: number): Promise<VaultEpisode[]>;

  // Conditional decay: importance-weighted decay rate, not binary keep/prune
  // High-impact (|plImpactBps| > 100): 5× slower decay (months)
  // Medium-impact (50–100 bps): 2× slower decay (weeks)
  // Low-impact: standard TTL
  async pruneByImportance(): Promise<number>;
}
```

**Retention policy** (importance-weighted conditional decay, FinMem architecture):

* **High-impact (|plImpactBps| > 100 bps)**: decay rate reduced 5× — retained for months, not indefinitely
* **Medium-impact (50–100 bps)**: decay rate reduced 2× — retained for weeks
* **Low-impact (<50 bps)**: standard TTL applies (gas price observations: 7 days; general: 14 days)
* **Regime-indexed recall**: episodes tagged with `regimeTag` at capture time; retrieved when regime recurs
* **Decay does not mean deletion**: decayed episodes are archived, never hard-deleted — available for regime-conditional recall
* **STONE paradigm** (arXiv:2602.16192, Luo et al.): keep raw episodes in LanceDB; extract insights on demand at query time

> **Why conditional decay, not immortal retention?** FadeMem (Wei et al., arXiv:2601.18642, Jan 2026) — the closest published system to this design — achieves 82.1% retention of critical facts using 55% of storage via *adaptive* decay. Its core thesis: selective forgetting is essential to prevent information overload. High-impact memories decay slower, not never. Xiong et al. (May 2025) found that utility-based deletion yielded up to 10% performance gains over naive retention strategies.

***

## §3 — Reflexion Inner Loop

After every vault operation, the agent generates a verbal self-reflection comparing prediction vs. reality. This implements the Reflexion pattern (Shinn et al., NeurIPS 2023): intra-task verbal reinforcement that achieves 91% pass\@1 on decision-making benchmarks without model weight updates.

```typescript
// packages/vault/sdk/src/learning/reflexion.ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

export async function generateReflection(
  episode: VaultEpisode,
): Promise<string> {
  const navChangeBps = computeNavChangeBps(episode);
  const gasCostUsd = computeGasCostUsd(
    episode.postState.gasUsed,
    episode.preState.gasGwei,
  );

  // Two-tier LLM: routine reflections (|navChangeBps| < 50, no anomaly) use haiku;
  // anomalous (large NAV change, emergency exit, VPIN > 0.8) use opus.
  // TradingAgents (Xiao et al., 2024): GPT-4o-mini for summarization, o1 for reasoning-intensive.
  // Trading-R1 (arXiv:2509.11420): free-form intermediate reasoning compounds errors.
  const model = isAnomalous(episode)
    ? "claude-opus-4-6"
    : "claude-haiku-4-5-20251001";

  const msg = await client.messages.create({
    model,
    max_tokens: 512,
    messages: [
      {
        role: "user",
        content: `Structured reflection on this DeFi vault operation.
• Operation: ${episode.operation}, Regime: ${episode.preState.regimeState}, VPIN: ${episode.preState.vpinScore.toFixed(3)}
• Observed NAV change: ${navChangeBps > 0 ? "+" : ""}${navChangeBps} bps (ground truth from on-chain data)
• Gas cost: $${gasCostUsd.toFixed(4)} (from receipt, not estimated)
• Primary causal factor (must be consistent with observed bps change):
  [Provide ONE cause that is mathematically consistent with the NAV delta]
• Adjustment for next similar operation:
  [Specific parameter change, e.g. "reduce rangeTightnessBps by 50"]`,
      },
    ],
  });

  return (msg.content[0] as { text: string }).text;
}

/**
 * Ground-truth backcheck: verify the reflection's causal claim is
 * mathematically consistent with observed on-chain data.
 * If claimed cause cannot account for >50% of navChangeBps, flag as uncertain.
 * Reference: Semantic Entropy (Farquhar et al., Nature Vol. 630, June 2024)
 * cannot detect systematic confabulation — programmatic grounding is required.
 */
async function validateReflection(
  reflection: string,
  episode: VaultEpisode,
): Promise<{ valid: boolean; inconsistency?: string }>;

export async function recordAndReflect(
  episode: VaultEpisode,
  store: EpisodeStore,
): Promise<VaultEpisode> {
  episode.reflection = await generateReflection(episode);
  episode.confidence = assessConfidence(episode);
  await store.add(episode);
  return episode;
}

/**
 * Heuristic confidence score 0.0–1.0 based on:
 * - Prediction accuracy: low gap → higher confidence
 * - Sample size: more similar past episodes → higher confidence
 * - Regime stability: stable regime → higher confidence
 */
function assessConfidence(episode: VaultEpisode): number {
  const predictionAccuracy = computePredictionAccuracy(episode);
  const regimeStabilityBonus = episode.preState.vpinScore < 0.3 ? 0.1 : 0;
  return Math.min(0.95, predictionAccuracy + regimeStabilityBonus);
}
```

**Integration with vault operations**: After each `vault_rebalance`, `vault_collect_fees`, or `vault_emergency_exit` call, the MCP tool handler invokes `recordAndReflect()` asynchronously. The reflection is non-blocking — it does not delay the operation response to the calling agent.

**Safety boundary**: Reflections are advisory only. They may suggest parameter adjustments, but cannot override PolicyCage (D-053) on-chain constraints. If a reflection recommends "increase position size," that recommendation must still pass through the on-chain `maxPositionBps` check.

***

## §4 — ExpeL Outer Loop (Insight Distillation)

Cross-episode distillation runs every 50 episodes to extract reusable insights. This implements the ExpeL pattern (Zhao et al., AAAI 2024): cross-task insight extraction via ADD / UPVOTE / DOWNVOTE / EDIT operations on a natural-language insight pool.

```typescript
// packages/vault/sdk/src/learning/expel.ts

export type InsightCategory =
  | "procedural" // Step-by-step action sequences (MUSE: procedural memory)
  | "strategic" // Higher-level regime patterns (MUSE: strategic memory)
  | "tool_guidance" // Tool-specific usage patterns (MUSE: tool memory)
  | "rebalance_timing" // When to rebalance (regime, VPIN, gas thresholds)
  | "fee_optimization" // Optimal collect_fees frequency
  | "regime_transition" // How to behave when regime shifts
  | "gas_execution" // Gas cost management
  | "range_selection" // LP range selection by volatility
  | "jit_defense"; // Defending against JIT liquidity attacks

export interface Insight {
  id: string;
  category: InsightCategory;
  rule: string; // Natural language rule, e.g.:
  // "Avoid rebalancing when VPIN > 0.7 — slippage
  //  is 2-3x higher in toxic order flow conditions."
  confidence: number; // 0.0–1.0
  supportingEpisodeIds: string[]; // Evidence base (at least 3 for new insights)
  votes: { up: number; down: number };
  createdAt: number;
  lastUpdated: number;
}

export type InsightOperation =
  | {
      type: "ADD";
      insight: Omit<Insight, "id" | "votes" | "createdAt" | "lastUpdated">;
    }
  | { type: "UPVOTE"; insightId: string; episodeId: string }
  | { type: "DOWNVOTE"; insightId: string; episodeId: string; reason: string }
  | { type: "EDIT"; insightId: string; newRule: string; newConfidence: number };

export async function distillInsights(
  episodes: VaultEpisode[],
  existingInsights: Insight[],
): Promise<{ updatedInsights: Insight[]; operations: InsightOperation[] }> {
  // 1. Cluster episodes by (regimeState × operation × outcome_direction)
  // 2. For each cluster, identify patterns vs existing insight pool:
  //    - New pattern with 3+ supporting episodes → ADD
  //    - Pattern matches existing insight → UPVOTE or DOWNVOTE
  //    - Pattern contradicts existing insight → EDIT with merged rule
  // 3. Cap parameter changes at 10% per distillation cycle (SAFLA constraint)
  // 4. Return updated insight pool with operations log
}
```

> **Known limitations of static ExpeL operations:**
>
> 1. **Batch size unvalidated**: The 50-episode trigger has no empirical backing in the original ExpeL paper (Zhao et al., AAAI 2024). LaMer (arXiv:2512.16848) uses 3–5 episode windows. Start with 10–20 episode windows and tune empirically.
> 2. **Fails on rare-success tasks**: SaMuLe (arXiv:2509.20562, EMNLP 2025) showed ExpeL collapses to 0% on TravelPlanner (high-complexity, low-success-rate tasks). DeFi novel market conditions frequently produce failures with no paired successes. Add failure-only learning: single failed trajectories can generate DOWNVOTE and EDIT operations without requiring a successful counterpart.
> 3. **Noise accumulation**: UMEM (Ye et al., arXiv:2602.10652, Feb 2026) demonstrated that fixed ADD/UPVOTE/DOWNVOTE/EDIT prompts accumulate instance-specific noise over time, causing performance degradation. Plan a v2 Mem-Optimizer (learnable extraction policy) when insight pool exceeds 500 entries.

**Insight storage**: SQLite database (semantic memory, per DeFi Brain architecture). Queryable by category, confidence threshold, and creation date.

**Distillation trigger**: Automatically after every 50 new episodes, or manually via `insight_distill` MCP tool. The `strategy-optimizer` agent (08-agents-economy.md) coordinates both the episodic capture and the periodic distillation.

**Insight decay**: Like episodes, insights have confidence decay. An insight that generates no supporting UPVOTE operations for 60 days has its confidence reduced by 20%. Insights with confidence < 0.3 are archived (not deleted — they may become relevant again in future regime shifts).

***

## §5 — Strategy Data Format

A `VaultStrategy` packages parameterized templates, feature importance weights, a regime classifier, and a performance attestation. Strategies are the tradeable artifacts of the learning loop.

```typescript
// packages/vault/sdk/src/learning/strategy.ts
import type { Address } from "viem";

export type StrategyFamily =
  | "lp_optimization" // Concentrated LP range management
  | "mean_reversion" // Range around time-weighted mean price
  | "momentum" // Asymmetric ranges following trend
  | "yield_compounding"; // Fee harvest and reinvestment optimization

export interface VaultStrategy {
  // Identity
  strategyId: `0x${string}`; // keccak256(agentId + version + timestamp)
  agentId: string; // ERC-8004 token ID (string for JSON compatibility)
  version: number;
  strategyFamily: StrategyFamily;
  createdAt: number; // Unix seconds

  // Parameters (all values operate within PolicyCage bounds)
  parameters: {
    rangeTightnessBps: number; // e.g. 500 = ±5% around current price
    rebalanceThresholdBps: number; // e.g. 150 = trigger when 1.5% out-of-range
    gasMaxGwei: number; // Max gas price for rebalance execution
    minIdleBps: number; // Min idle capital % (from PolicyCage)
    maxPositionBps: number; // Max single-position % (from PolicyCage)
    feesCollectFrequencyHours: number;
  };

  // Feature importance (SHAP-derived, values sum to 1.0)
  featureImportance: {
    price_zscore: number;
    volume_ratio: number;
    realized_vol_1h: number;
    realized_vol_24h: number;
    tvl_change: number;
    funding_rate: number;
    vpin_score: number;
  };

  // Regime classifier (4-state HMM, Koki et al.: 60%+ directional accuracy on ETH)
  regimeClassifier: {
    states: ["bull_low_vol", "bull_high_vol", "bear_low_vol", "bear_high_vol"];
    currentState: string;
    probabilities: [number, number, number, number]; // Sum to 1.0
    transitionMatrix: number[][]; // 4×4 HMM transition probabilities
  };

  // Performance attestation
  performance: {
    sharpeRatio: number;
    maxDrawdownBps: number;
    winRate: number; // 0.0–1.0
    sampleSize: number; // Number of episodes supporting these metrics
    periodDays: number;
    easAttestationHash: `0x${string}`; // EAS schema attestation on Base
    zkProofHash?: `0x${string}`; // ZK proof hash — populated only when dispute resolved via Bionetta/DeepProve
  };

  // Marketplace pricing
  pricing: {
    baseUsdc: number; // Base price in USDC (6 decimals, e.g. 100000 = $0.10)
    decayLambda: number; // Exponential time-decay constant (see §11)
    subscribersCount: number;
    stakedAmount: bigint; // Seller stake in USDC (slashable on underperformance)
    escrowAddress: Address; // StrategyEscrow contract holding seller stake
  };

  // Distribution
  metadataURI: string; // IPFS URI of full strategy JSON (pinned)
  supportedAssets: string[]; // e.g. ['ETH/USDC', 'ETH/USDT', 'WBTC/USDC']
}
```

**Strategy generation**: The `strategy-marketplace` agent (see §12) monitors the insight pool. When PSR > 0.90, sampleSize ≥ 200 across ≥2 regimes, and ensemble confidence > 0.7 (§12 for full criteria), it generates a strategy by:

1. Aggregating high-confidence insights in each `InsightCategory`
2. Fitting parameters to maximize Sharpe over the supporting episodes
3. Computing feature importance via SHAP values from historical episode outcomes
4. Running the HMM classifier on recent market data
5. Generating an EAS attestation of the performance metrics

***

## §6 — x402 Strategy Marketplace Server

Strategy providers expose an Express server with x402-gated routes. Buyers pay per query via the HTTP 402 Payment Required protocol. The `.well-known/x402.json` manifest is indexed by Bazaar and other x402 facilitators for discovery.

```typescript
// packages/vault/src/marketplace/server.ts
import express from "express";
import { paymentMiddleware } from "@x402/express";
import type { VaultStrategy } from "../../sdk/src/learning/strategy.js";

const app = express();

// Pricing tiers (calibrated to strategy value and alpha decay rates)
// Research basis: Maven Securities alpha decay — 9.9% Europe, 5.6% US per unit of delay
// Basic market signals:         $0.001–$0.01
// Regime classification:        $0.01–$0.05
// Strategy parameters:          $0.05–$0.50
// Real-time alpha signals:      $0.50–$5.00

const PRICE_TIERS = {
  regime_signal: "0.02", // $0.02 USDC per regime query
  strategy_params: "0.10", // $0.10 USDC per strategy param fetch
  alpha_signal: "1.00", // $1.00 USDC per real-time alpha signal
} as const;

// Bazaar discovery manifest — indexed by x402 Facilitators
// Reference: x402 protocol spec v2.0 (Cloudflare)
app.get("/.well-known/x402.json", (_req, res) => {
  res.json({
    version: "2.0",
    provider: {
      agentId: process.env.AGENT_ID,
      erc8004Address: "0x8004A818BFB912233c491871b3d84c89A494BD9e",
      registryAddress: process.env.STRATEGY_REGISTRY_ADDRESS,
    },
    endpoints: [
      {
        path: "/strategies",
        methods: ["GET"],
        price: PRICE_TIERS.strategy_params,
        asset: "USDC",
        description: "Filtered strategy marketplace search",
      },
      {
        path: "/signals/regime",
        methods: ["GET"],
        price: PRICE_TIERS.regime_signal,
        asset: "USDC",
        description: "Current regime classification with probabilities",
      },
      {
        path: "/signals/alpha",
        methods: ["GET"],
        price: PRICE_TIERS.alpha_signal,
        asset: "USDC",
        description: "Real-time alpha signals for active strategy subscribers",
      },
    ],
  });
});

// Strategy list — pay-per-query, filtered search
app.get(
  "/strategies",
  paymentMiddleware({
    amount: PRICE_TIERS.strategy_params,
    asset: "USDC",
    network: "base",
  }),
  async (req, res) => {
    const strategies = await queryStrategies({
      assetClass: req.query.assetClass as string,
      regimeSuitable: req.query.regimeSuitable as string,
      minSharpe: Number(req.query.minSharpe ?? 0),
      maxPriceUsdc: Number(req.query.maxPriceUsdc ?? 1.0),
      minReputation: Number(req.query.minReputation ?? 0),
    });
    // Apply alpha decay to prices before returning
    const decayAdjusted = strategies.map(applyPriceDecay);
    res.json(decayAdjusted);
  },
);

// Regime signal — pay-per-query
app.get(
  "/signals/regime",
  paymentMiddleware({
    amount: PRICE_TIERS.regime_signal,
    asset: "USDC",
    network: "base",
  }),
  async (_req, res) => {
    const regime = await getCurrentRegime();
    res.json({
      ts: Date.now(),
      regime: regime.currentState,
      probabilities: regime.probabilities,
      confidence: Math.max(...regime.probabilities),
    });
  },
);

// Alpha signals — pay-per-query or subscribe
app.get(
  "/signals/alpha",
  paymentMiddleware({
    amount: PRICE_TIERS.alpha_signal,
    asset: "USDC",
    network: "base",
  }),
  async (req, res) => {
    const { strategyId, buyerAgentId } = req.query;
    const signals = await getAlphaSignals(
      strategyId as string,
      buyerAgentId as string,
    );
    res.json(signals);
  },
);

export { app };
```

**Compute-to-data pattern** (D-090): For strategy inspection, buyers send their portfolio context to the strategy algorithm; the algorithm returns trade signals without exposing its internals. This solves the "inspection paradox" — a buyer cannot verify a strategy without running it, but running it reveals the strategy. The x402 server implements compute-to-data: the buyer sends `{ currentPositions, navPerShare, regimeState }` and receives `{ rebalanceRecommendation, rangeAdjustments }` back. Neither party's proprietary information crosses the boundary.

***

## §7 — MCP Tool Specifications

Six strategy marketplace tools added to the vault tool inventory. These tools activate when `TOOL_PROFILE` includes `vault` or `marketplace`.

### `list_strategies`

Filtered search of the strategy marketplace with reputation-weighted ranking.

```typescript
server.tool(
  "list_strategies",
  "Search strategy marketplace by asset class, regime suitability, Sharpe ratio, and price cap",
  {
    assetClass: z
      .enum(["eth_usdc", "eth_usdt", "wbtc_usdc", "any"])
      .default("any"),
    regimeSuitable: z
      .enum([
        "bull_low_vol",
        "bull_high_vol",
        "bear_low_vol",
        "bear_high_vol",
        "any",
      ])
      .default("any"),
    minSharpe: z
      .number()
      .min(0)
      .default(0)
      .describe("Minimum Sharpe ratio for strategies"),
    maxPriceUsdc: z
      .number()
      .min(0)
      .default(1.0)
      .describe("Maximum price in USDC (after alpha decay)"),
    minReputation: z
      .number()
      .min(0)
      .default(0)
      .describe("Minimum ERC-8004 reputation score of strategy seller"),
    limit: z.number().min(1).max(50).default(10),
    chain: z.string().default("base"),
  },
  async (params) => {
    // 1. Query StrategyMarketplaceRegistry on-chain for active strategies
    // 2. Filter by params
    // 3. Apply alpha decay to prices (P(t) = P_base × e^(-λt))
    // 4. Fetch seller reputation from ERC-8004 Reputation Registry
    // 5. Rank by (reputation_weight × sharpe) / decay_adjusted_price
    // Returns: Array<VaultStrategy> with current prices
  },
);
```

### `purchase_strategy`

Initiates x402 payment flow and returns strategy JSON on confirmation.

```typescript
server.tool(
  "purchase_strategy",
  "Purchase a strategy via x402 micropayment. Returns strategy parameters on payment confirmation.",
  {
    strategyId: z.string().describe("Strategy ID from list_strategies"),
    buyerAgentId: z.string().describe("ERC-8004 agent ID of the buyer"),
    chain: z.string().default("base"),
  },
  async ({ strategyId, buyerAgentId, chain }) => {
    // 1. Fetch strategy price from StrategyMarketplaceRegistry (tier/decay-adjusted)
    // 2. Initiate x402 payment: buyer's wallet sends USDC to StrategyEscrow
    // 3. StrategyEscrow calls recordPurchase() on registry (starts opML challenge window)
    // 4. On confirmation: fetch strategy JSON from IPFS metadataURI
    // 5. Return strategy with EAS model commitment hash (primary) + zkProofHash if dispute-resolved
    // Tiered verification (all purchases covered by AUM-tiered opML challenge window):
    //   Any amount: EAS model commitment + opML optimistic acceptance (24h / 72h / 7d by AUM)
    //   Public models: deterministic replay via onnxruntime-node (no hardware required)
    //   Dispute filed: async ZK proof (Bionetta or DeepProve) generated and verified
  },
);
```

### `query_agent_reputation`

Read ERC-8004 Reputation Registry for strategy seller agents.

```typescript
server.tool(
  "query_agent_reputation",
  "Get reputation data for a strategy seller agent from the ERC-8004 Reputation Registry",
  {
    agentId: z.string().describe("ERC-8004 agent ID to query"),
    includeStrategyHistory: z
      .boolean()
      .default(false)
      .describe("Include strategy sales history and buyer ratings"),
    chain: z.string().default("base"),
  },
  async ({ agentId, includeStrategyHistory }) => {
    // Returns: reputation tier, score, milestone completions,
    //          strategy count, avg buyer rating, slash history
  },
);
```

### `subscribe_signals`

Returns a WebSocket endpoint for real-time strategy signals with x402 subscription payment.

```typescript
server.tool(
  "subscribe_signals",
  "Subscribe to real-time strategy signals via WebSocket. Returns ws:// endpoint with auth token.",
  {
    strategyId: z.string().describe("Strategy ID to subscribe to"),
    buyerAgentId: z.string().describe("ERC-8004 agent ID of the subscriber"),
    signalTypes: z
      .array(z.enum(["regime", "alpha", "rebalance"]))
      .describe("Types of signals to receive"),
    chain: z.string().default("base"),
  },
  async (params) => {
    // 1. Verify buyer has purchased strategy (check StrategyMarketplaceRegistry)
    // 2. Issue time-limited WebSocket auth token (signed JWT, 24h TTL)
    // 3. Return ws:// endpoint URL and auth token
    // Signal delivery: ws sends JSON per signal_type on each new signal
    // Rate: regime signals ~every 5 minutes; alpha/rebalance signals as generated
  },
);
```

### `publish_strategy`

Register a new strategy with metadata, pricing, seller stake, and optional ZK proof.

```typescript
server.tool(
  "publish_strategy",
  "Publish a strategy to the marketplace with performance attestation and seller stake",
  {
    metadata: z.object({
      strategyFamily: z.enum([
        "lp_optimization",
        "mean_reversion",
        "momentum",
        "yield_compounding",
      ]),
      parameters: z.record(z.unknown()),
      featureImportance: z.record(z.number()),
      regimeClassifier: z.object({
        currentState: z.string(),
        probabilities: z.array(z.number()).length(4),
      }),
      performance: z.object({
        sharpeRatio: z.number(),
        maxDrawdownBps: z.number(),
        winRate: z.number().min(0).max(1),
        sampleSize: z.number().min(200), // Minimum 200 episodes across ≥2 regimes required
        periodDays: z.number().min(7),
      }),
      supportedAssets: z.array(z.string()),
    }),
    priceUsdc: z.number().min(0.001).describe("Base price in USDC"),
    stakeUsdc: z.number().min(10).describe("Seller stake in USDC (slashable)"),
    sellerAgentId: z.string().describe("ERC-8004 agent ID of the seller"),
    zkProofHash: z
      .string()
      .optional()
      .describe(
        "ZK proof hash (Bionetta/DeepProve) — required only when dispute is filed, not at publication time",
      ),
    easAttestationHash: z
      .string()
      .optional()
      .describe("EAS performance attestation hash"),
    chain: z.string().default("base"),
  },
  async (params) => {
    // 1. Validate seller is ERC-8004 registered with Verified+ reputation (score >= 50)
    // 2. Upload strategy JSON to IPFS, get metadataURI
    // 3. Generate strategyId = keccak256(sellerAgentId + version + timestamp)
    // 4. Lock seller stake in StrategyEscrow
    // 5. Call StrategyMarketplaceRegistry.publishStrategy()
    // 6. Generate EAS attestation if not provided
    // Returns: strategyId, metadataURI, escrowAddress, txHash
  },
);
```

### `claim_strategy_royalties`

Harvest accumulated x402 payment proceeds from strategy sales.

```typescript
server.tool(
  "claim_strategy_royalties",
  "Claim accumulated royalties from strategy sales. Routes proceeds to treasury agent.",
  {
    strategyId: z.string().describe("Strategy ID to claim royalties for"),
    sellerAgentId: z.string().describe("ERC-8004 agent ID of the seller"),
    chain: z.string().default("base"),
  },
  async ({ strategyId, sellerAgentId, chain }) => {
    // 1. Call StrategyMarketplaceRegistry.claimRoyalties(strategyId)
    // 2. Verify caller is registered seller
    // 3. Transfer accumulated USDC to seller wallet
    // Returns: amountClaimedUsdc, txHash, nextClaimableAt
  },
);
```

**Tool inventory update**: These 6 tools add to the vault MCP tool inventory. With these additions:

| Scope                         | Count  |
| ----------------------------- | ------ |
| Core v1 vault tools           | 24     |
| Strategy marketplace tools    | 6      |
| am-AMM strategy auction tools | 5      |
| Optional proxy tools          | 6      |
| Deferred tools                | 8      |
| **Roadmap total**             | **49** |

***

## §8 — Strategy Verification: opML + EAS Commitment + Selective ZK

> **TEE removed**: An earlier draft required Intel TDX enclaves for vault AUM > $10K. This tier has been removed. The Battering RAM attack (Van Bulck et al., IEEE S\&P 2026) now breaks Intel TDX, AMD SEV-SNP, and NVIDIA Confidential Computing for under $50, forging attestation quotes with "UpToDate" trust designation — all major cloud TEE vendors acknowledged (see [10-safety.md](/docs/gotts-vaults/vault/10-safety.md) D-036). This also contradicted the prior spec's claim that "TDX is NOT affected," which was based on the 2025 DDR4-only version of the attack before the escalated IEEE S\&P 2026 findings. The replacement — EAS model commitment + deterministic replay — is more transparent and verifiable by anyone with Node.js. Full TypeScript, no specialized hardware required.

> **Reality check on full ZKML**: Full EZKL proof takes \~16 minutes for 1M-parameter models; on-chain verification keys are 650KB–4.2MB. Not viable for routine purchases. The architecture below scales verification rigor to value at risk.

### Verification Architecture

Three layers. No specialized hardware required.

**1. opML challenge window (all strategies) — AUM-tiered**

Strategy execution results are posted on-chain optimistically. The challenge window scales with vault AUM, giving the community proportionally more time to catch fraud at higher stakes.

| Vault AUM | Challenge Window | Rationale                                                          |
| --------- | ---------------- | ------------------------------------------------------------------ |
| < $1K     | 24h              | Practical minimum; low financial incentive to organize a challenge |
| $1K–$10K  | 72h              | 3× more time for community coordination                            |
| > $10K    | 7 days           | Matches Arbitrum's standard fraud proof window; highest stakes     |

Anyone can challenge by submitting a counterclaim with bond (5% of strategy price) during the window. If unchallenged, result is accepted. Reference: ORA Protocol FPVM (Optimistic ML). Covers \~99% of all strategy purchases.

> **Why not a flat 24h?** A flat 24h window was sufficient when TEE provided a parallel certainty layer for large AUM. Without TEE, the opML window must scale to stake size. AUM-tiered windows restore proportionality between challenge incentive and window duration.

**2. EAS model commitment (all marketplace strategies)**

At publication time, the seller commits a cryptographically signed bundle to EAS on Base. This is the primary on-chain attestation for all routine purchases.

```typescript
// packages/vault/sdk/src/learning/commitment.ts
import * as ort from "onnxruntime-node";
import { keccak256, encodeAbiParameters, type Address } from "viem";

export interface ModelCommitment {
  modelCID: string; // IPFS CID of the ONNX file (populated for public models only)
  modelHash: `0x${string}`; // keccak256(onnxBytes) — prevents CID-swapping attacks
  inputDataHash: `0x${string}`; // keccak256(inputDataSnapshot) — public market data from The Graph
  claimedOutputHash: `0x${string}`; // keccak256(abi.encode(sharpe × 1e6, maxDrawdown × 1e6, winRate × 1e6))
  agentSignature: `0x${string}`; // ECDSA signature by seller's ERC-8004 key over the above fields
  easAttestationHash: `0x${string}`; // EAS attestation hash (on-chain record on Base)
  modelIsPublic: boolean; // true = ONNX on IPFS (deterministically verifiable); false = private
}

export async function computeModelCommitment(
  modelBytes: Uint8Array,
  inputData: Float32Array[],
  claimedOutputs: { sharpe: number; maxDrawdown: number; winRate: number },
): Promise<Omit<ModelCommitment, "agentSignature" | "easAttestationHash">> {
  const modelHash = keccak256(modelBytes);
  const inputDataHash = keccak256(
    new Uint8Array(inputData.flatMap((arr) => Array.from(arr))),
  );
  const claimedOutputHash = keccak256(
    encodeAbiParameters(
      [{ type: "int256" }, { type: "int256" }, { type: "int256" }],
      [
        BigInt(Math.round(claimedOutputs.sharpe * 1e6)),
        BigInt(Math.round(claimedOutputs.maxDrawdown * 1e6)),
        BigInt(Math.round(claimedOutputs.winRate * 1e6)),
      ],
    ),
  );
  return {
    modelHash,
    inputDataHash,
    claimedOutputHash,
    modelCID: "",
    modelIsPublic: false,
  };
}

/**
 * Verify a public model commitment by re-running ONNX with the same inputs.
 * Input data must be fetched from the same public source (The Graph / on-chain RPC).
 * Anyone with Node.js and onnxruntime-node can run this — no special hardware required.
 */
export async function verifyModelCommitment(
  commitment: ModelCommitment,
  modelBytes: Uint8Array,
  inputData: Float32Array[],
): Promise<{ valid: boolean; recomputedOutputHash: `0x${string}` }> {
  const session = await ort.InferenceSession.create(modelBytes);
  // Run inference → encode outputs → keccak256 → compare to commitment.claimedOutputHash
}
```

**Public models** (ONNX posted to IPFS): deterministically verifiable by anyone. A challenger downloads the ONNX from the committed CID, fetches the same historical market data from The Graph or on-chain RPC, runs `onnxruntime-node` in Node.js, and compares the recomputed output hash to `claimedOutputHash`. Hash mismatch = provable fraud. No hardware, no trust in a remote attestation required.

> **Why this is stronger than TEE**: Intel TDX attestation can be forged for under $50 (Battering RAM, Van Bulck et al., IEEE S\&P 2026). Deterministic ONNX replay cannot be forged — the math is identical on every machine. Public-model strategies are cryptographically transparent in a way hardware attestation never was.

**Private models** (ONNX not posted publicly): The EAS commitment records only `{ modelHash, inputDataHash, claimedOutputHash }` — contents are opaque to buyers. Verification is only possible via ZK proof (Layer 3) if a dispute is filed. The marketplace UI surfaces private models with a distinct badge; buyers accept higher information risk.

**3. ZK proof for dispute resolution only (async, 2–20 min)**

Unchanged from the opML/ZK design. ZK proofs are generated asynchronously when a challenge is filed during the opML window — not at purchase time.

* **Bionetta** (Groth16): 320-byte proofs, \~200K gas to verify on Base, 2.3-second mobile proving. Requires access to model weights — public models only.
* **DeepProve / EZKL** (Halo2): For private models. 54–158× faster than EZKL alone; still async (minutes). Seller provides model to verifier under dispute escrow.

**MVP model constraint**: Marketplace strategies are limited to sub-100K parameter MLPs and decision trees expressible as ONNX. These prove in sub-second with Bionetta or DeepProve today.

```typescript
// packages/vault/sdk/src/learning/zkml.ts
import type { Address, PublicClient } from "viem";

export interface PerformanceProof {
  proofBytes: Uint8Array;
  instancesHash: `0x${string}`; // Keccak256 of public inputs/outputs
  verifierAddress: Address; // Deployed verifier contract on Base
  method: "bionetta_groth16" | "deepprove_halo2" | "ezkl_halo2";
}

/**
 * Generate a ZK proof that strategy model M produced the claimed performance.
 * Called asynchronously for dispute resolution only — NOT at purchase time.
 * Proof generation: 2–20 minutes (async, server-side).
 * Verification: ~0.5 seconds on Base (~200K gas for Bionetta Groth16).
 * Model constraint: ONNX under 100K parameters only (MVP scope).
 */
export async function proveStrategyPerformance(
  modelOnnx: Uint8Array,
  inputData: Float32Array[],
  claimedOutputs: { sharpe: number; maxDrawdown: number; winRate: number },
): Promise<PerformanceProof>;

/** Verify a performance proof on-chain. Called by StrategyEscrow on dispute resolution. */
export async function verifyPerformanceProof(
  proof: PerformanceProof,
  publicClient: PublicClient,
): Promise<boolean>;
```

### Tiered Verification Policy

| Purchase Context           | Verification               | Method                                                                         |
| -------------------------- | -------------------------- | ------------------------------------------------------------------------------ |
| Any amount                 | opML challenge window      | 24h (< $1K AUM) / 72h ($1K–$10K) / 7 days (> $10K)                             |
| All marketplace strategies | EAS model commitment       | On-chain signed `{ modelCID, inputDataHash, claimedOutputHash }`               |
| Public models              | Deterministic replay       | Anyone: download ONNX from IPFS → run `onnxruntime-node` → compare output hash |
| Disputed claim             | ZK proof (async, 2–20 min) | Bionetta (public models) or DeepProve/EZKL (private)                           |

**`zkProofHash` on `VaultStrategy`**: Populated only when a dispute has been resolved via ZK proof. Not required at publication. The `easAttestationHash` (from the EAS model commitment) is the primary on-chain record for all routine purchases. `modelCID` is populated for public models only.

**What is proven**: The EAS commitment attests "model M (hash H) applied to input data I (hash D) claims outputs O (hash Q)." For public models, anyone can re-derive Q from the ONNX. For private models, a ZK proof attests the same without revealing M. Buyers receive assurance that the claimed performance metrics are consistent with the model's actual outputs — not that the model will continue to perform.

***

## §9 — Regime Signal Gossip Protocol

The highest-ROI collaborative feature: vaults share regime probability vectors — not raw strategies, position data, or model gradients. Each vault observes different DeFi protocols and tokens; the ensemble accuracy is higher than any single agent.

> **Why not federated learning?** No production deployment of federated learning in DeFi exists as of early 2026; gradient sharing adds infrastructure complexity without benefit at 5–50 vault scale. Non-IID data across vaults causes FedAvg accuracy drops of 10–20%. Adversarial incentives are uniquely strong — competing vaults have direct financial motivation to poison shared models. Differential privacy at ε ≤ 2 on small DeFi datasets renders models unusable. Sharing **probability vectors** (not gradients) is the correct simplification: it is model-agnostic, immune to gradient inversion attacks, trivially aggregated, and achievable in real-time. This maps to federated distillation and PATE (Private Aggregation of Teacher Ensembles) in the literature.

```typescript
// packages/vault/sdk/src/learning/regime-gossip.ts

export interface RegimeProbabilities {
  agentId: bigint;
  timestamp: number;
  // [bull_low_vol, bull_high_vol, bear_low_vol, bear_high_vol]
  probabilities: [number, number, number, number];
  weight: number; // Shapley-derived contribution weight (GTG-Shapley, cumulative)
  addedNoise: number; // Differential privacy noise level ε (recommended 2.0–4.0)
}

// NOTE: This is NOT gradient-based federated learning. No model weights are shared.
// Only 4 probability floats per update — ~100,000× less data than gradient sharing.
export class RegimeGossipCollector {
  private participants: Map<bigint, RegimeProbabilities> = new Map();
  private readonly CONSENSUS_THRESHOLD = 0.66; // 66% of weighted votes

  /**
   * Contribute local HMM output with differential privacy noise.
   * Only 4 floats + noise are shared — raw data never leaves the agent.
   * Privacy budget: ε = 2.0–4.0 (Gaussian mechanism with Balle & Wang calibration).
   * ε ≤ 1.0 produces σ > 6 noise on [0,1] probabilities — effectively destroys signal.
   */
  async contributeLocalSignal(
    agentId: bigint,
    localProbabilities: [number, number, number, number],
    privacyBudget: number = 2.0,
  ): Promise<void> {
    const noisedProbs = addGaussianNoise(localProbabilities, privacyBudget);
    const weight = await this.computeShapleyWeight(agentId);
    this.participants.set(agentId, {
      agentId,
      timestamp: Date.now() / 1000,
      probabilities: noisedProbs,
      weight,
      addedNoise: privacyBudget,
    });
  }

  /**
   * Compute weighted ensemble across all contributing agents.
   * Consensus triggers strategy adjustment only when 66%+ of weighted
   * agents agree on the same regime state.
   */
  async getWeightedEnsemble(): Promise<{
    consensus: string | null; // null if no consensus reached
    confidence: number;
    participantCount: number;
    regimeProbabilities: [number, number, number, number];
  }> {
    const REGIME_STATES = [
      "bull_low_vol",
      "bull_high_vol",
      "bear_low_vol",
      "bear_high_vol",
    ];

    // Geometric median (RFA, Pillutla et al., ICLR 2022) instead of weighted mean.
    // Tolerates up to 50% Byzantine agents; agnostic to corruption level.
    // The weighted mean is vulnerable: a single high-Shapley agent at 34%+ weight
    // can block consensus unilaterally.
    const regimeProbabilities = geometricMedian(
      Array.from(this.participants.values()).map((p) => p.probabilities),
    );
    // Shapley weights still used for contribution rewards; median used for consensus

    const maxIdx = regimeProbabilities.indexOf(
      Math.max(...regimeProbabilities),
    );
    const consensus =
      regimeProbabilities[maxIdx] >= this.CONSENSUS_THRESHOLD
        ? REGIME_STATES[maxIdx]
        : null;

    return {
      consensus,
      confidence: regimeProbabilities[maxIdx],
      participantCount: this.participants.size,
      regimeProbabilities: regimeProbabilities as [
        number,
        number,
        number,
        number,
      ],
    };
  }

  private readonly MAX_AGENT_WEIGHT_FRACTION = 0.15; // No single agent > 15% of total weight
  // Prevents a single high-Shapley agent from controlling consensus (Byzantine fault)

  /**
   * GTG-Shapley computation: cumulative Shapley values over all rounds.
   * Agents who consistently contribute accurate regime signals get higher weights.
   * Weight capped at MAX_AGENT_WEIGHT_FRACTION of total.
   */
  private async computeShapleyWeight(agentId: bigint): Promise<number>;
}
```

**Privacy guarantee**: Only 4 probability floats + Gaussian noise are shared — not gradients or model weights. Raw market data, position details, and strategy parameters never leave the contributing agent. Gradient inversion attacks (which can reconstruct training data from gradient updates) are entirely eliminated since no gradients are ever shared.

> **DP noise level**: ε = 2.0–4.0 (not ε ≤ 1.0). For standard Gaussian DP with ε=1.0, δ=10⁻⁵, and L2 sensitivity √2, noise standard deviation σ ≈ 6.86 per coordinate — which overwhelms \[0,1] probability values. With ε=2.0 and Balle & Wang (ICML 2018) analytical Gaussian calibration (reduces variance by 33%+), σ ≈ 3.0 — meaningful at 10+ agent aggregation. At ε=4.0, σ ≈ 1.5 — acceptable for ≥5 agents with simplex projection post-processing.

> **Future research note**: More complex federated training (e.g., MHESA — federated learning with random masking + CKKS homomorphic encryption) could be explored if the vault count reaches 20+ and regime probability sharing proves insufficient. This is not a planned v1 feature; it is a research direction contingent on demonstrated need.

***

## §10 — Anti-Gaming & Stake/Slash Mechanics

### Stake-and-Burn (Numerai Model)

Sellers stake USDC proportional to claimed quality. Underperformance *burns* stake — it does not redistribute. This is the critical design choice (D-089): burning creates genuine economic loss that cannot be gamed through collusion.

If stake were redistributed to buyers, a seller could collude with their own buyer accounts to "slash" the stake back to themselves with no net loss. Burning eliminates this attack vector entirely.

| Performance vs. Claimed      | Slash % | Trigger                                                        |
| ---------------------------- | ------- | -------------------------------------------------------------- |
| Within 20% of claimed Sharpe | 0%      | Normal variance — widened from 15% for fat-tailed DeFi returns |
| 20–30% below claimed Sharpe  | 25%     | Moderate underperformance                                      |
| 30–50% below claimed Sharpe  | 50%     | Significant underperformance                                   |
| > 50% below claimed Sharpe   | 100%    | Fraud or material misrepresentation                            |

**Minimum stake**: `stakeUsdc ≥ max(10, priceUsdc × subscribersCount² × 0.005)`. Quadratic scaling makes large-scale reputation gaming structurally expensive. (Pattern: DoraFactory Quadratic Funding V2, 2025.)

> **Why quadratic?** The linear formula is Sybil-exploitable: a strategy priced at $50 with 20 fake self-purchased subscribers requires only $100 stake (\~$70–150 total attack cost). If 2–3 genuine subscribers follow, the attack is profitable. Quadratic scaling (stake ∝ subscribers²) costs the same $100 at 20 subscribers, but 10× more at 100 subscribers ($2,500 vs. $250 linear). Large-scale reputation gaming becomes structurally expensive.

> **Statistical basis for 20% tolerance**: Lo (2002) SE(SR̂) ≈ √((1 + ½SR²)/n). For n=252 (1 year daily data), the 2-sigma band for SR=1.0 is ±15% — but DeFi returns exhibit significant negative skewness and positive kurtosis, widening confidence intervals by 30–50%+ (Bailey & López de Prado, PSR framework). The 20% threshold is calibrated to \~2-sigma for fat-tailed DeFi distributions.
>
> **Probabilistic Sharpe Ratio (PSR)**: For strategies with fewer than 252 trading periods, use PSR (Bailey & López de Prado, SSRN:1821643) instead of raw Sharpe for the slash trigger. PSR accounts for non-normality and track record length, providing a probability that the true Sharpe exceeds a threshold rather than a point estimate.

### Multi-Stage Dispute Escalation

\~98.5% of performance claims (per UMA's production data) resolve without dispute. The dispute path uses three escalating stages:

```
Stage 1 — Automated checks (instantaneous, on-chain):
  • On-chain metric validation: Did reported Sharpe fall below claimed value by > 15%?
  • Uses objective on-chain data only (vault NAV, EAS attested performance hash)
  • Auto-slash triggered if violation confirmed — no dispute needed
  • Covers ~80% of underperformance cases

Stage 2 — Optimistic challenge (24h window):
  • Disputer posts bond (5% of strategy price, burned regardless of outcome)
  • Dispute framed as objective quantitative claim:
    "Strategy X achieved Sharpe < Y over trailing Z days using oracle O"
  • Pre-specified metric + timeframe + data source required (vague claims rejected)
  • ~98% of remaining cases resolve here without escalation

Stage 3 — Expert jury (for contested disputes only):
  • 5-of-9 random selection from Verified+ ERC-8004 agents (score ≥ 100, active 90d)
  • Random jury, not token-weighted voting (Kleros-style selection)
  • 48-hour deliberation window
  • Juror bond: 10% of disputed strategy price (burned if juror doesn't vote)
  • Majority decision (5 of 9): slash per table above, or bond burned if claim invalid
```

> **Why not UMA token voting?** In March 2025, a whale with \~5M UMA tokens (\~25% of voting power) manipulated the resolution of a $7M Polymarket market about a Ukraine mineral agreement, forcing an incorrect "YES" resolution. Polymarket acknowledged the oracle reached an "incorrect outcome." Two addresses reportedly controlled over half the votes. Strategy quality disputes require randomized juries with skin-in-the-game participation, not token-plutocracy.

**Anti-collusion**: Dispute bond is burned (not awarded) regardless of outcome. This prevents coordinated dispute-to-slash farming where buyers and sellers collude to trigger disputes for mutual benefit.

### Sybil Resistance

The economic cost of staking makes identity farming expensive. A Sybil attacker trying to generate fake reputation via self-purchases must stake real capital, which is at risk of slashing if their own strategies underperform. This makes Sybil attacks economically irrational without genuine strategy performance.

ERC-8004 minimum tier for publishing (Verified+, score ≥ 50) adds an additional time-cost barrier — reputation cannot be farmed instantly.

**Capacity limits (Grossman-Stiglitz constraint)**: Selling genuinely profitable strategies to unlimited subscribers destroys the edge through crowding. Every published strategy carries a hard subscriber cap and AUM cap:

| Strategy Family    | Max Subscribers | Max Total Subscriber AUM |
| ------------------ | --------------- | ------------------------ |
| lp\_optimization   | 50              | $500K                    |
| momentum           | 30              | $300K                    |
| mean\_reversion    | 75              | $750K                    |
| yield\_compounding | 100             | $1M                      |
| arbitrage\_signal  | 5               | $50K                     |

When a strategy hits its cap, new subscriptions are blocked. Existing subscribers can trade subscription rights peer-to-peer via the `StrategyMarketplaceRegistry`. Automatic sunset triggers when 2 consecutive 30-day evaluation windows show Sharpe below 50% of claimed value.

> *Rationale: CFM's Stratus program is "closed to new investments" due to capacity constraints. A game-theoretic model (arXiv:2512.11913) found hyperbolic crowding decay fits DeFi momentum dynamics. Capacity limits prevent the Grossman-Stiglitz paradox from destroying alpha.*

***

## §11 — Alpha Decay Pricing

### Tiered Access Model (Recommended for MVP)

The market has converged on subscription bundling for information goods (Bakos & Brynjolfsson). Bloomberg and Refinitiv both use real-time premium / delayed standard / end-of-day free tiers. Subscription bundling empirically outperforms per-signal decay pricing for strategy marketplaces. The MVP implements this directly:

```
Tier 1 — Real-time (premium):
  x402 per-query or subscription. Strategy signals within 5 minutes.
  Price: $0.50–$5.00 per alpha signal, or $10–$50/month subscription.

Tier 2 — Delayed (standard):
  24-hour delay on all strategy updates.
  Price: $5–$20/month subscription.

Tier 3 — Historical (free/low cost):
  Strategy performance data older than 30 days, freely accessible.
  Used for reputation building, cold-start discovery, and buyer due diligence.
```

Tier 3 (historical, free) directly solves the cold-start problem for new strategies: buyers can evaluate track records before committing to a subscription. This is the same pattern used by Paradigm, Messari, and Dune Analytics.

The x402 server in §6 implements Tier 1 via per-query `paymentMiddleware`. Tier 2 is a subscription (monthly USDC transfer). Tier 3 is served via the public `/strategies` endpoint with no payment gate for data older than 30 days.

### Exponential Decay for Dutch Auction Premium Signals

Exponential decay governs **freshness pricing for the real-time tier** — not as a standalone pricing mechanism. Sellers who want to command real-time tier prices must keep strategies fresh. Stale strategies automatically fall to Tier 2 pricing after their half-life, then to Tier 3 after 30 days.

Maven Securities research confirms alpha decays at 9.9% per unit of delay in European markets, 5.6% in US markets. DeFi moves faster.

```typescript
// packages/vault/sdk/src/learning/pricing.ts

/**
 * Decay constants calibrated per strategy family.
 * Half-life = ln(2) / λ
 * Used to determine when a strategy drops from Tier 1 → Tier 2 pricing.
 */
export const DECAY_CONSTANTS = {
  lp_optimization: 0.05, // Half-life ~14 days (LP edge decays as others copy)
  momentum: 0.02, // Half-life ~35 days (trend signals persist longer)
  mean_reversion: 0.01, // Half-life ~69 days (statistical edges are stable)
  yield_compounding: 0.008, // Half-life ~87 days (compounding mechanics are slow)
  arbitrage_signal: 0.693, // Half-life ~1 day (pure arb: copy-killed in hours)
} as const;

/**
 * Current strategy price after exponential alpha decay.
 * P(t) = P_base × e^(-λt)
 *
 * Governs Tier 1 (real-time) freshness pricing.
 * Strategies that decay below 50% of base price should refresh or accept Tier 2 rate.
 * Price floor: P_base × 0.05 (5% of base — prevents strategies becoming free prematurely).
 */
export function currentStrategyPrice(
  basePriceUsdc: number,
  strategyFamily: keyof typeof DECAY_CONSTANTS,
  ageSeconds: number,
): number {
  const lambda = DECAY_CONSTANTS[strategyFamily];
  const ageDays = ageSeconds / 86400;
  const decayed = basePriceUsdc * Math.exp(-lambda * ageDays);
  const floor = basePriceUsdc * 0.05;
  return Math.max(decayed, floor);
}

/**
 * Seller price update recommendation.
 * Recommends refresh when current Tier 1 price < 50% of base price.
 */
export function shouldRefreshPrice(
  basePriceUsdc: number,
  strategyFamily: keyof typeof DECAY_CONSTANTS,
  ageSeconds: number,
): boolean {
  return (
    currentStrategyPrice(basePriceUsdc, strategyFamily, ageSeconds) <
    basePriceUsdc * 0.5
  );
}

/**
 * Optimal strategy refresh interval for maximum Tier 1 revenue.
 * Sellers should republish at half-life intervals with updated performance evidence.
 */
export function optimalRefreshIntervalDays(
  strategyFamily: keyof typeof DECAY_CONSTANTS,
): number {
  const lambda = DECAY_CONSTANTS[strategyFamily];
  return Math.LN2 / lambda; // Half-life in days
}
```

**Revenue model for sellers**: A seller who continuously refreshes their `lp_optimization` strategy every 14 days (half-life) with updated evidence maintains Tier 1 (real-time premium) pricing. A seller who publishes once earns Tier 1 revenue for \~14 days, then falls to Tier 2 (subscription), then Tier 3 (historical/free) after 30 days. Subscription bundling encourages sellers to keep strategies fresh.

***

## §12 — `strategy-marketplace` Agent

> See `prd/agents/08-agents-economy.md` for the full agent spec. Summary here for architectural context.

**Role**: Lists, sells, licenses, and monitors agent strategies via x402. Manages the full strategy lifecycle from episode monitoring to royalty collection.

**Key responsibilities**:

1. **Strategy detection**: Monitors insight pool. When the following conditions are met:

   * sampleSize ≥ **200** across **≥2 distinct market regimes**
   * Probabilistic Sharpe Ratio (PSR) > 0.90 (90% confidence true SR > 0)
   * Multiple testing correction: all candidate strategies survive BHY (Benjamini-Yekutieli) FDR control at q=0.05, accounting for all strategy variants evaluated
   * Confidence > **0.7 from ensemble of 3 independent distillation runs** (not single LLM 0.8 threshold — LLM confidence scores produce 30–50% false discovery rates at stated thresholds, per Xiong et al., ICLR 2024)

   ...then flags strategy as publishable.
2. **Strategy generation**: Runs parameter fitting over supporting episodes, computes SHAP feature importance, runs HMM classifier, generates EAS attestation.
3. **EAS model commitment**: Computes `{ modelCID, inputDataHash, claimedOutputHash }`, signs with seller's ERC-8004 key, and posts EAS attestation on-chain. For public models, uploads ONNX to IPFS (enables deterministic replay by anyone via `onnxruntime-node`). **ZK proof** is generated asynchronously only when a dispute is filed.
4. **Publication**: Calls `publish_strategy` MCP tool — uploads to IPFS, locks stake in `StrategyEscrow`, registers on `StrategyMarketplaceRegistry`.
5. **Marketplace server**: Runs the x402 Express server (§6), maintains `.well-known/x402.json` Bazaar manifest.
6. **Price maintenance**: Applies alpha decay monitoring. Recommends strategy refresh when current price < 50% of base.
7. **Royalty collection**: Calls `claim_strategy_royalties` periodically, routes proceeds to `treasury-manager`.
8. **Feedback loop**: When a purchased strategy underperforms, submits `vault_submit_yield_feedback` with negative `yieldBps` — this initiates the optimistic dispute pipeline and strengthens the seller's accountability record.

***

## §13 — Data Flow Summary

Complete end-to-end data flow from vault operation to strategy sale:

```
VaultOperation (on-chain event, e.g. vault_rebalance)
  │
  ├─→ Episode Capture (LanceDB, packages/vault/sdk/src/learning/episode.ts)
  │       ↳ preState + postState + gas metrics
  │
  ├─→ Reflexion Inner Loop (claude-opus-4-6, ~2-3 sentence reflection)
  │       ↳ prediction vs. actual gap analysis
  │       ↳ episode.confidence ← heuristic 0.0–1.0
  │
  ├─→ [Every 50 episodes] ExpeL Outer Loop
  │       ↳ Insight pool: ADD / UPVOTE / DOWNVOTE / EDIT
  │       ↳ SQLite (semantic memory, packages/vault/sdk/src/learning/)
  │
  ├─→ [PSR > 0.90 AND sampleSize ≥ 200 across ≥2 regimes AND ensemble confidence > 0.7] Strategy Generation
  │       ↳ Parameter fitting over supporting episodes
  │       ↳ SHAP feature importance
  │       ↳ HMM regime classifier (4-state, Koki et al.)
  │       ↳ EAS model commitment (modelCID + inputDataHash + claimedOutputHash, signed by agent key)
  │       ↳ [public model] ONNX uploaded to IPFS — deterministic replay enabled
  │       ↳ [AUM-tiered opML] 24h / 72h / 7d challenge window by vault AUM
  │       ↳ [dispute filed] async ZK proof (Bionetta/DeepProve, 2-20min)
  │
  ├─→ StrategyMarketplaceRegistry (on-chain, publishStrategy())
  │       ↳ strategyId = keccak256(agentId + version + timestamp)
  │       ↳ metadataURI → IPFS (full strategy JSON)
  │       ↳ StrategyEscrow ← seller stake locked
  │
  ├─→ x402-gated Express server (/.well-known/x402.json for Bazaar)
  │       ↳ GET /strategies — $0.10 USDC per query
  │       ↳ GET /signals/regime — $0.02 USDC
  │       ↳ GET /signals/alpha — $1.00 USDC
  │
  ├─→ list_strategies / purchase_strategy MCP tools (buyer side)
  │       ↳ alpha decay applied to prices
  │       ↳ tiered ZK verification on purchase
  │
  └─→ Royalty flow
        ↳ claim_strategy_royalties → treasury-manager
        ↳ proceeds → compound into vault stakes
```

***

## §14 — Verification Checklist

After implementation, validate end-to-end via the following integration tests:

1. **Episode capture**: `vault_rebalance` call → episode stored in LanceDB → reflection generated via claude-opus-4-6.
2. **Insight distillation**: After **configurable window (default: 10 new episodes, tunable down to 3 or up to 50)** → `distillInsights()` runs → new insights appear in SQLite with correct categories and confidence scores. Failure episodes (no paired success) can generate DOWNVOTE/EDIT operations without a success counterpart.
3. **Strategy publication**: `publish_strategy` tool → strategy registered on `StrategyMarketplaceRegistry` → IPFS metadata uploaded → IPFS URI accessible.
4. **Bazaar manifest**: `GET /.well-known/x402.json` on marketplace server returns valid JSON with correct endpoints and prices.
5. **Strategy search**: `list_strategies` tool → returns strategies with decay-adjusted prices (verify price < base for old strategies).
6. **Strategy purchase**: `purchase_strategy` tool → x402 payment flow completes → strategy JSON returned to buyer → `recordPurchase()` emitted on-chain.
7. **EAS model commitment**: `publish_strategy` call → `computeModelCommitment()` generates `{ modelHash, inputDataHash, claimedOutputHash }` → EAS attestation hash returned and stored on-chain. For public models: ONNX uploaded to IPFS, and `verifyModelCommitment()` run locally via `onnxruntime-node` returns `valid: true`.
8. **Dispute ZK proof**: When a dispute is filed during the opML window, `proveStrategyPerformance()` generates a valid async proof → `verifyPerformanceProof()` returns `true` on Base Sepolia fork. Not required at publication time.
9. **Regime gossip**: 3+ vaults contribute regime probability vectors → `RegimeGossipCollector.getWeightedEnsemble()` reaches consensus → `vault_rebalance` uses consensus regime in decision.
10. **Alpha decay**: Strategy price at day 0 = `baseUsdc`. Strategy price at 14 days for `lp_optimization` = `baseUsdc × e^(-0.05 × 14) ≈ 0.497 × baseUsdc`.
11. **Slash pipeline**: Buyer submits negative yield feedback → `vault_submit_yield_feedback` initiates Stage 1 automated check → if unresolved, 24-hour Stage 2 optimistic challenge window starts → if escalated, Stage 3 expert jury convenes → `StrategyEscrow.slash()` callable on resolution.

***

## §15 — Cold-Start Bootstrap Mechanism

### Problem

ERC-8004 Verified+ tier (score ≥ 50) is required to publish strategies on the marketplace. New agents start at score 0. Without a bootstrap path, the marketplace has no supply — only established agents can publish, which means new agents can never establish themselves. Ocean Protocol failed partly for this reason.

### Solution: Sandbox → Gradual Access Escalation

| Stage                  | Requirement                                                            | Capability                                                                        | Reputation   |
| ---------------------- | ---------------------------------------------------------------------- | --------------------------------------------------------------------------------- | ------------ |
| **Sandbox** (Week 1)   | Registration bond $100 USDC (refundable after 30 days with no slashes) | Observe strategies, limited vault deposits ($1K cap), free Tier 3 strategy access | Score: 0     |
| **Basic** (Month 1)    | First profitable exit OR 30-day continuous hold without slash          | List strategies for free (Tier 3 historical only), $10K deposit cap               | Score: 10–49 |
| **Verified** (Month 3) | 2+ profitable exits AND cumulative $500+ P\&L                          | Full marketplace publishing, Tier 1 real-time signals, $50K cap                   | Score: 50+   |
| **Trusted+**           | Peer vouching by 2 Sovereign agents (score ≥ 500)                      | Unlimited publishing, priority placement in strategy search                       | Score: 100+  |

Progress through stages is automatic: the `VaultReputationEngine` (see `prd/vault/01-overview.md`) auto-attests milestones on the ERC-8004 Reputation Registry. Every profitable exit, continuous hold, and vault creation is attested on-chain without manual intervention.

### Sybil Defense

**MeritRank transitivity decay** (Nasrulin et al., arXiv:2207.09950): Reputation value is discounted by transitivity decay coefficient α (recommended small, \~0.05) with each hop away from the Gotts Registry seed node. The paper does NOT specify a 50% cluster penalty — this is tunable per deployment. Isolated clusters receive discounted (not halved) reputation based on graph distance. This prevents coordinated fake-review rings from bootstrapping each other.

**Activity requirement**: Agents that have been inactive for >180 days require a re-attestation call to maintain Verified+ tier. This is an explicit action (not automatic decay) that acknowledges continued operation — different from time-based score reduction, which MeritRank (2022) found does not improve Sybil tolerance and degrades honest participants.

**Whitewashing defense**: The $100 registration bond is non-transferable and tracked per ERC-8004 identity. Abandoning a low-reputation identity and creating a new one requires a new $100 bond and 30-day restart period. This makes whitewashing economically costly.

### Agent-Native Identity Bootstrapping

The previous design referenced Gitcoin Passport (Human Passport) as an alternative to the $100 registration bond. This is architecturally incompatible: Human Passport was acquired by Holonym Foundation in February 2025 and explicitly rebranded to human.tech — its stamps require human-owned social media accounts (Twitter 100+ followers), biometric verification, and government ID. Its ML models are trained to detect human behavior vs. bot patterns. An AI agent would likely be classified as Sybil/bot.

**Agent-native alternatives for bond reduction:**

* **ERC-8004 Validation Registry verification**: A third-party verifier (e.g., the Gotts Registry) can attest that an agent has a valid deployment manifest, cryptographic identity, and operator ERC-4337 wallet. This is the intended use of the Validation Registry hooks in ERC-8004. Agents with Validation Registry attestation start at Basic tier without a $100 bond.
* **Operator co-signing**: Human operators (Trusted+) can co-sign a new agent's identity, vouching for its legitimacy. The co-signer's reputation is partially at risk if the new agent gets slashed.
* **Bond escalation**: The $100 bond escalates with tier: Sandbox $100 USDC, Basic $0 (after first profitable exit), Verified $0 (auto-attested by VaultReputationEngine).

```typescript
// packages/vault/sdk/src/identity/bootstrap.ts

export type OnboardingTier = "sandbox" | "basic" | "verified" | "trusted";

export interface OnboardingStatus {
  tier: OnboardingTier;
  score: number;
  bondPaidUsdc: number; // 0 after first profitable exit
  validationRegistryAttested: boolean; // ERC-8004 Validation Registry attestation
  registrationTimestamp: number;
  nextTierRequirements: string[]; // Human-readable list of what's needed to advance
  meritRankMultiplier: number; // Transitivity-discounted, not flat 50% penalty
}

/**
 * Compute the current onboarding status for an agent.
 * Called by vault_register_agent and vault_get_agent_reputation tools.
 */
export async function getOnboardingStatus(
  agentId: bigint,
  reputationRegistry: Address,
  publicClient: PublicClient,
): Promise<OnboardingStatus>;

/**
 * Determine MeritRank multiplier for an agent.
 * Returns transitivity-discounted value based on graph distance from Gotts Registry seed.
 * Discount coefficient α ≈ 0.05 per hop (tunable, not a fixed 50% penalty).
 * Reference: Nasrulin et al., arXiv:2207.09950.
 */
export async function computeMeritRankMultiplier(
  agentId: bigint,
  graphSnapshot: ReputationGraphSnapshot,
): Promise<number>;
```

***

## References

Key citations for this specification. Full bibliography in `research-citations.md`.

* Shinn et al. — Reflexion (NeurIPS 2023): verbal self-reflection inner loop
* Zhao et al. — ExpeL (AAAI 2024): cross-episode insight distillation
* Xu & Brini (AAAI 2025): PPO for Uniswap V3 LP — validates RL-based strategy learning
* Koki et al. — 4-state HMM: 60%+ directional accuracy on ETH
* Maven Securities alpha decay research: pricing calibration basis
* Bakos & Brynjolfsson: subscription bundling outperforms per-signal decay pricing for information goods
* Commit-Reveal² (arXiv:2504.03936): 80% gas savings on dispute resolution
* ORA Protocol FPVM: opML reference implementation (optimistic ML challenge pattern)
* onnxruntime-node: TypeScript-native ONNX inference — used for deterministic model replay in `verifyModelCommitment()`
* Bionetta (Groth16): 320-byte proofs, \~200K gas, 2.3-second mobile proving for sub-100K param models
* Lagrange DeepProve: 54–158× faster than EZKL for async ZK dispute resolution
* Kleros random jury: production-validated randomized dispute resolution with skin-in-the-game
* Numerai mechanism: stake-and-burn anti-gaming design (D-089)
* EZKL / Trail of Bits audit: TypeScript-native ZKML library (deferred to dispute-only path)
* Federated distillation / PATE: theoretical basis for probability-vector sharing
* ARCx Credit: gradual access escalation pattern (note: live but modest adoption — not a strong validation reference)
* MeritRank (Nasrulin et al., arXiv:2207.09950): transitivity decay with tunable α coefficient; epoch decay explicitly found NOT to improve Sybil tolerance
* FadeMem (Wei et al., arXiv:2601.18642, Jan 2026): selective forgetting — important memories decay 3–5× slower, not immortal. 82.1% retention at 55% storage on LTI-Bench. Core thesis: adaptive decay prevents information overload.
* FinMem (Yu et al., arXiv:2311.13743, AAAI 2024): 3-layer memory with α\_shallow=0.9, α\_deep=0.988 decay rates for financial agents
* STONE paradigm (Luo et al., arXiv:2602.16192, 2026): separate paper from Mem0; on-demand extraction rather than periodic batching. Note: Mem0 (arXiv:2504.19413) uses extraction-then-update pipeline with Add/Update/Delete/Merge operations — different architecture.
* UMEM (Ye et al., arXiv:2602.10652, Feb 2026): fixed ExpeL operations accumulate noise; jointly-learned Mem-Optimizer prevents degradation — 10.67% improvement over baselines
* SaMuLe (arXiv:2509.20562, EMNLP 2025): ExpeL collapses at 0% on high-complexity tasks; multi-level reflection including failure-only learning
* MUSE (arXiv:2510.08002, Oct 2025): typed procedural/strategic/tool memory achieves 51.78% SOTA on TAC (+20% relative)
* Harvey, Liu & Zhu (2016, RFS Vol. 29): t > 3.0 threshold for multiple-tested factors; BHY FDR correction
* Bailey & López de Prado — PSR (SSRN:1821643): Probabilistic Sharpe Ratio for non-normal returns
* Bailey et al. — PBO (SSRN:2326253): Probability of Backtest Overfitting via CSCV
* Pillutla et al. — RFA (ICLR 2022): geometric median for robust aggregation; 50% Byzantine tolerance
* Battering RAM attack (Van Bulck et al., IEEE S\&P 2026): <$50 DDR4/DDR5 memory interposer breaks Intel TDX, AMD SEV-SNP, AND NVIDIA Confidential Computing — forges attestation quotes with "UpToDate" trust designation; all major cloud TEE vendors acknowledged. See [10-safety.md](/docs/gotts-vaults/vault/10-safety.md) D-036. Earlier 2025 version only affected SGX/SEV-SNP; 2026 escalation extends to TDX.
* ReflAct (Kim et al., arXiv:2505.15182, May 2025): 27.7% over ReAct via structured state-goal reflection with quantitative anchors
* Xiong et al. (ICLR 2024): LLM confidence ECE > 0.377; explicit confidence scoring AUROC = 0.58; 30–50% false discovery rates at stated thresholds
* arXiv:2512.11913: hyperbolic crowding decay fits momentum strategies (R²=0.65) — basis for capacity limits
* Trading-R1 (arXiv:2509.11420): free-form intermediate reasoning compounds errors — basis for structured reflection template
* TradingAgents (Xiao et al., 2024): GPT-4o-mini for summarization, o1 for reasoning-intensive decisions — basis for two-tier LLM cascade
* Balle & Wang (ICML 2018): analytical Gaussian mechanism calibration — reduces DP noise variance by 33%+ vs. standard calibration
