> For the complete documentation index, see [llms.txt](https://gotts.gitbook.io/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://gotts.gitbook.io/docs/devenv/devenv/08-testing.md).

# Testing

> **Part of**: [Devenv PRD](/docs/devenv/devenv.md) | **Last Updated**: 2026-02-21

***

## Overview

This document is the comprehensive testing guide for the Gotts monorepo. It covers Solidity testing with Forge, TypeScript testing with Vitest, test file organization conventions, CI pipeline configuration, gas benchmarking, and coverage targets per implementation phase.

For vault-specific invariant properties and fuzzing parameter ranges, see [prd/vault/17-testing.md](/docs/gotts-vaults/vault/17-testing.md). This document covers the infrastructure and tooling layer.

***

## 1. Three-Tier Architecture

Tests are organized into three tiers based on chain dependency and execution speed.

| Tier | Name        | Chain              | Speed      | When to Run             |
| ---- | ----------- | ------------------ | ---------- | ----------------------- |
| 1    | Unit        | None (mocks only)  | < 5s total | On every file save      |
| 2    | Integration | Anvil (devenv)     | 30s–5min   | Pre-commit, PR CI       |
| 3    | E2E Swarm   | Anvil (full swarm) | 10–30 min  | Nightly CI, pre-release |

**Tier 1 (Unit)**: No blockchain calls. Mock viem clients, stub ERC-8004 registries, test pure SDK logic, validate tool parameter schemas, and verify error code mapping. Runs everywhere — no environment setup required.

**Tier 2 (Integration)**: Spins up a real Anvil instance with the full Gotts stack deployed. Uses `useGottsDeployment()` or `useDeployedUniswap()` fixtures to get a clean chain state before each test. Covers protocol interactions, MCP tool end-to-end behavior, and contract call sequences.

**Tier 3 (E2E Swarm)**: Full multi-agent simulation using `SwarmRunner` with 5 agent profiles running concurrent strategies. Not run in standard PR CI — too slow and flaky for that context. Run manually with `pnpm testnet:swarm` and in nightly CI.

***

## 2. Solidity Testing (Forge)

All Solidity tests live in `contracts/test/` within each package. Forge is configured via `foundry.toml` in the package root.

### 2.1 `foundry.toml` Configuration Reference

```toml
[profile.default]
src = "contracts/src"
out = "contracts/out"
libs = ["lib"]
optimizer = true
optimizer_runs = 200
via_ir = false           # Set true only for Permit2 (avoids 10-min compile)

[profile.default.fuzz]
runs = 1000              # Per test function in local runs
seed = "0x1"             # Deterministic; override with FOUNDRY_FUZZ_SEED

[profile.default.invariant]
runs = 500               # Campaign runs per invariant contract
depth = 256              # Max calls per run
call_override = false
fail_on_revert = false   # Allow reverts; only fail on invariant violations

[profile.ci]
fuzz = { runs = 10000 }
invariant = { runs = 1000, depth = 512 }

[profile.fork]
fork_block_number = 28_000_000  # Override with FOUNDRY_FORK_BLOCK_NUMBER
```

### 2.2 Test Naming Conventions (Normative)

| Prefix        | Usage                             | Example                                             |
| ------------- | --------------------------------- | --------------------------------------------------- |
| `test_`       | Unit test, expected to pass       | `test_depositMintsShares()`                         |
| `testRevert_` | Unit test, expected to revert     | `testRevert_depositBelowMinimum()`                  |
| `testFuzz_`   | Fuzz test                         | `testFuzz_depositWithdrawRoundtrip(uint256 amount)` |
| `testFork_`   | Fork test (requires `--fork-url`) | `testFork_depositWithRealERC8004()`                 |
| `invariant_`  | Invariant test function           | `invariant_sharePriceNeverDecreases()`              |

### 2.3 Unit Tests (No Fork)

```bash
# Run all unit tests (fast, no RPC)
forge test

# Run tests for a specific contract
forge test --match-contract AgentVaultCore -vv

# Run a specific test
forge test --match-test test_depositMintsShares -vvvv

# Run and show gas usage
forge test --gas-report
```

Standard test setup pattern:

```solidity
// contracts/test/AgentVaultCore.t.sol
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.26;

import { Test } from "forge-std/Test.sol";
import { AgentVaultCore } from "../src/AgentVaultCore.sol";
import { MockERC8004IdentityRegistry } from "./mocks/MockERC8004IdentityRegistry.sol";
import { MockERC4626Adapter } from "./mocks/MockERC4626Adapter.sol";
import { MockERC20 } from "./mocks/MockERC20.sol";

contract AgentVaultCoreTest is Test {
    AgentVaultCore vault;
    MockERC8004IdentityRegistry identityRegistry;
    MockERC20 usdc;
    address alice = makeAddr("alice");
    uint256 aliceAgentId;

    function setUp() public {
        usdc = new MockERC20("USD Coin", "USDC", 6);
        identityRegistry = new MockERC8004IdentityRegistry();

        vault = new AgentVaultCore({
            asset: address(usdc),
            name: "Test Vault",
            symbol: "TV",
            identityAdapter: address(identityRegistry),
            managementFeeBps: 100,
            performanceFeeBps: 1000,
        });

        // Register alice as a valid agent
        aliceAgentId = identityRegistry.register(alice, "ipfs://alice");
        usdc.mint(alice, 1_000_000e6);
        vm.prank(alice);
        usdc.approve(address(vault), type(uint256).max);
    }

    function test_depositMintsShares() public {
        vm.prank(alice);
        uint256 shares = vault.deposit(100e6, alice);
        assertGt(shares, 0);
        assertEq(vault.balanceOf(alice), shares);
    }
}
```

### 2.4 Fuzz Testing

```bash
# Run fuzz tests
forge test --match-test testFuzz_

# Run with more iterations (CI profile)
FOUNDRY_PROFILE=ci forge test --match-test testFuzz_
```

```solidity
function testFuzz_depositWithdrawRoundtrip(uint256 depositAmount) public {
    // Constrain to realistic range (1 USDC to 1M USDC)
    depositAmount = bound(depositAmount, 1e6, 1_000_000e6);

    vm.prank(alice);
    uint256 shares = vault.deposit(depositAmount, alice);

    vm.prank(alice);
    uint256 withdrawn = vault.redeem(shares, alice, alice);

    // Withdrawn amount is within 0.01% of deposit (rounding + fees)
    assertApproxEqRel(withdrawn, depositAmount, 0.0001e18);
}

function testFuzz_reputationTierCapEnforced(
    uint256 agentId,
    uint256 depositAmount,
    uint256 reputationScore
) public {
    agentId = bound(agentId, 1, 100_000);
    depositAmount = bound(depositAmount, 1, type(uint128).max);
    reputationScore = bound(reputationScore, 0, 999);

    vm.assume(identityRegistry.isRegistered(agentId));

    uint256 tierCap = vault.getTierDepositCap(reputationScore);
    if (depositAmount > tierCap) {
        vm.expectRevert(AgentVaultCore.TierLimitExceeded.selector);
    }

    vault.deposit(depositAmount, address(uint160(agentId)));
}
```

**Coverage-guided fuzzing (Foundry v1.3+):** Add `corpus_dir` to enable corpus persistence across runs. Foundry saves interesting inputs and reuses them, dramatically improving coverage over random inputs alone. Commit the corpus to git for CI reproducibility:

```toml
[fuzz]
runs = 1000
seed = "0x1"
corpus_dir = "fuzz-corpus"   # Foundry v1.3+: persist coverage corpus across runs
```

**`bound()` vs `vm.assume()`:** Prefer `bound()` over `vm.assume()` for constraining inputs. `vm.assume()` discards runs that don't satisfy the predicate, wasting fuzzer budget. `bound()` maps any input into the valid range without wasting a run. Only use `vm.assume()` for preconditions that cannot be satisfied by mapping (e.g., "this address must already be registered").

### 2.5 Invariant Testing

```bash
# Run invariant tests
forge test --match-test invariant_

# Run with CI parameters
FOUNDRY_PROFILE=ci forge test --match-test invariant_ -vvv
```

Invariant tests use handler contracts to bound the action space and track ghost variables that mirror expected accounting. See [prd/vault/17-testing.md §1](/docs/gotts-vaults/vault/17-testing.md) for the formal invariant statements (INV-1 through INV-14).

**Ghost variable handler pattern** — the `VaultHandler` maintains shadow accounting (`ghost_*` vars) that the invariant test uses to verify conservation of value:

```solidity
// contracts/test/invariants/handlers/VaultHandler.sol
pragma solidity ^0.8.26;

import { Test } from "forge-std/Test.sol";
import { AgentVaultCore } from "../../src/AgentVaultCore.sol";

contract VaultHandler is Test {
    AgentVaultCore public vault;
    uint256 public ghost_depositSum;
    uint256 public ghost_withdrawSum;
    address[] public actors;

    constructor(AgentVaultCore _vault, address[] memory _actors) {
        vault = _vault;
        actors = _actors;
    }

    modifier useActor(uint256 seed) {
        address actor = actors[bound(seed, 0, actors.length - 1)];
        vm.startPrank(actor);
        _;
        vm.stopPrank();
    }

    function deposit(uint256 amount, uint256 actorSeed) external useActor(actorSeed) {
        amount = bound(amount, 1e6, 10_000e6);
        address actor = actors[bound(actorSeed, 0, actors.length - 1)];
        deal(address(vault.asset()), actor, amount);
        vault.deposit(amount, actor);
        ghost_depositSum += amount;
    }

    function withdraw(uint256 amount, uint256 actorSeed) external useActor(actorSeed) {
        address actor = actors[bound(actorSeed, 0, actors.length - 1)];
        uint256 balance = vault.balanceOf(actor);
        if (balance == 0) return;
        amount = bound(amount, 1, balance);
        vault.withdraw(amount, actor, actor);
        ghost_withdrawSum += amount;
    }

    function advanceTime(uint256 seconds_) external {
        seconds_ = bound(seconds_, 1, 30 days);
        vm.warp(block.timestamp + seconds_);
    }

    // Expose for invariant checks
    function currentSharePrice() external view returns (uint256) {
        uint256 supply = vault.totalSupply();
        return supply == 0 ? 1e18 : (vault.totalAssets() * 1e18) / supply;
    }

    uint256 public lastSharePrice;
    uint256 public constant MAX_FEE_BPS = 500;

    function updateSharePriceSnapshot() external {
        lastSharePrice = this.currentSharePrice();
    }
}
```

Invariant test targeting only the handler:

```solidity
// contracts/test/invariants/SharePriceInvariants.t.sol
pragma solidity ^0.8.26;

import { StdInvariant } from "forge-std/StdInvariant.sol";
import { Test } from "forge-std/Test.sol";
import { VaultHandler } from "./handlers/VaultHandler.sol";

contract SharePriceInvariants is StdInvariant, Test {
    VaultHandler handler;

    function setUp() public {
        address[] memory actors = new address[](5);
        for (uint256 i; i < 5; i++) actors[i] = makeAddr(string(abi.encodePacked("actor", i)));

        AgentVaultCore vault = new AgentVaultCore(/* ... */);
        handler = new VaultHandler(vault, actors);

        targetContract(address(handler));
        bytes4[] memory selectors = new bytes4[](3);
        selectors[0] = handler.deposit.selector;
        selectors[1] = handler.withdraw.selector;
        selectors[2] = handler.advanceTime.selector;
        targetSelector(FuzzSelector({ addr: address(handler), selectors: selectors }));
    }

    /// @dev INV-1: Share price may only increase (modulo fee impact)
    function invariant_sharePriceMonotonic() public view {
        uint256 currentPrice = handler.currentSharePrice();
        uint256 lastPrice = handler.lastSharePrice();
        uint256 maxDecline = lastPrice * handler.MAX_FEE_BPS() / 10_000;
        assertGe(currentPrice + maxDecline, lastPrice);
    }

    /// @dev INV-2: Conservation of value — vault holds exactly what was deposited minus withdrawals
    function invariant_conservationOfValue() public view {
        assertEq(
            handler.vault().totalAssets(),
            handler.ghost_depositSum() - handler.ghost_withdrawSum()
        );
    }

    /// @dev INV-3: Dust deposit mints at least 1 share (inflation attack protection)
    function invariant_noInflationAttack() public {
        vm.prank(makeAddr("attacker"));
        uint256 shares = handler.vault().deposit(1, makeAddr("attacker"));
        assertGt(shares, 0);
    }
}
```

> **Critical caveat**: each `invariant_*` function gets a separate EVM executor. To check multiple properties against the same state snapshot, group them in a shared `assertInvariants()` helper called from each invariant function.

**DeFi-specific invariant categories for Gotts vaults:**

| Category       | Invariant                                                                       |
| -------------- | ------------------------------------------------------------------------------- |
| ERC-4626       | `totalAssets() >= sum(all user claims)`                                         |
| ERC-4626       | `convertToAssets(convertToShares(x)) <= x` (rounding never favors attacker)     |
| ERC-4626       | Zero total supply iff zero total assets                                         |
| Vault solvency | `vault.balance >= totalDeposits()`                                              |
| Fee cap        | `accruedFees / totalAssets <= MAX_FEE_BPS / 10_000`                             |
| Share price    | Share price monotonically non-decreasing (modulo max management fee per period) |
| Access         | Only registered ERC-8004 agents can hold positive balances                      |

**Beyond Foundry: Medusa v1 and Chimera/Recon**

For deep pre-audit campaigns, two tools complement Foundry's native fuzzer:

* **Medusa v1** (Trail of Bits, Feb 2025): parallel coverage-guided fuzzer written in Go on Geth. Multi-core, scales with hardware in ways Foundry's single-threaded fuzzer cannot. Benchmarks show 5–10x more coverage per hour than Foundry's fuzzer on complex DeFi protocols.
* **Chimera/Recon**: write invariants once, run with Foundry, Echidna, Medusa, and Halmos simultaneously. Prevents tool lock-in and ensures invariants are tool-agnostic.

```bash
# Bootstrap a Chimera-compatible project layout
forge init --template https://github.com/Recon-Fuzz/create-chimera-app
```

* **Halmos 0.3** (a16z): symbolic execution for formal verification; reuses existing Foundry fuzz tests. v0.3.0 added stateful invariant testing and a 32x faster EVM interpreter. Use for bounded formal verification of critical properties (e.g., `convertToShares` rounding always floors, never ceilings).

### 2.6 ERC-4626 Property Test Suite

Import and extend the a16z ERC-4626 property tests for all vault variants:

```solidity
// contracts/test/ERC4626Properties.t.sol
import { ERC4626Test } from "a16z-contracts/src/ERC4626Test.sol";

contract AgentVaultCoreERC4626Test is ERC4626Test {
    function setUp() public override {
        _underlying_ = new MockERC20("USDC", "USDC", 6);
        _vault_ = address(new AgentVaultCore({
            asset: address(_underlying_),
            // minimal config for property tests
        }));
        _delta_ = 0;
        _vaultMayBeEmpty_ = true;
        _unlimitedAmount_ = false;
    }
}
```

All a16z property tests must pass with zero delta for all AgentVaultCore configurations.

### 2.6a Property Testing Tools Comparison

| Tool         | Type                     | Speed    | Parallelism     | Best for                      |
| ------------ | ------------------------ | -------- | --------------- | ----------------------------- |
| Foundry fuzz | Stateless + stateful     | Fast     | Single-threaded | Day-to-day testing, CI        |
| Medusa v1    | Coverage-guided          | Fast     | **Multi-core**  | Pre-audit deep campaigns      |
| Echidna      | Grammar-based            | Moderate | Sequential      | Legacy / some unique features |
| Halmos 0.3   | Symbolic execution       | Slow     | N/A             | Bounded formal verification   |
| Certora      | Full formal verification | Slow     | N/A             | Critical protocol invariants  |

Recommended workflow: Foundry fuzz for daily development, Medusa for pre-audit sprints (1–2 weeks before external audit), Halmos for formal verification of ERC-4626 rounding properties, Certora for PolicyCage constraints and fee caps.

### 2.7 Fork Tests

```bash
# Run fork tests (requires BASE_RPC_URL)
forge test --match-test testFork_ --fork-url $BASE_RPC_URL

# Pin a specific block
forge test --match-test testFork_ --fork-url $BASE_RPC_URL --fork-block-number 28000000
```

```solidity
// contracts/test/fork/VaultCreation.fork.t.sol
pragma solidity ^0.8.26;

import { Test } from "forge-std/Test.sol";

contract VaultCreationForkTest is Test {
    // Real ERC-8004 Identity Registry on Base
    address constant IDENTITY_REGISTRY = 0x8004A818BFB912233c491871b3d84c89A494BD9e;

    function testFork_createVaultWithRealERC8004() public {
        // Fork test: uses real ERC-8004 registry at pinned block
        vm.createSelectFork(vm.envString("BASE_RPC_URL"), 28_000_000);

        // ...test body using real deployed contracts
    }
}
```

### 2.8 Gas Snapshots

```bash
# Take a gas snapshot
forge snapshot

# Check for regressions vs committed snapshot
forge snapshot --check       # Fails if any function exceeds committed values by >5%

# Generate detailed gas report
forge test --gas-report
```

`.gas-snapshot` is committed to the repository. Any PR that increases gas for a tracked function by more than 5% will fail the CI `gas-snapshot` job.

**Targeted gas measurement** excludes test setup overhead from gas readings using Foundry cheatcodes:

```solidity
// contracts/test/gas/VaultGas.t.sol

function test_gas_deposit() public {
    // Exclude expensive setup from the gas measurement
    vm.pauseGasMetering();
    usdc.mint(alice, 1_000_000e6);
    vm.prank(alice);
    usdc.approve(address(vault), type(uint256).max);
    vm.resumeGasMetering();

    // Named gas snapshot — measures only the deposit call
    vm.startSnapshotGas("deposit");
    vm.prank(alice);
    vault.deposit(100_000e6, alice);
    uint256 gasUsed = vm.stopSnapshotGas("deposit");

    // Assert against budget from §7 targets
    assertLt(gasUsed, 150_000, "deposit gas regression");
}

function test_gas_createVault() public {
    vm.pauseGasMetering();
    // ... factory setup ...
    vm.resumeGasMetering();

    vm.startSnapshotGas("createVault");
    agentVaultFactory.createVault(config);
    uint256 gasUsed = vm.stopSnapshotGas("createVault");

    assertLt(gasUsed, 500_000, "createVault gas regression");
}
```

Use `vm.pauseGasMetering()` / `vm.resumeGasMetering()` in all gas-sensitive tests to exclude setup overhead. Use `vm.startSnapshotGas("name")` / `vm.stopSnapshotGas("name")` for named measurements within a single test function.

### 2.9 Coverage

```bash
# Generate coverage report
forge coverage --report lcov

# Open HTML report
genhtml lcov.info --output-dir coverage/
open coverage/index.html
```

Coverage is measured per-contract and tracked against the phase targets in [§8](#8-coverage-targets-by-phase).

***

## 3. TypeScript Testing (Vitest)

### 3.1 `vitest.config.ts` Layout

Each package has its own `vitest.config.ts`. The structure is consistent across packages:

```typescript
// packages/vault/vitest.config.ts
import { defineConfig } from "vitest/config";
import { createAnvilPool } from "@gotts.ai/testnet/fixtures-parallel";
import { deployFullStack } from "./src/devenv/setup.js";

const anvilPool = createAnvilPool({
  preset: "fresh",
  poolSize: 4,
  setup: async (instance) => deployFullStack(instance.client, "from-scratch"),
  setupTimeout: 180_000,
});

export default defineConfig({
  test: {
    globalSetup: anvilPool.globalSetup,
    pool: "forks",
    poolOptions: {
      forks: { singleFork: false },
    },
    testTimeout: 60_000,
    hookTimeout: 180_000,
    reporters: ["verbose"],
    coverage: {
      provider: "v8",
      reporter: ["text", "lcov"],
      include: ["src/**/*.ts", "sdk/src/**/*.ts"],
      exclude: ["src/devenv/**", "**/__tests__/**"],
    },
    env: {
      DEVENV_MODE: "from-scratch",
    },
  },
});
```

### 3.2 Unit Tests (No Anvil)

Unit tests mock all blockchain interactions using `vi.mock`. They run without Anvil and complete in milliseconds.

**Mocking viem clients:**

```typescript
// packages/vault/sdk/src/__tests__/VaultClient.test.ts
import { vi, describe, it, expect, beforeEach } from "vitest";
import type { PublicClient } from "viem";
import { VaultClient } from "../VaultClient.js";

const mockPublicClient = {
  readContract: vi.fn(),
  simulateContract: vi.fn(),
  waitForTransactionReceipt: vi.fn(),
} satisfies Partial<PublicClient>;

describe("VaultClient", () => {
  beforeEach(() => {
    vi.clearAllMocks();
  });

  it("returns vault state from readContract", async () => {
    mockPublicClient.readContract.mockResolvedValueOnce({
      totalAssets: 100_000n * 10n ** 6n,
      totalSupply: 99_500n * 10n ** 18n,
      isPaused: false,
    });

    const client = new VaultClient({
      vaultAddress: "0x1234...5678",
      publicClient: mockPublicClient as unknown as PublicClient,
    });

    const state = await client.getState();
    expect(state.totalAssets).toBe(100_000n * 10n ** 6n);
    expect(state.isPaused).toBe(false);
  });
});
```

**Testing MCP tool parameter validation:**

```typescript
// packages/vault/src/__tests__/tools/vault-deposit.test.ts
import { describe, it, expect } from "vitest";
import { vaultDepositSchema } from "../../tools/vault-deposit.js";

describe("vault_deposit tool schema", () => {
  it("rejects negative amounts", () => {
    const result = vaultDepositSchema.safeParse({
      amount: "-100",
      vaultAddress: "0x...",
    });
    expect(result.success).toBe(false);
  });

  it("rejects invalid address", () => {
    const result = vaultDepositSchema.safeParse({
      amount: "100",
      vaultAddress: "not-an-address",
    });
    expect(result.success).toBe(false);
  });

  it("accepts valid inputs", () => {
    const result = vaultDepositSchema.safeParse({
      amount: "1000",
      vaultAddress: "0x1234567890123456789012345678901234567890",
      chainId: "8453",
    });
    expect(result.success).toBe(true);
  });
});
```

**Mocking the MCP server:**

```typescript
// packages/safe/src/__tests__/tools/get-pool-info.test.ts
import { vi, describe, it, expect } from "vitest";
import { createMockMcpServer } from "../helpers/mock-mcp-server.js";
import { registerTool } from "../../tools/get-pool-info.js";

describe("get_pool_info tool", () => {
  it("registers without error", () => {
    const server = createMockMcpServer();
    expect(() => registerTool(server)).not.toThrow();
  });

  it("validates required pool_address parameter", () => {
    const server = createMockMcpServer();
    registerTool(server);
    const tool = server.getRegisteredTool("get_pool_info");
    const result = tool.schema.safeParse({});
    expect(result.success).toBe(false);
  });
});
```

### 3.2a Property Testing with fast-check (TypeScript SDK)

For TypeScript SDK logic — encoding, fee math, share price rounding — use [fast-check](https://fast-check.dev/) via its Vitest integration. This is the TypeScript analog to Foundry's fuzz testing, but for off-chain code only. Do not use fast-check to test Solidity contracts; use Foundry's native fuzzer for that.

```typescript
// packages/vault/sdk/src/__tests__/ERC4626Math.prop.test.ts
import { test } from "@fast-check/vitest";
import * as fc from "fast-check";
import { previewDeposit, previewRedeem } from "../ERC4626Math.js";

// Property: converting to shares and back should never yield more assets than deposited
// (rounding always floors, never ceilings — protects existing shareholders)
test.prop([
  fc.bigInt({ min: 1n, max: 1_000_000n * 10n ** 6n }),
  fc.bigInt({ min: 0n, max: 10_000_000n * 10n ** 6n }),
  fc.bigInt({ min: 0n, max: 1_000_000n * 10n ** 18n }),
])(
  "convertToShares round-trip never profits attacker",
  (depositAmount, totalAssets, totalSupply) => {
    // Skip degenerate states (empty vault with no shares)
    fc.pre(!(totalAssets === 0n && totalSupply > 0n));

    const shares = previewDeposit(depositAmount, totalAssets, totalSupply);
    const redeemed = previewRedeem(
      shares,
      totalAssets + depositAmount,
      totalSupply + shares,
    );
    expect(redeemed).toBeLessThanOrEqual(depositAmount);
  },
);

// Property: fee calculation always stays within caps
test.prop([
  fc.bigInt({ min: 0n, max: 1_000_000n * 10n ** 6n }),
  fc.integer({ min: 0, max: 500 }), // management fee bps (capped at 500)
])("management fee never exceeds cap", (totalAssets, feeBps) => {
  const fee = computeManagementFee(totalAssets, feeBps);
  expect(fee).toBeLessThanOrEqual((totalAssets * BigInt(feeBps)) / 10_000n);
});
```

Install: `pnpm add -D @fast-check/vitest fast-check`

fast-check is for **TypeScript SDK logic only** (encoding, fee math, rounding, input validation). For anything that touches Solidity state, use Foundry's native fuzzer.

### 3.3 Integration Tests (With Anvil)

Integration tests use the `useGottsDeployment()` fixture (or `useDeployedUniswap()` for safe-only tests). Anvil starts once per test file via `createTestFixture`, takes a snapshot after setup, and reverts before each test.

> **Snapshot isolation caveat**: Anvil snapshots are **consumed on revert** — you cannot reuse a snapshot ID. `createTestFixture()` handles this automatically by re-taking the snapshot after each revert. If you implement `beforeEach`/`afterEach` manually, you must re-take the snapshot after every revert:
>
> ```typescript
> let snapshotId: `0x${string}`;
>
> beforeAll(async () => {
>   await deployContracts();
>   snapshotId = await testClient.snapshot();
> });
>
> afterEach(async () => {
>   await testClient.revert({ id: snapshotId });
>   snapshotId = await testClient.snapshot(); // Must re-take — previous ID is now invalid!
> });
> ```

```typescript
// packages/vault/src/__tests__/factory.integration.test.ts
import { describe, it, expect } from "vitest";
import { useGottsDeployment } from "../devenv/fixtures.js";
import { parseUnits, getAddress } from "viem";

describe("AgentVaultFactory", () => {
  const { getClient, getDeployment, getAccounts } = useGottsDeployment();

  it("creates a vault for a registered agent", async () => {
    const client = getClient();
    const { gotts } = getDeployment();
    const [, alice] = getAccounts();

    const hash = await client.writeContract({
      address: gotts.agentVaultFactory,
      abi: AGENT_VAULT_FACTORY_ABI,
      functionName: "createVault",
      args: [
        {
          name: "Test Vault",
          symbol: "TV",
          asset: getDeployment().mockTokens.usdc,
          agentId: 1n, // alice's agent ID
          managementFeeBps: 100n,
          performanceFeeBps: 1000n,
          autoPool: false,
          hookAddress: "0x0000000000000000000000000000000000000000",
        },
      ],
      account: alice,
    });

    const receipt = await client.waitForTransactionReceipt({ hash });
    expect(receipt.status).toBe("success");

    // Extract vault address from VaultCreated event
    const log = receipt.logs.find((l) => l.topics[0] === VAULT_CREATED_TOPIC);
    expect(log).toBeDefined();
  });

  it("rejects vault creation for unregistered agent", async () => {
    const client = getClient();
    const { gotts } = getDeployment();
    const unregistered = getAccounts()[10]; // account without ERC-8004 registration

    await expect(
      client.writeContract({
        address: gotts.agentVaultFactory,
        abi: AGENT_VAULT_FACTORY_ABI,
        functionName: "createVault",
        args: [
          /* config with unregistered agentId */
        ],
        account: unregistered,
      }),
    ).rejects.toThrow(/AgentNotRegistered/);
  });
});
```

### 3.4 MCP Tool Integration Tests

Test MCP tools by calling their handler functions directly (not via MCP transport). Inject a `TestClient` that wraps the Anvil `PublicClient`.

```typescript
// packages/vault/src/__tests__/vault-tools.integration.test.ts
import { describe, it, expect } from "vitest";
import { useGottsDeployment } from "../devenv/fixtures.js";
import { createVaultToolHandler } from "../../tools/vault-deposit.js";

describe("vault_deposit MCP tool", () => {
  const { getClient, getDeployment, getAccounts } = useGottsDeployment();

  it("deposits and returns share count", async () => {
    const client = getClient();
    const deployment = getDeployment();
    const [, alice] = getAccounts();

    const handler = createVaultToolHandler({
      publicClient: client,
      walletClient: client,
    });

    const result = await handler({
      vaultAddress: deployment.gotts.seedVaults[0].address,
      amount: "1000",
      chainId: "31337",
      agentId: "1",
    });

    expect(result.content[0].type).toBe("text");
    const parsed = JSON.parse(result.content[0].text);
    expect(parsed.shares).toBeDefined();
    expect(BigInt(parsed.shares)).toBeGreaterThan(0n);
  });
});
```

### 3.5 Fixture Composition

Stack fixtures from bottom to top. Each fixture can add setup on top of the previous:

```typescript
// Custom fixture: vault with a specific state for CCA testing
export function useCCATestSetup() {
  const base = useGottsDeployment();

  let ccaVaultAddress: Address;

  beforeAll(async () => {
    const client = base.getClient();
    const deployment = base.getDeployment();

    // Create CCA-configured vault on top of base deployment
    ccaVaultAddress = await deployTestCCAVault(client, deployment);
  });

  return {
    ...base,
    getCCAVault: () => ccaVaultAddress,
  };
}
```

**Canonical Vitest 3.x `test.extend()` pattern** — an alternative to the `useX()` hook pattern above, providing Playwright-style dependency injection with automatic cleanup:

```typescript
// packages/vault/src/__tests__/fixtures.ts
import { test as baseTest } from "vitest";
import { createTestClient, http } from "viem";
import { foundry } from "viem/chains";

export const test = baseTest.extend<{
  deployment: GottsDeployment;
  snapshot: `0x${string}`;
}>({
  deployment: async ({}, use) => {
    const pool = Number(process.env.VITEST_POOL_ID ?? 1);
    const rpcUrl = `http://127.0.0.1:8545/${pool}`;
    const deployment = await loadDeployment(rpcUrl);
    await use(deployment);
    // No teardown needed — Anvil pool handles cleanup
  },

  snapshot: async ({ deployment }, use) => {
    const id = await deployment.client.snapshot();
    await use(id);
    await deployment.client.revert({ id });
    // Note: snapshot consumed — next test will get a fresh fixture call
  },
});

// Usage:
test("deposits mint shares", async ({ deployment, snapshot }) => {
  const { client, contracts, accounts } = deployment;
  // snapshot is already taken; Anvil reverts after this test
  const shares = await client.writeContract({
    address: contracts.agentVaultCore,
    abi: VAULT_ABI,
    functionName: "deposit",
    args: [1000n * 10n ** 6n, accounts[1].address],
    account: accounts[1],
  });
  expect(shares).toBeGreaterThan(0n);
});
```

Use `test.scoped()` (Vitest 3.1+) to override fixture values per `describe` block — e.g., testing the same vault in different states (empty vs. seeded with liquidity).

### 3.6 Mock Patterns

Standard mock helpers used across all packages:

```typescript
// packages/vault/src/__tests__/helpers/mocks.ts

// Minimal viem PublicClient stub
export function createMockPublicClient(
  overrides: Partial<PublicClient> = {},
): PublicClient {
  return {
    readContract: vi.fn().mockResolvedValue(undefined),
    simulateContract: vi
      .fn()
      .mockResolvedValue({ result: undefined, request: {} }),
    waitForTransactionReceipt: vi
      .fn()
      .mockResolvedValue({ status: "success", logs: [] }),
    getBlock: vi.fn().mockResolvedValue({
      number: 1000n,
      timestamp: BigInt(Date.now() / 1000),
    }),
    ...overrides,
  } as unknown as PublicClient;
}

// MCP server that captures registered tools
export function createMockMcpServer() {
  const tools = new Map<string, { schema: z.ZodType; handler: Function }>();
  return {
    tool: (
      name: string,
      _desc: string,
      schema: z.ZodType,
      handler: Function,
    ) => {
      tools.set(name, { schema: z.object(schema), handler });
    },
    getRegisteredTool: (name: string) => {
      const t = tools.get(name);
      if (!t) throw new Error(`Tool not registered: ${name}`);
      return t;
    },
    getAllTools: () => [...tools.keys()],
  };
}

// TypeScript stub for ERC-8004 registry interactions
export class MockERC8004Registry {
  private agents = new Map<
    bigint,
    { wallet: Address; score: number; frozen: boolean }
  >();
  private nextId = 1n;

  register(wallet: Address): bigint {
    const id = this.nextId++;
    this.agents.set(id, { wallet, score: 0, frozen: false });
    return id;
  }

  setScore(agentId: bigint, score: number): void {
    const agent = this.agents.get(agentId);
    if (!agent) throw new Error(`Agent ${agentId} not registered`);
    agent.score = score;
  }

  isValid(agentId: bigint): boolean {
    const agent = this.agents.get(agentId);
    return agent !== undefined && !agent.frozen;
  }
}
```

### 3.7 Skill Tests with Promptfoo

Skills are tested with [Promptfoo](https://promptfoo.dev/) rubric scoring to verify the LLM invokes the right tools and produces correct outputs.

```yaml
# packages/vault/skills/tests/vault-strategy.yaml
description: Test vault-strategy skill invokes correct tools
providers:
  - id: anthropic:claude-sonnet-4-6
    config:
      tools: [vault_get_state, vault_rebalance, vault_get_risk_metrics]

tests:
  - description: Vault with high idle capital should trigger rebalance suggestion
    vars:
      prompt: "My vault has $500k idle. Should I rebalance?"
    assert:
      - type: llm-rubric
        value: "Response includes a recommendation to rebalance and references current idle capital percentage"
        threshold: 0.85
      - type: tool-calls
        value:
          - tool: vault_get_state
          - tool: vault_get_risk_metrics

  - description: Manager without Sovereign tier should be warned about deposit limits
    vars:
      prompt: "I want to accept a $200k deposit into my vault"
    assert:
      - type: llm-rubric
        value: "Response mentions reputation tier limits and what tier is required for deposits above $100k"
        threshold: 0.85
```

```bash
# Run skill tests in CI
pnpm promptfoo eval --config packages/vault/skills/tests/vault-strategy.yaml

# Pass threshold: 85%
```

***

## 4. Test File Organization (Normative)

```
packages/vault/
├── contracts/test/                      # Forge unit + fuzz + invariant tests
│   ├── AgentVaultFactory.t.sol
│   ├── AgentVaultCore.t.sol
│   ├── VaultHook.t.sol
│   ├── NAVAwareHook.t.sol
│   ├── LaunchFeeHook.t.sol
│   ├── FeeModule.t.sol
│   ├── OnboardRouter.t.sol
│   ├── mocks/
│   │   ├── MockERC8004IdentityRegistry.sol
│   │   ├── MockERC8004ReputationRegistry.sol
│   │   ├── MockPriceOracle.sol
│   │   ├── MockERC4626Adapter.sol
│   │   └── MockERC20.sol
│   ├── invariants/
│   │   ├── SharePriceInvariants.t.sol   # INV-1 through INV-4
│   │   ├── AccountingInvariants.t.sol   # INV-5 through INV-8
│   │   ├── AccessInvariants.t.sol       # INV-9 through INV-11
│   │   ├── CircuitBreakerInvariants.t.sol  # INV-12 through INV-14
│   │   └── handlers/
│   │       ├── VaultHandler.sol
│   │       ├── FactoryHandler.sol
│   │       └── ProxyHandler.sol
│   └── fork/
│       ├── VaultCreation.fork.t.sol
│       ├── CircuitBreaker.fork.t.sol
│       └── ERC8004Integration.fork.t.sol
│
├── sdk/src/__tests__/                   # Vitest unit tests for TypeScript SDK
│   ├── VaultClient.test.ts
│   ├── StrategyEngine.test.ts
│   ├── ReputationReader.test.ts
│   ├── FeeCalculator.test.ts
│   └── helpers/
│       └── mocks.ts
│
└── src/__tests__/                       # Vitest integration tests (with Anvil)
    ├── helpers/
    │   └── mock-mcp-server.ts
    ├── factory.integration.test.ts
    ├── deposit-withdraw.integration.test.ts
    ├── rebalance.integration.test.ts
    ├── proxy.integration.test.ts
    ├── reputation.integration.test.ts
    ├── onboard-router.integration.test.ts
    └── share-pool.integration.test.ts

packages/safe/src/__tests__/
├── helpers/
│   └── mock-mcp-server.ts
├── tools/                               # MCP tool unit tests (no Anvil)
│   ├── get-pool-info.test.ts
│   ├── get-token-price.test.ts
│   ├── vault-deposit.test.ts
│   ├── vault-withdraw.test.ts
│   └── vault-get-state.test.ts
└── integration/                         # Integration tests against devenv
    ├── vault-tools.integration.test.ts
    └── data-tools.integration.test.ts

packages/agent-proxy/
├── contracts/test/                      # Forge tests
│   ├── AgentProxy.t.sol
│   ├── AgentProxyFactory.t.sol
│   └── invariants/
│       └── ProxyDelayInvariants.t.sol
└── sdk/src/__tests__/
    ├── ProxyClient.test.ts
    └── MonitorBot.test.ts
```

***

## 5. `vitest.config.ts` Patterns per Package

### `packages/vault/vitest.config.ts`

Full Gotts deployment for integration tests. Parallel pool of 4 Anvil instances for concurrent test file execution.

```typescript
import { defineConfig } from "vitest/config";
import { createAnvilPool } from "@gotts.ai/testnet/fixtures-parallel";
import { deployFullStack } from "./src/devenv/setup.js";

const anvilPool = createAnvilPool({
  preset: "fresh",
  poolSize: 4,
  setup: (instance) => deployFullStack(instance.client, "from-scratch"),
  setupTimeout: 180_000,
});

export default defineConfig({
  test: {
    globalSetup: anvilPool.globalSetup,
    pool: "forks",
    testTimeout: 60_000,
    hookTimeout: 180_000,
    include: ["sdk/src/__tests__/**/*.test.ts", "src/__tests__/**/*.test.ts"],
    reporters: process.env.CI ? ["json", "verbose"] : ["verbose"],
  },
});
```

### `packages/safe/vitest.config.ts`

Uniswap devenv only (no Gotts vault layer). Smaller pool — safe integration tests are faster.

```typescript
import { defineConfig } from "vitest/config";
import { createAnvilPool } from "@gotts.ai/testnet/fixtures-parallel";
import { deployUniswapStack } from "@gotts.ai/devenv";

const anvilPool = createAnvilPool({
  preset: "fresh",
  poolSize: 2,
  setup: (instance) => deployUniswapStack(instance.client, "from-scratch"),
  setupTimeout: 120_000,
});

export default defineConfig({
  test: {
    globalSetup: anvilPool.globalSetup,
    pool: "forks",
    testTimeout: 30_000,
    hookTimeout: 120_000,
    include: ["src/__tests__/**/*.test.ts"],
  },
});
```

### `packages/agent-proxy/vitest.config.ts`

Lightweight — mostly unit tests. Integration tests use a minimal Anvil with no Uniswap or Gotts layer.

```typescript
import { defineConfig } from "vitest/config";

export default defineConfig({
  test: {
    pool: "forks",
    testTimeout: 30_000,
    hookTimeout: 60_000,
    include: ["sdk/src/__tests__/**/*.test.ts"],
    // Integration tests in separate suite via --project proxy-integration
  },
});
```

### Vitest Projects: separating unit from integration within one config

Use Vitest's `projects` feature to run fast unit tests (no Anvil) and Anvil-dependent integration tests from a single config, with separate `--project` flags in CI:

```typescript
// packages/vault/vitest.config.ts — projects-based separation
import { defineConfig } from "vitest/config";
import { createAnvilPool } from "@gotts.ai/testnet/fixtures-parallel";
import { deployFullStack } from "./src/devenv/setup.js";

const anvilPool = createAnvilPool({
  preset: "fresh",
  poolSize: 4,
  setup: (instance) => deployFullStack(instance.client, "from-scratch"),
  setupTimeout: 180_000,
});

export default defineConfig({
  test: {
    projects: [
      {
        test: {
          name: "unit",
          include: ["sdk/src/__tests__/**/*.test.ts"],
          // No globalSetup — runs without Anvil; completes in < 5s
        },
      },
      {
        test: {
          name: "integration",
          include: ["src/__tests__/**/*.integration.test.ts"],
          globalSetup: anvilPool.globalSetup,
          pool: "forks",
          testTimeout: 60_000,
          hookTimeout: 180_000,
        },
      },
    ],
  },
});
```

```bash
# Run fast unit tests only (no Anvil, no network)
vitest --project unit

# Run integration tests only (starts Anvil pool)
vitest --project integration

# Run all
vitest
```

This is the recommended config for `packages/vault` and `packages/safe` to keep pre-commit feedback loops fast.

***

## 6. CI Pipeline (GitHub Actions)

### Trigger

All jobs run on:

* Push to `main`
* Pull request (any branch → main)

### Job Graph

```
lint-typecheck ──► unit-tests ──► integration-tests
                               └► forge-tests ──► forge-invariants
                                              └► gas-snapshot
                               └► fork-tests (if BASE_RPC_URL available)
coverage (collects from unit + forge + integration)
skill-tests (nightly only)
e2e-swarm (nightly only)
```

### Job Definitions

**`lint-typecheck`** (no chain, \~1 min)

```yaml
- run: pnpm lint
- run: pnpm typecheck
```

**`unit-tests`** (no chain, \~2 min)

```yaml
- run: pnpm test --run --reporter=junit --outputFile=test-results/unit.xml
  env:
    VITEST_SKIP_INTEGRATION: "true" # Skip any test requiring Anvil
```

**`forge-tests`** (no fork, \~3 min)

```yaml
- run: forge test -vv --no-match-test "testFork_"
  working-directory: packages/vault
```

**`forge-invariants`** (\~10 min in CI)

```yaml
- run: FOUNDRY_PROFILE=ci forge test --match-test "invariant_" -vvv
  working-directory: packages/vault
  env:
    FOUNDRY_FUZZ_RUNS: "10000"
    FOUNDRY_INVARIANT_RUNS: "1000"
    # Semi-deterministic seed per PR: same inputs across retries, different per PR
    FOUNDRY_FUZZ_SEED: "0x${{ github.event.pull_request.number || '0' }}"
```

**`integration-tests`** (\~8 min)

```yaml
- name: Cache Anvil state
  uses: actions/cache@v4
  with:
    path: |
      packages/devenv/anvil-state.json
      packages/devenv/.devenv-deployment.json
      packages/vault/.devenv-gotts.json
    key: anvil-gotts-${{ hashFiles('packages/devenv/src/deployer/**', 'packages/vault/src/devenv/**') }}

- run: pnpm test:integration --run
  env:
    DEVENV_MODE: from-scratch
```

**`fork-tests`** (\~5 min, skipped if no `BASE_RPC_URL` secret)

```yaml
- if: secrets.BASE_RPC_URL != ''
  run: forge test --match-test "testFork_" --fork-url $BASE_RPC_URL --fork-block-number 28000000
  env:
    BASE_RPC_URL: ${{ secrets.BASE_RPC_URL }}
```

**`gas-snapshot`** (\~3 min)

```yaml
- run: forge snapshot --check
  working-directory: packages/vault
  # Fails if any function gas exceeds the committed .gas-snapshot by >5%
  env:
    # Semi-deterministic fuzz seed: same inputs per PR number, reducing false-positive
    # gas regressions caused by different random inputs across CI machines.
    FOUNDRY_FUZZ_SEED: "0x${{ github.event.pull_request.number || '0' }}"
```

**Gas diff on PRs**: The `foundry-gas-diff` GitHub Action (Rubilmax/foundry-gas-diff, \~208 stars) is the current best practice for posting gas diff comments on PRs. It stores gas reports as GitHub Artifacts per branch, computes the diff, and posts a markdown table comment. Wire this in at Phase 3+ alongside the `gas-snapshot --check` gate:

```yaml
# Phase 3+: add to gas-snapshot job
- run: forge test --gas-report > gasreport.ansi
  working-directory: packages/vault
  env:
    FOUNDRY_FUZZ_SEED: "0x${{ github.event.pull_request.number || '0' }}"

- uses: Rubilmax/foundry-gas-diff@v3
  id: gas_diff
  with:
    sortCriteria: avg,max
    sortOrders: desc,desc

- uses: marocchino/sticky-pull-request-comment@v2
  if: github.event_name == 'pull_request'
  with:
    header: gas-diff
    delete: ${{ !steps.gas_diff.outputs.markdown }}
    message: ${{ steps.gas_diff.outputs.markdown }}
```

**`coverage`** (\~10 min, runs after unit + integration)

```yaml
- run: forge coverage --report lcov --report-file coverage/forge-lcov.info
- run: pnpm test --run --coverage
- run: |
    # Merge lcov reports
    lcov --add-tracefile coverage/forge-lcov.info \
         --add-tracefile coverage/vitest-lcov.info \
         --output-file coverage/merged-lcov.info
    # Check thresholds (see §8)
```

### Caching Strategy

| Cache                   | Key                                                                                               | Path                                             |
| ----------------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------------ |
| Foundry build artifacts | `foundry-${{ hashFiles('foundry.toml', 'packages/*/foundry.toml') }}`                             | `~/.foundry/cache/`, `packages/*/contracts/out/` |
| Anvil devenv state      | `anvil-gotts-${{ hashFiles('packages/devenv/src/deployer/**', 'packages/vault/src/devenv/**') }}` | `*.json` deployment files                        |
| Fork RPC block cache    | `fork-base-28000000` (static key, block is pinned)                                                | `~/.foundry/cache/rpc/base/`                     |
| pnpm store              | `pnpm-${{ hashFiles('pnpm-lock.yaml') }}`                                                         | `~/.pnpm-store/`                                 |

### Secrets

| Secret            | Usage                         | Fallback           |
| ----------------- | ----------------------------- | ------------------ |
| `BASE_RPC_URL`    | Fork tests, fork mode devenv  | Skip fork tests    |
| `ALCHEMY_API_KEY` | Alternative to `BASE_RPC_URL` | Use `BASE_RPC_URL` |

Fork tests are **advisory** (not blocking) until Phase 3 (when mainnet deployment is imminent).

### PR Gates by Phase

| Phase | Required to Merge                                          |
| ----- | ---------------------------------------------------------- |
| P0    | `lint-typecheck`, `unit-tests`                             |
| P1    | + `forge-tests`, `forge-invariants`                        |
| P2    | + `integration-tests`                                      |
| P3    | + `fork-tests` (if `BASE_RPC_URL` present), `gas-snapshot` |
| P5    | + `coverage` (must meet thresholds in §8)                  |

***

## 7. Gas Benchmarking Targets

These are normative targets for the `gas-snapshot` CI gate (enforced at Phase 3+). Targets represent the maximum acceptable gas for the specified operation under standard conditions (no competing transactions, warm storage slots).

| Operation                      | Contract                      | Gas Budget            | Notes                     |
| ------------------------------ | ----------------------------- | --------------------- | ------------------------- |
| `createVault()`                | `AgentVaultFactory`           | ≤ 500,000             | Includes hook deployment  |
| `createVault()` with auto-pool | `AgentVaultFactory`           | ≤ 700,000             | Includes V4 pool init     |
| `deposit()` first depositor    | `AgentVaultCore`              | ≤ 200,000             | Cold storage              |
| `deposit()` subsequent         | `AgentVaultCore`              | ≤ 150,000             | Warm storage              |
| `withdraw()`                   | `AgentVaultCore`              | ≤ 200,000             | Full redemption           |
| `redeem()` partial             | `AgentVaultCore`              | ≤ 180,000             | Partial shares            |
| `rebalance()`                  | `AgentVaultCore`              | ≤ 300,000 per adapter | Does not include swap gas |
| `report()`                     | `AgentVaultCore`              | ≤ 80,000              | Fee accrual snapshot      |
| `proxy.announce()`             | `AgentProxy`                  | ≤ 80,000              | Stores tx hash + delay    |
| `proxy.execute()`              | `AgentProxy`                  | ≤ 50,000 + target gas | Target call gas excluded  |
| `proxy.cancel()`               | `AgentProxy`                  | ≤ 30,000              | Mark as cancelled         |
| `register()`                   | `MockERC8004IdentityRegistry` | ≤ 100,000             | Includes NFT mint         |
| `hook.beforeSwap()`            | `NAVAwareHook`                | ≤ 30,000              | Per-swap overhead         |
| `hook.beforeSwap()`            | `VaultHook`                   | ≤ 25,000              | Identity check overhead   |

**Regression threshold**: Any function that exceeds its committed `.gas-snapshot` value by more than **5%** blocks merge (CI `gas-snapshot --check` fails). Improvements (lower gas) always pass.

***

## 8. Coverage Targets by Phase

Coverage is measured separately for Solidity (Forge `--coverage`) and TypeScript (Vitest `--coverage`). Targets apply to the code that exists at the time of each phase gate; pre-existing tests are not retroactively required.

### Phase 1 (Core Vault Contracts)

| Target                                  | Metric     | Threshold           |
| --------------------------------------- | ---------- | ------------------- |
| `AgentVaultFactory` line coverage       | Forge      | ≥ 90%               |
| `AgentVaultCore` line coverage          | Forge      | ≥ 90%               |
| `IdentityRegistryAdapter` line coverage | Forge      | ≥ 85%               |
| `FeeModule` line coverage               | Forge      | ≥ 85%               |
| ERC-4626 property tests                 | a16z suite | All pass, delta = 0 |

### Phase 2 (SDK + MCP Tools)

| Target                      | Metric           | Threshold                     |
| --------------------------- | ---------------- | ----------------------------- |
| TypeScript SDK (`sdk/src/`) | Vitest line      | ≥ 85%                         |
| All 21 vault MCP tools      | Integration test | ≥ 1 test each                 |
| V4 Hooks line coverage      | Forge            | ≥ 80%                         |
| Error code exhaustiveness   | Unit test        | All error codes have ≥ 1 test |

### Phase 3 (Safety + Proxy)

| Target                        | Metric           | Threshold      |
| ----------------------------- | ---------------- | -------------- |
| `AgentProxy` line coverage    | Forge            | ≥ 95%          |
| `AgentProxy` branch coverage  | Forge            | ≥ 90%          |
| Circuit breaker paths         | Forge            | ≥ 95%          |
| `OnboardRouter` line coverage | Forge            | ≥ 85%          |
| Proxy MCP tools (5 tools)     | Integration test | ≥ 2 tests each |

### Phase 5 (Pre-Mainnet)

| Target                    | Metric           | Threshold                          |
| ------------------------- | ---------------- | ---------------------------------- |
| All v1 contracts combined | Forge line       | ≥ 95%                              |
| All v1 contracts combined | Forge branch     | ≥ 85%                              |
| TypeScript (all packages) | Vitest line      | ≥ 85%                              |
| INV-1 through INV-14      | Invariant tests  | All pass, 10k runs, 512 depth      |
| Skill tests               | Promptfoo rubric | ≥ 85% pass rate                    |
| Fork test suite           | All scenarios    | All pass against pinned Base block |

***

## 9. Running the Full Test Suite Locally

```bash
# Tier 1: Unit tests (no chain, fast)
pnpm test --run                         # All packages, unit only

# Forge unit tests (no fork)
cd packages/vault && forge test         # Unit + fuzz (default profile)

# Forge invariants (slower)
FOUNDRY_PROFILE=ci forge test --match-test invariant_

# Tier 2: Integration tests (requires Anvil)
pnpm test:integration --run             # All packages with Anvil fixtures

# Forge fork tests (requires BASE_RPC_URL)
forge test --match-test testFork_ --fork-url $BASE_RPC_URL

# Gas snapshot
forge snapshot                          # Update .gas-snapshot
forge snapshot --check                  # Fail on regression

# Coverage
forge coverage --report lcov
pnpm test --coverage

# Tier 3: E2E swarm (manual, ~15 min)
cd packages/vault && pnpm testnet:swarm

# Full CI simulation (sequential)
pnpm lint && pnpm typecheck
pnpm test --run
cd packages/vault && forge test && forge test --match-test invariant_
pnpm test:integration --run
```
