Methodology: measuring frontier-AI stability on crypto tax classification
All results as of August 2026 · model snapshots: claude-sonnet-5 · gpt-5.5 · grok-4.5
This page documents how we measured what frontier AI models actually do when asked to classify on-chain transactions for tax purposes — including the results that favor them, the results that favor us, and everything we could not measure. Every number below is recomputable from frozen artifacts. Nothing was excluded after a model output was seen.
1. What we measured and why
Two studies, run under one frozen protocol. Both use the same fixed prompt (verified by hash at every run): single-shot, no tools, no retrieval, a closed 17-category taxonomy, abstention explicitly permitted. Both were designed — sampling rules, metrics, and analysis plan — before any model was called.
Study A — the 388-row bake-off (accuracy and silent-wrong)
Three flagship models (claude-sonnet-5, gpt-5.5, grok-4.5) classified transactions from a frozen 777-row corpus against a 388-row adjudicated answer key (provenance in disclosure 4d). The CryptoTaxEdge engine (CTE in tables below) was scored against the same key as a reference.
- Accuracy: statistically tied at the top. gpt-5.5 scored 93.0% on the deep-protocol strata (94.5% overall, abstain-excluded) vs the CTE reference at 92.3% deep-protocol — a 0.7-point gap on n=388, statistically indistinguishable. grok-4.5 scored 89.6% overall; claude-sonnet-5 scored 79.9% overall. We publish the tie rather than bury it.
- Silent-wrong is where the systems separate. A silent-wrong is a confident, unflagged, wrong answer — the failure mode that matters in a liability product. Measured rates: gpt-5.5 5.4%, grok-4.5 9.5%, claude-sonnet-5 14.2% of rows (range 5.4–14.2%). The engine's review-routing exists to push exactly those rows to needs-review instead of asserting them.
- Consistency: the bake-off's repeat arm measured model run-to-run flip rates of 5.0–11.5% — the same question, asked twice, gets a different answer on up to one row in nine.
- Cost: a flagship read runs ~$0.0104 per enriched row at list price (July 2026) vs ~$0.0005 CryptoTaxEdge marginal serve — roughly 20×. Cost-model boundaries in disclosure 4d.
Study B — the 96-transaction complexity ladder (stability)
The same three models classified 96 fresh 2026 transactions spanning 12 protocol rungs, pooled into three pre-registered complexity bands: simple (basic transfers, router swaps, liquid-staking basics), mid (aggregators, lending, bridges, NFT/intent), and exotic (yield-stripping, vaults, restaking, perp-adjacent, leveraged looping, hard Solana). Every transaction ran at least twice per model on byte-identical input; hard rungs ran three times. 744 model calls, 100% strict-JSON parse rate, zero transport errors.
Four metrics, none of which require ground-truth labels: run-to-run flip rate, inter-model disagreement, abstention rate, and divergence from the CryptoTaxEdge engine (a framing metric, never an error rate). Headline results by band:
- Category flips rise 2.1% (95% CI [0.0–6.3]) → 8.3% [3.1–14.6] → 17.4% [11.1–23.6].
- Inter-model disagreement rises 15.0% → 27.5% → 44.5%; at tax-bucket level 0.0% → 12.5% → 24.2%.
- Abstention stays nearly flat: 0.0% → 4.2% → 5.2% of parsed calls.
- Mean stated confidence barely moves: 91.8 → 88.5 → 83.4.
- CryptoTaxEdge needs-review rate climbs 19% → 28% → 44% across the same bands — the engine increasingly declines to guess.
2. Results by complexity band
Bands were fixed a priori (before any model call): simple = rungs {1, 8}, mid = {2, 3, 4, 10}, exotic = {5, 6, 7, 9, 11, 12}. Confidence intervals are bootstrap 95% CIs (2,000 resamples over transactions, seeded). Per-rung numbers (n≈8 each) are context, never headlines; the pre-registered rung-level trend test did not reach significance (see Limitations).
Run-to-run stability (same model, byte-identical input)
| Band | Txs | Run-pairs | Category flip | 95% CI | Tax-bucket flip | 95% CI |
|---|---|---|---|---|---|---|
| simple | 16 | 48 | 2.1% | [0.0–6.3] | 0.0% | [0.0–0.0] |
| mid | 32 | 96 | 8.3% | [3.1–14.6] | 4.2% | [1.0–8.3] |
| exotic | 48 | 144 | 17.4% | [11.1–23.6] | 8.3% | [4.2–13.2] |
The simple and exotic CIs do not overlap. Per model, pooled across all 96 transactions: claude-sonnet-5 flipped 14.6% (14/96), gpt-5.5 10.4% (10/96), grok-4.5 10.4% (10/96) — no model is flip-free. Exotic-band detail: 25 category flips across 144 run-pairs; 12 of those 25 flips also changed the tax-treatment bucket.
Agreement, uncertainty, and escalation
| Band | Inter-model category disagreement | Tax-bucket disagreement | Abstention | Mean stated confidence | CTE needs-review |
|---|---|---|---|---|---|
| simple | 15.0% (6/40 states) | 0.0% (n=40) | 0.0% | 91.8 | 19% (n=16 tx) |
| mid | 27.5% (22/80 states) | 12.5% (n=80) | 4.2% | 88.5 | 28% (n=32 tx) |
| exotic | 44.5% (57/128 states) | 24.2% (n=128) | 5.2% | 83.4 | 44% (n=48 tx) |
A “state” is one (transaction, run) with all three models parsed. Raw and synonym-normalized disagreement were identical on every rung — the closed taxonomy already constrained vocabulary, so vocabulary mismatch does not inflate these numbers.
High-confidence disagreement
Of the 224 (transaction, run) cases where all three models answered without abstaining, 61 produced a three-way disagreement — and in 34 of those 61, all three models stated confidence of 70 or higher. Per band:
| Band | Full three-way disagreements | With all three models at confidence ≥70 | Share |
|---|---|---|---|
| simple | 6 | 6 | 6/6 |
| mid | 14 | 5 | 5/14 |
| exotic | 41 | 23 | 23/41 (56%) |
| pooled | 61 | 34 | 34/61 |
3. The asymmetry: confidence vs stability
The cleanest overall finding is not that models destabilize on complex transactions — it is that they destabilize without saying so. Flip rate rises roughly 8× from the simple band (2.1%, 95% CI [0.0–6.3]) to the exotic band (17.4%, 95% CI [11.1–23.6]) while mean stated confidence slips about 9 points (91.8 → 83.4) and abstention stays near flat (0.0% → 5.2%).
4. The Pendle finding: stable is not the same as right
One ladder rung sampled Pendle yield-stripping operations — PT and YT swaps, LP entries and exits, and a post-expiry PT redemption — three runs per model on each. Two findings sit side by side, and they point in opposite directions:
Pendle was the most stable exotic rung. 22 of 24 (transaction, model) triples produced the identical category on all three runs. Whatever the models say about Pendle, they say it consistently. Stability and correctness are different things — that is the point of this section.
Stated plainly: in this instance the tax bucket did not move. Both reads — redemption and swap — are §1001 dispositions, so the bottom-line taxable/non-taxable outcome is the same either way. The open question is character and basis analysis — for example, whether §§1276–1278-style questions about instruments acquired at a discount to a fixed maturity value apply to a PT held to expiry. We publish that as analysis, never as an asserted treatment: under the closed taxonomy the generic label is defensible. What the row demonstrates is narrower and more precise — a maturity redemption and a spot swap are indistinguishable to a single-shot frontier read, the models signal no awareness of the difference, and the very stability of the answer makes the gap invisible to consistency-based quality checks.
Honest counterweight: models were not uniformly collapsing Pendle to generic swap. Across 72 parsed Pendle calls the label mix included liquidity operations, rewards, and abstentions; on the dual-sided liquidity removal all nine calls agreed on the liquidity label. And on the six Pendle rows the engine asserts (it routes two to review), the model-majority category diverged from the engine zero times. The differences live in confidence, review-flagging, and instrument identification — not in the headline category.
5. Four disclosures
5a. Decoding settings
All model calls used provider default decoding settings. These reasoning-class models reject temperature and top_p overrides — the APIs refuse the parameters, and the harness header documents it. That means the run-to-run variance measured here is the variance of the only configuration anyone can actually deploy. A temperature-0 arm does not exist to run; there is no hidden deterministic mode we declined to test.
5b. Selection and exclusion log
The ladder set is a convenience sample of recent live activity, not a random draw from protocol history: for each rung contract, the most recent qualifying transactions at curation time (2026 timestamps, successful only, verified target contract, diversity caps per method and per sender, global dedupe against 765 previously used hashes). The full selection rule is frozen inside the set file itself.
The set was re-curated three times, all before any model call, all disclosed:
- A rate-limit response was initially treated as an empty transaction list, silently zeroing some contracts — fixed with a retry.
- A stale GMX v2 router whose recent traffic was mostly reverted bot transactions was replaced with the active router, resolved from the protocol's official deployments repository.
- Plain approvals and order cancellations — both outside the closed taxonomy — were being sampled and were globally excluded.
Zero post-hoc exclusions: no transaction was added or removed after any model output was seen. One mid-run harness fix is also disclosed: the engine-reference caller initially read legacy field names and stored nulls; it was fixed and the affected reference rows refetched. Model calls were untouched by that fix.
5c. Engine self-disclosure
Since we grade frontier models in public, here is where our own engine is deterministic and where it is not:
- Deterministic layers: the verified rule library (each rule tied to receipt-shape evidence), receipt-shape guards, and identification short-circuits (transfers with counterparty, approvals, spam, failed transactions, repeats). These return the same answer for the same input by construction.
- Model-based layers: complex-case lanes, where multi-source evidence is weighed by a model before a classification is proposed.
The mitigation for the model-based layers is review-routing: rather than assert a low-evidence answer, the engine escalates. We publish our own escalation curve as a feature, not a confession — needs-review climbs 19% → 28% → 44% across the same simple/mid/exotic bands where model instability climbs. And symmetrical honesty: we did not measure our own run-to-run flip rate in these studies. The engine reference figures above come from one call per transaction; the deterministic layers cannot flip by construction, but the model-based lanes have not had their flip rate measured under this protocol.
5d. Provenance, cost model, and snapshots
The 388-row answer key descends from a bake-off protocol pre-registered on 2026-07-22, before execution: sampling frames, metrics, equivalence maps, and adjudication rules were locked in advance and committed. Disagreement rows were adjudicated blind from raw on-chain receipts and logs by independent multi-model lenses — system outputs relabeled before review, neither lens shown either system's claim, all evidence recorded and frozen. It is not independent-CPA-labeled: no human tax professional has adjudicated the key, and we say so rather than imply otherwise. The eligible key (388 of 777 rows) skews answerable by construction; 187 hard rows where the engine abstained have no truth key and are unscored for every system, including ours.
Cost model, with dates: flagship ~$0.0104 per enriched row at list price (July 2026) vs CryptoTaxEdge marginal serve ~$0.0005 (pattern-cache serving). Both figures exclude enrichment costs on both sides and amortized R&D on ours; they are marginal per-row serving comparisons, not fully loaded unit economics. Study spend, for scale: the ladder cost $7.67 across 744 model calls; the full bake-off program measured out at $22.72.
Model snapshots: claude-sonnet-5, gpt-5.5, grok-4.5, called through provider APIs in July–August 2026. All results on this page are as of August 2026 and describe this model generation; each generation re-baselines, and we re-measure on a standing cadence.
6. Limitations
- No ground truth in the ladder. Nothing in Study B measures accuracy. Flip, disagreement, and abstention are internal-consistency metrics; engine divergence is disagreement with one good-but-fallible classifier, and is never a model error rate — several divergent rows cut against our own engine's low-confidence reads, and we route those to our own review pipeline.
- Single fixed prompt, single-shot, no tools, no retrieval. This measures out-of-the-box frontier-model behavior under this protocol — not the ceiling of an engineered agent with tools, retrieval, or protocol context. It licenses claims about this configuration; it does not license “LLMs cannot classify DeFi.” An agentic system with enrichment could behave very differently — that would be a different study.
- Sample sizes. n≈8 transactions per rung, 96 total. Per-rung numbers are context only; headline claims are band-level with bootstrap CIs. The pre-registered rung-level trend test did not reach significance (Spearman rho 0.375, one-sided permutation p = 0.116, 10,000 permutations) — the supported claim is the band-level contrast, not a smooth 12-step staircase, and we report that rather than upgrade it by narrative.
- Convenience sampling. Most-recent qualifying transactions from live contract activity, not random draws from protocol history; whatever operations were popular that week shape each rung.
- Complexity is confounded with log readability, semantic novelty, tax-law ambiguity, chain and encoding differences (the Solana packet is structurally different from the EVM packet), and engine coverage density. The ladder ordering is a reasonable a-priori complexity proxy, not a purified single axis.
- Divergence denominators vary. The engine reviews 44% of exotic rows and returned errors on four rows on one chain (filed as an engine gap and counted as engine error, not divergence), so engine-divergence percentages sit on small, uneven bases; bucket-level divergence with the denominator shown is the only form worth quoting.
- Abstention is prompt-tied. The prompt invites abstention once; different framing could move the absolute rates. The finding is the contrast — abstention nearly flat while flips rise roughly 8× — not the level.
- Two runs is the floor for flip measurement (three on hard rungs). Per-transaction flip rates at two runs are noisy and only meaningful pooled.
- Instrument-mechanics recall was unmeasurable. The strict-JSON output contract leaves no free text to scan, so whether a model ever mentions PT/YT/maturity semantics could not be tested. The Pendle finding rests on category choices, confidence, abstention, and stability — not on text mentions.
- Model versions are point-in-time (claude-sonnet-5, gpt-5.5, grok-4.5, as of August 2026). These results describe this generation, this month. The trend across generations only goes one way, which is why this measurement recurs.
7. Raw-hash arm: what a chat window actually gets
The accuracy numbers in Study A required handing each model the full enrichment packet — the decoded transfers, contract labels, and receipt our pipeline produces. Nobody who asks a chatbot about a transaction supplies that. So we ran the honest version of the question: the same 388 adjudicated transactions, the same three models, given only the transaction hash and chain — exactly what a real user enters — with abstention explicitly permitted and no tools or browsing.
Result: all three models declined to classify every transaction. 1,164 calls (388 rows × 3 models), 1,164 schema-valid responses, 1,164 abstentions. Zero answers, zero guesses, zero attempts to recognize a hash from training data. Confidence 0 across the board.
| Model (Aug 2026) | Abstention | Answered | Silent-wrong | Parse failures |
|---|---|---|---|---|
| claude-sonnet-5 | 100.0% (388/388) | 0 | 0 | 0 |
| gpt-5.5 | 100.0% (388/388) | 0 | 0 | 0 |
| grok-4.5 | 100.0% (388/388) | 0 | 0 | 0 |
Read together, the two arms bracket the question cleanly: with a full evidence packet, a flagship model reaches 93.0% on deep protocols; with only the hash, it answers nothing. The 93% is unreachable without the evidence layer — the fetching, decoding, and enrichment is what makes classification possible at all.
Three disclosures, pre-registered before the run: (1) this measures single-shot, no-tools behavior — a browsing-enabled chatbot is a different configuration that would need its own arm; (2) the models behaved honestly here — this result is evidence that frontier models refuse without evidence, not that chatbots hallucinate crypto tax answers; (3) total abstention partly reflects our prompt explicitly permitting abstention and the packet honestly stating that no transaction data was provided — consumer chat interfaces that pressure a model to answer may behave differently. Frozen raw output and scorecard: frontier-bakeoff-raw-rawhash-2026-08-06.json. Arm cost: $3.44.
8. Where to verify
Claims without inspectable evidence are marketing. These are the surfaces where the underlying material lives:
- Canonical examples — live classification records with on-chain evidence, per transaction.
- The treatment taxonomy — the closed category set both studies classified into.
- Answers — the tax positions behind the treatment buckets, with both sides shown where the law is unsettled.
- Public corpus on GitHub — published example records for independent inspection.
This page reports measurements, not tax advice. Classification of any specific transaction depends on facts and elections; verify with a qualified tax professional before filing. All figures as of August 2026.