Home › Research

Research

Every number CryptoTaxEdge states publicly traces to a published study with frozen artifacts. This hub collects the studies, the open datasets behind them, and the methodology that governs how they were run. Claims without inspectable evidence are marketing; these pages exist so ours are not.

Benchmarks · 2026 H1

Frontier-AI stability on crypto tax classification

The 388-row adjudicated bake-off (silent-wrong rates and the raw-hash refusal arm) and the 96-transaction complexity ladder (run-to-run stability): study descriptions, the labelled 388-row corpus, public transaction-hash lists, the adjudication rubric, and aggregate results with confidence intervals. As of August 2026.

Open data

The datasets are published under CC BY 4.0, no gate and no form: take them, score your own system, disagree with us in public. The 388-row answer key is open because it is spent, not because it is safe. Our own engine was later tuned against it, so it can no longer measure us, and section 5 of the benchmark hub says so plainly. Comparative rankings appear on this site only from pre-registered, multi-run designs, on corpora none of the compared systems was tuned against, with deltas outside measured run-to-run noise.

bakeoff-labels.json · 388 labelled rows bakeoff-labels.csv ladder-hashes.json · 96 txs

Labels are model-adjudicated, not independent-CPA-labelled. A category and its default treatment are a starting point for professional review, not a tax position. Citation details: how to cite.

Protocol

Methodology

How the studies were designed and run: the frozen protocol, four disclosures, the full limitations list, and where to verify every figure.

These pages report measurements, not tax advice. Verify with a qualified tax professional before filing.