Why use CryptoTaxEdge instead of Claude, ChatGPT, or Grok?
An earlier version of this post led with a single-run accuracy comparison between the best frontier model and our engine, scored once against the July adjudicated key. We removed that comparison, and every paraphrase of it, from this post and from every page on this site, in both directions. The reason is a measurement criterion, not the result: our own published run-to-run data shows a model's verdict on byte-identical input flips between 2.1% and 17.4% of the time depending on complexity, a band wider than the gap that table showed, so a single run cannot rank the systems. The retired table is preserved unchanged, numbers intact, in the benchmark archive. The instability, disagreement, refusal, and silent-wrong findings below are from the same program and are unchanged. Comparative rankings appear on this site only from pre-registered, multi-run designs with deltas outside measured run-to-run noise.
Honest answer first: a single-run accuracy score cannot decide this question, and we no longer publish one (see the editor's note above). What decides it is measured behavior, from tests we ran ourselves under a published, frozen methodology, with the models given the same decoded transaction data our engine works from. On exotic DeFi, asking a model the identical question twice changes the answer 17.4% of the time. And given only what a real user has, a hash and a chain, all three models refused every call: 1,164 out of 1,164.
- No single-run accuracy ranking. Run-to-run flips on byte-identical input (2.1% simple to 17.4% exotic) are wider than the deltas a single-run accuracy table showed, so we retired the comparison in both directions; the original is archived unchanged.
- Stability is not. Identical re-runs change the model's answer 2.1% of the time on simple transactions and 17.4% on exotic DeFi, and 12 of the 25 hardest-band changes also changed the tax treatment.
- Evidence is the product. Given only a transaction hash and chain, the models refused all 1,164 calls. Classification starts with the evidence layer, not the model.
If raw accuracy on a single classification were the whole job, you could ask a chatbot. It isn't the whole job. Here is the measurement that shows why: 96 fresh transactions arranged from simple transfers up to exotic DeFi, three frontier models (Claude Sonnet, ChatGPT's GPT-5.5, and Grok, August 2026 snapshots), every transaction run at least twice with input identical down to the byte.
The same question, different answers
On simple transactions, identical runs change the answer 2.1% of the time. On exotic DeFi (yield-stripping, restaking, leveraged looping) that climbs to 17.4%, roughly one answer in six. The changes are not cosmetic: in the hardest band, 12 of the 25 also changed the tax treatment. An answer that changes on a second ask is not a work product a practitioner can put a name on.
Confidence that doesn't track reality
Across that same climb, stated confidence slips only nine points (92 to 83) and the models decline to answer under 6% of the time. Run all three on the same exotic transaction and they disagree on the transaction type 44.5% of the time; in over half of those disagreements (23 of 41), every model reports confidence of 70 or higher. Instability rises roughly eightfold; the uncertainty signal barely moves.
Our engine inverts that curve: its needs-review rate climbs from 19% on simple transactions to 44% on exotic ones, escalating to a human instead of guessing. That asymmetry, instability that outruns the uncertainty signal versus escalation you can see, is the difference between an AI answer and evidence a professional can work from.
The failure you can't catch by double-checking
The sharpest single finding: a redemption of a Pendle principal token (PT) after maturity. All three models confidently labeled it a generic swap in every run, confidence 82 to 94, no flag. The deciding fact, whether the position had matured, lives in the protocol's own state, not the transaction record. And because the wrong answer is stable, double-checking catches nothing. The tax bucket happened to survive here (both reads are dispositions), but the character and basis analysis a practitioner would want flagged was invisible to every model, every time.
What happens when you actually just ask
Everything above required handing the models our enrichment packet (the decoded transfers, contract labels, and receipt our pipeline produces), which nobody who asks a chatbot supplies. So we ran the honest version: the same 388 transactions, the same three models, given only what a real user enters in a chat window, the hash and the chain. All three declined every single transaction: 1,164 calls, 1,164 refusals. No answers, no guesses, no hallucinated classifications, only honest refusal for lack of evidence.
If a chatbot has ever given you transaction details from just a hash, one of three things happened: it browsed, fetching a block-explorer page and reading it back; it recognized a famous transaction from its training data; or it guessed, plausible details with no evidence behind them. None of these contradicts this result. To say anything true about a transaction it has never seen, a model first needs an evidence layer, and a block-explorer page is a partial one: transfers, but no instrument semantics, protocol state, or verified taxonomy. And the ladder above shows what happens once a model has good evidence: that is where the 17.4% changed answers and the confident disagreements were measured. Evidence is where the instability starts, not where it ends. To the models' credit, refusing is the right behavior. It is also the whole argument: classification is unreachable without the evidence layer; fetching, decoding, and enrichment make it possible at all. (Measured single-shot with no tools or browsing, abstention explicitly permitted; a browsing-enabled agent is a different configuration, and the methodology page carries the full disclosures.)
Evidence and economics
A classification you can file on needs receipts: which rule matched, what the decoded transfers show, and a replay you can run yourself. Our published reference examples are re-verified against the live engine nightly, and every API response carries its evidence. A chat answer carries a paragraph.
The economics land in the same place. At API list prices, a flagship-model read of an already-prepared transaction costs about a cent*, and what comes back still has to be verified by hand. For roughly the same cent, CryptoTaxEdge returns a finished classification: category, US tax treatment, confidence score, the on-chain evidence behind it, and an honest review flag on the calls that belong with you, ready to drop into a workpaper.
*Flagship API list-price tokens per enriched row: ~$0.0104, measured July 2026. CryptoTaxEdge Pro: $249/month for 25,000 billable classifications, about $0.010 each, evidence and review-routing included.
What we're not claiming
We tested the chatbots as any user gets them: one call at a time, one fixed prompt, no browsing or tools. A custom-built agent could do better; we did not test one. We did not measure our own engine's run-to-run stability in this study, so we won't claim it here. And we are not claiming chatbots hallucinate crypto tax answers. Under honest no-tools conditions we observed the opposite: they refuse. A browsing-enabled agent is a different configuration we haven't measured. The full methodology, including the rung-level trend test that did not reach significance (p=0.116) and the limits of the sample, is published, with confidence intervals, on the methodology page. If you want to check our work, that's the point: you can.
Deterministic guardrails, honest escalation, and evidence are the product. Raw model accuracy is the commodity.
If you have a client wallet full of unknown DeFi transactions, enter one hash in the Classification Explorer and see what a finished answer looks like. Developers can start on the free tier of the API.