Why use CryptoTaxEdge instead of Claude, ChatGPT, or Grok?
Give Claude, ChatGPT, or Grok nothing but a transaction hash and a chain, which is all a real user has, and ask for the US tax treatment. We did: 388 transactions, three frontier models, 1,164 calls. We got 1,164 refusals. That is the right behavior, and it is the first half of the answer to this page's title: a model with no evidence will not classify anything. So we removed the excuse. We handed all three models the decoded transfers and contract labels our own pipeline produces, and asked each one the identical question twice. On exotic DeFi, the answer changed 17.4% of the time, roughly one in six. Those two numbers, measured under a published, frozen methodology, are what this post is about.
- A bare hash gets zero answers. Given only a transaction hash and chain, all three frontier models refused all 1,164 calls. Right behavior, and the reason classification starts with evidence, not a model.
- Perfect evidence does not buy stability. Asked the identical question twice, the models changed their answer 2.1% of the time on simple transactions and 17.4% on exotic DeFi; 12 of the 25 hardest-band changes also changed the tax treatment.
- Escalation is the product. Model confidence slips nine points while instability rises roughly eightfold. The engine inverts that curve: needs-review climbs from 19% to 44% with complexity, routing hard calls to a human instead of guessing.
Run the experiment yourself
Open any chatbot and enter a transaction hash and its chain, the entire input a real user has. Across our 388-transaction adjudicated corpus, single-shot, with no browsing or tools and abstention explicitly permitted, that produced refusals on every one of the 1,164 calls: no answers, no guesses, no hallucinated classifications. If a chatbot has ever read you transaction details from a bare hash, it browsed an explorer page, recognized a famous transaction from its training data, or guessed. And an explorer page is partial evidence at best: transfers, with no instrument semantics, protocol state, or verified taxonomy behind them. Fetching, decoding, and enrichment are what make classification possible at all.
Same evidence, same question, different answer
The 17.4% comes from a ladder of 96 fresh transactions arranged from simple transfers up to exotic DeFi (yield-stripping, restaking, leveraged looping), run against the same three models (Claude Sonnet, GPT-5.5, and Grok, August 2026 snapshots), every transaction at least twice with input identical down to the byte, every run supplied with the decoded evidence nobody who asks a chatbot actually provides. On simple transactions, identical runs change the answer 2.1% of the time. On exotic DeFi that climbs to 17.4%, and the changes are not cosmetic: in the hardest band, 12 of the 25 also changed the tax treatment. An answer that changes on a second ask is not a work product a practitioner can put a name on.
Confidence that doesn't track reality
Across that same climb, stated confidence slips only nine points (92 to 83) and the models decline to answer under 6% of the time. Run all three on the same exotic transaction and they disagree on the transaction type 44.5% of the time; in over half of those disagreements (23 of 41), every model reports confidence of 70 or higher. Instability rises roughly eightfold; the uncertainty signal barely moves.
The failure you can't catch by double-checking
The sharpest single finding is a redemption of a Pendle principal token after maturity. All three models labeled it a generic swap in every run, confidence 82 to 94, no flag. The deciding fact, whether the position had matured, lives in the protocol's own state, not the transaction record. The tax bucket happened to survive (both reads are dispositions), but the character and basis analysis a practitioner would want flagged was invisible to every model, every time. And because the wrong answer is stable, double-checking catches nothing.
The engine inverts the curve
The failure mode that matters in a liability product is the silent wrong: a confident, unflagged, wrong answer. Across our published 388-row corpus, the models' silent-wrong rates span 5.4% to 14.2% of adjudicated rows, with abstentions honored, never scored as wrong; the per-model figures are on the methodology page. Our engine is built to route exactly that class of row to a human instead: its needs-review rate climbs from 19% on simple transactions to 44% on exotic ones, escalating as complexity rises rather than guessing through it. We did not measure the engine's own run-to-run stability in this study, so we claim none here. What we can show is the routing, and every flagged row arrives with its evidence and the reason it needs you.
Evidence and economics
A classification you can file on needs receipts: which rule matched, what the decoded transfers show, and a replay you can run yourself. Our published reference examples are re-verified against the live engine daily, and every API response carries its evidence. A chat answer carries a paragraph.
The economics land in the same place. At API list prices, a flagship-model read of an already-prepared transaction costs about a cent*, and what comes back still has to be verified by hand. For roughly the same cent, CryptoTaxEdge returns a finished classification: category, US tax treatment, confidence score, the on-chain evidence behind it, and an honest review flag on the calls that belong with you, ready to drop into a workpaper.
*Flagship API list-price tokens per enriched row: ~$0.0104, measured July 2026. CryptoTaxEdge Pro: $249/month for 25,000 billable classifications, about $0.010 each, evidence and review-routing included.
Deterministic guardrails, honest escalation, and evidence are the product. Raw model accuracy is the commodity.
Scope. We tested the chatbots as any user gets them: one call at a time, a fixed prompt, no browsing or tools; a purpose-built agent is a different configuration we have not measured. The full methodology, confidence intervals, and limitations, including the rung-level trend test that did not reach significance, are published on the methodology page, and checking our work is the point.
If you have a client wallet full of unknown DeFi transactions, classify a transaction in the Explorer and see what a finished answer looks like: category, treatment, confidence, evidence, and an honest flag when a call belongs with you. Developers can start on the free tier of the API.