ai
Harvey's Margin Swing Is a Tokenomics Story, Not Only a Legal-AI Story
Harvey’s gross margin sat near 50% early in 2026, fell to about minus 50% by June as customer usage of its AI agents spiked on rented frontier models, and turned positive again after August, when the company released Harvey Tenet, its first in-house open-weight model, post-trained on Moonshot’s Kimi K3 with Fireworks. Bloomberg, via The Next Web, reported that arc; Harvey declined to comment on specific financials. Harvey’s own research post describes Tenet as improving long-horizon legal-agent performance and cost-efficiency. The stack remains hybrid. Anthropic has publicly reminded investors that Harvey still needs Opus for its hardest tasks.
I run Object Edge, a bootstrapped firm that sells services and a company-brain stack. I read Harvey as a unit-economics case about knowledge-work volume on metered APIs, and about what happens when that volume finds a cheaper open-weight path.
Power users and the meter
Harvey was a power user of frontier APIs for knowledge work. Anthropic and OpenAI are confirmed customers. Token usage moved from about one trillion tokens a month in January 2026 toward a twelve-to-thirteen-trillion monthly pace by May, according to CEO commentary, on the order of a twentyfold rise across 2026. The dollar amount of the API bill is not disclosed. I will not invent millions or tens of millions. The operator lesson does not require that conversion.
Development work and knowledge work both consume tokens, often on the same model price list. Anthropic Pro, Max, and Enterprise seats have rate limits; enterprise deals often combine seats with usage priced at API rates. Coding tools such as Claude Code can share those limits. I will not claim a law that coding tokens are always cheaper than knowledge tokens. The useful distinction is operating shape. Heavy knowledge and agentic production volume cannot live on consumer seats alone. Once agents run all day on real matters, the meter is the business.
Concentration is the uncomfortable question
How much of Anthropic’s revenue comes from power users like Harvey versus a long tail? Only public texture belongs here. Reuters reported in October 2025 that enterprise was about 80% of Anthropic revenue. A Ramp economist, via PYMNTS, estimated that the top 1% of customers account for about 80% of enterprise revenue for both OpenAI and Anthropic, based on card-spend estimates. In August 2025, Cursor and GitHub Copilot were reported as roughly a quarter of Anthropic at a then roughly $4–5 billion milestone. That is a historical snapshot. It should not be read as a current split. There is no public coding-versus-knowledge-work revenue split for Anthropic. Harvey is not disclosed as a named top-percent Anthropic revenue account.
Anthropic’s reported revenue run-rate climbed through late 2025 and 2026. Use that ladder as texture only. The thesis is not a growth cheer. It is concentration risk plus open-weight displacement.
Open weights as COGS strategy and as API risk
Harvey’s path is part of a wider move. Application companies under gross-margin pressure are post-training and routing onto open or smaller fine-tuned models. In the same reporting cycle, Abridge is building on open weights, Decagon routes about 80% of queries through models of its own, Ramp is weighing training for the first time, and Thomson Reuters has discussed building “Thomson.” For those apps, open weights are a cost-of-goods answer. For Anthropic and OpenAI, the same move is a revenue risk when power users stop renting every token.
Anthropic looks more exposed than OpenAI on composition if it lacks a ChatGPT-scale consumer long tail to monetize. That is a directional reading from the enterprise-heavy mix, not a quantified hedge. OpenAI’s enterprise API is also concentrated. Exact exposure is not public.
Two operating models we see at Object Edge
At Object Edge we work across development and knowledge-work lanes. We see knowledge-work API bills behave differently from coding seats once agents are in production, without inventing our own dollar or token figures. When every seat is a private frontier meter, the bill and the silo arrive together. Shared company context, with Hive as the company brain, Sayya as the harness, and Mentat as the brain for revenue, can reduce redundant expensive calls by putting common memory under the seats. That is amplification of the people and the models you already pay for, not a claim that open weights replace frontier models for every hard task. Harvey still needs Opus at the top of the stack. Many enterprises will too.
How AI-native delivery rewrites services and API P&Ls is still an open question. Harvey’s swing from plus fifty to minus fifty and back is simply evidence that the question now has a P&L attached.
People do careful work when the cost of judgment is visible on their own books. They get surprised when the meter belongs to someone else until the margin flips sign. Rented intelligence has unit economics. Open weights are now how serious operators answer part of that bill, while still paying frontier prices for the work that still needs them.
Sources: Bloomberg reporting via The Next Web (21 Sep 2026), “AI model costs are pushing startups towards cheaper open weights”; Harvey, “Harvey Tenet Research Preview” (20 Aug 2026); Reuters on Anthropic enterprise mix (Oct 2025); Ramp economist via PYMNTS on top-1% enterprise concentration; Anthropic and secondary run-rate commentary as attributed in market reporting. See also: Faster, cheaper tokens: why your AI bill won’t fall like you think.
Related articles
-
ai
AI-Native Services Are a Real Market. Enterprise Still Needs a Company Brain.
Greg Isenberg's $100B AI-native services thesis is right about the market. The operator question is whether the enterprise keeps a durable record under its agents. Without one, a thousand AI seats become a thousand silos.
-
ai
Enterprise AI Has Two Operating Models: Development Work and Knowledge Work
Enterprise AI is splitting into two operating models — development work and knowledge work. Why coding-copilot playbooks stall on operational workflows, and which engineering disciplines transfer.
-
ai
Tokenomics Is the New AI Efficiency Frontier — and Here's How We're Winning It
AI tokenomics is the discipline of managing token consumption at enterprise scale. Learn how semantic infrastructure, context-aware retrieval, and agent budgeting cut AI costs without sacrificing quality.