Guest post: The AI bill many law firms haven’t priced in

By James Harrison

Profile image

 

 

 

 

Most firms buying legal AI are buying on a per-seat licence. That unit is on its way out.

Legora has moved its Agent Pro product to consumption pricing and Harvey’s co-founder, Gabe Pereyra, has been direct about where that leads: “customers are going to start getting these consumption bills of $10 million, and they’re going to ask, ‘What did my agent do that cost me $10 million?’” Harvey reports that token usage is up fourteenfold in six months. Pereyra cites Uber burning through its annual coding token allocation in just four months.

These numbers are easy to check. At today’s rate, a drafting query will cost around £15 worth of tokens. Reviewing 100,000 contracts runs to around £15,000.

A seat licence allows access and an agent consumes compute. Those two stop being the same product when a tool is allowed to run unsupervised across a dataset, and every vendor roadmap I’ve found so far points that way.

I have never known law firm boards this engaged in what is ostensibly an IT transformation. It’s a different level of interest compared with PMS and case management projects. With AI, they can’t get enough.

The reasons are not mysterious, and I keep hearing the same three things from peers, interviewers and recruiters alike: clients are demanding it, there is a great deal of FOMO in the market, and lawyers understand that they are knowledge workers and are therefore exposed. None of those is a bad reason to act. However, all three reward speed and none of them rewards good cost discipline.

The economics behind the meter

The frontier labs are not charging what inference costs them. OpenAI is projected to burn $14bn in 2026, up from $8–9bn in 2025. Anthropic moved from roughly -94% margins in 2024 to +40% in 2025 and remains under pressure. Microsoft supplies OpenAI compute below market rate. These are the conditions of a land grab, financed by investors who will want their money back. May Habib, chief executive of the enterprise AI firm Writer, put it plainly: these companies will go public and they will raise prices, because they have to.

This is now a well-established playbook. Uber entered cities at fares no operator could sustain and covered the gap with investor money. It waited for the market to reorganise around it and then put prices up. The subsidy was customer acquisition, not generosity. The people who paid for it were the ones who had already sold the car.

Legal AI is early in the same sequence. Today’s pricing is set by vendors competing for reference clients in a market that will consolidate, not by what the compute costs. A firm rebuilding its operating model around that pricing is selling the car.

Uber, incidentally, is now on the other side of it. The client Pereyra describes burning a year’s token allocation in four months is Uber’s own engineering team.

There is a real counterargument that also needs stating. Per-token prices have fallen steadily as inference gets more efficient. Competition between labs is genuine. Open-weight models put a ceiling on what anyone can charge for commodity work. On unit price alone, the pessimists may be wrong.

However, that is not where the exposure sits. Unit price and total spend are different variables. Cheaper tokens make it rational to use more of them. Harvey’s fourteenfold increase is not a pricing story, it is a consumption story. A firm that halves its cost per token and multiplies its volume by fourteen has a larger bill and a more difficult conversation with its board. This is Jevons paradox. Jevons, a Victorian economist, noticed that as steam engines got better at burning coal, Britain burned more of it, not less. Efficiency made coal cheaper to use, so people found more uses for it.

So, the planning assumption should not be “AI will get cheaper.” It should be “our AI bill will grow, it will become variable, and someone will have to justify it.”

Start with the boring work

The industry conversation has settled on drafting and contract review. Both are legitimate uses, but both are also where the work is highest-value, most partner-visible and hardest to verify at scale, which makes them the worst place to learn what this technology costs and where it fails.

It makes far more sense to start where volume is high, legal risk is low, and the work is still very manual. In most firms that means file opening, conflict checks, AML and KYC document handling, medical records sorting and chronology building, disclosure de-duplication and first-pass relevance, bundling and pagination, court form population, time-narrative cleanup against client billing guidelines, and inbound enquiry triage.

None of that is sexy, but all of it is measurable. The cost per unit is known today because a paid employee is currently doing it, which means a baseline exists before any AI is deployed. Errors surface quickly and are recoverable. And the work is repetitive enough that consumption per output is stable and forecastable, which is exactly what you need before signing anything metered.

Drafting and review then become the second wave, deployed by a team that has already learned what its own token curve looks like (or will look like). Even if the product is not on a token model yet, you should plan and use it as if it were.

There is also a governance argument for the same sequencing. These processes sit on the firm’s most sensitive data, but in the least structured part of its estate. Classification, retention and access controls must be right before an agent is pointed at them. Doing that work on file-opening and disclosure first is cheaper than doing it on live matter drafting.

Nobody knows if it works

The measurement gap is documented, although the best data is American. BARBRI published research in August, based on ten leaders across nine US firms from Am Law 100 down to smaller tech-led practices. Firms rated their own technology rollouts seven out of ten. They can identify who has activated a tool, but almost none can identify who has changed how they work, and not a single firm in the study had an AI competency framework for its associate pipeline.

There is no reason to think UK firms are any further ahead. The tooling is the same, the vendors are the same, and the billable hour remains the problem in both markets.

And that gap matters more once the meter starts running. Licence spend is a fixed line cost that was probably approved a year ago. Every firm likely knows its E3, E5 or even E7 cost to the penny, but even Microsoft has now split the model. A user-based Copilot licence is a fixed cost. Copilot’s agents, however, are billed in credits, pooled at tenant level and metered from June this year. The means the most predictable vendor has already moved the goalposts, and most firms probably aren’t aware because the credits come out of a pool that was purchased in advance.

Useful measures may not be sexy, but they are sensible:

· Matter profitability by type, measured before the tool goes in and after it has bedded in. Not time saved in the abstract, and not self-reported hours.

· Consumption per completed output. How many tokens it takes to produce one finished chronology, one bundle, one conflicts check. This is the number that tells you whether a use case survives a price rise, because you can double it and see whether the work still makes sense.

· Rework and escalation rate. How often the output needs a second pass, and whether that pass is done by someone more expensive than the person the tool replaced. Work that moves from a paralegal’s desk to an associate’s is not a saving, it is a transfer.

· Cycle time on the step the tool touched, not the matter end to end. Matter duration is driven by the other side, the court list and the client, none of which your tool controls.

· Write-off rate on the work the tool touched.

The billable hour sits behind all of this and the BARBRI work names it as the primary obstacle. An efficiency gain with no pricing change never reaches the P&L. Firms that have not decided where AI savings go, whether to margin, to capacity, or back to the client as a lower price, will not be able to answer the board’s question when the invoice arrives.

There is a second structural problem sitting next to that one. Capital investment reduces the profit distributed to equity partners in the year it is made, which is why firms have preferred to rent AI as a subscription rather than own any part of it. That preference was rational while the subscription was cheap and fixed. It becomes expensive once the subscription turns into a meter.

Buy yourself room

The contractual position is where a CIO can still act.

Price protection is the first ask. Rate cards should be fixed for a defined term, with notice periods on repricing long enough to migrate, as well as a cap on year-on-year increases. Vendors currently competing for logos will concede more now than they will in eighteen months.

Track AI spend against the client and matter from day one, the same way you already track a disbursement. If you cannot say which client, which matter and which task a cost belongs to, you cannot bill it on, you cannot knowingly decide to absorb it, and you cannot justify it if it is challenged on assessment.

Keep the model layer replaceable. The commercial risk is concentration, not capability. Once your prompts, templates, matter history and workflows live inside a vendor’s platform, your switching cost is a number they can estimate. That number, and not the market, sets your renewal. They know what they can charge you before they tell you what they’re charging you.

Latham & Watkins has gone further than anyone. The FT reported on the 11th of September that the firm has bought its own Nvidia GPU servers and is fine-tuning open-weight Nemotron models in leased data centre space only its own staff can access. Its chief information officer, Rene Mendoza, gave two reasons. Some client information is too sensitive to put with any cloud vendor. And with “a lot of consumption costs coming” in how AI usage is priced, the firm wanted flexibility. His summary of the strategy is the argument in one sentence: “We are not hitching our wagon to one particular company.”

Most firms cannot follow. The FT puts the cost of running hardware at that scale in the tens of millions of dollars a year, and Latham employs more than 900 technology specialists against $8.3bn of revenue. Artificial Lawyer has argued the economics do not stack up below a certain size, and at UK revenue levels that is probably right. Latham is also the outlier rather than the trend. Kirkland is working with Palantir, A&O Shearman with Harvey, Freshfields with Anthropic. The signal still matters more than the strategy. The largest buyer in this market has looked at where token pricing is going and decided it is worth owning the hardware.

You need to work out whether each use case pays for itself before you scale it. Anything that is marginal at today’s token prices is a loss at tomorrow.

And settle the question of who pays with clients now. Pereyra thinks firms will start billing AI costs back: “Law firms already pass through costs for other legal tech so it’s not a completely crazy idea.” Clients who are themselves buying AI will have views on that, and those conversations go better before the charge appears on a bill than after.

Where I stand

This is not an argument against AI in legal services. The capability is real, the direction is settled, and firms that sit it out will be at a structural disadvantage.

It is an argument about sequencing and cost discipline. The current price of AI reflects a competitive land grab and not the true cost of production. Firms deploying against today’s pricing, on today’s seat licences, with no way of telling which client or matter the spend belongs to, and no settled position on who pays, are underwriting a variable they have not modelled.

Buy the boring use cases first. Measure matter profitability, not how many people switched the tool on. And above all, factor in an inevitable and hefty price rise now.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top