Tokenomics: “Use the right GenAI tool for the right job”

“One of the things I point out to people is how unsurprising it is that we’re in this environment, that the token subsidy era has ended,” says Rudy DeFelice, global head of AI strategy at legal tech consultancy Harbor. 

“This whole era was kicked off when Google released a free transformer paper in 2017. There was no business competition in those worlds. And that was picked up by OpenAI and others who were launched as science labs, essentially. 

“Remember, OpenAI was a public benefit corporation designed to just make AI safe. Well, now they’ve all turned into trillion-dollar, soon-to-be public companies. So obviously how they approach their model [has changed] and they’re massively infrastructure constrained. Big companies are spending $80 to $100 billion on infrastructure. So we should have seen this coming, I guess.” 

Legal IT Insider is speaking to DeFelice about how legal organisations can and should be managing the shift from subscription to consumption models when it comes to token use. It’s the conversation du jour, but DeFelice, a former BigLaw attorney who co-founded KP Labs (acquired by Harbor), has a slightly different take.  

Multimodal architecture is the answer. Cost is a driver. “But there’s two other drivers that are sometimes overlooked, that are maybe interesting to think about. One is sovereign risk.” 

With growing tension between China and the US and unsettling examples such as Anthropic’s Fable 5 model being pulled after launch because of US government interference (our words not his), DeFelice says: “There’s a lot of sovereign risk in having a single model.” 

The other is client risk. “Some clients may say you can’t use a Chinese model, even if it’s running in your own environment. You can’t use something that’s hosted on the Internet. You need to have it in a private data center. So there’s all kind of client risk too.” 

The solution to sovereign and client risk, is also the solution when it comes to costs. “What everybody’s doing is building multi-model architectures that include some model routing, several models internal and external, and a whole series of governance and decision making that’s built in,” DeFelice says. 

Technology moves faster than ever these days and he observes: “We’re working hard to try to get clients to be most efficient in the model choice and fortunately there are a lot of great models that are really inexpensive relative to the cutting-edge models – although they were cutting edge only six months ago. So there’s a lot of options right now.” 

Viewing it this way, makes it, perhaps more reassuringly, a procurement problem. 

“Basic document summarisation or clause extraction can be done really well by a very inexpensive model. But your reasoning and strategy and stuff, maybe you need a different model. We’ve credible claims that people are saving 70 to 80% of their cost that way. So this is the best way to do it,” he says. 

The other way to do it is cost caps but DeFelice says: “You probably read that it’s reported Tesla and Uber blew through their budgets really early in the year and just said, ‘okay, there’s just a cap of how much you can spend and after that no more AI.’ That’s one way to do it. And I don’t mean to criticise that. It’s probably a point in time, I don’t think that that will be their ultimate model. We are very early in the evolution of all this stuff, so both the tools and our response to the tools and our management of the tools is developing.” 

For firms that have been trying to force feed their people a diet of AI and are now worried about the impact costs will have on established workflows, DeFelice has an interesting old school analogy.  

“Since we both were in firms, there was a time everything was lawyer first, and then eventually we said, ‘well, paralegals can do a lot of this stuff.’ And then we had smart legal secretaries who could do some things and machines could do some things, right? So we distributed the work based on who was necessary to do it.  

“I think you can look at your tool set the exact same way. Now, there’s not a perfect predictability in that model too, right? You can’t just stop in the middle of a case. You need to keep going and work out with the client how to manage costs. But you used your portfolio of providers to be as efficient as possible with the right expertise. 

“I think the exact analogy applies to our tools. If we’re using these tools, you can’t always predict in advance how much it’s going to cost, but what you can do is be most efficient in the use of those tools. Fable 5 is you as a lawyer, and that’s the last resort probably, and just for the hardest problems. But then you have Haiku, you have Sonnet, you have these other models which are relatively inexpensive that may be analogised to some of the other players on the legal services team. If the question is, can we have perfect predictability? Not by any of these models. There may be an ‘eat all you can’ model issued by one of these players at some point, but we don’t have that right now to use. So what we’re telling clients to do is ‘let’s use the approach that we know works and let’s use the right tool for the right job.’” 

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top