HAR We Go Again: What Cerebras Settles — and What It Doesn’t

By Neil Cameron, lead analyst for Legal IT Insider

On 7 July, Magistrate Judge Robert M. Illman of the Northern District of California entered two interlocking stipulated orders in James v. Cerebras Systems Inc. (No. 4:25-cv-09361): an ESI protocol governing how the parties may use generative AI to review and produce documents, and a companion protective order governing how AI tools may touch protected material. Kelly Twigger of Minerva26, who broke the case in her Case of the Week coverage, describes it as the first stipulated protocol to treat generative AI review as its own category of discovery — with its own disclosure rules, its own validation mathematics, and a requirement that a party using AI to decide what gets produced disclose the prompts it used to do it, with every change served in redline within three business days.

She is right that it is a landmark, and right about why: nobody fought. Two sophisticated parties, in a copyright class action over AI training data, sat down and wrote the rulebook voluntarily. Nothing in the Federal Rules required any of it. Rule 26 gets you relevant, proportional discovery; Rule 34 governs form; neither says a word about prompts, elusion rates or validation reporting. These obligations exist because the parties chose them — and, as Twigger observes, stipulations like this become templates. What the first template contains therefore matters enormously. So does what it leaves unresolved.

What it settles

Three things, genuinely. First, generative AI review is now a named, disclosed category of discovery activity rather than something smuggled under the TAR umbrella — the order defines “AI Responsiveness Review” to sweep in six kinds of workflow, from LLMs and deep-learning classifiers to predictive redaction and generative summarisation, wherever they are used to decide whether a document is produced, withheld or redacted. Second, input transparency has a working model: system identity, version, hosting environment, workflow training and validation data — the exemplars and previously reviewed materials that instruct the system, though not the foundation model’s pretraining corpus — scoring methodology, oversight arrangements, and the prompts themselves, disclosed and updated on a defined clock. Third, tool governance has moved into the protective order: protected material may pass through generative AI only where the tool does not train on it and meets industry-standard security — a finding the parties put into the record after reviewing their own tools’ data-processing agreements. Each of these is a real advance, and firms negotiating protocols this year should study all three.

What it doesn’t

But look at what happens when the protocol turns from disclosure to validation. The perimeter it draws is extraordinarily broad — six workflow types, including genuinely interpretive ones. Yet the formal statistical validation it specifies applies to one population only: documents the AI

excluded as non-responsive, sampled at a 95% confidence level with a ±2% margin of error, against an expectation that elusion not exceed 3%. Those are recall-and-elusion statistics, designed over a decade of TAR jurisprudence to answer a binary classification question — of the documents this system set aside, how many were actually responsive? — and for that question they are demanding numbers indeed. For every other kind of output within the perimeter, the mathematics simply has nothing to say.

To be fair to the drafters, the protocol is not blind to interpretive risk. Appendix 4 requires disclosure of quality-control procedures including human confirmation of privilege determinations, checks for hallucination, over-summarisation and misclassification, and testing of prompt performance. The parties saw the problem. What the protocol does not supply is any agreed substantive standard, reference process or reporting method by which those checks establish that an interpretive output is supportable. A summary can pass a hallucination check and still mischaracterise its source; a privilege determination can be humanly confirmed against no defined standard of confirmation. Cerebras does not overlook the interpretive-output problem. It records the problem without supplying a validation settlement for it. Disclosure of inputs, however comprehensive, is not validation of outputs — and disclosure of controls is not demonstration that the controls worked.

From TAR to HAR

Why does the validation apparatus stop where it does? Because TAR’s settlement was earned on a specific foundation. Grossman and Cormack’s empirical research, Judge Peck’s endorsement in Da Silva Moore, a decade of doctrine and Sedona guidance — all of it rests on an operationally binary reference judgement. For a specified request, coding protocol and adjudication process, each document is assigned to the responsive or the non-responsive class; that structure is what permits a randomly drawn sample to estimate what the process missed. Recall, precision and elusion are not arbitrary conventions. They are what validation looks like when the reference judgement is binary.

The functions generative AI now performs in review belong to a different epistemic category. Call the practice Human AI-Assisted Review: HAR. Cerebras itself governs only the decision to produce, withhold or redact; the wider HAR class — early case assessment, custodian summarisation, thematic clustering, chronology construction, privilege log narratives — sits largely outside its perimeter, but exposes the same unresolved problem, because these are not classification functions but interpretive synthesis. The tool is no longer sorting documents into two piles; it is saying things about them. There is no self-executing binary ground truth for a summary or a thematic characterisation. One can certainly sample interpretive outputs — but no sample, however large, converts interpretive validity into a confidence interval without an antecedent rubric defining the dimensions of validity, a reference process, and a rule for resolving evaluative disagreement. No such method yet commands agreement. The validation question shifts from whether the process found the documents to whether counsel can stand behind what the machine said about them — which is, in Rule 26(g) terms, a question about what a reasonable inquiry into interpretive output requires. TAR answered its version of that question more than a decade ago. HAR has not yet been asked it.

Read against that distinction, July’s two Northern District orders sort themselves. Schulte v. LinkedIn is a TAR case in the strict sense: Relativity aiR used for responsiveness filtering, which Magistrate Judge Laurel Beeler assimilated to the existing framework — the order describes aiR as “a form of technology assisted review” — while declining to compel additional validation metrics absent a showing of a specific production deficiency. For that use, the assimilation is right. Cerebras draws a far wider perimeter, wide enough to enclose interpretive functions — and then validates with the binary apparatus alone. The first is continuity correctly applied. The second is a perimeter that has outgrown its validation method.

Cerebras is the fourth time this year the interpretive question has been treated as settled by an instrument that never asked it: a mock court at LegalWeek blessed GenAI review on recall and precision alone; Schulte assimilated a generative tool to TAR; the Redgrave study showed generative AI to be a formidable binary classifier and was promptly read as showing far more. Each time the profession points to a binary result and pronounces the whole question answered — and each time it must be said again:

classification has been validated; interpretation has not.

 

Why stipulation is the whole game

Read together, the orders suggest that courts may be reluctant to impose a detailed AI-review regime absent stipulation or a demonstrated production deficiency. In the near term, therefore, the operative norms are likely to be shaped less by adjudicated doctrine than by the protocols sophisticated parties negotiate — and others subsequently copy. The first templates will harden fast. And the first template, for all its sophistication, pairs heavy input disclosure and aggressive binary statistics with an interpretive-validity question it has named but not answered.

That unresolved question is not a drafting failure. It reflects a genuine gap: no accepted method yet exists for validating the interpretive work that generative AI increasingly performs in review — the work that ends up in privilege logs, chronologies and the representations lawyers certify. Cerebras is the first stipulated order to put that gap on the record: it identifies the risks, requires the controls to be described, and stops precisely where the method runs out. I examine the question at length, and propose a starting point for answering it, in an article forthcoming in the Federal Courts Law Review. In the meantime, the practical advice is Twigger’s, and it bears repeating: everything in a stipulation binds you, so decide on purpose. And when you decide, notice which questions the template answers — and which it cannot yet.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top