JuriScripta

How it works

Law First, AI Second

A language model left to itself writes the answer first and supplies citations afterwards. JuriScripta will not draft a sentence until it has the opinions in hand, and it will not hand you the draft until a model from a different company has checked it against them.

The run

  1. Identify candidate authorities A model proposes the cases likely to govern the question. At this point they are only names.
  2. Confirm each one in the court record Every proposed citation is looked up in CourtListener. One that cannot be found is dropped — and listed, so you can see what was refused.
  3. Find later opinions A search for decisions that cite those cases, including ones filed after the model's training data ends.
  4. Read the opinions The full text of each confirmed opinion is fetched. This is the only material the draft may use.
  5. Draft Written from the fetched text. A citation that was not confirmed cannot appear.
  6. Review, twice A quality review, then a peer review by a model from a different provider, which does not share the first one's blind spots. The draft is revised against both.
  7. Check every claim Each statement the draft makes about a case is tested against that case's text and given a verdict.

Three verdicts, all of them shown.

Grounded — the opinion's text supports the claim.

Not supported — the case is real, but it does not say what the draft says it does. This is the error a citation check alone never catches.

Could not verify — not enough of the opinion's text was available to decide. Reported as its own count, never folded into the grounded ones.

The totals always appear together. A score that hides how many claims were checked is easy to flatter.

When it finds nothing, it says so.

If no real authority can be confirmed for a question, the run stops and tells you that. It does not compose a plausible answer around a citation that does not exist. That refusal is the system working.

How we measure

No accuracy number yet, on purpose.

Every brief already carries its own measurement: each claim, its verdict, and the counts, published with the draft. That is the figure that matters for the brief in front of you.

For the system as a whole, we built a 115-question benchmark on the question taxonomy Stanford researchers published for testing legal AI. We will publish results when a full run completes under a protocol written down and dated before it starts, stating what was counted and out of how many. Until then we would rather show you the method than quote a number from a partial run.

Stanford has not tested JuriScripta. The benchmark is our own work, following their published methodology.

The limits

What this does and does not establish.

Request access