How it works
A language model left to itself writes the answer first and supplies citations afterwards. JuriScripta will not draft a sentence until it has the opinions in hand, and it will not hand you the draft until a model from a different company has checked it against them.
The run
Grounded — the opinion's text supports the claim.
Not supported — the case is real, but it does not say what the draft says it does. This is the error a citation check alone never catches.
Could not verify — not enough of the opinion's text was available to decide. Reported as its own count, never folded into the grounded ones.
The totals always appear together. A score that hides how many claims were checked is easy to flatter.
If no real authority can be confirmed for a question, the run stops and tells you that. It does not compose a plausible answer around a citation that does not exist. That refusal is the system working.
How we measure
Every brief already carries its own measurement: each claim, its verdict, and the counts, published with the draft. That is the figure that matters for the brief in front of you.
For the system as a whole, we built a 115-question benchmark on the question taxonomy Stanford researchers published for testing legal AI. We will publish results when a full run completes under a protocol written down and dated before it starts, stating what was counted and out of how many. Until then we would rather show you the method than quote a number from a partial run.
Stanford has not tested JuriScripta. The benchmark is our own work, following their published methodology.
The limits