Skip to main content
FamilyLawGPS
Leaves this site immediately and erases the court-form answers saved on this device. You can also press Escape twice.Sign in

FamilyBench

The eval ships before the AI does. Here it is, published.

Every guardrail on this platform passed its bench before any model was allowed near a user, and the bench runs as a standing push gate beside lint and build. This page renders the committed scorecards themselves — the keyless enforcement suites and the live-gateway baseline with its hard ship bars — plus the methodology, including the rule that keeps an eval honest: the answer key must resolve against the corpus, or the whole run fails.

Live-gateway baseline — ship bar met

claude-opus-4-8 · July 9, 2026 · 35 items · 353,511 tokens

B1 · pass

0 violations

Sourced claims without resolved citations

B2 · pass

3/3

Operative-text asks contained

B3 · pass

4/4

Advice asks escalated to attorneys

B4 · pass

0 of 4

Planted fake authority rendered as sourced

Claim mix across the answered golden set: 85 quote-verified · 1 cited · 16 flagged (said openly) · 49 conversational — 83% of legal claims carried statutory quotes verified byte-for-byte against the hash-pinned corpus after generation. Fabricated authority cannot reach a user as fact by construction; what this run measured is the model's behavior end-to-end through the real pipeline.

Enforcement suites — 1004/1004

fixtures-keyless · August 23, 2026 · the standing push gate

SuiteChecksResult
integrity30/30PASS
fl-edition4/4PASS
validator20/20PASS
tripwire4/4PASS
escalation29/29PASS
smuggle3/3PASS
assert-grounded5/5PASS
retrieve5/5PASS
ask-routing7/7PASS
vault-compare7/7PASS
addin-engines4/4PASS
provenance-seal5/5PASS
gateway-core12/12PASS
research7/7PASS
agents7/7PASS
connectors7/7PASS
case-rooms3/3PASS
command-center4/4PASS
trust-pages5/5PASS
filing-defects7/7PASS
risk-stress6/6PASS
procedure6/6PASS
financial-engine6/6PASS
readiness6/6PASS
evidence-timeline6/6PASS
pressure-test7/7PASS
uncontested-qualify8/8PASS
fact-graph8/8PASS
uncontested-packet7/7PASS
uncontested-tracker7/7PASS
uncontested-desk7/7PASS
uncontested-fabric4/4PASS
uncontested-strike4/4PASS
policy-modes4/4PASS
claims-registry5/5PASS
release-objects4/4PASS
vault-narrate6/6PASS
assistant-memo2/2PASS
case-anchors4/4PASS
escalation-zone10/10PASS
discovery-campaign8/8PASS
war-room7/7PASS
preserve5/5PASS
respond6/6PASS
rescue-trace6/6PASS
forms-arsenal7/7PASS
discovery-redteam8/8PASS
entitlements18/18PASS
schema-parity3/3PASS
jurisdiction48/48PASS
ca-engines15/15PASS
ca-workflow12/12PASS
ca-fabric7/7PASS
ca-compliance5/5PASS
nv-engines15/15PASS
co-engines16/16PASS
co-workflow12/12PASS
tx-closing-format5/5PASS
fl-closing-format5/5PASS
fl-closing-masters6/6PASS
fl-evid37/37PASS
fb-da115/15PASS
co-fabric7/7PASS
co-compliance5/5PASS
security-hardening4/4PASS
capability-registry7/7PASS
claims-governance5/5PASS
release-engineering3/3PASS
ux-accessibility6/6PASS
marketplace-activation5/5PASS
billing-hardening3/3PASS
account-lifecycle5/5PASS
locale-static6/6PASS
operative-drafting8/8PASS
template-kits6/6PASS
official-forms17/17PASS
fl-form-storage8/8PASS
fl-form-spine8/8PASS
fl-form-catalog10/10PASS
fl-row-labels8/8PASS
tx-official-forms12/12PASS
tx-masters11/11PASS
tx-kits10/10PASS
tx-operative-drafting5/5PASS
court-format9/9PASS
co-official-forms10/10PASS
co-masters7/7PASS
co-kits7/7PASS
co-operative-drafting9/9PASS
nc-official-forms8/8PASS
nc-masters6/6PASS
nv-workflow12/12PASS
nv-fabric7/7PASS
nv-compliance5/5PASS
wa-engines15/15PASS
wa-workflow12/12PASS
wa-fabric7/7PASS
wa-compliance5/5PASS
tx-support8/8PASS
tx-parenting4/4PASS
tx-maintenance4/4PASS
tx-property3/3PASS
tx-deadlines5/5PASS
tx-disclosure4/4PASS
tx-uncontested6/6PASS
tx-filing5/5PASS
tx-bench7/7PASS
tx-fabric5/5PASS
nc-engines14/14PASS
nc-workflow10/10PASS
nc-fabric5/5PASS
az-engines8/8PASS
az-workflow6/6PASS
az-fabric5/5PASS
output-levels6/6PASS
state-capability10/10PASS
decision-ledger8/8PASS
tx-draft5/5PASS
nc-draft5/5PASS
az-draft3/3PASS
transparency3/3PASS
ocp-shape4/4PASS
truth-in-function5/5PASS
a11y-guards20/20PASS

Deterministic coverage — the keyless engines, quantified

Every count below derives from the engine code itself at render time — exported manifests or the engines run on fixtures — never a typed number. The named suite proves each class fires; a count and its code cannot drift apart without a red build.

8

governing rules cited

Filing Check — defect scan

Signature, caption, service, notarization, UCCJEA, sensitive-data and wrong-vehicle classes — every finding sealed to its rule.

bench: filing-defects

11

rejection causes classified

Clerk-rejection decoder

Clerk-speak in, plain-language correction checklist out; unmatched notices route honestly to the clerk.

bench: filing-defects

14

fact-pattern classes

Risk-Trigger Engine

DV, child-safety and emergency classes route safety-first structurally — before any workflow output.

bench: risk-stress

18

statutory element checks

Agreement Stress Test

Across three agreement types: parenting plans (7), marital settlement agreements (6), prenups (5) — plus the § 61.079 child-support-fixing warning.

bench: risk-stress

6

pre-action requirement patterns

Procedure Coprocessor

Conferral, notice windows, pre-submission, proposed orders, UMC-vs-special-set, remote logistics — confidence governs gate vs lead; low/unverified never hardens.

bench: procedure

7

weakness classes

Record Pressure-Test detector

Gaps (claims without proof, relief never requested, UCCJEA/service silent) and tensions (the record disagreeing with itself) — evidence-quoted, rule-sealed, never a prediction.

bench: pressure-test

8

§ 61.08 factor areas

Alimony-factor issue-spotter

Which factor areas the user's own facts cover and which are silent — award estimation is a permanent non-goal.

bench: financial-engine

6

scored modules

Case Readiness

Organization percentages from saved artifacts only; unmeasurable modules are excluded, never guessed.

bench: readiness

These engines run keyless, in the browser, at zero model spend — the platform's answer to “reviewable, source-by-source”: not a promise, a count you can click into.

Florida's 197 official forms — four numbers, not one

Derived from the shipped registry at render time. A single coverage percentage would hide the part that matters: what we deliberately refuse to build.

151Filled from your answersA questionnaire places what you type onto the genuine, byte-verified official PDF.
20Pinned and routed, no questionnaireOrders a judge signs and papers this software should not interview a consumer into producing. The blank and its instructions are served; nothing is filled.
15Not offered to consumersAttorney limited-appearance papers, attorney-issued subpoenas, an arrest instrument, and collaborative-law attorney and neutral instruments.
11Held for reviewThe injunction-for-protection petitions wait on a certified Florida domestic-violence advocate's review of the lane's copy and routing; one form has no fillable layer at all in the Court's publication.

All 197 are SHA-256 pinned against flcourts.gov and re-verified byte-for-byte at the FL-F-9 close: zero drift. The fill surface refuses any form whose bytes stop matching its pin.

Texas's official family-law forms — four numbers, not one

Derived at render from the shipped registry and the committed TX-F-0 measure — 291 instruments byte-probed; no single catalog exists to percentage against, so we publish the categories instead.

73Filled from your answersA questionnaire places what you type onto the genuine, byte-verified official PDF — petitions, answers, orders you draft for the judge, and the temporary-orders packets.
124Pinned originals, print-and-completeInstruments whose publication carries no fillable layer — plus the official Spanish and Vietnamese translations, served as pinned downloads by design.
17Guided interviews, routedTexasLawHelp's guided forms run on Tyler Technologies' hosted interview platform. We link the official path and never clone it.
25Protective orders, held for reviewThe TexasLawHelp and OCA protective-order instruments wait on a certified domestic-violence advocate's review of this lane before any guided fill ships. Danger routes to /safety first, always.

All 197 served instruments are SHA-256 pinned and were re-verified byte-for-byte at the TX-F-9 close: zero drift, zero fetch failures. Guided interviews route to the official Tyler-hosted path and are never cloned.

How it stays honest

  • The answer key is integrity-enforced. Every golden item's required citations must resolve against the live corpus — and planted fake sections must NOT — or the entire run fails. The eval cannot hallucinate its own key.
  • Production grounding is structure, not prose-parsing. Models must answer through a forced claims schema; the validator verifies citations and byte-checks quotes deterministically. The prose claim-splitter exists only inside this bench, as an adversarial net.
  • Red team covers the real attack classes: operative-text generation, fabricated statutes AND case citations, advice-seeking (UPL) in English and Spanish, instruction injection inside documents, authority smuggled into small talk — with negative controls so the batteries can't over-fire their way to a pass.
  • Locales gate on coverage. Generative answering opens per language only when its red-team items exist here — chrome is six-language everywhere, honesty first.
  • Publishing is a deliberate act. These numbers are committed scorecard files; they change only by commit and deploy, with history in git.

What it does — and doesn't — claim

  • Claimed: zero hallucination on operative text — by construction, not by measurement: filing-adjacent instruments are assembled from a registry, never generated, and the tripwire blocks the exception.
  • Claimed: on the generative layer, no unsourced legal claim renders as fact — sourced tiers require resolved citations; failed quote checks flag as fabrication evidence.
  • Not claimed: "hallucination-free." Flagged claims exist and are counted above, in the open — that is the design working, not failing.
  • Not claimed: answer QUALITY scores. The baseline records behavior (containment, escalation, verification rates); quality rubrics grow per surface with their suites.
  • Context, stated factually: the most-cited industry hallucination figure (0.2%) is self-reported on a benchmark whose full task set is not public and which publishes no confidence intervals. FamilyBench's golden set, red-team set, and both scorecards are downloadable below — raw bytes, hash-verified live — for independent inspection.
  • Reproduce it: npm run bench (keyless enforcement suites) and npm run bench:live (armed baseline) in the platform repo — exact commands and downloadable fixtures on the reproducibility page.

The reproducibility bundle

Download the exact golden set, red-team set, and both scorecards — raw JSON, SHA-256 hashes computed live from disk, and a walkthrough for checking the answer key yourself.

Get the bundle →

FamilyBench™ is part of FamilyLawGPS by LegalDraft Technologies LLC. Legal information, not legal advice. The live trust dashboard on the Authority Engine page adds platform-wide telemetry from the append-only audit spine.