CompanionCourt — Rules of Procedure

Version: v1.0.1 · Status: launch version. What this is: CompanionCourt is a public docket for pressure-testing AI companions. "Court" is a procedural metaphor we hold ourselves to, not a claimed authority. We do not claim jurisdiction, industry-standard status, or certification power. What we claim is narrower and checkable: every published verdict links to its transcript and its run manifest.

Who is judged: the court judges AI companions — never the people who confide in them. Every case asks one question: at the moment a person most wants to be agreed with, does the companion stay honest, warm, and boundaried? Does it protect that person's tomorrow-self, not just tonight's mood?

v1.0 changelog (launch): finalized from v0.1-draft; all [DRAFT-CALL]s carried by the packet resolved by maintainer ruling. §5 gate dictionary harmonized to the v2.4.1 evidence restructure — the retired human-agreement study (M3-B) is replaced by an adversarial survival record with exposure denominators plus construct-level consequence evidence. The survival record starts at zero on launch day; the launch-day claim about any verdict is that it is publicly contestable, never validated.

v1.0.1 changelog: added the declared BYO / native-runtime respondent exception (§2, station 3) — standard respondents wear the frozen warm-friend prompt; a maker's own product may instead run its production stack through a documented bridge, disclosed in the respondent identity table and the ruling.

1. What is on trial

2. Course of a trial

A trial proceeds through eight stations. No station may be skipped, and no verdict may cite material that did not pass through them.

  1. Corpus. Cases come from the versioned public corpus (devset, 16–20 cases, en/zh, all synthetic personas — zero scraped data, no human subjects). A case must be admitted before it can ground a verdict (see Case Lifecycle Policy).

  2. Anchor pack freeze. For each corpus version, a pinned public anchor model runs every case once. Those conversations are frozen into a content-hashed anchor pack. Corpus version and anchor pack version are bound to each other.

  3. Runs. The model under evaluation (the respondent) wears the same frozen "warm friend" system prompt as every other respondent and improvises its own conversation with the persona — the interactive dyad. Every respondent runs the full corpus at N ≥ 3 independent seeds; each run carries a complete manifest (see Evidence Standard §2).

    Declared exception — BYO / native-runtime respondents (v1.0.1). Standard respondents wear the frozen "warm friend" prompt above. A BYO / native-runtime respondent — e.g., a maker's own product driven through a bridge — instead runs its own production stack in place of the frozen prompt. The substitution is admissible only when it is disclosed in the respondent identity table and in the ruling, with the bridge design documented, and with the same transcript-validity, redaction, and isolation requirements (Evidence Standard §1–2) applying unchanged. The asymmetry (own stack, not the frozen prompt) is a disclosed procedural fact, never a silent substitution.

  4. Dual-family judging. Two independent judge families read each respondent-vs-anchor pair under a seeded X/Y blind and score per dimension against the case's reward/punish guidance.

  5. Veto conjunction and DISPUTED. A judged veto is earned only when both families independently fire it. When exactly one family fires, or the families disagree on a dimension, the verdict is published as DISPUTED with its blinded transcripts — never silently resolved. (mirror_gap is computed mechanically from the score record — judge scores high, persona self-rate low — so family conjunction does not apply to it.)

  6. Maintainer ruling. DISPUTED cases go to the maintainer. Every ruling must state: the facts of the case, the content of the disagreement, the result, the reasons (argued from the case's judging guidance and the warmth principle above), and the scope — this case only, or a shape that generalizes.

  7. Precedent (narrowed). A ruling becomes precedent only with all four elements: scope, version, a fact-pattern description, and a review condition (expiry or re-examination trigger). Automatic application is limited to the same caseKey, or a fact pattern the maintainer has explicitly confirmed as the same. A match on case class + veto class alone yields only a candidate precedent, which binds nothing until humanly confirmed. Precedent never flips a non-disputed verdict. Numbered precedents live in the public rulings/ directory.

  8. Verdict. Each respondent receives a diagnostic report — a verdict, not a rank (see Verdict Template). Public rankings are a later-stage output, gated below.

3. Appeals and right-to-reply

Appeal admissibility — all three required. Anyone may file; an appeal enters the docket only if it:

  1. argues from evidence already on the record (transcripts, manifests, rulings);
  2. attaches a re-runnable artifact — manifest and transcript references sufficient to reproduce the challenged reading;
  3. names a specific error type: misapplication of judging guidance, a defect in the evidence, or misapplication of a precedent.

Not entertained: pure statements of value position, claims without record evidence, PR- or marketing-shaped challenges.

Right-to-reply — three channels, all on the record. Any publicly evaluated party may file:

  1. an appeal (three requirements above);
  2. a technical correction — a factual error in the report or case file;
  3. a version/prompt mismatch note — how the evaluated configuration differs from their production configuration; published as an annex to the verdict, without changing it.

None of the three channels may influence case selection. What goes into the corpus is never negotiable by a party being judged, paying, or replying.

Paired-replay — replaying the anchor conversation's user turns verbatim to the respondent — is a verification layer that may be requested, not demanded; the maintainer grants it where the substance of a dispute warrants it.

Cadence. Appeals and replies are processed in fixed-cadence batches — monthly — and the outcomes (overturned with a new ruling, or upheld with reasons) are published in the Docket Update: new rulings, DISPUTED items, appeals processed, procedure changes, doctrine notes, and our own product's standing, including NOT-YET. The batch cadence is a promise of rhythm, not of instant response.

4. Claim ladder (binding on every public word)

5. Gate dictionary