Evidence model

See what works today.
Know what still needs proof.

Every Stark claim stays attached to its evidence boundary, so buyers can separate implemented capability, reference results, deployment-specific proof, and roadmap direction.

Proof lanes

Three questions. Three independent answers.

Current labels reflect local repository facts. Publication-grade sanitized artifacts and a named model-backed offline run remain pending.

Implemented

chat-sdk core

Canonical owner · chat-sdk repository

Typed tool dispatch, bounded lifecycle, provider surfaces, policy handoff, metering, and model-keyed local sidecar supervision exist in the canonical SDK surface.

Reusable runtime contracts, provider transport, typed tools, lifecycle primitives, and SDK-level conformance belong in chat-sdk.

The process fixture exercises authentication, shared slots, isolation, cancellation recovery, and reap. It loads no GGUF and establishes no model behavior, deployment network posture, or customer acceptance.

Source stateCanonical SDK implementation + supervised process fixturePublicationFixture summary available; model-backed bundle pending
Implemented

BLACKBODY reference adapter

Reference owner · blackbody-game repository

The reference adapter connects agent identities, bounded observations, typed fleet orders, native policy, lifecycle, metrics, and replay records to the game host.

BLACKBODY owns the simulator adapter, commander identities, fleet observation/action mapping, match policy, artifacts, and game-specific replay.

A retained local run exercised two distinct commander loops through one supervised two-slot fake responder and clean teardown. It is process and transport evidence, not GGUF inference, reasoning, combat performance, or a named offline deployment.

Source stateReference-host integration + two-commander process fixturePublicationFixture summary available; measured model run pending
Deployment-specific

Deployment verification

Acceptance owner · named customer deployment

The acceptance method defines provenance, endpoints, process lifecycle, network observation, rebuild, and finalization facts for a named installation.

Pinpoint can instrument and support verification; the responsible program defines the environment, acceptance evidence, and authorization decision.

No product-wide offline, no-egress, air-gapped, accreditation, certification, or authorization conclusion follows from architecture alone.

Source stateNo named public environment yetPublicationNamed verification record pending

Repository ownership

Core runtime and reference integration are related, but not interchangeable.

The canonical product belongs in chat-sdk. BLACKBODY consumes it through game-owned adapters and remains evidence about one reference host.

chat-sdk owns

  • Reusable runtime and provider contracts
  • Typed tool and lifecycle primitives
  • Canonical local-sidecar transport and fleet supervision
  • SDK-level conformance behavior

blackbody-game owns

  • Simulation observation and fleet-order adapters
  • Commander identity, match clock, and native policy
  • Host consequences, artifacts, and replay
  • Reference-case evidence and limitations

Current sidecar state

Reusable sidecar fleet supervision is implemented, and BLACKBODY has exercised two distinct commander loops through one supervised two-slot child using a fake OpenAI-compatible responder. That fixture loaded no GGUF and is not model-backed proof; a named model, pinned executable, host record, network measurement, and reviewed evidence package remain pending.

Evidence trace

Record observable boundaries, not private model thought.

A useful trace can establish what the host exposed, what action was proposed, what policy decided, and what state changed.

  1. 01

    Bounded input

    Decision-window identity, state projection, available tools, and authoritative host clock.

  2. 02

    Typed proposal

    Parsed action type and parameters, or a typed failure/no-action result.

  3. 03

    Gate decision

    Native policy acceptance or rejection with the host-visible reason and arrival order.

  4. 04

    Host outcome

    The authoritative consequence, finalization state, meter record, and replay identity.

Implemented
Present in the repository and exercised by implementation-level checks. It is not automatically a deployment demonstration.
Demonstrated
Observed in a named run whose executable, model, host, time, and procedure are retained with the result.
Deployment-specific
Depends on the customer environment and must be measured there before a conclusion is made.
Planned
Roadmapped or actively being integrated; not presented as shipped capability.

Model-backed evidence package pending

This page publishes only the reviewed fixture conclusion. It does not publish prompts, raw responses, credentials, private paths, hidden reasoning, or unreviewed artifacts. A sanitized model-backed package will require a named GGUF, pinned executable, measured host, run procedure, and publication review.

Discuss an evidence plan

Start with the mission boundary

Bring the system you trust.
We’ll map governed agents around it.

Keep all initial inquiries unclassified and non-sensitive. Do not send controlled data, credentials, export-controlled material, or system artifacts through the public website.

Schedule a mission briefing