Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Atmospheric Evidence Discovery — A Live Labeler

draft first cut, 2026-09-21.

A signed record on its own only answers “who said this, when.” It doesn’t answer whether an acknowledgement of a claim is a genuine attestation or a pro-forma “nice work” that never actually engaged with what it’s acknowledging. That second question needs judgement, not just signatures — and breadcrumbs now has a live, self-hosted service that makes exactly that judgement, publishes it as a real signed label, and is running today at labeler.breadcrumbs.run.

What’s actually making the judgement

TypeSafe’s Jev is a “System One” model: given some state and a question, it returns a typed judgment or probability — not generated text, not a chat reply. Two of its primitives do the work here:

  • Noul — a bare probability that a condition holds, no separate confidence band. Used to ask: does this acknowledgement genuinely attest to the claim, or is it templated, bot-like, or not actually engaging with the specific thing being acknowledged?
  • Score — a probability-weighted position on a small set of described levels. Used to judge an evaluation’s own justification text against the claim it evaluates, rather than trusting a bare pre-normalized number with no visibility into why it landed where it did.

Both replace a cheap, gameable proxy — a raw acknowledgement count, a bare numeric score — with a judgement of the evidence’s actual content. Neither generates anything; they only judge what’s already there.

Three ways evidence could reach it, one way it’s live today

The full design covers three discovery pipelines feeding one judgement call — see the diagram for the complete picture, including which are built and which are still gaps:

  • Record track — a formal org.hypercerts.context.acknowledgement record on the acknowledging party’s own PDS. The shape breadcrumbs’ signal-reading logic already expects; real record parsing isn’t built yet.
  • Outbound engagement — someone’s own Bluesky post naming a crumb. Needs real full-text Bluesky search to discover, which doesn’t exist yet — a known, named gap, not an oversight.
  • Inbound reaction — likes, reposts, and replies landing on a trail record’s own AT-URI, discoverable via Constellation’s backlink index once trail records are themselves interactable content.

What’s live today is a fourth, more direct path that proves the judgement layer end to end rather than waiting on all three: a reply to one of breadcrumbs’ own Chatto-room confirmations (see Physical Crumbs for the badge-claiming flow this extends) gets judged by Jev, thresholded into one of two label values — genuine-acknowledgement or pro-forma-acknowledgement — and published as a real, signed com.atproto.label record. The raw probability stays visible in the reply for a human to see; it never becomes the label value itself. That distinction isn’t cosmetic: the AT Protocol label spec explicitly recommends against encoding scores or confidence values directly in a label, precisely because a label is meant to be drawn from a small, fixed, inspectable vocabulary — real deployments (Bluesky’s own Automod included) collapse continuous judgement down to a handful of named tokens before anything gets published, and labeler.breadcrumbs.run follows the same convention.

A real identity, not a bare API

Labels are signed, attributed records — src names the issuing DID, sig is a signature over the label, verifiable against that DID’s registered signing key. None of that works without a real, resolvable identity behind it, so labeler.breadcrumbs.run is exactly that: a real did:plc account, its handle verified, self-publishing its own app.bsky.labeler.service record the moment it starts up. Anyone — a Bluesky client, a future breadcrumbs viewer, anyone at all — can subscribe to it and independently verify what it says, the same way any other labeler on the network works. Building a bare, unsigned HTTP endpoint instead would have been less work and would have thrown away the actual point of using the label mechanism: native rendering in real atproto clients, and a verdict that isn’t just breadcrumbs’ own word for it.

What’s genuinely still open

  • Outbound and inbound discovery (above) aren’t built. The judgement layer is proven against one real evidence source; the wider atmospheric-discovery breadth the diagram describes is still ahead.
  • How a judged verdict becomes a trust score is deliberately unresolved, not quietly decided. Breadcrumbs’ trust bandit (arms.rs) picks a single winning signal per claim on purpose — never a blend — specifically so a later outcome can be attributed back to the one signal that was actually bet on. A verdict from this labeler can either strengthen that one signal directly, or become a separate, blended number shown alongside it (e.g. on a trail page) without ever feeding the bandit’s own reward calculation. Both are real options; neither is picked yet. The diagram marks this fork explicitly rather than presenting a settled answer.
  • Which failure modes get their own label, not just a yes/no. A badge’s issuance context, for instance, has at least three real outcomes worth distinguishing — plausible, self-dealt, reciprocal exchange — currently collapsed into one bare probability. Worth a richer judgement later, not solved here.

Why this is the concrete version of an accountability argument

Why This Fits AI for Science & Safety argues that as AI systems increasingly participate in research, the attestation layer has to identify and hold what made a claim accountable, not only verify humans. This labeler is that argument made real rather than aspirational: an AI judgement, individually attributable to a real signed identity, independently verifiable by anyone, publishing a narrow, inspectable vocabulary instead of an opaque internal score. The same discipline this project already applies to human claims and evidence, applied to what an AI system itself asserts.