Atmospheric Evidence Discovery — A Live Labeler
draft first cut, 2026-09-21.
A signed record on its own only answers “who said this, when.” It
doesn’t answer whether an acknowledgement of a claim is a genuine
attestation or a pro-forma “nice work” that never actually engaged with
what it’s acknowledging. That second question needs judgement, not just
signatures — and breadcrumbs now has a live, self-hosted service that
makes exactly that judgement, publishes it as a real signed label, and
is running today at labeler.breadcrumbs.run.
What’s actually making the judgement
TypeSafe’s Jev is a “System One” model: given some state and a question, it returns a typed judgment or probability — not generated text, not a chat reply. Two of its primitives do the work here:
- Noul — a bare probability that a condition holds, no separate confidence band. Used to ask: does this acknowledgement genuinely attest to the claim, or is it templated, bot-like, or not actually engaging with the specific thing being acknowledged?
- Score — a probability-weighted position on a small set of described levels. Used to judge an evaluation’s own justification text against the claim it evaluates, rather than trusting a bare pre-normalized number with no visibility into why it landed where it did.
Both replace a cheap, gameable proxy — a raw acknowledgement count, a bare numeric score — with a judgement of the evidence’s actual content. Neither generates anything; they only judge what’s already there.
Three ways evidence could reach it, one way it’s live today
The full design covers three discovery pipelines feeding one judgement call — see the diagram for the complete picture, including which are built and which are still gaps:
- Record track — a formal
org.hypercerts.context.acknowledgementrecord on the acknowledging party’s own PDS. The shape breadcrumbs’ signal-reading logic already expects; real record parsing isn’t built yet. - Outbound engagement — someone’s own Bluesky post naming a crumb. Needs real full-text Bluesky search to discover, which doesn’t exist yet — a known, named gap, not an oversight.
- Inbound reaction — likes, reposts, and replies landing on a trail record’s own AT-URI, discoverable via Constellation’s backlink index once trail records are themselves interactable content.
What’s live today is a fourth, more direct path that proves the
judgement layer end to end rather than waiting on all three: a reply to
one of breadcrumbs’ own Chatto-room confirmations (see
Physical Crumbs for the badge-claiming flow this
extends) gets judged by Jev, thresholded into one of two label values —
genuine-acknowledgement or pro-forma-acknowledgement — and published
as a real, signed com.atproto.label record. The raw probability stays
visible in the reply for a human to see; it never becomes the label
value itself. That distinction isn’t cosmetic: the AT Protocol label
spec explicitly recommends against encoding scores or confidence values
directly in a label, precisely because a label is meant to be drawn from
a small, fixed, inspectable vocabulary — real deployments (Bluesky’s own
Automod included) collapse continuous judgement down to a handful of
named tokens before anything gets published, and labeler.breadcrumbs.run
follows the same convention.
A real identity, not a bare API
Labels are signed, attributed records — src names the issuing DID,
sig is a signature over the label, verifiable against that DID’s
registered signing key. None of that works without a real, resolvable
identity behind it, so labeler.breadcrumbs.run is exactly that: a real
did:plc account, its handle verified, self-publishing its own
app.bsky.labeler.service record the moment it starts up. Anyone —
a Bluesky client, a future breadcrumbs viewer, anyone at all — can
subscribe to it and independently verify what it says, the same way any
other labeler on the network works. Building a bare, unsigned HTTP
endpoint instead would have been less work and would have thrown away
the actual point of using the label mechanism: native rendering in real
atproto clients, and a verdict that isn’t just breadcrumbs’ own word for
it.
What’s genuinely still open
- Outbound and inbound discovery (above) aren’t built. The judgement layer is proven against one real evidence source; the wider atmospheric-discovery breadth the diagram describes is still ahead.
- How a judged verdict becomes a trust score is deliberately
unresolved, not quietly decided. Breadcrumbs’ trust bandit
(
arms.rs) picks a single winning signal per claim on purpose — never a blend — specifically so a later outcome can be attributed back to the one signal that was actually bet on. A verdict from this labeler can either strengthen that one signal directly, or become a separate, blended number shown alongside it (e.g. on a trail page) without ever feeding the bandit’s own reward calculation. Both are real options; neither is picked yet. The diagram marks this fork explicitly rather than presenting a settled answer. - Which failure modes get their own label, not just a yes/no. A badge’s issuance context, for instance, has at least three real outcomes worth distinguishing — plausible, self-dealt, reciprocal exchange — currently collapsed into one bare probability. Worth a richer judgement later, not solved here.
Why this is the concrete version of an accountability argument
Why This Fits AI for Science & Safety argues that as AI systems increasingly participate in research, the attestation layer has to identify and hold what made a claim accountable, not only verify humans. This labeler is that argument made real rather than aspirational: an AI judgement, individually attributable to a real signed identity, independently verifiable by anyone, publishing a narrow, inspectable vocabulary instead of an opaque internal score. The same discipline this project already applies to human claims and evidence, applied to what an AI system itself asserts.