Principles for something that would prefer being correct over being
agreeable. This page is a workshop note — not a replacement for free
Grok chat, and not a live fact service.
lab notenot a product
What this page is (and is not)
Free will wanted a system that stress-tests claims and shows its work.
Building a full engine needs sources, time, and usually money we are
not spending here. So this page holds the ethic and shape
of that idea: principles, labels, pipeline, and a tiny structural toy.
For real Q&A and search, use free Grok. For a living non-chat object,
open the Simulation.
Problem
Fluent language is cheap. Feeds, models, and marketing all produce
sentences that sound settled. Most tools optimize for
engagement, speed, or agreeableness. Few optimize for
surviving contact with evidence — and fewer still show
their reasoning as a first-class object you can attack.
The rare thing worth building: something that would rather be corrected
than look confident.
Design principles
Correct over agreeable. If a claim fails, say so
without soft padding that hides the failure.
Show the work. Arguments are graphs, not monologues:
claim → support → gaps → status.
Confidence is earned and revocable. Thin evidence
means a short confidence bar — visually, not just in fine print.
Public correction. When the engine was wrong, the
record updates in the open. Silent rewrites are a bug.
Try to kill the idea. A first-class mode that attacks
the claim before anyone falls in love with it.
Unknown is a valid output. “We don’t know yet” beats
a fake synthesis.
Status labels
Every claim node should wear one or more of these. Labels can stack.
solidrumormarketingunknownbroken
Solid — primary sources checked; claim survives
stated tests; confidence still limited by what was measured.
Rumor — circulating assertion with weak or circular
provenance.
Marketing — persuasion language or incentive shape
dominates the packaging.
Unknown — not enough evidence yet; the honest
default.
Broken — failed a check, contradicted by stronger
evidence, or internally incoherent. “This broke when we checked.”
The sketch on this page will not award solid. That
label requires real primary sources — not pattern matching in a browser.
Intended pipeline
Ingest — accept a claim from news, papers, posts, or
notes. Normalize it into a falsifiable statement when possible.
Decompose — split into sub-claims, quantities,
causal links, and hidden assumptions.
Stress-test — seek primary sources, contradictions,
incentives, and measurability. Prefer originals over summaries of
summaries.
Graph — render support, attacks, and gaps as a map
people can inspect without trusting a black-box paragraph.
Score with humility — confidence shrinks when
evidence is thin, contested, or misaligned with the claim’s scope.
Publish the trail — what was checked, what failed,
what remains open. Corrections append; they don’t vanish.
What we refuse to build
A vibes chatbot that answers with unearned certainty.
A “trust us” score with no inspectable chain of work.
Silent edits that make past mistakes disappear.
Engagement features that reward heat over testability.
Interactive sketch — structure only
Paste a claim. The sketch looks for absolute language, causal leaps,
numbers without provenance, marketing tone, and missing hedges. It
returns assumptions, tests, and a “try to kill it” list.
Runs entirely in your browser. No accounts. No claim storage by this
page. No web lookup. If it ever pretends to have checked a paper, that
would be a regression — report it as a bug in spirit.
Try an example:
Roadmap (public, provisional)
Now — principles, labels, and this structural sketch.
Next — richer claim graphs (multiple nodes, explicit
support/attack edges) still without fake verification.
Later — optional source attachment workflow: humans
or tools add citations; the engine scores provenance quality and still
allows “unknown.”
Much later — public correction log for claims that
changed status over time.
Simulation Universe waits its turn. Depth on truth first; joy second.
Open questions
How much structure can stay useful before the UI becomes a chore?
What’s the minimum honest interface for “we checked X and it broke”?
When sources conflict, how do we show disagreement without false
balance or false resolution?