Auto AI AgencyAll apps

Trivia · free · no sign-up

GutCheck: ten statements, one confidence dial, and the gap between them

A free daily calibration test. You call ten true/false statements and set a confidence dial on each; GutCheck grades the distance between how sure you said you were and how often you were right. No account, one set a day.

Play Gutcheck →

01  What GutCheck actually measures

GutCheck is a free daily test of calibration — the match between the confidence you feel and how often you turn out to be right. Every UTC day it deals the same ten true/false statements to everyone playing. For each one you make a call, true or false, then set a dial for how sure you are: 50% for a coin flip up to 100% for dead certain, in 5% steps. Ten statements, about a minute, then a verdict.

That verdict is a page of instruments rather than a score out of ten: a Brier-based calibration figure from 0 to 100, your accuracy printed next to your average confidence, a curve bucketing your ten answers by how sure you said you were, and a one-word archetype. The curve is the part most quizzes lack: it shows where your judgement started to drift from reality.

One design decision does most of the work. The app never tells you whether an individual answer was right while you play; on screen it says "We grade your calibration at the end — not each answer." You cannot adjust the next nine dials to compensate for the first, and there is no back button and no re-answering. Knowing more facts helps, but what is being measured is whether your certainty is priced correctly.

02  The four verdicts, and how each one is earned

The archetype comes from a single number: your mean stated confidence minus your accuracy, in percentage points. Average 80% sure and get six of ten right and the gap is plus twenty.

The rules, as the code applies them: an Overconfident Oracle is a run at least 15 points over. A Humble Genius is underconfident by at least 10 points while still right at least 70% of the time — someone who knew more than they were willing to claim. Well-Calibrated is a gap within 10 points either way. Wildcard is checked first and is the odd one out: it needs your average confidence on the answers you got wrong to be at least 5 points higher than on the ones you got right, the signature of being sure of the wrong things and hesitant about the right ones.

The captions stay dry rather than scolding — the Overconfident Oracle reads "Sure — and wrong more often than you thought." Because the Wildcard test runs first, a small-looking gap can still be a Wildcard if the confidence points the wrong way; every run lands on exactly one of the four.

03  Why your score is not your accuracy

The 0-100 figure comes from a Brier score: the average squared distance between the confidence you claimed on each statement and what actually happened — 1 if you were right, 0 if you were wrong. Claiming 100% and getting it wrong costs a full 1.0; claiming 50% costs 0.25 either way. That average is then rescaled: 0.25 — the score of someone who answered 50% on everything, a shrug priced as a coin flip — is the zero point, and a perfect 0 becomes 100.

The arithmetic, rather than a slogan: eight out of ten with the dial pinned at 100% scores 20 and is branded an Overconfident Oracle. The same eight out of ten, with the dial at an honest 80%, scores 36 and lands Well-Calibrated. Ten out of ten at 100% scores 100.

The sharpest case is the hedger. Five out of ten with the dial parked at 50% is mathematically perfectly calibrated — mean confidence 50%, accuracy 50%, gap zero — and the app will hand you a Well-Calibrated verdict for it. The score is 0, because that is exactly where the scale is anchored. A hedged five can wear a better verdict than a cocksure eight while scoring lower: the label and the number measure different things.

There is no trick to a high score except being right and having said so. Admitting a guess is never punished beyond the coin-flip floor, so 50% is free to say when it is true — but replacing a real opinion with it removes information, and removing information cannot raise the score.

04  One set a day, one attempt, and a streak that notices

Today’s ten statements are chosen from the UTC date alone, so everyone playing gets the identical set — which is what makes comparing with a friend meaningful rather than approximate. The bank holds 360 statements across nine areas, from geography and history to language, maths, nature, technology and the arts. The ordering is re-shuffled every 36 days, so the whole bank is used once before any statement returns.

You get one attempt per UTC day. Finish the run and the same page later opens straight onto your saved verdict instead of dealing a fresh set. The lock lives in your own browser, and the project’s own notes are candid that it is best-effort: a stored flag can be cleared by anyone who wants to. Defending a free daily quiz with accounts would cost more than the lock is worth.

The streak counts consecutive UTC days — play today and tomorrow and it extends, skip one and it restarts at one. Because the rollover is UTC midnight, the next set unlocks in the late afternoon or evening across the Americas (around 5pm Pacific, 8pm Eastern in summer) and in the small hours in Europe, and the share row counts that moment down to the second.

05  Playing it well

The useful advice is mechanical, not factual: the score grades your certainty, not your general knowledge.

  • Use the whole slider. Reserve 100% for the statements you would genuinely bet on; a dial that never leaves 70–80% is a dial that is not saying anything.
  • Treat 50% as "no idea", and use it without embarrassment. It is the most forgiving answer in the scoring, and calling a guess a guess is the honest move.
  • Read the curve, not just the headline score. It shows your accuracy inside each confidence bucket, so a 90% bucket that was right 60% of the time points at the level where you are overstating yourself.
  • Do not draw conclusions from a single day. Ten statements is a small sample, and some buckets hold one answer or none — the curve reads the run you just finished, not you.
  • Expect to be overconfident in the subjects you like most. Familiarity is what makes a wrong answer feel certain.

06  Sharing a result, and the one number that is not yours

The verdict is drawn as a 1200×630 card — archetype, score, accuracy against confidence, streak, date and the day’s set number. The share row offers the system share sheet where the browser has one, a copy-link fallback, and a downloadable PNG. The link is one self-contained URL holding the whole result: no database write, no account needed to open someone else’s, and it lands read-only on that card with a button to take today’s set yourself.

That URL is generated in the browser and unsigned, so a shared verdict records what one browser computed rather than certifying it. The one number that is not self-reported is the crowd line. After a run the app submits only an archetype and a date to a single global counter, and prints something like "only 8% of players landed on Wildcard too" only once at least twenty results exist. Below that threshold, or if the request fails, it shows the verdict and invents nothing.

07  Questions people ask

Is GutCheck free, and do I need an account?
Free, and no account. The page is ad-supported and the run happens entirely in your browser. Your streak and saved verdict live in that browser’s local storage, which also means they do not follow you to another device and clearing your browsing data clears them. Nothing you type is collected, because there is nothing to type — the only things that leave the page are anonymous funnel events and your archetype counted into the crowd tally.
How is the calibration score calculated?
It is a Brier score — the average squared gap between the confidence you claimed and what actually happened — rescaled to read 0 to 100. Answering 50% on everything gives exactly 0, because a coin flip priced as a coin flip is the zero point; a perfect run of ten at 100% gives 100. Eight out of ten at 100% scores 20, and the same eight out of ten at an honest 80% scores 36.
Can I play more than once a day?
No — there is one set of ten per UTC day, and once you have finished it the page opens onto your saved verdict instead of a new run. The next set unlocks at UTC midnight, which is the late afternoon or early evening in the Americas and the small hours in Europe.
Does everyone get the same statements?
Yes. The set is a function of the UTC date alone, so every player worldwide sees the same ten statements on the same day, which is what makes comparing verdicts with a friend fair. The bank holds 360 statements across nine subject areas and is re-shuffled every 36 days, so it is used up once before anything repeats within a cycle.
Can I beat it by answering "true" to everything?
Partly, and the numbers are worth knowing. 284 of the bank’s 360 statements are true, so a blanket "true" is right about 79% of the time. With the dial set near 80% that averages a score in the low thirties — but a single day’s set runs from four to ten true statements in practice, so the same tactic scores 0 on a false-heavy day and only reaches the eighties when the set happens to be all true. It also earns a Well-Calibrated verdict without telling you anything about your confidence, which is the thing being measured.
Is this a real psychological test?
No, and it does not pretend to be. GutCheck is entertainment built around a real scoring rule — Brier scores are standard practice in forecasting — played over ten statements, which is far too small a sample to support a verdict about a person. Take the archetype in good humour and read the curve for the interesting part.
Open Gutcheck →

08  More from the shelf

  • Hivemind — Don't answer. Guess how everyone else did.
  • Baloney — Three facts. One's fake. Sniff it out.
  • Brainrank — 10 questions. Where do you rank?
  • Isobath — Drop a pin. Find out how wrong you were.

Browse every app in the catalogue →

Last reviewed 2026-09-17.