Exam Readiness: How Do You Know You’re Actually Ready to Sit?
Am I ready? Every candidate asks it in the last fortnight before an exam, and the usual answer is a number. Score 80% on practice tests and you are good to go. That rule is everywhere. As far as I can establish, nobody has ever published evidence for it.
The harder problem sits underneath the number. Judging your own readiness means judging your own knowledge, and people are measurably bad at that in specific, well-documented ways.
Why can't you just ask yourself?
Karpicke and Roediger (2008) ran a vocabulary learning study that is usually cited for the testing effect. It contains a second result that gets less attention. Students were asked to predict how much they would remember a week later. Their predictions bore no relationship to what they actually recalled. The people who went on to remember 80% and the people who went on to remember a third were not distinguishable by their own forecasts.
Koriat and Bjork (2005) explain part of the mechanism. When you judge how well you know something, you do it while the answer is in front of you. At the exam the answer is absent and you have to produce it. They call the resulting error a foresight bias. In their paired-associate experiments, items where studying activated a strong link that would not be available at test produced predictions of 75.7 against actual recall of 60.3.
Be careful how far that result travels. The authors are explicit that overconfidence is not a general property of self-assessment. In their words, "By and large, JOLs do not exhibit an overconfidence bias and, in fact, for many of the items used in this study, JOLs were very well calibrated." For a different set of items in the same experiment, predictions of 78.1 sat against recall of 78.8. The bias is selective. It shows up when the conditions of study flatter you in a way the exam will not.
The third finding is the one that matters most for exam timing. Koriat, Bjork, Sheffer and Bar (2004) asked people to predict their recall when they expected to be tested immediately, after a day, or after a week. The predictions barely moved. Actual recall fell steeply. At the one-week delay participants predicted better than 50% and recalled under 20%. People can estimate what they know now. They do not spontaneously subtract what they are going to forget between now and the exam.
That is the core of the readiness problem. Your sense of readiness is a reading taken today, and you sit the exam in three weeks.
What is wrong with the 80% rule?
Three things, and they compound.
First, a practice score is a fact about a question bank, not about the exam. Most certification exams report a scaled score rather than a percentage. CompTIA passes Security+ at 750 on a 100 to 900 scale and does not publish how many questions that takes. Scoring 80% in an app does not convert into 750.
Second, nobody has calibrated the bank against the real cut score. A third-party question bank has not been through a standard-setting process. Its difficulty is whatever its authors chose, so 80% on a hard bank and 80% on a soft one mean different things.
Third, a single score is a single sample. Sit the same bank on a different day and you will get a different number. One reading tells you very little about the next one. I have written about score prediction in more detail in why you fail practice tests but pass the real exam.
What actually counts as evidence you are ready?
Better signals exist. None of them is a single percentage.
- You produced the answer, rather than recognised it. Multiple choice lets you work backwards from the options. If you can state the answer before you read the choices, you know it. If you can only pick it out of a line-up, you may not.
- You got it right on material you have not touched for a fortnight. Recent study inflates everything. Performance on cold material is closer to what exam day will look like, which is the whole argument for spacing your reviews out.
- You scored consistently across several sessions. Three sessions in the same range beats one good session by a distance. The variation between your sessions is itself information.
- You covered the whole blueprint, not the comfortable parts. Most candidates drift toward domains they enjoy. Check your weakest domain separately, because a composite score hides it.
- You were right when you felt sure. Being correct matters. So does whether your confidence tracked your correctness.
What does confidence calibration measure?
Calibration is the gap between how sure you felt and how often you were right. Ask for a confidence rating on every question and you end up with four groups instead of two.
Confident and correct is settled knowledge. Unsure and wrong is a known gap, and it is the easiest kind to fix because you already know it is there. The two mixed cases carry the information. Unsure and correct usually means a lucky guess or a half-remembered rule, and it will not survive a harder version of the same question. Confident and wrong is the dangerous one. You will not revise it, because you have no reason to think anything is broken.
A plain accuracy score merges all four. Two candidates at 78% can be in completely different positions, one of them mostly confident and right, the other carrying a pile of confident errors.
Every Meridian Labs app asks how confident you were on each question and reports those groups back to you. That is a measurement, not a promise. It tells you where your self-assessment and your performance disagree. Whether it raises anyone's pass rate is untested. What the research above supports is the narrower point that unaided self-assessment is unreliable, so measuring the disagreement beats trusting the feeling.
Calibration is one of four inputs to the readiness score those apps display. Accuracy carries 35% of it, coverage of the blueprint 30%, calibrated confidence 20%, and consistency of study 15%. The Exam Ready certificate sits behind further gates, among them answering about 80% of the question bank, 80% overall accuracy, a minimum in every domain, and a two-week study streak.
Those weights are a design judgement. They were not derived from exam outcomes, and they could not have been. Progress is stored on your device and never sent anywhere, so no dataset exists linking what anyone did in an app to whether they passed. Read the number as a structured summary of several signals, not as a forecast.
How do you run a readiness check?
Something like this, over about a week.
Take a full-length timed session under exam conditions. No notes, no pausing, no looking anything up. Record the score and set it aside, because the single number is the least useful part.
Then break the result down by domain and find your worst one. Work that domain for several days and leave the rest alone. Take a second full session and compare the domain breakdown rather than the totals. What you want to see is the weak domain moving and the others holding.
Last, go back through the questions you got right while feeling unsure. Try to answer each one again from memory without the options in front of you. The ones you cannot reconstruct were never secure.
What can a readiness check not tell you?
It cannot give you a probability of passing. Nobody outside the testing body can, because the cut score, the item difficulties and the scaling are not public. Any app or site that hands you a pass likelihood is modelling its own bank, not the exam.
Exam day is outside its reach. Time pressure, an unfamiliar centre and nerves all cost something, and the amount varies by person.
It cannot tell you about the questions you never see. A bank samples the blueprint, and your result is a statement about that sample.
What a readiness check gives you instead is a better class of information than a feeling. The feeling is uncorrelated with performance, indifferent to how long you have left, and inflated exactly where study conditions flattered you.
Readiness Tracking in Every Meridian Labs App
All 28 apps ask for a confidence rating on every question and report accuracy and calibration by domain, so you can see where your self-assessment and your results disagree.