Every paper graded by the AI displays a reliability score, expressed as a percentage. It helps you quickly spot papers that deserve a closer review.
What the score measures
The reliability score combines two dimensions:
Transcription confidence: how confident the AI is that it correctly read what the student wrote (handwriting legibility, scan quality).
Grading confidence: how confident the AI is in its evaluation relative to the grading scale or competency grid.
Where it appears
The reliability score is available at two levels:
At the whole-paper level, for a quick overview.
At the individual question level, to precisely target the answers to check before validating the grading.
How to use it
A high score generally means you can trust the suggested grading. A lower score doesn't mean the grading is wrong, but that it deserves your attention: hard-to-decipher handwriting, an ambiguous answer, or a borderline case relative to the grading scale.
In practice, a good approach is to sort or browse your papers by ascending reliability score, so you focus your review where it adds the most value, rather than rechecking every question one by one.