Is AI marking accurate enough to trust?
It is a fair question and it deserves a straight answer rather than a statistic. This page sets out how the grade is actually produced, what the engine deliberately refuses to do, and where its limits are.
The rule the engine is built on
One rule shapes everything else: the part of the system that decides your grade never calls an AI model.
AI reads your answer and describes it. AI writes the feedback afterwards. In between, the grade is computed by ordinary code from those descriptions and the exam board's published weightings. No prompt, no model, nothing that can be talked round.
This is a design rule. A language model asked to produce a grade is doing pattern completion on text that looks like marking, and it will hand the same essay a 7 and then a 5 depending on how the request was phrased. Taking the model out of that one step is what makes the rest defensible.
What the grade is made of
Every grade carries its own audit trail. When you open a marked answer you can see each of these.
| Recorded | What it tells you |
|---|---|
| Level matched | Which band of the board's published level descriptors the response met. |
| Mark band | The real mark range that level carries on the board's own mark scheme. |
| Descriptor wording | The published sentence the response was judged against, not a paraphrase. |
| AO breakdown | How the assessment objectives were weighted for that board, paper and question type. |
| What prevents a higher grade | The specific thing standing between this response and the next band. |
The point of publishing the reasoning is that you do not have to take the number on faith. If the engine has misread something, the trail shows you where.
What the engine refuses to do
Some of the design is about what it will not produce. These are enforced as rules that run on every change, only what the words on the page do.
| Situation | What happens |
|---|---|
| No verified answer key for a question | Reported as unmarked and excluded from the total. Never scored zero. |
| Answer appears copied from the source text | Detected as copying and graded accordingly, rather than rewarded for reading fluently. |
| Answer does not address the question | Two independent relevance signals must agree before the grade is floored. One alone only flags. |
| Marks would exceed the question's tariff | Impossible by construction; marks are bounded to the real maximum on every path. |
| A paper's content could not be extracted properly | The paper is flagged rather than published with questions missing. |
Determinism, and why it matters
Submit the same answer twice and you get the same grade twice. That sounds like a small thing until you have used a tool that does not do it.
If marking drifts, none of the feedback means anything. You cannot tell whether last week's 5 and this week's 6 reflect real improvement or the model having a different morning. Progress tracking built on a drifting grade is decoration.
Because the grading step is arithmetic over stored measurements, reproducing a grade is not a matter of luck. Each grade is also stamped with the version of the rules that produced it, so an old result can be understood in terms of the engine that actually marked it.
The limits, stated plainly
A page arguing for trustworthiness that only listed strengths would not deserve any.
There is no examiner-agreement figure. We do not publish one because we cannot support one. Measuring it honestly needs a substantial body of real scripts marked independently by two examiners, which is not something we have been able to obtain. Any competitor quoting a precise percentage should be asked where the scripts came from, who marked them and when.
Coverage is uneven. Where a board publishes a full level-descriptor grid, grading is anchored to it. Where a board does not, the engine falls back to assessment-objective weightings alone, which is a weaker signal.
Creative writing is the hardest case. Judging originality is genuinely difficult, and the engine is more cautious there by design.
It is not your teacher. It marks what is in front of it. It does not know what you were aiming for, and it cannot tell you that this is the best thing you have written all year.
Common questions
- Is AI marking accurate?
- Ask a sharper question: accurate against what? There is no published figure for how often Marked agrees with a human examiner, because measuring that honestly needs a body of real double-marked scripts that we do not have and will not fabricate. What we can show is what the grade was built from: the level it matched, the mark band it came from, the descriptor wording behind it, and the measurements that positioned it. You can check the reasoning rather than trust a number.
- Is it as good as a teacher?
- No, and it is not trying to be. A teacher knows you, knows what you did last term and knows what you are capable of on a good day. Marked does one narrow thing: it applies a mark scheme to the words in front of it, at two in the morning, for the twentieth essay this week, without getting tired. Those are different jobs.
- Will it give the same answer the same grade twice?
- Yes. The step that decides the grade is arithmetic over stored measurements, so identical input gives identical output. That is checked automatically on every change, never assumed.
- What stops it being fooled by fluent waffle?
- Fluency and correctness are scored separately, and relevance to the question is measured by two independent signals that must agree before an answer is marked down as off-topic. An answer copied out of the source text is detected as copied rather than rewarded for reading well. A confident paragraph that does not answer the question will not score as though it did.
- Which exam boards does the engine know?
- AQA, Edexcel (Pearson), OCR and WJEC/Eduqas, for English Language and English Literature. Each has its own assessment-objective weightings and, where the board publishes them, its own level descriptors.
For the mechanics of a single submission, see how AI marking works.
Put one answer through it
The free tier needs no card. Paste an answer you have already written and see the grade, the mark band and the reasoning behind both.