Every AI marking tool competing for UK schools' attention claims to be accurate. So we went looking for the published evidence behind those claims — the correlations, the error rates, the methodologies — for every major tool on the market, and recorded exactly what we found. Much of the table is empty — and much of what isn't doesn't survive a second reading. This audit sets out, vendor by vendor, what is actually published, what is claimed but not evidenced, and what that gap should mean for your procurement decision.
A note on bias, before anything else. We build one of the tools audited here (Top Marks AI), so you should read our findings with appropriate scepticism — and hold our section to a harder standard than any other. Every claim below about a competitor is limited to what their own published material did or did not contain when we checked it on 18 August 2026, and we've included a standing invitation at the end: if any vendor publishes evidence that meets the bar described here, we will update this page and say so.
The AI marking market has produced a wave of "best tools" listicles in 2026 — most of them written by vendors, including us. What the market has not produced is a single page that answers the question a head of department actually needs answered before trusting a machine with a student's mock grade: what evidence exists that this tool's marks are right?
Our guide for school leaders argues that the first question to ask any vendor is "where is your published accuracy data?" This audit simply asks that question of the whole market at once — including ourselves — and writes down the answers. The six tools were chosen for visibility, not convenience: they are the tools marketed to UK schools and students for GCSE and A Level marking that buyers most often meet in this year's comparison articles and search results; adjacent tools aimed at universities or vocational training are out of scope. Tools are listed alphabetically. Nothing here is inferred from demos or hearsay: every finding comes from what each vendor has published on its own website, checked on 18 August 2026.
Marking accuracy has an established measurement tradition, because exam boards have spent decades checking human markers against senior examiners. The standard instruments are simple:
Equally important is what does not count. "Aligned to AQA, Edexcel and OCR mark schemes" describes an input — the mark scheme was given to the model — not an outcome. "Trained on real GCSE answers" is a statement about how the system was built, not about how well it performs. "Accuracy within normal human variation" is a claim shaped like a result, but without a number attached it cannot be checked, compared, or falsified. And "saves teachers 50% of marking time" is a workload finding — valuable, but it measures speed, not correctness.
GradeDrive is a UK-focused platform for marking handwritten exam papers, including a "Level of Response" mode aimed at extended writing — which puts it squarely in the territory where accuracy evidence matters most.
Its published material argues, reasonably, that AI marking should be measured against human marking rather than an imagined perfect answer, and states that its marking falls "within normal human inter-rater variation on the large majority of responses". Its accuracy article goes further, explicitly declining to give figures on the grounds that an honest answer is "more useful than a headline percentage". Yet its homepage carries exactly such a headline: a stat tile reading "98% accuracy vs manual marking", with no methodology, sample size, dataset, or definition of "accuracy" published anywhere on the site — including in its own write-up of how the platform was tested, which names the exam boards covered but contains no numbers. A figure that cannot be interpreted — 98% of what, measured how, against whom? — sits in tension with the article's own, better argument.
On the substance, we'd put the counter-argument this way: it is true that two experienced teachers won't always agree — that is precisely why the discipline benchmarks against chief-examiner standardisation marks, and why the known human baseline (~0.65 correlation, Fowles 2009) exists as the number to beat. The answer to an imperfect human baseline is to publish your distance from the standard, not to decline to measure it — and not to headline a number without showing the measurement.
GradeOrbit is a UK tool for marking handwritten GCSE and A Level mock papers: teachers photograph or scan scripts, select the exam board, level, and subject, and the system transcribes and marks against the relevant scheme.
Its marketing describes AI marking as "exceptionally, undeniably accurate" for technical elements such as spelling, punctuation and factual content, and explains the approach as semantic understanding of rubric language rather than keyword matching. The only accuracy numbers on its site are generic: a claim that "the best AI marking tools" achieve exact grade matches around 60–70% of the time with the rest within one grade boundary — presented as an industry expectation, uncited, and never as a measured GradeOrbit result. Beyond that, its accuracy articles invite teachers to trial the platform and judge for themselves, even proposing a do-it-yourself parallel-marking test. A trial is genuinely good advice — our procurement guide recommends exactly that — but an invitation to self-test is a substitute for evidence, not a form of it. A school running a ten-script trial cannot detect the difference between a 0.65-correlation tool and a 0.90 one; that takes standardisation materials and a proper sample, which is the vendor's job to have done first.
MarkMe is a student-facing revision product: pupils submit GCSE practice answers (typed or photographed) and receive instant marks and examiner-style feedback. It supports AQA, Edexcel and OCR, with a focus on essay-based humanities subjects. It's a different category from school-deployed marking infrastructure — a student practising at home, not a department processing mock scripts — and should be judged as such.
Its site's accuracy language is qualitative — "accurate marking, tailored to your exam boards" — and its fuller claims appear in syndicated directory listings and its own replies to reviewers, which describe models "fine-tuned on real student responses" and trained to follow UK exam-board mark schemes. All of that may well be true — fine-tuning on real answers is a sensible approach — but these are statements about how the system was built. We found no published figures showing how the resulting marks compare with examiner marks, which matters even for a revision tool: a student calibrating their revision against a marker of unknown accuracy is navigating with an uncalibrated compass.
PaperAce is a student-facing GCSE revision platform: AI essay marking with instant grades and examiner-style feedback, alongside practice questions, timed mocks, predicted grades, and revision plans. It covers AQA, Edexcel, OCR and Eduqas across 20+ subjects, and reports over 10,000 student users. Like MarkMe, it belongs to the revision category rather than school-deployed marking infrastructure.
Its core accuracy claim, from its own FAQ, is "We train on real AQA, Edexcel, OCR and Eduqas mark schemes" — which, per the decoder below, describes an input, not an outcome. To its credit, its own small print is unusually plain: the platform "uses AI to simulate GCSE marking for revision guidance only", and its terms add that AI marking "does not guarantee the same result as official examiner marking". That is an honest description of an unbenchmarked tool. We found no published accuracy figures, no methodology, and no independent validation; the evidence offered is student testimonials about grade improvements, which are outcomes of revision, not measurements of marking accuracy.
ReMarkAble AI offers instant essay feedback aligned to AQA, Edexcel, OCR and WJEC mark schemes, including handwritten submission via OCR, with a free GCSE essay marker aimed at students as well as teachers.
To its credit, its published material is more candid than most: its guide for parents states plainly that AI marking "is not infallible, and it is not a replacement for teacher judgement", and frames the category as sufficiently accurate for formative feedback on practice attempts rather than high-stakes assessment. That candour is worth acknowledging — it is the right way to describe an unbenchmarked tool. The accuracy numbers its articles do quote are all category-level statistics from external, unlinked research — a trial of eleven AI models on 150 AQA GCSE scripts said to have "matched human grading" while halving marking time, an unnamed study's "0.94 correlation with official grades", a "0.7–0.85" correlation range for structured essays — none of them a measurement of ReMarkAble's own marking. (Its own reference list attributes the eleven-model trial to a comparison study by Marking.ai — a different marking vendor.) For its stated formative use-case the missing first-party evidence is less dangerous than it would be for summative marking; it does mean a school cannot currently know how close its marks land to an examiner's.
Top Marks AI is a school- and MAT-facing marking platform with 400+ individually calibrated tools across 40+ subjects, each benchmarked against exam-board standardisation materials before release. Because this is our product, this section applies the audit's standard at its strictest — here is what we publish, and, just as importantly, what we don't.
What we publish: more than thirty per-question accuracy studies on our accuracy blog, each with correlation, error, and sample described. Headline figures: 0.94 Pearson correlation on AQA GCSE English Language, 0.91 on OCR English Literature, and 0.90 on Edexcel IGCSE English, with ~84% of marks within board tolerance against ~45% for experienced human markers (Fowles, 2009). In a head-to-head test on 51 Edexcel A Level Politics standardisation essays, our calibrated tool achieved 0.84 correlation with a mean absolute error of 2.55 marks, against 0.48 and 5.0 for a rubric-on-an-LLM competitor.
Independent verification: our accuracy findings have been independently corroborated by Ark Schools — one of the UK's largest multi-academy trusts — and by Community Schools Trust: organisations that ran their own checks on our marking rather than taking our word for it. This is the third-party verification this audit asks of every vendor, and as of August 2026 we are the only tool on this page whose accuracy has been independently checked against exam-board standardisation materials. A study on AQA GCSE Shakespeare essays found 93% agreement with human markers across 30 handwritten scripts.
Where we fall short of the ideal: our studies are conducted and published in-house rather than in peer-reviewed journals. The independent corroboration from Ark and Community Schools Trust closes much of that gap — it is exactly the external check a school should demand — but peer-reviewed publication remains a higher bar still, and until we clear it we won't claim to have.
One row per tool, alphabetically; every cell reflects what was publicly available on 18 August 2026. "None found" means exactly that — we looked and could not find it, and we will correct any cell a vendor can show to be wrong.
| Tool | Category | Numeric Accuracy Figures | Methodology Described | Independent Validation |
|---|---|---|---|---|
| GradeDrive | School — handwritten exam marking | Headline "98%" tile, unexplained; no correlation/MAE | No | None found |
| GradeOrbit | School — handwritten mock marking | None found | No | None found |
| MarkMe | Student — GCSE revision | None found | Training approach only | None found |
| PaperAce | Student — GCSE revision | None found | No | None found |
| ReMarkAble AI | Student/teacher — formative essay feedback | None for own marking | No | None found |
| Top Marks AI | School/MAT — calibrated batch marking | Yes — Pearson, MAE, tolerance, 30+ studies | Yes, per study | Ark Schools, Community Schools Trust |
None of the phrases below is dishonest. Each is simply an answer to a different question than the one you're asking. When you see one, translate it — then ask for the number.
A standing invitation. This page describes the state of published evidence on 18 August 2026. If you are one of the vendors named here and you publish accuracy data — correlations, error rates, tolerance percentages, with sample and method described — email us at info@topmarks.ai and we will update this page to reflect it, with a dated correction note. The market improves fastest when publishing evidence becomes the norm, and we would genuinely rather compete on data than on its absence.
This audit is deliberately narrow, and it is worth being explicit about why. Accuracy is not the only criterion that matters when choosing AI marking. The ability to mark at scale — whole class sets, handwritten scripts, results flowing into your MIS — is a real criterion. The pedagogical quality of the feedback is a real criterion. The ability to manage adoption as a department, a school, or a trust is a real criterion. We audit accuracy first, and alone, because it gates all of them: batch marking doesn't dilute a marking error, it multiplies it by the size of the cohort; a beautifully scaffolded piece of feedback justifying the wrong mark is confidently wrong pedagogy; and a trust-level dashboard is only as useful as the marks flowing into it. Evidence of accuracy is the entry ticket — the other criteria decide the game among tools that hold one.
For those criteria, we've written the comparisons separately: our comparison of the wider tool market assesses handwriting support, batch marking, and MIS integration; our guide for multi-academy trusts covers procurement, governance, and rollout at trust level; and the ScaMP framework sets out the pedagogy behind feedback that is scaffolded, modelled, and precise rather than merely generated.
If you are choosing an AI marking tool this term, the audit reduces to a simple procedure. First, decide what the marks will be used for: formative practice feedback tolerates uncertainty that target-setting, moderation, and reports to parents do not. Second, for anything of consequence, ask each candidate vendor the question this page asked: where are your published figures against board standardisation materials? Third, run a real pilot with blind moderation — our guide for school leaders sets out the full framework, and our comparison of the wider tool market covers the general-purpose options (ChatGPT, Grammarly, Gradescope and others) not audited here.
And hold us to it too. Our numbers are published precisely so that you can check them against your own scripts before believing them.
Every accuracy claim in this audit about Top Marks AI links to a published study. Browse them all, or book a demo and we'll walk through the data for your subjects — and show you how to run the same evaluation on any vendor.
As of August 2026, of the AI marking tools most visible to UK schools, only Top Marks AI publishes accuracy data benchmarked against exam-board standardisation materials — Pearson correlations, mean absolute error, and tolerance rates across 30+ per-question studies, independently corroborated by Ark Schools and Community Schools Trust. GradeOrbit, MarkMe, PaperAce, and ReMarkAble AI publish no measured accuracy figures for their own marking; GradeDrive's only figure is an unexplained "98%" headline claim on its homepage.
As of 18 August 2026, GradeDrive's homepage displays one headline figure — "98% accuracy vs manual marking" — with no methodology, sample size, or definition published anywhere on its site, while its own accuracy article argues against headline percentages. GradeOrbit publishes no measured accuracy figures for its tool; the only accuracy numbers on its site are an uncited claim that "the best AI marking tools" achieve 60–70% exact grade matches. Neither vendor publishes correlations, error rates, sample sizes, or benchmarking methodology. If either publishes such data, we will update this audit to reflect it.
ReMarkAble AI's published position — set out in its guide for parents — is that AI marking "is not infallible", is not a replacement for teacher judgement, and is sufficiently accurate for formative feedback on practice attempts rather than high-stakes assessment: a candid framing for a tool without published benchmarks. As of August 2026 it publishes no correlations or error rates for its own marking — the figures its articles quote come from external, category-level research — so its distance from examiner standards cannot currently be verified either way.
The benchmark is the human baseline: research on experienced GCSE examiners (Fowles, 2009) found typical correlations of around 0.65 with chief-examiner marks, with only ~45% of marks within board tolerance. A credible AI marking tool should publish figures at or above that baseline against board standardisation materials. Top Marks AI's published studies show correlations of 0.90–0.94 across Humanities subjects with ~84% of marks within tolerance.
No. Time saved is a workload result, not an accuracy result — a tool can halve marking time while producing marks far from examiner standards. Workload and accuracy should be evidenced separately: a time-saving figure tells you the tool is fast, and only a correlation or error-rate figure benchmarked against examiner marks tells you it is right.
Ask for published Pearson correlations, mean absolute error, and percentage-within-tolerance benchmarked against exam-board standardisation materials, with sample sizes and methodology described — then run a pilot with blind moderation: have teachers and the AI mark the same held-back scripts independently and compare both against each other and the mark scheme. If a vendor can supply neither published figures nor a credible pilot protocol, treat their accuracy claims as unverified.
We use cookies for analytics and marketing to improve your experience — these are only set if you accept. Decline and we'll only use cookies that are strictly necessary. (Live chat is always available either way.) Learn more in our Cookie Policy.