Top Marks AI is an AI marker built for the essays UK schools actually set: GCSE and A Level responses for AQA, Edexcel, OCR, Eduqas and WJEC, handwritten or typed, marked one script or a whole class set at a time. Every one of its 400+ marking tools is calibrated against exam-board standardisation materials before release, so the mark it gives tracks what a senior examiner would give, and the feedback tells the student how to move up a band.
Pearson correlation with examiners
AQA GCSE English Language
Marks within board tolerance
vs ~45% for experienced human markers
An AI marker reads a student's response and produces a mark and written feedback against a mark scheme. The important distinction is how it arrives at the mark. A general-purpose model given a rubric will produce something that reads like examiner feedback, but it has no way of knowing where the band boundaries fall across real student work. A calibrated marker has been tested against standardisation scripts with known chief-examiner marks, and adjusted until its marks land where an examiner's do. Top Marks AI is the second kind. Here is what that looks like for a teacher:
The honest baseline is human marking. Research on experienced GCSE English examiners marking against chief-examiner scores (Fowles, 2009) found a Pearson correlation of around 0.65, with only about 45% of marks inside the board's tolerance. Any AI marker should be measured against that, and should publish the measurement.
Top Marks AI's headline figures: 0.94 correlation on AQA GCSE English Language, 0.91 on OCR English Literature, and 0.90 on Edexcel IGCSE English. On a 30-mark GCSE English question the average error is 1.75 marks against 4.0 for experienced human markers, with around 84% of marks within tolerance. In a head-to-head test on 51 Edexcel A Level Politics standardisation essays, the calibrated tool reached 0.84 correlation with a mean absolute error of 2.55 marks, against 0.48 and 5.0 marks for a competitor that feeds the rubric to a general-purpose model.
Those findings have been independently corroborated by Ark Schools, one of the UK's largest multi-academy trusts, and by Community Schools Trust, both of which ran their own checks. A study on AQA GCSE Shakespeare essays found 93% agreement with human markers across 30 handwritten scripts. Our studies are published in-house rather than peer-reviewed, and we say so; in August 2026 we audited what every major UK AI marking tool publishes, and as of that date ours was the only one with measured, independently checked figures.
Check the numbers yourself. Every study is on our accuracy blog, with correlation, error and sample described. Bring your own standardisation scripts to a trial and compare.
The platform covers 40+ subjects across GCSE, IGCSE, A Level, IB, IELTS and KS3, on AQA, Edexcel, OCR, Eduqas, WJEC, CCEA and Cambridge. The papers below are those with a published accuracy study you can read before you trial the tool.
Paper 1 Question 2, Paper 1 Question 3, Paper 1 Question 5 creative writing, Paper 2 Question 2, Paper 2 Question 3, Paper 2 Question 4, Paper 2 Question 5 persuasive writing, and the 0.94 correlation overview.
Shakespeare (and a handwritten-script study), nineteenth-century prose, post-1914 drama and prose, poetry anthology and unseen poetry.
English Language Paper 2 Question 3, Paper 2 Question 6 and the 40-mark transactional writing question; English Literature Shakespeare extract, post-1914 40-mark question, nineteenth-century extract and poetry anthology; IGCSE English and literary heritage.
OCR GCSE English Literature Shakespeare and nineteenth-century novels, and OCR GCSE Religious Education; Eduqas GCSE English Literature and English Language Component 2 Question 2.
Edexcel GCSE History 12-mark and 16-mark questions; Edexcel A Level Politics source and no-source 30-mark questions; AQA A Level Politics 25-mark extract; Edexcel A Level English Language 45-mark question; AQA A Level Sociology 10-mark outline. Geography, Economics, Psychology, Business, Philosophy, Drama, PE and further subjects are covered on the platform; ask for the studies for your subjects at a demo.
Most mocks are still handwritten, so an AI marker that only takes typed text solves a small part of the problem. Top Marks AI's handwriting-to-text conversion is built in: photograph or scan the scripts, upload them as one PDF, and each script is transcribed and marked. Teachers can assemble complete exam papers on the platform with Assignment Packs, combining several question types into one paper that students sit under exam conditions, then batch-upload and mark the lot. Heads of department can then moderate a batch by describing the adjustment in plain English; the platform turns that into explicit rules, re-marks the affected scripts and regenerates the feedback to match, keeping the original marks for audit.
Results flow out the way schools need them: Word and Excel downloads for departments, and direct export to the school MIS through Wonde or Bromcom, which also imports student data so there is no re-keying. For schools that pay for external marking or moderation during peak assessment periods, this replaces that cost. A UCL study puts the average teacher's marking at around 230 hours a year; Top Marks AI estimates a 55% reduction, roughly 125 hours returned per teacher per year.
A mark on its own does not move a student. Every Top Marks AI tool has a bespoke feedback engine written by subject specialists, following our ScaMP framework: Scaffolded, Modelled and Precise. Feedback is organised by Assessment Objective, quotes the mark scheme, and includes a worked example showing what the next band up looks like for this student's own response. Whole-cohort feedback surfaces the patterns across a class so the next lesson can target them.
ChatGPT will mark an essay if you ask it to, and its commentary can be useful. What it cannot do is place a response in the right band with any reliability: it has no access to current mark schemes, no standardisation scripts to calibrate against, and research consistently finds general models more generous and less consistent than trained examiners. The same applies to any tool that pastes a mark scheme into a general-purpose model. If you are weighing options, our ranked comparison of the best AI marking tools covers seven of them on published evidence.
Schools and trusts including Merchant Taylors', City of London School, Weydon Multi Academy Trust, AIM Academies Trust, Corvus Learning Trust and Community Schools Trust. Teachers at UTCN reported a 50% reduction in marking load after adopting the platform. Our guide for multi-academy trusts covers procurement and rollout across several schools, and our KS3 guide covers starting earlier than GCSE.
"We've had a lot of success with, and positive feedback about, Top Marks, and in our experience it is the most accurate, with the most impact on workload, compared to others we have tried."
— Head of Sociology, Weald of Kent Grammar
"Top Marks AI exceeded my expectations. I went into the process sceptical of how well AI could respond to students' Literature exams, but I was very pleasantly surprised. I would recommend Top Marks AI as a reliable and time effective way of marking summative assessments."
— Head of English, Pocklington School
School and MAT plans are priced to the institution, with credits shared across all staff, dedicated support and onboarding training. Free trials are available, and the trial we recommend is the demanding one: bring your own standardisation scripts with known marks, run them through the tools for your papers, and compare. Our vendor-neutral guide to choosing AI marking software sets out how to run that pilot for any vendor, including us. Full product documentation, from setting up a school to sharing feedback, is in the Top Marks AI Help Centre.
Book a demo and we will walk through the accuracy data for your subjects, then mark a set of your scripts live.
An AI marker is software that reads a student's response, typed or handwritten, and produces a mark and written feedback against a mark scheme. The useful distinction is between general-purpose models given a rubric and tools calibrated against exam-board standardisation materials so that their marks track what a senior examiner would give. Top Marks AI is the second kind: each of its 400+ tools is benchmarked against board standardisation scripts with known chief-examiner marks before it is released.
A calibrated one can. Top Marks AI reaches a 0.94 Pearson correlation with examiners on AQA GCSE English Language, 0.91 on OCR English Literature and 0.90 on Edexcel IGCSE English, with around 84% of marks within board tolerance against roughly 45% for experienced human markers (Fowles, 2009). These findings have been independently corroborated by Ark Schools and Community Schools Trust. Accuracy varies enormously between tools, so ask any provider for their published figures.
Yes. Handwriting-to-text conversion is built in. Photograph or scan the scripts, upload them as a single PDF for the whole class, and each script is transcribed and marked. Teachers can review and correct every transcript before marking begins. A study on AQA GCSE Shakespeare essays found 93% agreement with human markers across 30 handwritten scripts.
AQA, Edexcel, OCR, Eduqas, WJEC, CCEA, Cambridge IGCSE and CIE across GCSE, IGCSE, AS and A Level, plus IB, IELTS and KS3, in 40+ subjects including English Language, English Literature, History, Geography, Economics, Psychology, Sociology, Politics, Business, Philosophy, Drama, PE and Religious Studies. Each question type on each board has its own tool rather than a generic essay marker.
ChatGPT can comment usefully on a typed essay, but it has no access to current UK mark schemes, nothing to calibrate against, and research consistently finds general models more generous and less consistent than trained examiners. A calibrated AI marker has been tested against standardisation scripts and adjusted until its marks land where an examiner's do, and it publishes the measurement.
Through their school. Teachers set assignments and students submit through school-managed accounts; it is not a standalone consumer app. Students whose school does not yet use it can point their English or Humanities lead to the published accuracy studies.
School and MAT plans are priced to the institution, with credits shared across all staff, dedicated support and onboarding training. Free trials are available so a school can evaluate accuracy against its own scripts before committing.
Bring your own standardisation scripts with known senior-examiner marks, run them through the tools for your specific papers, and compare the AI's marks with the known marks and with your own teachers' blind marks. A classroom-sized trial can tell you whether a tool feels plausible; standardisation scripts tell you whether it is calibrated. Book a demo and we will run this with you.
We use cookies for analytics and marketing to improve your experience — these are only set if you accept. Decline and we'll only use cookies that are strictly necessary. (Live chat is always available either way.) Learn more in our Cookie Policy.