Adoption of AI marking has been rapid. Best practice has not. Every demo looks immaculate, and from a leadership view a flood of intuitive, low-cost tools can look like progress — while underneath, each teacher quietly picks their own tool, applies their own standards, and generates feedback their own way. This guide is a vendor-neutral framework for cutting through the polish: what to require, where the human belongs, and the questions to ask before you let any tool near a student's work.
A note on bias. We build an AI marking platform (Top Marks AI), so apply appropriate scepticism to any example we give — and hold us to exactly the same questions below. The framework matters more than the vendor; if it's any good, it should make your decision harder for us, not easier.
AI marking touches the most sensitive parts of how a school runs — grades, feedback, data, and safeguarding. The benefits are real, but the risks are rarely visible from a demo, and it is easy to leap when you're looking for a lifeboat. "I'll just drop their essays into ChatGPT and have it marked by lunchtime" is a sentence heard in a lot of staffrooms; it is also exactly the leap this guide is designed to slow down.
The fix isn't caution for its own sake. It's a shared playbook — a short set of standards every tool has to meet, so that the decision is made once, at leadership level, rather than a thousand times by individual teachers. Three "muddles" are where schools most often get caught: accuracy, the role of the teacher, and governance.
Every demo looks immaculate. Three things separate the tools that hold up in moderation from the ones that don't.
1. Accurate against what? "Accurate" on its own is meaningless. The gold standard is the exam board's own standardisation materials — scripts with known chief examiner marks — not a single teacher's opinion and not an internal dataset you can't inspect. Ask for three numbers:
2. What does "fair" actually mean? Hold any tool to three qualities working together: consistency (similar work earns similar marks across cohorts, across days, across markers), objectivity (judgements rest on clear criteria, not bias or fatigue), and alignment (marks reflect the official rubric, not an internal house style).
3. Is the feedback credible, or just convincing? Large language models are very good at producing text that sounds authoritative — which is not the same as text that helps a student improve. The gap between the two is where the value of a paid tool actually lives. Press any vendor on the pedagogical principles behind their feedback model, whether it's tied to the board's assessment objectives rather than generic prose, what stops it inventing facts to justify praise, and — bluntly — what makes their paid output better than the free version of ChatGPT.
The trap to remember: most of the time, these tools are right. That is the problem. Deceptive accuracy gives you no way to know in advance which script the tool will fail on — which is why published, auditable evidence matters more than any demo.
AI should replace repetition — not expertise, and not the wider human role of being a teacher in a school. An AI mark is best understood as a fast, consistent first draft; the professional judgement, and the accountability, stays with the teacher. A workable rule: no AI mark of consequence without a teacher's eye on it. In practice that means a tool should surface its own uncertainty rather than bury it, flag outliers for moderation, and let staff sample efficiently rather than forcing them to re-mark everything.
There's a question most vendors haven't considered: what happens when a student discloses something in their writing? A teacher reading a script might catch a safeguarding signal an algorithm would miss entirely. If the AI is the first — or only — reader, who is responsible for the catch, and how is it logged? A platform built only to award marks has misunderstood the wider role of the person it's standing in for. Ask whether the model can flag concerning content, where that flag goes and how fast, whether it routes to your designated safeguarding lead, and whether both the trigger and the response are logged.
Governance is data, workflow and oversight — how you prove, to yourself, to parents, and to inspectors, that you've done this properly. It's also where a slick pilot meets reality.
Will it survive a wet Tuesday in November? The conditions a tool meets in a real mock season are nothing like a demo. Ask how it handles each of these:
| A real-world pitfall | A mitigation worth asking for |
|---|---|
| Crossed-out text, marginal additions, arrows | Confidence thresholds that flag messy pages for teacher review before marking |
| Two answers spilling onto the same page | Pre-printed booklets with one question per page and per-student barcodes |
| A scanner that can't handle 10+ pages; an admin off sick | A simple SOP, MIS integration, and a vendor that helps you set the habits |
| Staff with no time and no training to learn yet another tool | A vendor with a difficult-pilot story — and an honest answer for what went wrong |
Whose data is it, and where is it going? Three non-negotiables:
If an inspector asked you tomorrow, could you answer? A solid local framework needs five things:
A pilot is how you replace a vendor's claims with your own evidence. A credible one runs roughly like this:
The single test that matters: if you can't measure both correlation and time saved, you don't have a pilot — you have a vibe.
Every vendor will give you a polished demo. Your job is to cut through the polish. These are the six things worth asking before you let any tool near a student's work — copy or print them and bring them with you:
For what it's worth, applying these questions to the market is how we ended up building Top Marks AI the way we did — with published accuracy studies, a teacher-in-the-loop workflow, and exam-board-based training rather than student data. If it's useful to see how the criteria play out across the tools schools actually consider, our AI marking software comparison and secondary-schools guide work through them in detail, and there's a dedicated procurement & rollout guide for multi-academy trusts. And if the question on your desk is really about workload rather than tooling, we've set out why the marking-workload debate keeps asking the wrong question. But the framework above stands on its own, whichever vendor you choose.
Bring the six questions to your next vendor meeting — including ours. If you'd like to see how Top Marks AI answers them, with the accuracy data for your specific subjects, book a demo and we'll walk you through it.
Decide once, at leadership level, against a shared set of standards — rather than letting every teacher pick their own tool. Require accuracy evidence benchmarked against exam-board standardisation materials (correlation with chief examiners, mean absolute error in marks, and percentage of scripts within the board's tolerance), insist that a teacher stays in the loop for any mark of consequence, and check the governance basics: data protection, safeguarding, and an inspector-ready local framework. Then run a pilot that measures both correlation and time saved.
Three numbers, all benchmarked against the exam board's own standardisation materials rather than a single teacher or an internal dataset you can't inspect: correlation with chief examiner marks, mean absolute error expressed in marks (not as a percentage), and the percentage of scripts that fall within the board's own tolerance. If a vendor can't provide published, auditable figures, treat "accurate" as a marketing claim rather than a metric.
No. AI should replace the repetition in marking, not the expertise. The reliable model is to treat an AI mark as a fast, consistent first draft, with professional judgement and accountability remaining with the teacher — no AI mark of consequence without a teacher's eye on it. A vendor offering to remove the teacher from the loop entirely is a red flag.
Ask whether student work is anonymised before processing (names, candidate numbers and class IDs stripped); whether the vendor trains on student data — the only acceptable answer is an unambiguous "no", stated in the Data Processing Agreement; and where data is processed (UK, EU or US), whether a DPA is in place, and whether a Data Protection Impact Assessment (DPIA) has been completed.
Sign-off and named owners first (DPIA, DPA, SLT approval; a teacher lead and an admin lead), then a small calibration on 5–10 scripts to check transcription and feedback quality, then a blind-moderation phase where the AI and several teachers mark a held-back sample independently and every outlier is investigated, and finally a decision and a written policy. The test of a real pilot is that it measures both correlation with the standard and time saved — if it measures neither, it's a vibe, not a pilot.
A teacher reading a script may catch a disclosure or safeguarding signal an algorithm would miss. Ask whether the tool can detect and flag concerning content, where that flag is escalated and how quickly, whether it routes to your designated safeguarding lead, and whether both the trigger and the response are logged for audit. A platform that only awards marks, with no safeguarding provision, has misunderstood the role it's standing in for.
We use cookies for analytics and marketing to improve your experience — these are only set if you accept. Decline and we'll only use cookies that are strictly necessary. (Live chat is always available either way.) Learn more in our Cookie Policy.