How to Choose AI Marking Software: A Best-Practice Guide for School Leaders

Adoption of AI marking has been rapid. Best practice has not. Every demo looks immaculate, and from a leadership view a flood of intuitive, low-cost tools can look like progress — while underneath, each teacher quietly picks their own tool, applies their own standards, and generates feedback their own way. This guide is a vendor-neutral framework for cutting through the polish: what to require, where the human belongs, and the questions to ask before you let any tool near a student's work.

A note on bias. We build an AI marking platform (Top Marks AI), so apply appropriate scepticism to any example we give — and hold us to exactly the same questions below. The framework matters more than the vendor; if it's any good, it should make your decision harder for us, not easier.

Key Takeaways
  1. "Accurate" is a marketing line, not a metric. Demand evidence benchmarked against board standardisation materials: correlation with chief examiners, mean absolute error in marks, and the percentage of scripts within the board's own tolerance.
  2. The danger isn't a tool that's obviously wrong — it's one that's right often enough to feel reliable, and wrong often enough that you meet the failures one parent phone-call at a time. Require evidence that is published, transparent and auditable.
  3. AI should replace repetition, not expertise. No AI mark of consequence without a teacher's eye on it. A vendor offering to take the teacher out of the loop is a red flag.
  4. Governance is what you'll actually be asked about: data protection (anonymisation, no training on student data, DPA/DPIA/jurisdiction), safeguarding when a student discloses something in their writing, and an inspector-ready local framework.
  5. Run a real pilot and measure both correlation and time saved. If you can't measure both, you don't have a pilot — you have a vibe.

It's still the wild west

AI marking touches the most sensitive parts of how a school runs — grades, feedback, data, and safeguarding. The benefits are real, but the risks are rarely visible from a demo, and it is easy to leap when you're looking for a lifeboat. "I'll just drop their essays into ChatGPT and have it marked by lunchtime" is a sentence heard in a lot of staffrooms; it is also exactly the leap this guide is designed to slow down.

The fix isn't caution for its own sake. It's a shared playbook — a short set of standards every tool has to meet, so that the decision is made once, at leadership level, rather than a thousand times by individual teachers. Three "muddles" are where schools most often get caught: accuracy, the role of the teacher, and governance.

The accuracy muddle

Every demo looks immaculate. Three things separate the tools that hold up in moderation from the ones that don't.

1. Accurate against what? "Accurate" on its own is meaningless. The gold standard is the exam board's own standardisation materials — scripts with known chief examiner marks — not a single teacher's opinion and not an internal dataset you can't inspect. Ask for three numbers:

  • Correlation with chief examiners — how closely the tool tracks the official standard across a cohort, not against one marker.
  • Mean absolute error, in marks — when it's wrong, how far out is it? Reported in the units the board uses, never as a flattering percentage.
  • Percentage of scripts within tolerance — the same tolerance the board applies to its own moderators. If a vendor can't tell you, the conversation is effectively over.

2. What does "fair" actually mean? Hold any tool to three qualities working together: consistency (similar work earns similar marks across cohorts, across days, across markers), objectivity (judgements rest on clear criteria, not bias or fatigue), and alignment (marks reflect the official rubric, not an internal house style).

3. Is the feedback credible, or just convincing? Large language models are very good at producing text that sounds authoritative — which is not the same as text that helps a student improve. The gap between the two is where the value of a paid tool actually lives. Press any vendor on the pedagogical principles behind their feedback model, whether it's tied to the board's assessment objectives rather than generic prose, what stops it inventing facts to justify praise, and — bluntly — what makes their paid output better than the free version of ChatGPT.

The trap to remember: most of the time, these tools are right. That is the problem. Deceptive accuracy gives you no way to know in advance which script the tool will fail on — which is why published, auditable evidence matters more than any demo.

The teacher muddle

AI should replace repetition — not expertise, and not the wider human role of being a teacher in a school. An AI mark is best understood as a fast, consistent first draft; the professional judgement, and the accountability, stays with the teacher. A workable rule: no AI mark of consequence without a teacher's eye on it. In practice that means a tool should surface its own uncertainty rather than bury it, flag outliers for moderation, and let staff sample efficiently rather than forcing them to re-mark everything.

There's a question most vendors haven't considered: what happens when a student discloses something in their writing? A teacher reading a script might catch a safeguarding signal an algorithm would miss entirely. If the AI is the first — or only — reader, who is responsible for the catch, and how is it logged? A platform built only to award marks has misunderstood the wider role of the person it's standing in for. Ask whether the model can flag concerning content, where that flag goes and how fast, whether it routes to your designated safeguarding lead, and whether both the trigger and the response are logged.

The governance muddle

Governance is data, workflow and oversight — how you prove, to yourself, to parents, and to inspectors, that you've done this properly. It's also where a slick pilot meets reality.

Will it survive a wet Tuesday in November? The conditions a tool meets in a real mock season are nothing like a demo. Ask how it handles each of these:

A real-world pitfallA mitigation worth asking for
Crossed-out text, marginal additions, arrowsConfidence thresholds that flag messy pages for teacher review before marking
Two answers spilling onto the same pagePre-printed booklets with one question per page and per-student barcodes
A scanner that can't handle 10+ pages; an admin off sickA simple SOP, MIS integration, and a vendor that helps you set the habits
Staff with no time and no training to learn yet another toolA vendor with a difficult-pilot story — and an honest answer for what went wrong

Whose data is it, and where is it going? Three non-negotiables:

  • Anonymise before processing. Names, candidate numbers and class IDs should be stripped before anything is uploaded; the persistent identifiers stay with the school.
  • No training on student data. You want an unambiguous "no", in writing, in the Data Processing Agreement. Vendors should rely on exam-board or synthetic training sets.
  • DPA, DPIA, jurisdiction. Where is data processed — UK, EU, US? Is there a Data Processing Agreement in place? Has a Data Protection Impact Assessment been completed?

If an inspector asked you tomorrow, could you answer? A solid local framework needs five things:

  • A named owner — a real person, typically a deputy or assistant head, who owns AI marking decisions.
  • A use policy — when AI marks can be used unmoderated, and (more importantly) when they cannot.
  • An audit trail — a log of AI-assisted marking decisions, including moderation and overrides.
  • A review cycle — re-validate accuracy each term, and again whenever the board updates its mark scheme.
  • A disputes process — a clear route when a parent or student challenges an AI-assisted mark.

What a credible pilot looks like

A pilot is how you replace a vendor's claims with your own evidence. A credible one runs roughly like this:

  • Week 0 — Sign-off & owners. DPIA, DPA and SLT approval. Name a teacher lead and an admin lead.
  • Week 1 — Small calibration. Five to ten scripts. Validate transcription quality and feedback tone before scaling.
  • Weeks 2–4 — Blind moderation. Take a held-back sample; the AI and several teachers mark independently; investigate every outlier.
  • Post-pilot — Decide & set policy. Refine settings, build CPD, and agree explicitly where AI marks may inform predictions and where they may not.

The single test that matters: if you can't measure both correlation and time saved, you don't have a pilot — you have a vibe.

The six questions to take into any vendor meeting

Every vendor will give you a polished demo. Your job is to cut through the polish. These are the six things worth asking before you let any tool near a student's work — copy or print them and bring them with you:

The Six-Question Checklist
  1. How do you measure accuracy — against what? Correlation with the chief examiner, MAE in marks, % within the board's tolerance. Not "are you accurate" — but how do you know.
  2. What does your feedback actually look like? Bring your own script with edge cases. Watch for plausible generalities versus specific, AO-aligned, actionable advice.
  3. Do you train on student data? The only acceptable answer is an unambiguous "no" — and you want it explicit in the Data Processing Agreement.
  4. Where is data processed, and what's your data-protection framework? UK, EU, US? Is there a DPA? A DPIA? Does it align with what your exam board permits?
  5. What are the practical pain points? Ask for a difficult-pilot story. The vendors worth working with will have one — and a plan for the friction.
  6. What does your safeguarding provision look like? No detection, no escalation, no human-review trigger? That should be a dealbreaker.

For what it's worth, applying these questions to the market is how we ended up building Top Marks AI the way we did — with published accuracy studies, a teacher-in-the-loop workflow, and exam-board-based training rather than student data. If it's useful to see how the criteria play out across the tools schools actually consider, our AI marking software comparison and secondary-schools guide work through them in detail, and there's a dedicated procurement & rollout guide for multi-academy trusts. And if the question on your desk is really about workload rather than tooling, we've set out why the marking-workload debate keeps asking the wrong question. But the framework above stands on its own, whichever vendor you choose.

Putting the framework to work

Bring the six questions to your next vendor meeting — including ours. If you'd like to see how Top Marks AI answers them, with the accuracy data for your specific subjects, book a demo and we'll walk you through it.

Frequently Asked Questions

How should a school choose AI marking software?

Decide once, at leadership level, against a shared set of standards — rather than letting every teacher pick their own tool. Require accuracy evidence benchmarked against exam-board standardisation materials (correlation with chief examiners, mean absolute error in marks, and percentage of scripts within the board's tolerance), insist that a teacher stays in the loop for any mark of consequence, and check the governance basics: data protection, safeguarding, and an inspector-ready local framework. Then run a pilot that measures both correlation and time saved.

What accuracy evidence should an AI marking vendor provide?

Three numbers, all benchmarked against the exam board's own standardisation materials rather than a single teacher or an internal dataset you can't inspect: correlation with chief examiner marks, mean absolute error expressed in marks (not as a percentage), and the percentage of scripts that fall within the board's own tolerance. If a vendor can't provide published, auditable figures, treat "accurate" as a marketing claim rather than a metric.

Should AI replace teacher marking?

No. AI should replace the repetition in marking, not the expertise. The reliable model is to treat an AI mark as a fast, consistent first draft, with professional judgement and accountability remaining with the teacher — no AI mark of consequence without a teacher's eye on it. A vendor offering to remove the teacher from the loop entirely is a red flag.

What data-protection questions should schools ask AI marking vendors?

Ask whether student work is anonymised before processing (names, candidate numbers and class IDs stripped); whether the vendor trains on student data — the only acceptable answer is an unambiguous "no", stated in the Data Processing Agreement; and where data is processed (UK, EU or US), whether a DPA is in place, and whether a Data Protection Impact Assessment (DPIA) has been completed.

What does a credible AI marking pilot look like?

Sign-off and named owners first (DPIA, DPA, SLT approval; a teacher lead and an admin lead), then a small calibration on 5–10 scripts to check transcription and feedback quality, then a blind-moderation phase where the AI and several teachers mark a held-back sample independently and every outlier is investigated, and finally a decision and a written policy. The test of a real pilot is that it measures both correlation with the standard and time saved — if it measures neither, it's a vibe, not a pilot.

What safeguarding questions should schools ask about AI marking tools?

A teacher reading a script may catch a disclosure or safeguarding signal an algorithm would miss. Ask whether the tool can detect and flag concerning content, where that flag is escalated and how quickly, whether it routes to your designated safeguarding lead, and whether both the trigger and the response are logged for audit. A platform that only awards marks, with no safeguarding provision, has misunderstood the role it's standing in for.

Alex Chapman

Alex Chapman

Chief Operating Officer, Top Marks AI

Alex leads operations at Top Marks AI, working with Multi-Academy Trusts and schools across the UK to embed AI marking infrastructure at scale. This guide is adapted from a CPD session he delivered to school leaders on navigating best practice in AI assessment.

We use cookies for analytics and marketing to improve your experience — these are only set if you accept. Decline and we'll only use cookies that are strictly necessary. (Live chat is always available either way.) Learn more in our Cookie Policy.