The National Student Survey says feedback needs work. Leadership says “use AI.” Here’s the part of the problem that’s actually yours to solve, and the kind of tool that solves it.

The NSS results land in July. You scan for the red.

And there it is, sitting roughly where it sat last year: assessment and feedback. Year on year it inches upward, but it’s still one of the lowest-scoring themes in the whole survey.  At the most selective universities it scores lowest of all, with positive responses running up to ten percentage points behind less selective institutions. Students are delighted with the library. They rate the staff who explain things. It’s the marking and the feedback that lag.

By the afternoon, there’s an email from someone senior. Subject line, more or less: AI?

So now it’s yours. The marking problem, handed over with a hunch attached: this AI thing is everywhere, surely it can take care of this.

It can. But not the way the headlines imply, and not before we’re honest about which part of “marking” we actually mean.

“Marking” is several problems wearing one coat

When a senior leader says marking, they usually mean a whole tangle of things at once. Getting assessments authored and delivered. Making sure the right person sits them unaided. Checking the work is the student’s own. And then, at the end of all that, the actual act of reading a script and deciding what it’s worth.

Every one of those is a real problem, and most of them have mature tools built to solve them. (We’ll come back to that, because it matters more than you’d think.)

This article is about the last one. The reading-and-judging. Sitting down with a pile of scripts and working out, fairly and consistently, what each one has earned. It’s the hardest of the lot to add technology to, and the easiest to underestimate.

Why marking is difficult

Marking is difficult, because it matters.

A summative mark is a promise. When your institution awards a grade, it’s vouching to an employer, a professional body, sometimes the public, that this person can do this thing. Get marking wrong and, depending on the discipline, the cost isn’t an awkward email; it’s a risk to someone’s safety down the line. And formative marking is the teaching itself. Feedback is how a learner finds out what to do next. Educators don’t mark because they enjoy the evenings; they mark because feedback is where the learning actually happens.

Marking is difficult, because it doesn’t scale.

The pile grows with every extra student you admit. The number of people available to mark it does not. It’s a faintly Sisyphean business. Clear one stack of scripts, and the next is already waiting. The large cohorts, the hundred-plus modules, feel it most acutely.

Marking is difficult, because it’s slow.

Consider that it takes, lets say, 30 minutes a script to mark a substantial piece. Multiply that by a cohort, then by a department, and you’re looking at staff hours nobody is getting back: time and energy that might have gone into designing better assessments, or teaching.

Marking is difficult, because consistency is genuinely hard.

Put five markers on one cohort and you’ll get five slightly different ideas of what a distinction looks like. Moderation and calibration exist precisely because human judgement drifts, between people and across a long marking session.

Marking is hard, because it isn’t algorithmic.

Marking isn’t arithmetic you can hand to a computer. It’s pedagogical judgement, built over years. And that, more than anything, is why it has resisted automation for so long.

So what would actually help?

Here’s a useful exercise. Forget the products for a moment and just spec the tool you’d want. If you could order it from a catalogue, what would it have to do?

Six things, probably:

  • Learn from your best markers: replicate the judgement your educators already have, rather than importing some opinion of its own.
  • Get faster the more work there is: so the punishing large cohorts become the ones that benefit most, not least.
  • Work on any unit of work: typed or handwritten, a maths proof or an essay, three marks or thirty.
  • Stay consistent: mark the five-hundredth script the way it marked the first.
  • Produce educator-quality feedback: not a grade in a vacuum, but comments a learner can actually use.
  • Be auditable: every decision visible, explainable, and defensible at an appeal.

This is the bar you should measure a marking tool by.

Enter Inspera Graide

Inspera Graide was built by educators who’d done the marking. They wanted a platform like this back in 2019, couldn’t find one, and went and built it. So it’s designed around how marking actually works, not how a software team imagines it works.

The heart of it is the Inspera Graide classification AI. Inspera Graide observes how one of your educators marks a given type of response (the standard they apply, the feedback they give) and then replicates that judgement across every similar answer in the cohort. It doesn’t bring its own view of what a good answer looks like. It learns yours, and applies it consistently.

That’s also why it gets faster as the pile grows. The more you mark, the more it can recognise and replicate. So a 400-student module, the kind that used to ruin a fortnight, is exactly where the time saving compounds.

And you stay in the chair. Inspera Graide is human-in-the-loop by design: it proposes grades and feedback, and the educator reviews, adjusts, or overrides them. Every one of those decisions is logged, so you get a full audit trail: the explainability a high-stakes grade demands, and that “the computer said so” never survives.

Not all “AI” is the AI your leadership is picturing

This is really important because it’s where most of the scepticism and most of the genuine risk lives.

When leadership says “use AI,” they’re usually picturing a generative model: a ChatGPT-style system that writes plausible prose. Those tools are remarkable, but for high-stakes marking they carry three problems you can’t wave away:

  1. They can be biased.
  2. They aren’t consistent.
  3. They can’t reliably explain why they reached a verdict.

For a grade that has to stand up at an appeal, none of that is acceptable.

Inspera Graide uses classification AI instead, a more reliable variant. Rather than inventing an answer, it learns to recognise and reproduce how your markers grade, on relatively little data, in a way that can be explained and audited. AI as a partner to the educator, not a replacement for their judgement.

How good is Inspera Graide?

In research with the University of Birmingham, marking physics work on Inspera Graide rather than on paper cut median marking time from 11.2 minutes a script to just 1.2 while the feedback students received grew more than sevenfold, from an average of 23 words per assessment to 166. 

84% of the Birmingham users agreed or strongly agreed that the platform was easy to use.

In a pilot with Oxbridge, an online learning provider, Inspera Graide marked 235 student responses on Othello, The Great Gatsby, and a poetry comparison, reaching up to 99–100% congruence with the human pre-marked grades. The provider also saw a 15% rise in learner satisfaction in a single quarter.

Through a partnership with Aptem, tutors in apprenticeships and vocational training are saving up to half their marking time.

No system is an island

Now, back to that tangent we set aside.

Marking is one beast. Authoring, delivery, proctoring, and originality are others, and solving the marking beautifully doesn’t help much if it leaves you stitching five unrelated tools together by hand. 

Which is why Inspera has acquired Graide. Inspera has spent more than 25 years building one of the world’s most trusted digital assessment platforms, used by universities, awarding bodies, and governments. It covers authoring, secure delivery, proctoring, and originality. 

Today you can use Inspera Originality embedded directly inside Inspera Graide, so integrity checking and intelligent marking live in one connected environment rather than two disconnected ones. 

Soon Inspera Graide will be embedded within Inspera Assessment, so it takes care of the assessment around the marking. Inspera Graide takes care of the marking itself. Together, that’s something close to the end-to-end marking ecosystem.

Word can only do so much, see it in action

The problem that landed on your desk has a shape now. Delivery, proctoring, and originality is solved by Inspera. The stubborn part, the sitting-down-and-marking, has a tool built specifically for it, by people who used to do the marking themselves.

The prize isn’t just “faster marking” .The prize is what your educators do with the hours they get back. The mechanical, repetitive, do-it-a-hundred-times part gets delegated. The human part of teaching, the bit no model should ever touch, gets the time it deserves.

That’s the version of “use AI” worth saying yes to.

Want to see it on your own assessments? Book a demo with the Inspera Graide team and bring a real script: a STEM question, an essay, or an apprenticeship portfolio. Watching the AI work is far more convincing than any blog post.