Skip to main content

 

This Isn’t the AI You Know.
This One Is Different.

Generative AI vs. Classification AI in Marking:
What Higher Education Needs to Know

 

Authored by Manjinder Kainth

You’re right to be tired of AI.

You’ve sat through the keynotes. You’ve deleted the vendor emails. You’ve rewritten your assessment policies twice since November 2022, and somewhere along the way you’ve probably been told, by someone who has never marked a script in their life, that a chatbot is coming for your job.

So here’s an opening you won’t hear at a conference: your scepticism is correct.

It is also in the peer-reviewed literature. When Flodén (2025) compared ChatGPT with human graders on higher education exam responses, exact agreement came in at around 30% and adjacent accuracy came in at 45% (British Educational Research Journal).

But here is the thing the keynotes keep missing. The AI you’re tired of has a quieter, older sibling. It doesn’t write essays, invent citations, or improvise. It has been trusted with credit decisions, cancer screening, and the sorting of the world’s post for decades. It’s called classification AI, and almost nobody in education is talking about it.

This paper is about that sibling.

In a nutshell

THESIS

Marking is a classification problem, and treating it like a generative problem doesn't work.

The AI making headlines is generative: it produces new content, brilliantly and unpredictably. The AI that has quietly run fraud detection since 1992 and mammography screening since 1998 is classification: it recognises and categorises, deterministically and auditably.

Marking is a categorisation task, not a content-generation task, and the architecture you choose determines whether AI-assisted marking can ever be consistent, explainable, and defensible to a student, an external examiner, or a regulator.

The distinction is not an implementation detail. It is the whole question.

Where the mistrust came from

"AI" became shorthand for Generative AI, and it turns out educators were right not to trust it.

Generative large language models arrived so loudly in late 2022 that they swallowed the word “AI” whole.

Understandably so: ChatGPT reached the public faster than any technology in history, and it landed on educators’ desks not as a tool but as a crisis.

What followed was three years of preaching. AI will transform education. AI will destroy education. AI is inevitable. Exhausting, all of it, and largely beside the point, because the objections educators raised were soon validated by research and regulations.

Consider the evidence. Warr, Oster, and Isaac (2025) ran a controlled experiment in which a large language model assigned different scores to identical student writing when only the demographic descriptors were varied: experimental proof of implicit bias in AI grading (Journal of Research on Technology in Education). 

England’s qualifications regulator, Ofqual, put it plainly in its January 2026 working paper Principles of AI use in marking: current large models function as black boxes, exhibit variability and unpredictability in outputs, and produce confidence measures that do not necessarily reflect real-world accuracy (Ofqual, 2026).

Generative AI is not useless in education. It is, however, a technology whose grading behaviour depends heavily on the task, the prompt, and whose failure modes (bias, inconsistency, opacity) are precisely the ones that marking and feedback cannot tolerate.

Your instincts were right. They were just aimed at one side of the technology family.

Three benchmarks

Consistency, bias, and explainability separate classification AI from generative AI, and no amount of prompting closes that gap.

Classification AI is the older sibling. Where generative AI creates, classification AI recognises. It looks at a new thing and answers one question: which of the categories I have been shown does this belong to?

In marking terms, it works like this. Your team grades a set of responses to a question. The model learns which features of a response correspond to which mark. 

From then on, it suggests marks for new responses to that same question. Nothing is generated. Nothing is improvised. The model’s entire knowledge of the question is the marking your educators supplied, and when a response doesn’t resemble anything it has seen before, a well-built classification system says so and routes the script to a human.

Whether any marking technology deserves a place in assessment, this kind or the headline-grabbing kind, comes down to three benchmarks:

01

Consistency

Does the same answer get the same mark, every time, for every student?

02

Bias

Where can unfairness creep in, and can you see it, measure it, and correct it?

03

Explainability

Could you tell a student, an appeals panel, or an external examiner why this mark?

Hold both families up against those three benchmarks and the differences turn out to be structural, not cosmetic.

When Ofqual’s researchers tabulated the properties of marking architectures in their 2026 working paper, they also touched on these three principles, but in more detail.

Property Inspera Graide Generative AI with Mark Scheme
Transparent to subject experts ✓ Yes ✗ No
Decisions based on construct-relevant features Controlled by designers ✗ No
Same response, same output, every time ✓ Yes ✗ No
Training data fully specifiable ✓ Yes ✗ No
Local training data required Yes, per question None
Main source of potential bias Local human marking Web-scale pre-training data

When we map the table back to the three benchmarks we see the following: 

Consistency lives in the determinism row: a grade that changes between runs is indefensible to a student appeal or an external examiner, and non-replicability is part of the design of generative AI.

Bias lives in the training data rows: a model trained only on your institution’s marking can inherit only your institution’s biases, and those are visible and correctable through the moderation and calibration processes you already run. Biases inherited from an undisclosed slice of the internet are not.

Explainability lives in the top row: decision criteria that subject experts can inspect are decision criteria you can defend.

A road already travelled

Medicine, finance, and the postal service have trusted classification AI with high-stakes decisions for decades, with a person always checking its work.

Would education be taking a reckless leap by trusting classification AI with decisions that matter? Hardly. It would be arriving roughly last to a party that started half a century ago.

Medicine has been using classification AI for decades. The first hospital-based computer system for routine ECG interpretation went live at Glasgow Royal Infirmary around 1971, and automated ECG analysis traces its roots to Washington DC around 1960 (Macfarlane & Kennedy, 2021, Hearts). Fifty years of a machine suggesting and a clinician confirming. In June 1998, the US Food and Drug Administration approved the first computer-aided detection system for mammograms, R2 Technology’s ImageChecker, marketed then and ever since as “a second pair of eyes” for the radiologist. In trials it lifted radiologists’ detection from roughly 80 cancers in 100 to almost 88 (CancerNetwork, 1998); by 2004, more than 6.8 million women had had mammograms read with its help (Hologic, 2004). A second pair of eyes. Not a replacement pair.

Finance bet billions on it. Fair Isaac introduced the FICO score in 1989 as the first industry-standard credit scoring system, adopted by all three major bureaus by 1991 (Capital One). Three years later, in 1992, FICO’s Falcon began screening payment card transactions in real time; its consortium now spans more than 9,000 banks feeding it roughly 20 billion transaction records a month (FICO; Machine Learning Times). Every time your card works abroad without a hitch, a classification model vouched for you.

Society trusted it with our post. In 1989 the US Postal Service handed Yann LeCun’s team 9,298 scanned handwritten ZIP codes; the resulting neural network hit a 95% success rate and paved the way for deployment across the postal network in the 1990s (Guinness World Records). The handwritten address interpretation system that grew from that research has run in most USPS mail processing centres since first being deployed in 1997–98 (CEDAR, University at Buffalo). 

Now notice the shape all three sectors share. An expert’s judgement, learned and replayed at scale. A human for review. The system knowing when it doesn’t know is the design principle.

Nobody is proposing education skip those steps. The proposal is to walk the same road with better safeguards.

The regulatory picture

Every major regulator has converged on the same four demands: transparency, fairness, accountability, human oversight, and classification AI meets all four.

Wherever you teach, your regulator has been busy. I’ve looked at the frameworks from the UK, Europe, Australia, New Zealand, India, and South East Asia side by side and they all ask for the same four things: transparency, fairness, accountability, and human oversight.

United Kingdom

Ofqual has been explicit that using AI as the sole marker of student work does not comply with its regulations, citing the requirement for human judgement, and it operates under the UK's five principles: safety, security and robustness; appropriate transparency and explainability; fairness; accountability and governance; and contestability and redress (Ofqual, 2024).

European Union

The European Union's AI Act classifies systems used to evaluate learning outcomes as high-risk, attracting obligations around transparency, data governance, and human oversight (Regulation (EU) 2024/1689). UNESCO calls for a human-centred approach and for institutions to validate generative tools before use (Miao & Holmes, 2023).

New Zealand

NZQA requires schools with consent to assess to hold an authenticity policy covering acceptable AI use, and generative AI is simply not permitted in NCEA external assessment (NZQA).

India

India published the India AI Governance Guidelines in November 2025, built on seven "sutras" including People First, Accountability, and Understandable by Design, and insisting that humans retain final control over AI systems, which in education means the teacher stays the ultimate decision-maker (MeitY, 2025).

Australia

Australia has gone furthest on assessment specifically. TEQSA published Assessment Reform for the Age of Artificial Intelligence in 2023, followed it with Enacting Assessment Reform in a Time of Artificial Intelligence in 2025, and on 3 June 2024 issued a request for information requiring every registered Australian higher education provider to submit a credible action plan, overseen by appropriate governance, addressing the risk generative AI poses to award integrity. TEQSA received a 100% response rate (TEQSA, Gen AI strategies for Australian higher education: Emerging practice, November 2024). Nationally, the Voluntary AI Safety Standard (2024) and its successor Guidance for AI Adoption (2025) set out guardrails for transparency and accountability across the AI supply chain (Department of Industry, Science and Resources).

South East Asia

Across Southeast Asia, the ASEAN Guide on AI Governance and Ethics (February 2024) sets seven principles for its ten member states, transparency and explainability, fairness and equity, and human-centricity among them, with human oversight required as risk rises (ASEAN, 2024). Malaysia's national AIGE guidelines (September 2024) echo the same principles (MOSTI, 2024). Revealingly, ASEAN found generative AI different enough to warrant an entirely separate expanded guide in 2025.

Transparency, fairness, accountability, and human oversight are running throughout these policies worldwide. Regulators are not anti-AI. They are anti-unauditable.

A deterministic model, trained only on specifiable local data, wrapped in human review, is that consensus expressed as an architecture.

The evidence, at scale

Over 50,000 real customer marking decisions show Inspera Graide is accurate, its confidence can be trusted, and educators stay in control.

Most evidence on AI grading comes from small studies. We would rather test our claims against real data, and lots of it: over 15,000 AI feedback comment suggestions and nearly 40,000 AI-graded rubric essays, drawn from everyday use of Inspera Graide across six institutions.

These are the key findings from our analysis.

1.

The model's confidence is a valid indicator: suggestions rated 80–100% confident were accepted 91% of the time, and each confidence band matched a corresponding level of marker acceptance across the full range.

2.

Markers accepted between 77.6% and 89.5% of AI suggestions, high enough to show the suggestions are genuinely useful, while the remainder shows markers exercising real judgement.

3.

High-confidence predictions agree with educators more closely than educators typically agree with each other: a QWK of up to 0.93, against the typical level of agreement between two educators of 0.70 to 0.80.

Confidence is a valid indicator

First, the model’s confidence is a valid indicator. The confidence score exists to route work deliberately: it decides which suggestions can be confirmed quickly and which deserve a closer look from an educator, so it has to be trustworthy. It is.

Across 15,400 suggestions, each confidence band matched a corresponding level of marker acceptance, with suggestions rated 80–100% confident accepted 91% of the time and acceptance tracking confidence across the full range, at around 58% for suggestions rated 0–20%. The 60–80% band contains the fewest suggestions, and that is by design: confidence scores in that middle range are easy to over-read, so the system deliberately holds back suggestions there rather than present markers with signals that are hard to interpret.

Graide - Graph 1

The suggestions are genuinely useful, and markers stay in control

Second, the suggestions are genuinely useful, and markers stay in control. Markers accepted between 77.6% and 89.5% of AI suggestions, and an acceptance rate that high only happens when the suggestions are doing real work.

The remainder is just as informative: acceptance varied by institution, stayed below 100% even for the model’s most confident suggestions, and the institution with the highest acceptance rate was not the one where the AI was most involved in marking. That is what active educator judgement looks like in the data: every suggestion reviewed, and every accepted mark a decision the educator made.

Graide - Graph 2

High-confidence predictions out-agree human-to-human agreement

Third, high-confidence predictions agree with educators more closely than educators typically agree with each other. On the largest essay rubrics (six to nine bands), the full model achieved a quadratic weighted kappa (QWK) of 0.76 to 0.81, and the high-confidence subset reached 0.85 to 0.93.

For context, QWK measures agreement between two graders on a scale where 0 is chance and 1 is perfect, penalising large disagreements more heavily than near-misses. Two trained educators grading the same essays typically agree at around 0.70 to 0.80, and Williamson, Xi, and Breyer (2012) set 0.70 as the minimum for automated scoring to be considered acceptable. The model clears both benchmarks, and where it is confident, it clears them comfortably.

In short, the deployment data matches the design. The confidence scores mean what they say, the suggestions save time because they are usually right, and educators keep the final word on every mark.

Graide - Graph 3

Due diligence

If a vendor can't answer three simple questions about judgement, consistency, and confidence, walk away.

Marks go out under your name and your institution’s authority, so the question that actually matters here is where responsibility sits once AI enters the workflow.

Let’s be precise about it: every AI suggestion passes through human review before anything reaches a student. High-confidence suggestions get a fast confirm-or-amend; low-confidence ones get detailed review; and every accepted, edited, or overridden decision retrains the model.

Graide - How it works

If you take one practical tool from this paper, make it this. Whatever system lands on your desk next, from us or anyone else, ask the vendor three questions:

01
Where does the judgement come from?
Your educators' marking, or an unspecifiable pre-training corpus?
02
Does the same answer get the same mark twice?
If not, explain that to your external examiner.
03
Can it tell you when not to trust it, and what happens then?
A validated confidence signal routing work to humans, or an answer for everything?

Each question maps directly onto a regulatory criterion: specifiable training data, determinism, and human oversight. If a vendor can’t answer all three, you haven’t found a marking partner.

Marking replicates judgement, it doesn't generate it. Classification AI is built for exactly that.

For three years the sector has been stuck on the wrong question. “AI in marking: yes or no?” was always a false binary, and the fatigue you feel is the sound of that binary going nowhere.

The real question, the one regulators from London to Canberra are now asking, is: which kind of AI, for which task, under whose control?

Marking is a judgement-replication task. Classification AI replicates judgement; generative AI improvises it. Finance and medicine settled the analogous question decades ago, with human-in-the-loop classification, evidence before autonomy, and uncertainty routed to experts.

Education can walk the same road, and classification AI offers that path forward. By focusing on pattern recognition rather than content generation, it mirrors the proven, human-in-the-loop systems that medicine, finance, and global logistics have relied on for decades. It doesn’t replace the educator, it amplifies their judgment, maintains strict determinism, and ensures that every decision remains transparent, auditable, and firmly under institutional control.

Every design decision in Inspera Graide follows from one rule: the educator’s judgement is core, and the technology’s only job is to carry it further. A suggestion is never a verdict, and an unsure model says so out loud.

The benefits of Classification AI are not just theoretical. Graide’s Classification AI has supported marking across six Institutions, expanding on feedback 200,000 Responses, all while being accountable and repeatable.

To discover how Graide can transform your marking workflow, contact us.

About the Author

Manjinder Kainth - Photo

Dr. Manjinder Kainth

Dr Manjinder Kainth is the Global Director of AI at Inspera. A theoretical physicist by training, he was named among the world's top young physicists by the Lindau Nobel Laureate Committee. After years of university teaching, marking, and giving feedback, he became convinced the process needed rethinking and co-founded Graide (now Inspera Graide) to do it. He now leads the use of classification AI to help educators return better, more consistent feedback, faster.

Bibliography

Sources

Peer-reviewed and regulatory

  1. Flodén, J. (2025). Grading exams using large language models. British Educational Research Journal, 51(1), 201–224. https://doi.org/10.1002/berj.4069
  2. Warr, M., Oster, N. J., & Isaac, R. (2025). Implicit bias in large language models. Journal of Research on Technology in Education, 57(6), 1324–1349. https://doi.org/10.1080/15391523.2024.2395295
  3. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x
  4. Ofqual (2024; 2026). Ofqual’s approach to regulating the use of artificial intelligence in the qualifications sector; Principles of AI use in marking (working paper). https://www.gov.uk/government/publications/principles-of-ai-use-in-marking
  5. European Union (2024). Artificial Intelligence Act, Regulation (EU) 2024/1689. http://data.europa.eu/eli/reg/2024/1689/oj
  6. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535

Historical precedents

  1. Capital One (2023). When did credit scores start? capitalone.com
  2. FICO (2021). Explainable AI in fraud detection. fico.com
  3. CancerNetwork (1998). FDA approves ImageChecker computer system. cancernetwork.com
  4. Macfarlane, P. W., & Kennedy, J. (2021). Automated ECG interpretation: a brief history. Hearts, 2(4). mdpi.com
  5. Guinness World Records: first neural network to identify handwritten characters. guinnessworldrecords.com
  6. CEDAR, University at Buffalo. Handwritten address interpretation. cedar.buffalo.edu

Regulation and guidance in worldwide markets

  1. TEQSA. Gen AI knowledge hub; Assessment Reform for the Age of Artificial Intelligence (2023); Enacting Assessment Reform in a Time of Artificial Intelligence (2025). teqsa.gov.au
  2. TEQSA (2024). Gen AI strategies for Australian higher education: Emerging practice (confirms the 3 June 2024 RFI and 100% provider response rate). teqsa.gov.au
  3. Department of Industry, Science and Resources (Australia). Voluntary AI Safety Standard (2024); Guidance for AI Adoption (2025). industry.gov.au
  4. NZQA. Guidance on the acceptable use of artificial intelligence. nzqa.govt.nz
  5. MeitY (2025). India AI Governance Guidelines. pib.gov.in
  6. ASEAN (2024). Guide on AI Governance and Ethics. asean.org
  7. MOSTI, Malaysia (2024). National Guidelines on AI Governance & Ethics. mastic.mosti.gov.my