INDEPENDENT RESEARCH · UIB MATHEMATICS
Can AI be trusted to grade student mathematics assignments?
In autumn 2025, independent researchers at the Mathematical Institute and STEM Education Research Centre at UiB, Kasper Troøyen, Therese Saltskår, and Sehoya Cotner, tested AI-assisted assessment on two introductory mathematics courses (MAT101 and MAT111), using a Lectora prototype. In MAT101, the comparison with the course professor's scores yielded R² = 0.68 and 87.7% pass/fail agreement at the 3.0/6 boundary. In MAT111, inter-rater ICC fell from 0.87 without AI to 0.61 with AI assistance. Students in MAT111 did not rate the usefulness of human-curated AI feedback significantly differently from human-only feedback (p = 0.45); this does not establish equivalence. These findings support further course-specific evaluation, not a general conclusion that AI assessment is ready for deployment at scale. The page below reports the findings, including where AI assistance performed worse, alongside the published validation summary.
About the UiB AI grading research project
The project — "Assessment in Introductory Mathematics", supported by UiB insentivmidler — set out to find evidence-based solutions to the challenge of giving effective feedback in large (>400 students) introductory mathematics courses. The research team led by Kasper Troøyen, Therese Saltskår, and Sehoya Cotner ran two parallel studies in autumn 2025, one on each of UiB's two large first-year math courses: MAT111 (mathematics-intensive programmes, ~450–500 students) and MAT101 (less mathematics-intensive programmes, ~450–500 students). Lectora was the AI tool the researchers evaluated; we provided the prototype but had no role in study design, data collection, or analysis.
Study 1 (MAT111) tested AI as a grading partner on mandatory assignment 1 — 394 students submitted, graded over a two-day retreat by 11 student-assistant graders on a 0–17-point scale (students received pass/fail). Each paper was randomly assigned to either human-only grading or human+AI grading (AI drafted, human curated); students didn't know which condition they were in. 20 random papers were graded four times each (twice human-only, twice human+AI) to measure inter-rater reliability under each condition. A follow-up student survey (150 respondents, 138 read the feedback) measured perceived usefulness of feedback within each condition — a controlled-group design since each student saw only one feedback type.
Study 2 (MAT101) tested AI as one of three parallel feedback sources on milestone-check submissions (1,051 paired AI/professor scores from 356 students across milestone checks 0–4, on a 0–6-point scale). Each student received instructor feedback, peer feedback from two assigned peers, and AI feedback from Lectora, all side-by-side on the same submission. A follow-up student survey (127 respondents of 848 in the perception cohort that covered milestone checks 2–4) measured which feedback students read and how useful they found each source — a side-by-side comparison design where each student saw all three sources together. The findings below are reproduced from the published presentation in full, including where the AI underperformed.
The handwritten mathematics example uses fictional work, authored example feedback and illustrative cohort data. It is not a record of MAT101 or MAT111 students or a measurement of results from this research.
What did the MAT101 study find on AI grading agreement with the professor?
On the statistical-agreement question — does the AI's draft score line up with what the course professor would have given? — the MAT101 milestone dataset gives a clear, mixed answer. Across 1,051 paired (AI, professor) scores from 356 students on a 0–6 scale, Lectora's drafts agreed with the professor at R² = 0.68 (Pearson r = 0.82), with a mean absolute error of 0.64 points. 61% of paired scores landed within ±0.5 pt of the professor, 85% within ±1.0 pt, and 93% within ±1.5 pt. Exact match — AI and professor on the same quarter-point — happened on 26% of submissions.
On the pass/fail decision at the 3.0/6 boundary, the AI agreed with the professor on 87.7% of submissions (Cohen's κ = 0.65). Of the 129 disagreements: 52 papers where the AI said pass but the professor said fail, and 77 papers where the AI said fail but the professor said pass. Both clusters sit near the boundary itself. With the professor's scores as the reference, sensitivity to passes was 91% and specificity to fails was 77%. The reference scores are the comparison standard, not proof that every individual decision was correct.
The AI's mean bias was −0.24 pt: on average, Lectora drafted slightly lower than the professor, concentrated in the high-score band where the AI used scores 5 and 6 less often (89 AI sixes vs 120 professor sixes on the same submissions). A lower draft score is not inherently safer; it can also disadvantage a student if accepted without review. R² = 0.68 describes the fitted linear relationship with the professor's scores on this dataset. We have no published human-vs-human R² for mathematics to set beside it, and R² alone does not establish correct individual scores or reduced review workload.
How consistent was the AI agreement across the four module milestones?
The dataset covers milestone checks 0–4 in MAT101 (5 of the 6 milestone checks in the syllabus). For per-module analysis, the researchers' Analytics sheet groups submissions into four time-window buckets by submission date (the labels "Module 1–4" below refer to these analytic groupings, not the course's milestone-check numbering). Per-module R² ranged from 0.62 (Module 2, n = 319) to 0.77 (Module 4, n = 171), with Modules 1 and 3 at 0.69 and 0.70 respectively. The MAE ranged from 0.50 pt (Module 3, the tightest) to 0.76 pt (Module 1, the loosest). Pass/fail agreement at the 3.0/6 boundary ran 85–93% across the four buckets, with Module 3 the highest at 92.9%.
These four time-window groupings show variation in the observed comparisons. They do not isolate topic, score distribution or calibration as the cause. R² depends on the score distribution as well as the fitted relationship, so it should be read alongside absolute errors and boundary disagreements. The breakdown can guide closer inspection of submissions; it does not by itself establish which topics the system verifies reliably or how much manual review they require.
How did using AI affect grader consistency on MAT111?
This is the finding most likely to give institutional buyers pause, and the reason it's reported here. In the MAT111 study, 20 random papers were each graded four times: twice by a human alone and twice by a human using AI assistance, distributed across 11 student-grader assistants. The inter-rater reliability — measured as ICC(2,1), the standard intra-class correlation for absolute agreement — dropped from 0.874 (human-only) to 0.608 (human + AI). On the same 20 papers, the share of total grade variance attributable to which grader handled the paper rose from 3.7% (human-only) to 26.3% (human + AI). An ANOVA found that AI use significantly increased grader-attributable variance (p = 0.0098).
No significant mean-grade shift was detected: graders using AI gave on average 0.23 points lower than graders without AI on the 20-paper subset (p = 0.199, lmer) and 0.10 points lower across the full population (p = 0.58, lmer). Those results do not prove the absence of a systematic difference. On the full population, an ANOVA comparing variance by method gave p = 0.058: below α = 0.1, but not below α = 0.05. This is weaker evidence than the 20-paper finding and should not be presented as conclusive replication across the full cohort.
Two possible explanations deserve investigation. The graders were using an early prototype with a course rubric that had not been tuned beforehand, so familiarity and rubric interpretation may have mattered. Reviewing AI suggestions may also change how graders arrive at a score. The study does not separate these explanations or establish that onboarding and tuning would recover the observed loss of agreement.
How did students perceive the AI feedback compared to instructor and peer feedback?
Two study designs, two answers. In MAT101 (n = 127 survey respondents from the 848-student perception cohort), students received instructor, peer, and AI feedback side-by-side on the same submission, and AI rated lowest on both measures: read rates instructor 94% / peers 89% / AI 57%; usefulness on a 1–7 scale instructor 5.72 (variance 0.73) / peers 4.96 (variance 0.79) / AI 4.38 (variance 1.80); all pairwise differences significant (Wilcoxon, Bonferroni-adjusted, p < 0.01). Free-text comments split roughly as ~55% "unreliable or inaccurate", ~30% "inconsistent with the instructor", ~20% "not mature enough to deploy", ~15% "useful for exam practice". On a 58-student subset that read all three sources, peers slightly outperformed AI on item-by-item agreement with the professor (peers MAE 0.52, AI MAE 0.74).
In MAT111, each student received only one feedback condition (Human+AI or Human-only) and did not know which. No significant difference in usefulness was detected (coefficient 0.20, p = 0.45 on linear regression; mean usefulness 5.1/7 across the cohort, 138 of 150 surveyed read the feedback). A non-significant difference does not establish equivalence. The AI-assisted condition included human curation, so it is not a test of unreviewed AI feedback.
Both findings answer specific questions. MAT101 asks how students rate three feedback sources presented together; AI rated lowest. MAT111 compares human-curated AI feedback with human-only feedback in a single-condition design; it detected no significant usefulness difference. Course, task, presentation and human curation differ, so the studies cannot identify presentation as the cause of the contrast. A controlled three-arm comparison of AI-only, peer-only and instructor-only feedback could help separate these effects. Outcomes on later assignments and educator review workload are also questions for follow-up research.
What state was the Lectora prototype in at the time of the study, and what has changed since?
The autumn 2025 studies evaluated an early Lectora prototype. The MAT111 grading retreat was the documented cohort's first exposure to AI-assisted assessment; MAT101 ran during the same semester. These results describe that prototype in those courses. They do not establish how the current version performs, and subsequent development is not evidence that the observed problems have been resolved.
The MAT111 grader-variance finding remains relevant to an institution considering a pilot. Course-specific evaluation should examine both the draft and how different graders use it. Training is a reasonable part of that evaluation, but its effect needs to be measured. Prediction intervals discussed in the validation methodology belong to a retrospective MED12 simulation and the site's simulator; they are not a claim that the current assessment interface provides calibrated score intervals.
How does Lectora read handwritten math and verify the calculations today?
Mathematics submissions can include scanned pages, images and typed working. Lectora's document workflow supplies submission content for assessment against the course rubric. Faint scans and ambiguous symbols can be misread, so the educator should check the original material when reviewing the draft. The research reported here does not measure recognition accuracy by line or establish a calibrated handwriting-confidence display.
For quantitative work, the assessment workflow has access to code execution and instructions to use it for checks such as numerical answers, algebra, calculus and matrix operations. Executed checks can support the assessment, but access to the tool is not a guarantee that every line is executed or that every arithmetic error is caught. The code's inputs and assumptions also need to be correct.
A numerical check and a mathematical proof serve different purposes. Computing an integral or matrix operation can help check a result; testing an induction step at representative values does not prove it for all values. The rubric governs the proposed credit, and the examiner reviews both the reasoning and the draft score.
How does Lectora handle partial credit, symbolic equivalence, and notation correctness?
Partial credit is where rubric design matters. A student who identifies the base case and inductive hypothesis but makes an algebraic error in the inductive step has shown different work from a student who omits the induction entirely. If the rubric assigns separate marks for those elements, Lectora can use them when drafting a score and explanation. The examiner decides whether the proposed method marks and deductions are justified; the published studies do not establish correct partial-credit allocation for every proof.
Equivalent expressions may deserve the same credit when the rubric allows either form. For example, 2(x+1) and 2x+2 are equivalent; e^(ln x) and x require the domain condition x > 0 over the reals. Lectora can consider such expressions in its assessment draft, but this is not a claim of a universal symbolic verifier. Domain assumptions and requirements for simplified or factored form must be checked against the task and rubric.
Notation can be a separate criterion when the rubric defines one. Variable definitions, set-membership symbols, and the distinction between equality, equivalence and implication may then be assessed separately from method. Lectora drafts against those criteria, and the examiner checks the proposed deductions and edits the explanation before it reaches the student.
What should an institution evaluating AI grading tools take from this research?
The autumn 2025 UiB research gives institutions reasons to evaluate AI-assisted assessment carefully. MAT101 reported R² = 0.68 and 87.7% pass/fail agreement with the professor at the 3.0/6 boundary, alongside individual score differences. MAT111 found lower inter-rater agreement with AI assistance, while its feedback survey detected no significant usefulness difference (p = 0.45). Together, these findings do not establish readiness for deployment at scale, equivalent feedback quality, or performance of other assessment tools.
A course-specific evaluation can compare drafts with previously assessed submissions under the same rubric, examining R² alongside absolute errors, boundary decisions and agreement between several graders using the tool. A set of 50–200 papers is a possible starting point to discuss, not a validated minimum or a guarantee of sufficient evidence. The published methodology explains the score-comparison measures. Grader-calibration training is also worth evaluating: the MAT111 ICC drop from 0.87 to 0.61 gives reason to examine how graders use a draft, but the study did not include a trained comparison arm and cannot show that training closes the gap.
If you want to evaluate Lectora with your course's submissions and rubric, request access to discuss a pilot. The same conversation can cover follow-up research: separating feedback-source and presentation effects, measuring outcomes on later assessments, or examining educator review workload. The study design and Lectora's role should be agreed with the institution before work begins.