Faculty of Medicine
University of Bergen
R² = 0.85
against the course teacher
The published anchor: 12 sittings of the final clinical exam, 889 candidate-exam pairs, scored blind against the same rubric.
Read the case study →VALIDATION
Lectora has not been validated in a single one-shot study. It has been built with three academic partners over several years, and every partnership re-runs the same comparison against that course's own teacher's prior grading. The published anchor is the final clinical exam at the Faculty of Medicine at the University of Bergen: twelve sittings, 889 candidate-exam pairs of long-form written exams, about 43,700 item-pair comparisons, run blind. Across them, Lectora's draft agreed with the course teacher at R² = 0.85, a stronger fitted linear relationship than the R² = 0.64 reported between two human graders on the same material. The question every partnership asks is the same one: is the draft close enough to review, or are you starting from scratch?
Across twelve sittings of the final clinical exam at the Faculty of Medicine at the University of Bergen, 889 candidate-exam pairs and roughly 43,700 item-pair comparisons, Lectora's draft agreed with the course teacher at R² = 0.85, reaching 0.86 on the summative exam. The two independent human graders who had already scored the same papers agreed with each other at R² = 0.64. Lectora was run blind, against the same rubric, with no access to either grader's scores.
Lectora (0–100) ↑
Course teacher (0–100) →
12 exams · 889 candidate-exam pairs · 43,683 item-pair comparisons
Dots are illustrative.
In a retrospective simulation, the targeted manual scoring method used a stratified calibration set of fourteen candidates and flagged candidates whose 99% prediction intervals crossed one pass/fail threshold. Across the twelve sittings, it would have required manual scores on 12,694 of the 41,733 examiner-task pairs: 69.6% fewer, rounded to 70%, with about 40% to about 84% per exam. This measures simulated task-pair reductions, not minutes saved or performance on a letter-grade scale. The method and the simulator below are not a feature activated at installation; a qualified examiner approves every grade in the product workflow.
RETROSPECTIVE SIMULATION
Exam 01 · 78 candidates · 48 short-answer tasks
4. Manually score the risk block
Simulated task reduction
73.1%
2 736 of 3 744 examiner-task pairs omitted in the simulation, not measured time saved
Manual scoring
26.9%
Risk block
7
Calibration R²
0.882
Full-cohort R²
0.833
Selected task pairs
1 008
Full-pass task pairs
3 744
Prediction interval
Three academic partnerships, each re-running the same validation against its own course teacher's prior grading, including the parts where AI underperformed.
In the two independent mathematics studies at the Department of Mathematics at the University of Bergen, grading agreement held: R² = 0.68 against the professor (Pearson r = 0.82), mean absolute error 0.64 points on a 0 to 6 scale, and 87.7% pass/fail agreement at the boundary. The feedback result split by study design. In the milestone-check study, where students saw instructor, peer and AI feedback side by side, they read the AI feedback least (57% against 94% for the instructor) and rated it significantly less useful (4.38 against 5.72 on a 7-point scale). In the grader-reliability study, where each student received either human-plus-AI or human-only feedback without knowing which, no significant difference in usefulness was detected (p = 0.45); the case study states that this is not proof of equivalence. When students could see the source they preferred the human; when they could not see it, no difference was detected. The same researchers also found inter-rater reliability falling from ICC 0.874 to 0.608 with the early prototype, which is on the cards below rather than left out.
University of Bergen
R² = 0.85
against the course teacher
The published anchor: 12 sittings of the final clinical exam, 889 candidate-exam pairs, scored blind against the same rubric.
Read the case study →University of Bergen
R² = 0.68
milestone-check study, measured by UiB researchers
Not our measurement. The same researchers found AI assistance dropped inter-rater agreement from ICC 0.87 to 0.61 on the early prototype, reported here because it is part of the evidence. On feedback, students who could see the source preferred the human.
Read the full findings →NHH, Norwegian School of Economics
—
Pass/fail screening on long-form analytical hand-ins. Per-cohort figures are not published yet — the per-pilot summary goes to the course coordinator.
Read the pilot scope →Read the full methodology → — what R² measures in grading, how the human baseline was produced, and the targeted-scoring math in full.
QUESTIONS & CASES GRADED ACROSS DEPLOYMENTS
…and growing every week.
20+ courses
IN PRODUCTION AND VERIFIED
Across medicine, mathematics, finance and more.
12 sittings · 889 candidates
UIB MEDICINE TARGETED-SCORING PILOT
~43,700 AI-vs-examiner item-pair comparisons via the targeted manual scoring workflow.