VALIDATION

How has Lectora been validated?

Lectora has not been validated in a single one-shot study. It has been built with three academic partners over several years, and every partnership re-runs the same comparison against that course's own teacher's prior grading. The published anchor is the final clinical exam at the Faculty of Medicine at the University of Bergen: twelve sittings, 889 candidate-exam pairs of long-form written exams, about 43,700 item-pair comparisons, run blind. Across them, Lectora's draft agreed with the course teacher at R² = 0.85, a stronger fitted linear relationship than the R² = 0.64 reported between two human graders on the same material. The question every partnership asks is the same one: is the draft close enough to review, or are you starting from scratch?

How closely does Lectora agree with the course teacher?

Across twelve sittings of the final clinical exam at the Faculty of Medicine at the University of Bergen, 889 candidate-exam pairs and roughly 43,700 item-pair comparisons, Lectora's draft agreed with the course teacher at R² = 0.85, reaching 0.86 on the summative exam. The two independent human graders who had already scored the same papers agreed with each other at R² = 0.64. Lectora was run blind, against the same rubric, with no access to either grader's scores.

Where two graders typically land0.50.60.70.80.91.0R² = 0.64Two human graders, same examR² = 0.85Lectora vs. course teacherAGREEMENT (R²) — FURTHER RIGHT IS CLOSER TO THE TEACHER

Lectora (0–100)

Course teacher (0–100)

Lectora vs. course teacher
R² = 0.85
Two human graders
R² = 0.64

12 exams · 889 candidate-exam pairs · 43,683 item-pair comparisons

Dots are illustrative.

What did the retrospective workload simulation find?

In a retrospective simulation, the targeted manual scoring method used a stratified calibration set of fourteen candidates and flagged candidates whose 99% prediction intervals crossed one pass/fail threshold. Across the twelve sittings, it would have required manual scores on 12,694 of the 41,733 examiner-task pairs: 69.6% fewer, rounded to 70%, with about 40% to about 84% per exam. This measures simulated task-pair reductions, not minutes saved or performance on a letter-grade scale. The method and the simulator below are not a feature activated at installation; a qualified examiner approves every grade in the product workflow.

Full-cohort pass41,733Targeted manual scoring12,69470%less manual scoringexaminer-task pairs scored by handACROSS 12 FINAL CLINICAL EXAM SITTINGS · ~40% TO ~84% PER EXAM

RETROSPECTIVE SIMULATION

Exam 01 · 78 candidates · 48 short-answer tasks

4. Manually score the risk block

Switch exam:
010203040506070809101112
00252550507575100100PASS/FAIL THRESHOLD · 57.7AI-drafted score (% of max)Examiner score (% of max)

Simulated task reduction

73.1%

2 736 of 3 744 examiner-task pairs omitted in the simulation, not measured time saved

Manual scoring

26.9%

Risk block

7

Calibration R²

0.882

Full-cohort R²

0.833

Selected task pairs

1 008

Full-pass task pairs

3 744

Prediction interval

90%95%99%

Where does the evidence come from?

Three academic partnerships, each re-running the same validation against its own course teacher's prior grading, including the parts where AI underperformed.

In the two independent mathematics studies at the Department of Mathematics at the University of Bergen, grading agreement held: R² = 0.68 against the professor (Pearson r = 0.82), mean absolute error 0.64 points on a 0 to 6 scale, and 87.7% pass/fail agreement at the boundary. The feedback result split by study design. In the milestone-check study, where students saw instructor, peer and AI feedback side by side, they read the AI feedback least (57% against 94% for the instructor) and rated it significantly less useful (4.38 against 5.72 on a 7-point scale). In the grader-reliability study, where each student received either human-plus-AI or human-only feedback without knowing which, no significant difference in usefulness was detected (p = 0.45); the case study states that this is not proof of equivalence. When students could see the source they preferred the human; when they could not see it, no difference was detected. The same researchers also found inter-rater reliability falling from ICC 0.874 to 0.608 with the early prototype, which is on the cards below rather than left out.

Published

Faculty of Medicine

University of Bergen

R² = 0.85

against the course teacher

The published anchor: 12 sittings of the final clinical exam, 889 candidate-exam pairs, scored blind against the same rubric.

Read the case study →
Independent study

Department of Mathematics

University of Bergen

R² = 0.68

milestone-check study, measured by UiB researchers

Not our measurement. The same researchers found AI assistance dropped inter-rater agreement from ICC 0.87 to 0.61 on the early prototype, reported here because it is part of the evidence. On feedback, students who could see the source preferred the human.

Read the full findings →
Pilot in progress

Finance

NHH, Norwegian School of Economics

Pass/fail screening on long-form analytical hand-ins. Per-cohort figures are not published yet — the per-pilot summary goes to the course coordinator.

Read the pilot scope →

Read the full methodology → — what R² measures in grading, how the human baseline was produced, and the targeted-scoring math in full.

~248,500

QUESTIONS & CASES GRADED ACROSS DEPLOYMENTS

…and growing every week.

20+ courses

IN PRODUCTION AND VERIFIED

Across medicine, mathematics, finance and more.

12 sittings · 889 candidates

UIB MEDICINE TARGETED-SCORING PILOT

~43,700 AI-vs-examiner item-pair comparisons via the targeted manual scoring workflow.

Across twelve sittings of the final clinical exam at the Faculty of Medicine at the University of Bergen, Lectora agreed with the course teacher at R² = 0.85 over 889 candidate-exam pairs and roughly 43,700 item-pair comparisons, and 0.86 on the summative exam. It was run blind, against the same rubric. That is score agreement on long-form written exams. For comparison, two independent human graders on the same material agreed at R² = 0.64. The fitted score relationship was stronger for Lectora versus the course teacher in this comparison; R² alone does not measure the size of individual score differences. The UiB Medicine case study carries the per-exam breakdown. The same loop is in pilot at the Department of Mathematics at UiB and at NHH Finance.