VALIDATION · METHODOLOGY

How is Lectora's score agreement actually measured?

This is the long-form methodology behind the headline figures on the validation summary. It covers what R² actually measures when two people grade the same exam, how the R² = 0.64 baseline between two human graders on the final clinical exam at the Faculty of Medicine at the University of Bergen was produced, why Lectora's R² = 0.85 against the course teacher sits where it does, and how targeted manual scoring, in a retrospective simulation at one pass/fail boundary, would have converted that agreement into roughly 70% fewer manually scored examiner-task pairs. Every figure here is the same one reported on the validation summary: this page is the working, not a separate claim.

What does R² mean in the context of grading?

R² (the coefficient of determination) describes the proportion of variation accounted for by a fitted model. In the linear score comparisons reported here, it describes how closely scores follow a fitted line. It does not measure the percentage of correct grades or the size of the differences between individual scores.

R² = 1.0 means the observations lie on the fitted line; it does not require the two graders to give identical scores. A systematic offset or difference in scale can remain. R² = 0.0 means the fitted model accounts for none of the variation around the mean in this comparison, not that a grader is making random decisions.

The comparison on the final clinical exam at the Faculty of Medicine at the University of Bergen reports R² = 0.85 for Lectora against the course teacher and R² = 0.64 between the two human graders. These are results on that dataset, not a universal range for human grading. Evaluating agreement also requires examining score differences and the consequences at the assessment's decision boundaries.

What evidence is there that Lectora is accurate?

Evidence comes from three ongoing academic partnerships. The published anchor is the twelve sittings of the final clinical exam at the Faculty of Medicine at the University of Bergen. Separate studies cover handwritten math at the Department of Mathematics at UiB, written up in the math case study, while NHH Finance is an ongoing pilot on long-form analytical answers. The studies have different tasks, designs and outcome measures.

The medicine set covers 889 candidate-exam pairs from twelve six-hour sittings of the final clinical exam at UiB, producing roughly 43,700 item-pair comparisons. Two independent human graders had already scored the papers, giving the baseline R² = 0.64. Lectora was run blind on the same papers and rubric without access to either human grader's scores. The comparison with the course teacher's reference scores produced R² = 0.85, with the summative exam reaching 0.86.

The published math case study reports the milestone-check study's score comparisons and the grader-reliability study's inter-rater findings, including where AI assistance performed worse. Finance per-cohort R² results are not published yet; the pilot summary is shared with the course coordinator. The medicine result does not establish the result for a different course, rubric or assessment format.

Lectora (0–100)

Course teacher (0–100)

Lectora vs. course teacher
R² = 0.85
Two human graders
R² = 0.64

12 exams · 889 candidate-exam pairs · 43,683 item-pair comparisons

Dots are illustrative.

What's the inter-rater agreement among human graders?

The two independent human graders on the final clinical exam dataset at the Faculty of Medicine at the University of Bergen had R² = 0.64. The UiB Medicine case study explains that comparison. It is a measured baseline for these exams, not a general benchmark for all human graders.

In the separate grader-reliability study at the Department of Mathematics at UiB, eleven student graders scoring the same twenty papers had an inter-rater ICC of 0.874; with AI assistance, ICC dropped to 0.608 on the early prototype, and the share of grade variance attributable to the grader rose from 3.7% to 26.3%. ICC and R² are different statistics and cannot be ranked against each other as if they measured the same thing. Subject, rubric, study design and grader calibration all matter when interpreting these findings.

Differences in rubric interpretation can produce different scores for the same answer. An institution evaluating Lectora should examine those differences on its own assessment material, alongside the aggregate statistics.

Is Lectora more consistent than human graders?

The medicine comparison found a higher R² for Lectora versus the course teacher (0.85) than between the two human graders (0.64). That is stronger linear association in this comparison; R² alone does not establish that individual grades are correct, that review takes less time, or that the result transfers to another course.

The mathematics grader-reliability counterstudy is also part of the evidence: AI assistance reduced inter-rater ICC on the tested prototype. Its companion milestone-check study found the grading agreement holding at R² = 0.68 against the professor, mean absolute error 0.64 points on a 0 to 6 scale, and 87.7% pass/fail agreement at the boundary, so the two results have to be read together rather than one at a time. Institutions should assess the draft quality and review workload in their own setting.

In the product workflow, a qualified examiner reviews and approves every grade before publication. The published studies inform that evaluation; they do not replace examiner review or validate every Lectora capability.

What does this mean for workload? Targeted manual scoring at the pass/fail boundary

Targeted manual scoring was evaluated in a retrospective simulation across twelve sittings of the final clinical exam at the Faculty of Medicine, University of Bergen. It asks how many examiner-task pairs would have needed manual scoring to support one pass/fail boundary. It is a published method and a simulator on this site, not a feature activated when Lectora is installed.

In the simulation, the method used a stratified calibration set of fourteen candidates: eight from the low end of the AI-drafted distribution, four around the pass/fail boundary and two from the top. A fitted regression linked the AI draft to examiner scores and supplied a 99% prediction interval for each remaining candidate. Candidates whose intervals crossed the pass/fail threshold were selected for additional manual scoring; the remaining candidates used the calibrated estimate within the simulation.

Across the twelve sittings, the method would have required manual scores on 12,694 of the 41,733 examiner-task pairs: 69.6% fewer, rounded to 70%. The per-exam reduction ranged from about 40% to about 84%. Cohort size and the concentration of candidates near the threshold affected the result; the fourteen-candidate calibration is a larger share of a smaller cohort.

These are simulated task-pair reductions, not measured minutes saved or a result for a letter-grade scale. The prediction intervals determine selection for manual scoring within the model; they do not guarantee correct decisions for a new cohort. The product's requirement for qualified examiner approval of every grade remains.

RETROSPECTIVE SIMULATION

Exam 01 · 78 candidates · 48 short-answer tasks

4. Manually score the risk block

Switch exam:
010203040506070809101112
00252550507575100100PASS/FAIL THRESHOLD · 57.7AI-drafted score (% of max)Examiner score (% of max)

Simulated task reduction

73.1%

2 736 of 3 744 examiner-task pairs omitted in the simulation, not measured time saved

Manual scoring

26.9%

Risk block

7

Calibration R²

0.882

Full-cohort R²

0.833

Selected task pairs

1 008

Full-pass task pairs

3 744

Prediction interval

90%95%99%