PILOT CASE STUDY · MEDICINE
What twelve clinical exam sittings at UiB show about AI-assisted assessment
The MED12 validation at the Faculty of Medicine, University of Bergen, covers twelve six-hour clinical exam sittings: 889 candidate-exam pairs and roughly 43,700 item-pair comparisons. On these long-form written exams, Lectora's draft had a fitted score relationship of R² = 0.85 with the course teacher, compared with R² = 0.64 between two human graders on the same material. A separate retrospective simulation examined how many examiner-task pairs would need manual scoring at one pass/fail boundary. It did not measure time saved or establish that examiner review can be omitted.
About the UiB Medicine validation
The MED12 comparison concerns the final integrated clinical exam at the Faculty of Medicine, University of Bergen. It covers twelve sittings, 889 candidate-exam pairs and roughly 43,700 item-pair comparisons on long-form written exams. The reported R² was 0.85 for Lectora versus the course teacher and 0.64 between the two human graders on the same material; the summative exam comparison reached 0.86.
The human graders had already scored the papers. Lectora was run blind against the same papers and rubric, without access to either grader's scores. This makes it a comparison with existing examiner reference scores, not evidence that the faculty replaced its grading process with the simulation described below.
The written clinical questions address areas such as differential diagnosis, history-taking, treatment-plan justification, test ordering and ethical or legal reasoning. The rubric and examiner guidance establish what each answer should demonstrate. The study measures score agreement; it does not separately validate every type of clinical reasoning or the quality of the generated feedback.
The simulator below presents candidate scores using per-exam codes. The separate medical assessment example is a fictional teaching case, not an extract from this validation dataset.
How does the targeted manual scoring simulation work?
The retrospective method proceeds in five stages. First, it starts with a draft score for every candidate and previously available examiner reference scores. Second, it selects a stratified calibration set of fourteen candidates: eight from the low end of the AI-score distribution, four around the pass/fail boundary and two from the top. This samples several parts of the distribution rather than just its centre.
Third, a linear regression links the AI draft to the examiner scores in the calibration set. At the default 99% level, its prediction intervals reflect the fitted model's residual variation and each candidate's position relative to the calibration data. Fourth, candidates selected by the threshold rule receive additional manual scores in the simulation, and the regression is refitted on the expanded reference set.
Fifth, the optional benchmark stage reveals the remaining examiner scores to compare the calibrated estimates with the full reference data. This is an analysis step in the simulator. It does not make formal examiner approval optional, and neither the benchmark nor a prediction interval guarantees correct pass/fail decisions for another cohort.
This staged view uses 99% prediction intervals and a default threshold of 70% of the cohort's P90. The comparative simulator on the validation page also offers 90% and 95% interval levels. These are assumptions for exploring the retrospective method, not recommended institutional pass marks or model-certainty controls activated in the product. Wider intervals can bring more candidates into the selected review group. An institution defines its own assessment boundary.
RETROSPECTIVE SIMULATION
Exam 01 · 78 candidates · 48 short-answer tasks
Stage
The simulation starts with AI draft scores and existing examiner reference scores. The plot reveals the references as you step through the analysis. It does not show live grading or measured time saved.
Manual scoring
—
Simulated task reduction
—
Calibration R²
—
Risk block
—
Simulation stages
Retrospective simulation of examiner-task counts at one pass/fail boundary. This is not measured time saved or a replacement for qualified examiner approval of every grade.
What reduction in manual scoring did the simulation find?
Across the twelve sittings, the method would have required manual scores on 12,694 of the 41,733 examiner-task pairs in a full pass: 69.6% fewer, rounded to 70%. A task pair is one examiner-task scoring unit in this analysis. It is not a minute of work, and this count does not measure the time needed to read, review, correct or approve an assessment.
Per-exam reductions ranged from about 40% to about 84%. Cohort size and the concentration of candidates around the pass/fail threshold affected the result. A fixed fourteen-candidate calibration set represents a larger share of a smaller cohort; a dense group near the threshold requires more additional reference scores. The simulator lets readers inspect these differences rather than assume that every sitting has the same result.
The simulation uses calibrated estimates for candidates outside its selected manual-scoring group. That modelling choice is not a claim that these candidates can safely skip review in practice. The result applies to the studied data and one pass/fail boundary, not a letter-grade scale or a guarantee for future exams. In Lectora's product workflow, a qualified examiner reviews and approves every grade before publication.
How can Lectora support assessment of written clinical reasoning?
The examiner guidance can distinguish the student's conclusion from the reasoning offered for it. Differential diagnosis, history-taking, treatment-plan justification, test ordering and ethical or legal reasoning can each contribute to an assessment where the rubric specifies them. Lectora drafts scores and written explanations from the supplied material; the examiner judges whether those proposals match the answer and the course's standards.
For differential diagnosis, an answer might need to compare alternatives and justify a priority using the case information. The examiner determines which alternatives are acceptable and whether an unusual but well-supported argument earns credit. The AI draft can help organize that review, but it does not establish the clinical validity of an answer.
For a written history-taking question, the guidance can specify which questions, observations and missing information the student should address. An examiner can use the draft to examine omissions and the student's explanation of relevance. This does not extend the published evidence to live clinical performance, patient care or OSCE stations.
Ethical and legal questions can distinguish naming a principle from applying it to the given facts. The supplied rubric governs the proposed credit. Clinical and legal references still need examiner checking; the model is not an independent authority on the applicable rules.
What does the comparison with human graders establish?
On long-form written exams in the MED12 dataset, Lectora versus the course teacher had R² = 0.85, compared with R² = 0.64 between two human graders on the same material. That is a stronger fitted linear relationship in this comparison. R² alone does not measure the size of individual score differences, prove that a grade is correct or establish a general advantage over human graders.
The scope matters when considering another course. The medical setting makes the comparison relevant to written clinical assessment, but it does not supply one accuracy figure for all clinical curricula, feedback quality or assessment formats. A new evaluation should use the institution's own questions, rubric and reference scores, with explicit checks at its decision boundaries.
See the full validation methodology for the score-comparison limits and the retrospective targeted-scoring analysis.