PILOT CASE STUDY · FINANCE
How the Norwegian School of Economics (NHH) used Lectora to review multi-page regression analyses and long-form finance exams
The Norwegian School of Economics (NHH) pilots Lectora across finance courses that demand multi-page analytical answers: factor models, regression interpretation and valuation reasoning. This page documents the pilot scope and assessment approach; finance-specific score-agreement results have not been published.
About the NHH Finance pilot
The Norwegian School of Economics (NHH) is a Norwegian business school. The pilot scope covers finance courses where exam answers run two to five pages: multi-step regression interpretation, factor model derivations, valuation problems and policy reasoning.
These courses don't fit the single-question / single-score model. They require the examiner to follow an argument, weigh how the student used the data and judge whether the conclusion follows. That makes them a demanding setting in which to evaluate assessment drafts.
The finance showcase uses fictional work, authored example feedback and illustrative cohort data. It is not a record of NHH students or a measurement of results from this pilot.
How does Lectora handle multi-page financial analyses?
Long-form finance exams combine charts, regression output, tabulated calculations and written interpretation over several pages. Lectora can use those submissions alongside the assignment, solution and examiner guidance to draft an assessment. The examiner reviews the reading of the document as well as the proposed score.
For each sub-question, the rubric can distinguish the calculation in the table, the interpretation in the prose and the connection between them. For example, examiner guidance can award credit for a correct regression calculation while requiring a separate assessment of whether the student interprets R² appropriately. Lectora drafts against that guidance; the examiner checks the proposed allocation and explanation.
Following cross-references can require moving between a table on one page and an interpretation on another. The draft can help focus that review, but it is not an automatic reconciliation of every figure. Check important values, labels and conclusions against the original submission.
Can Lectora grade regression output and hypothesis tests?
The directional setup of the hypothesis (one-sided vs two-sided), choice of null, interpretation of the test statistic and p-value framing can each be specified in the examiner guidance. Lectora uses that guidance when drafting scores and feedback for review.
For example, a student may choose a two-sided test where the assignment requires a directional hypothesis. A draft can give the examiner a starting point for checking that choice and its consequences. This is an illustrative assessment question, not a reported finding about the NHH cohort.
For five-factor model regressions and similar analyses, the guidance can ask whether the student justifies factor attribution, discusses coefficient stability across sub-periods and interprets the residual appropriately. Those are criteria for the draft and the examiner's review, not checks that are guaranteed to succeed automatically.
How does Lectora score R² interpretation and explained-variance framing?
Interpreting R² requires distinguishing explained variance from claims about causality, predictive power or the strength of an individual coefficient. A rubric can ask the student to make those distinctions and relate the statistic to the specific regression.
For example, examiner guidance could distinguish four levels: incorrect (R² treated as causality or a significance proxy), surface (fit described without qualification), applied (explained variance interpreted in context), and full (the interpretation connected to the question and the model's limitations). These are illustrative rubric levels, not a fixed classifier built into Lectora. Your guidance determines the point allocations; the examiner reviews the draft feedback and score.
How accurate is AI grading on finance exams compared to human graders?
The published R² of 0.85 (Lectora vs course teacher) and 0.64 (two human graders against each other) concerns score agreement on long-form written clinical exams at UiB Medicine, not finance exams. Those figures do not establish accuracy on financial calculations, chart reading or finance-specific assessment.
We have not published a finance-specific score-agreement result or a human-on-human R² for finance. To evaluate a finance pilot, compare Lectora's drafts with examiner assessments on representative submissions using the same assignment and marking guidance. That comparison is needed before drawing conclusions about performance in the course.
See the validation page for the full study breakdown and the methodology used for per-course re-validation.