Eindhoven
University of
Technology

AI Grader for Mathematical Assessments

Background Information

Grading mathematical assessments is notoriously time-consuming and scales proportionally with the number of students enrolled in a course. Due to a significant lack of resources, including time and teaching staff, educators are increasingly forced to resort to using simple multiple-choice questions in their continuous assessments and final exams. Unfortunately, these multiple-choice formats fail to fully evaluate a student’s true academic abilities because they focus exclusively on the final output rather than assessing the underlying problem-solving process. This logistical issue has become particularly acute in rapidly growing academic environments. For instance, student enrollment in the fundamental 4DB00 course recently surged by twenty percent to over five hundred students, making traditional grading extremely burdensome and necessitating the hiring of numerous teaching assistants. Consequently, teachers are severely constrained in their ability to utilize long-answer questions that better test deep conceptual understanding and complex mathematical logic. Furthermore, manual grading of these large student cohorts often introduces long delays, sometimes taking up to four weeks to return assignments. This delay significantly hinders the prompt delivery of the actionable, formative feedback that is absolutely essential for driving self-directed student learning. There is also the persistent challenge of subconscious human bias affecting teaching assistants during the evaluation of extensive written reports. As student populations continue to expand, finding a robust technological solution that can handle complex mathematical reasoning while alleviating the administrative strain on teaching staff has become a critical priority.

Aim of the Project

The primary objective of this project is to pilot a third-party artificial intelligence grading tool, specifically mathgrader.ai, to rigorously evaluate its speed and efficacy in grading both multiple-choice and complex, long-answer mathematical questions. By utilizing this advanced technology, the project aims to dramatically reduce assignment grading times from four weeks to just a couple of days, thereby significantly shortening vital feedback loops for students. The pilot will initially test the tool in a small, ten-student course, 4SC060, by having the software assess a final take-home exam. The efficacy of the artificial intelligence will be meticulously determined by comparing its output directly to manually graded exams to ensure the resulting grades are normally distributed and align closely with human evaluation. Subsequently, the project will conduct a larger-scale historical analysis using student data from the extensive 4DB00 course, comparing previous teaching assistant grades against the artificial intelligence grader's automated results. Mechanically, this process involves the software scanning handwritten submissions using optical character recognition, converting complex mathematical symbols into a machine-readable format, and evaluating the semantic syntax tree. A large language model then acts as the grader, comparing the student's step-by-step logic against a known rubric to pinpoint exactly where conceptual errors occurred. Ultimately, the project seeks to understand teacher and student perceptions of automated assessment while determining if the software can provide targeted, high-quality, and unbiased feedback. A successful pilot would motivate the future creation of a dedicated, in-house artificial intelligence grader.

This project is still ongoing.


For more information, please contact:

University Lecturer
Matthew James
Mechanical Engineering