Educators create learning tasks to help students develop knowledge, skills at application and critical skills set out as learning objectives. We create assessment tasks to evaluate how successful students have been at meeting those learning objectives.
Large language models (LLMs) typically used as chatbots such as ChatGPT, Claude, etc., enable students to outsource the work required to complete those tasks in all except the most closely supervised settings.
Understanding how well LLMs are able to complete the tasks we set enables us to make better decisions about which tasks we set for students, in what settings to run the tasks, and what guidance should accompany the tasks.
However, students use a wide range of different LLM based tools (a recent Postit note exercise in two modules generated a list of more than 20 different LLM based chatbots and other applications used by the students). It is unreasonable to expect that educators would copy-paste dozens of learning and assessment tasks into dozens of different tools to understand the range of responses they might generate.
AI Validator is a prototype app to help automate the process.
The app can import MCQ test banks, in standard formats, and educators can add text based tasks (mostly copy-paste) and then batch test these against a selection of LLMs. The app exports data in standard formats to make analysis and review easy.
These are a series of screen shots of the key panels, showing a list of tasks imported into the app and the panels for editing the tasks, selecting the LLMs to test, reviewing the LLMs' responses and exporting files for more detailed analysis and reporting.
I have generated two sample reports, one on a batch of 15 MCQs (from a textbook I wrote and which I have used in formative in-class assessments in previous years) and the other on a single essay style task I used as a discussion point in a seminar.
The MCQ report shows the results from testing 15 MCQs against six different LLMs, from Anthropic (Claude), OpenAI, Baidu (Ernie) and Alibaba (Qwen). The report is in two sections:
In practical terms,
The sample report (HTML and PDF) are hosted here: https://edumeme.com/validator-reports/sample-mcq/
I am part way through a larger study of MCQ test banks in marketing and economics. So far, some of the test banks are are only being answered correctly less than half the time. In such cases, there is much more scope for using the questions in different assessment settings, but much less scope for LLMs to be used to support learning of the material.
The essay style task report. The task is to critically engage with the methods applied in a very recent article on workslop [Niederhoffer, et al, 2024] and was sent to six different LLMs twice: first just the tasks was sent, then the task was sent with a copy of the text of the article.
The report skips the analysis of the responses because the app does not mark tasks except for MCQs which have an answer key. Domain experts are the only ones qualified to evaluate domain-specific LLM output.
Instead it reproduces the task and two sets of outputs. The first generated by a simple copy-paste of the task. Five of the models return statements acknowledging that the task cannot be addressed without a copy of the article text and instead provide sound generic guidance on how to address the task. One LLM goes off track.
When the article text is included, the responses are sound but vary widely in how they are presented.
A copy of the report is hosted here: https://edumeme.com/validator-reports/sample-essay/
There are a number of useful applications of the tools and approach documented above:
Kosmyna, N. et al. (2025) ‘Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task’. arXiv. Available at: https://doi.org/10.48550/arXiv.2506.08872.
Niederhoffer, K. et al. (2025) ‘AI-Generated “Workslop” Is Destroying Productivity’, Harvard Business Review. Available at: https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity (Accessed: 30 September 2025).