How are AI models evaluated in the SAAI Research Program?
Answer:
Models are evaluated using suitable metrics, confusion matrices, error analysis, cross-experiment comparison and a written explanation of limitations.
Students are taught that one headline accuracy score is not enough to establish that a model works well. Evaluation begins by selecting metrics that match the research question and the structure of the dataset. Learners may use accuracy, precision, recall, F1 score, class-wise results, confusion matrices and other suitable measures. They compare performance across training runs, models, vectorisers, datasets or experimental conditions. Error analysis is used to identify which examples are being misclassified and why. Students also document uncertainty, sample-size limitations, label quality, bias and cases where the model should not be trusted. The final conclusion must be supported by the recorded evidence and should distinguish between what the experiment demonstrates and what remains unknown.
Topic: AI Model Evaluation for Students · Audience: Students, AI Learners, Parents and Educators