📊
AI Model Evaluation Guide — How Do You Know Your Model Is Good?
Accuracy, BLEU, CER, SDR — different metrics per task and why each is used
Why Evaluation Is Needed
You ran trainer.train(). But how do you know your model is actually good? "Looks OK when I try it" is not evaluation. You need numbers.
Core Concept: Test Data with Known Answers
Training data: model learns from (study materials)
Test data: model is evaluated on (exam questions)
Using study materials as the exam is meaningless. Evaluate on data the model has never seen.
Metrics by Task
| Task | Metric | Direction | Code |
|---|---|---|---|
| Classification | Accuracy, F1 | Higher ↑ | load("accuracy") |
| Translation | BLEU | Higher ↑ | load("bleu") |
| Summarization | ROUGE | Higher ↑ | load("rouge") |
| OCR | CER, WER | Lower ↓ | load("cer") |
| Source separation | SDR | Higher ↑ | mir_eval |
All available via pip install evaluate → load("metric_name").
Fine-tuning + Evaluation = Set
Add compute_metrics to Trainer → see accuracy per epoch during training.
Key Concepts
1
pip install evaluate — install HuggingFace evaluation library
2
metric = load("accuracy") — load metric matching your task
3
metric.compute(predictions, references) — compare predictions to answers → one number
4
Add compute_metrics to Trainer — auto-evaluate every epoch during training
Use Cases
Fine-tuning quality check — track accuracy/CER/BLEU per epoch to prevent overfitting
Model comparison — compare multiple models on same test data
Deployment decision — set criteria like "deploy if CER below 3%"