Evaluation¶
Table of Contents
Offline Evaluation¶
Golden Set Curation¶
Criteria for selection
Coverage: Includes all relevant feature distributions.
Accuracy: Labels verified by experts.
Diversity: Edge cases, rare conditions.
Update frequency?
Periodically (e.g., quarterly) or when drift is detected.
How to balance representation?
Maintain real-world distribution while oversampling rare cases.
Metric¶
Basic Building Blocks
TP = True Positives
FP = False Positives
FN = False Negatives
TN = True Negatives
From these:
Precision = TP / (TP + FP)
Recall = TP / (TP + FN)
F1-score = 2 × (Precision × Recall) / (Precision + Recall)
Binary Classification¶
ROC-AUC: Measures ability to distinguish classes across all thresholds; useful when class balance is not extreme.
PR-AUC: Focuses on positive class performance (precision vs recall); useful when positives are rare.
When to prefer ROC-AUC vs PR-AUC?
ROC-AUC: When positives and negatives are balanced.
PR-AUC: When positives are rare (e.g., fraud detection, rare disease prediction).
Multi-class Classification (1 label/sample)¶
\(n_i\) is the number of true instances for class i, and \(N\) is the total number of samples.
Accuracy = (Number of correct predictions) / (Total samples)
Macro Precision = average of per-class precision
- \[\text{MacroPrecision} = \frac{1}{C} \sum_{i=1}^{C} \frac{TP_i}{TP_i + FP_i}\]
Macro Recall = average of per-class recall
- \[\text{MacroRecall} = \frac{1}{C} \sum_{i=1}^{C} \frac{TP_i}{TP_i + FN_i}\]
Macro F1 = average of per-class F1
- \[\text{MacroF1} = \frac{1}{C} \sum_{i=1}^{C} \text{F1}_i\]
Weighted F1
- \[\text{WeightedF1} = \sum_{i=1}^{C} \frac{n_i}{N} \cdot \text{F1}_i\]
Multi-label Classification (multiple labels/sample)¶
Let \(Y_i\) be the true labels and \(\hat{Y}_i\) the predicted labels for instance i.
Subset Accuracy (Exact Match)
\[\text{Accuracy} = \frac{1}{N} \sum_{i=1}^{N} \mathbf{1}(Y_i = \hat{Y}_i)\]Micro Precision / Recall / F1 (aggregate TP/FP/FN across all labels)
- \[\text{MicroPrecision} = \frac{\sum_{l} TP_l}{\sum_{l} (TP_l + FP_l)}\]
- \[\text{MicroRecall} = \frac{\sum_{l} TP_l}{\sum_{l} (TP_l + FN_l)}\]
- \[\text{MicroF1} = 2 \cdot \frac{\text{MicroPrecision} \cdot \text{MicroRecall}}{\text{MicroPrecision} + \text{MicroRecall}}\]
Macro F1 (average across labels)
- \[\text{MacroF1} = \frac{1}{L} \sum_{l=1}^{L} \text{F1}_l\]
Multi-task Binary Classification (1 binary label per task)¶
Let there be T tasks. Each task is a binary classification problem.
For each task, compute:
\(Accuracy\_t = (TP + TN) / (TP + FP + TN + FN)\)
\(AUC\_t\), \(F1\_t\) (as per binary classification)
MacroAUC / MacroF1 (across tasks)
\[\text{MacroF1}_{\text{tasks}} = \frac{1}{T} \sum_{t=1}^{T} \text{F1}_t\]
Slice-based Performance Evaluation¶
How to choose slices for evaluation?
Numerical features: Quantile-based bins (e.g., age groups).
Categorical features: Stratify by value distribution.
Temporal features: Time-based slices (e.g., recent vs past).
Edge cases: Identify rare but critical scenarios.
When is a model ready for production?
Stable performance across test & validation sets.
Performs better than baseline (existing model or heuristic).
Low failure rate in stress tests (edge cases, adversarial inputs).
Model Evaluation Beyond AUC¶
Calibration: Platt scaling, isotonic regression.
Expected Calibration Error (ECE): Ensuring confidence scores are well-calibrated.
Robustness Testing: Adversarial robustness, stress testing with synthetic data.
Online Evaluation¶
A/B testing - Interleaving vs non-interleaving
Bayesian A/B testing
Metrics
Resource: [transparency.meta.com] How Meta improves
LLM App Evaluation¶
Practical¶
[github.com] The LLM Evaluation guidebook
[confident.ai] LLM Evaluation Metrics: The Ultimate LLM Evaluation Guide
[confident-ai.com] How to Evaluate LLM Applications: The Complete Guide
[arize.com] The Definitive Guide to LLM App Evaluation
[arize.com] RAG Evaluation
[guardrailsai.com] Guardrails AI Docs
Academic¶
[arxiv.org] The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
[arxiv.org] Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
Evaluation of instruction tuned/pre-trained models
MMLU
Big-Bench