All field notesEvaluation · Field note

Accuracy is only part of the decision.

Why false alarms, missed events and the review workflow belong in the same conversation.

Start with the cost of being wrong.

A model score is not a product outcome. In an operational setting, a false alarm consumes somebody’s time. A missed event can matter much more. Which trade-off is acceptable depends on the task and on what happens after the prediction.

The evaluation set should reflect that operating environment: ordinary conditions, difficult examples and the transitions that cause trouble. An impressive aggregate result can conceal a weak spot that dominates the real experience.

Separate the stages of success.

Receiving data, producing a prediction, delivering an alert and helping someone act are different events. Each needs evidence. A healthy data pipeline does not establish model quality; a correct prediction does not establish successful delivery.

Useful evaluation follows the whole path. Keep the model metrics, delivery evidence and human outcome visible without collapsing them into a single success flag.

Use a threshold to explore the trade-off.

The small example below uses invented, labelled scores. It is a teaching aid, not a benchmark of any project or model. Move the threshold to see which kinds of error change.

The right operating point cannot be chosen from the score alone. It depends on review capacity, the cost of a miss and the evidence available to the person making the decision.

Interactive example / Synthetic data

What changes
when the bar moves?

20 invented examples. 8 actual events.
A score at or above the threshold raises an alert.

More alertsFewer alerts
0.96Hit
0.91Hit
0.86False alarm
0.81Hit
0.76Hit
0.71False alarm
0.66Hit
0.61False alarm
0.56Hit
0.51False alarm
0.46Miss
0.41Quiet
0.36Quiet
0.31Miss
0.26Quiet
0.21Quiet
0.16Quiet
0.11Quiet
0.06Quiet
0.01Quiet
6Events detected
4False alarms
2Events missed
60%Precision
75%Recall

Precision: the share of raised alerts that are real events. Recall: the share of real events detected. Precision is undefined when no alerts are raised.