#370 — recovering the EU no-result tail with a span-rescore
370 — recovering the EU no-result tail with a span-rescore
Measuring the parser — metrics, golden sets, parity scorecards, calibration, and eval discipline.
View all tags370 — recovering the EU no-result tail with a span-rescore
"How good is the matcher at deduplication?" sounds like a question with one
Every span the parser emits carries a number when the model says it's 94% sure, is it right 94% of the time?
734's collision-risky lever
Mailwoman's eval methodology learned its most important lessons the hard way — from shipping two model versions that regressed on headline F1 but told a different story when the failures were examined properly. This article documents the discipline: what to measure, what not to trust, and how to read a model release report.
Cross-model parser failure comparison over golden:dev — which address classes each candidate trades, and what stays beyond reach.
Cross-model parser failure comparison over parity-corpus.triaged.jsonl — which address classes each candidate trades, and what stays beyond reach.
Cross-model parser failure comparison over golden:dev — which address classes each candidate trades, and what stays beyond reach.
Cross-model parser failure comparison over golden:dev — which address classes each candidate trades, and what stays beyond reach.
A geocoder turns an address into a point, and the obvious thing to ask is whether the point is right. That question has no answer, because "right" assumes a single tolerance and there isn't one. A point that is right enough to tell you which county regulates this facility can be uselessly wrong for routing a truck to its loading dock. The question that does have an answer is the one the operator keeps coming back to, and the one this page is about: close enough for what?
The addresses in a test fixture are clean. The addresses a geocoder receives are
How Mailwoman development runs — every change is a hypothesis with a pre-registered gate, negative results ship with mechanisms, and the eval suite is an executable bug log.
This is the repeatable recipe for giving a new country a retrieval-augmented decode prior — a
Date: 2026-05-29
If you've evaluated geocoders before, you've seen the coverage page: a world map,