How mail delivery actually works
Before you can judge a parser, you need to know what it is parsingaddress parsingThe process of decomposing a free-text postal address string into structured components — house number, street name, locality, region, postcode, and country — so a geocoder can resolve them to coordinates. for. Address parsingaddress parsingThe process of decomposing a free-text postal address string into structured components — house number, street name, locality, region, postcode, and country — so a geocoder can resolve them to coordinates. exists to route physical objects to physical places, through a system designed in the 19th century and patched ever since.
This article walks through what happens when you mail a letter, because the mechanics show something important: the postal system is already fuzzy. It handles ambiguity through human intervention, local knowledge, and layered fallbacks, and a parser that reports its own uncertainty is more faithful to the real world than one that returns a confident wrong answer.
Step 1 — Collection
You drop a letter in a blue box or hand it to a carrier. The letter enters the USPS network at the nearest processing and distribution center (P&DC). There are about 250 of these in the United States. Each one serves a regionregionThe first-level administrative subdivision of a country — a US state, a French region, a province. The component between country and locality. roughly the size of a congressional district.
At this point, nobody has read your address. The letter is in a bin with thousands of others, sorted only by rough destination zone. The destination ZIP code — not the streetstreetThe named linear feature along which house numbers are ordered. Decomposes into a name plus street affixes; one of the Tier 2 fine labels. address — determines which P&DC it goes to next. This is the first hint that the postal system cares about routing codes more than precise addresses.
Step 2 — OCR and the Remote Encoding Center
At the destination P&DC, the letter passes through an optical character reader at about 10 letters per second. The OCR tries to read the address block and resolve it to an 11-digit delivery point codedelivery point code (DPBC — Delivery Point Barcode). The full 11-digit USPS code (5-digit ZIP + 4-digit extension + 2 delivery-point digits) that uniquely identifies a single mailbox, encoded as the barcode on an envelope for automated sorting. — ZIP+4 plus the last two digits of the street numberhouse numberThe numeric or alphanumeric identifier of a building on a street. Mailwoman's house_number component; its position relative to the street name flips between locales. or PO boxPO boxA numbered mailbox at a post office used as a delivery address instead of a physical street location. Mailwoman tags it as the po_box component; structurally the same family as a subpremise.. If it succeeds, the letter gets a barcode sprayed on the envelope and proceeds to sorting. About 85-90% of machine-printed mail passes OCR on the first pass.
If the OCR cannot read the address — handwriting, smudged ink, non-standard formatting, foreign scripts — the system takes a photograph of the envelope and sends it to a Remote Encoding Center (REC). A human operator at a keyboard sees the photograph for about two seconds and types the ZIP code. Another human sees the street numberhouse numberThe numeric or alphanumeric identifier of a building on a street. Mailwoman's house_number component; its position relative to the street name flips between locales. and streetstreetThe named linear feature along which house numbers are ordered. Decomposes into a name plus street affixes; one of the Tier 2 fine labels. name. Between them, they produce enough information for the sorting machines to route it.
The RECs handle roughly 200 million piecesECE (Expected Calibration Error). A metric that measures how well a model's confidence scores align with its actual accuracy. Lower is better. Mailwoman's held-out ECE drops from 0.067 (raw) to 0.0035 (calibrated). per year. That is 200 million piecesECE (Expected Calibration Error). A metric that measures how well a model's confidence scores align with its actual accuracy. Lower is better. Mailwoman's held-out ECE drops from 0.067 (raw) to 0.0035 (calibrated). where the machine gave up and a human squinted at bad handwriting and figured it out anyway. This is the postal equivalent of graceful degradation: the system escalates bad input instead of failing on it.
Step 3 — ZIP+4 resolution
The 5-digit ZIP code gets the letter to the right post office. The +4 extension gets it to the right carrier routecarrier routeThe delivery path one postal carrier covers on a shift. A ZIP code is fundamentally a set of carrier routes, not a geographic polygon, which is why ZIP boundaries are fuzzy and shift over time. or building. The full 11-digit delivery point codedelivery point code (DPBC — Delivery Point Barcode). The full 11-digit USPS code (5-digit ZIP + 4-digit extension + 2 delivery-point digits) that uniquely identifies a single mailbox, encoded as the barcode on an envelope for automated sorting. gets it to the right mailbox.
ZIP codes are not polygons. They are carrier routescarrier routeThe delivery path one postal carrier covers on a shift. A ZIP code is fundamentally a set of carrier routes, not a geographic polygon, which is why ZIP boundaries are fuzzy and shift over time. — the sequence a postal carrier walks or drives. A ZIP code boundary follows streetsstreetThe named linear feature along which house numbers are ordered. Decomposes into a name plus street affixes; one of the Tier 2 fine labels., not census blocks. It can change when a carrier retires and routes get redrawn. The US Census Bureau publishes ZIP Code Tabulation Areas (ZCTAsZCTA (ZIP Code Tabulation Area). The US Census Bureau's polygon approximation of a ZIP code, generalized from census blocks rather than USPS delivery routes. Explicitly not USPS ground truth and often wrong in rural areas.) as an approximation, but ZCTAsZCTA (ZIP Code Tabulation Area). The US Census Bureau's polygon approximation of a ZIP code, generalized from census blocks rather than USPS delivery routes. Explicitly not USPS ground truth and often wrong in rural areas. are not USPS ground truthground truthThe correct answer for an example, used as the standard a prediction is graded against. Mailwoman's ground truth is the hand-labeled golden set; its quality caps achievable accuracy. and USPS explicitly disclaims them. If you are geocodinggeocodingThe process of converting an address into geographic coordinates (latitude and longitude). Mailwoman geocodes in a multi-tier cascade: exact address-point match → street interpolation → locality centroid. Each tier is progressively coarser but more widely available. by ZIP centroid and calling it "the address," you are off by anywhere from a few hundred feet to several miles, and the error is systematic for rural areas.
The +4 extension narrows this: it maps to a block faceblock faceOne side of a street between two intersections. Several postal systems specify delivery to block-face granularity — e.g. the USPS ZIP+4 extension. (one side of one streetstreetThe named linear feature along which house numbers are ordered. Decomposes into a name plus street affixes; one of the Tier 2 fine labels. between two cross streetsstreetThe named linear feature along which house numbers are ordered. Decomposes into a name plus street affixes; one of the Tier 2 fine labels.), a single large building, or a PO boxPO boxA numbered mailbox at a post office used as a delivery address instead of a physical street location. Mailwoman tags it as the po_box component; structurally the same family as a subpremise. section. The delivery point codedelivery point code (DPBC — Delivery Point Barcode). The full 11-digit USPS code (5-digit ZIP + 4-digit extension + 2 delivery-point digits) that uniquely identifies a single mailbox, encoded as the barcode on an envelope for automated sorting. narrows further: the specific mailbox.
Step 4 — Carrier delivery
The letter arrives at the destination post office, sorted into trays by carrier routecarrier routeThe delivery path one postal carrier covers on a shift. A ZIP code is fundamentally a set of carrier routes, not a geographic polygon, which is why ZIP boundaries are fuzzy and shift over time. and walk sequence. The carrier loads their vehicle (or satchel) and delivers.
Carriers possess local knowledge that no database has. They know that the blue house on Elm Street is 142 Elm, even though the mailbox says 140. They know that the new apartment building at the end of the block uses the address of the demolished warehouse that stood there before, because USPS hasn't updated the delivery sequence yet. They know that the occupant of 15 Main Street moved to Florida six months ago and mail should be forwarded.
This is the deepest layerlayerOne transformer block — attention plus a feed-forward network, with normalization and residual connections — applied to every position. Stacking layers lets the model build up richer representations; Mailwoman's encoder has 6. of the postal system's fuzziness: the human at the end of the chain who overrides the machine's decisions.
Step 5 — What happens when the address is wrong
If the address is malformed enough that even the REC operators cannot resolve it, the letter goes to dead letter processing. The Mail Recovery Center in Atlanta tries to identify the sender or recipient from the contents. If it cannot, the letter is destroyed or auctioned.
But usually, wrong addresses get handled earlier:
- Return to sender. If the address exists but the recipient has moved, the letter gets a yellow sticker and goes back.
- Forwarding. If the recipient filed a change-of-address, the letter gets a new barcode and continues to the new destination. Forwarding orders last 12 months for first-class mail; after that, mail is returned.
- Carrier correction. The carrier knows the delivery point. If the address is close but not quite right ("142 Elm" when the mailbox says "140"), the carrier corrects it without returning it. This is friction, but it works often enough that the system tolerates it.
What this means for a parser
The postal system has four layerslayerOne transformer block — attention plus a feed-forward network, with normalization and residual connections — applied to every position. Stacking layers lets the model build up richer representations; Mailwoman's encoder has 6. of ambiguity resolution, escalating from cheapest to most expensive:
| LayerlayerOne transformer block — attention plus a feed-forward network, with normalization and residual connections — applied to every position. Stacking layers lets the model build up richer representations; Mailwoman's encoder has 6. | Mechanism | Cost | Success rate |
|---|---|---|---|
| 1. OCR + delivery point barcode | Machine | Near zero | ~85-90% |
| 2. Remote Encoding Center | Human typing | ~$0.03/piece | Most of the remaining ~10% |
| 3. Carrier local knowledge | Human walking | ~$0.50/piece (amortized) | Most of the rest |
| 4. Dead letter / return | Disposal | ~$1.00+ | The rest |
A parser that returns "I'm 70% sure this is Springfield IL, 25% Springfield MA, 5% Springfield MO" is mimicking LayerlayerOne transformer block — attention plus a feed-forward network, with normalization and residual connections — applied to every position. Stacking layers lets the model build up richer representations; Mailwoman's encoder has 6. 2 — human triage — before the letter ever leaves the sender's computer. It tells the downstream system: "you need more information before you can route this." That serves the consumer better than returning "Springfield IL" with 98% confidence because the gazetteergazetteerA geographical index that maps place names and postcodes to real-world coordinates. Mailwoman uses a custom-built Who's On First (WOF) SQLite database as its gazetteer — the 'atlas' half of the grammar/atlas architecture. has a population-weighted default.
The design principle for Mailwoman's confidence modelneural classifierThe machine learning model at the core of Mailwoman's parser — a transformer encoder (~30M parameters) trained from scratch to do BIO token classification over addresses. It learns the 'grammar' of address formats; the gazetteer supplies the 'atlas.': a low-confidence correct answer is better than a high-confidence wrong one. The postal system already works this way. The carrier would rather get an envelope that says "maybe 142 Elm St?" than one that says "definitely 140 Elm St" when the correct address is 142. An uncertain machine gets a second look from a human; a falsely confident one does not.
See also
- How humans break addresses — the failure taxonomy the parser must handle
- The database fallacy — why there is no master list of all addresses
- The 90% trap — why the 90% that works isn't enough