Skip to content

A word list is not a dictionary: we recounted Holt's with two keys

Counting by the label a human reads understated contamination in Holt's dictionary by two to three times. With a second independent key the honest figure came back at about 18% and 14%. After the merge the learner-visible dictionary was unchanged.

holt · evals · vocabulary · case

Counting duplicates by the name a human reads flattered us, and we can say by how much: counted by the display label, our duplicate metric had been understating contamination by two to three times.

Holt is our language app, and the number on its shelf is counted in senses. Before adding the next language pair, we re-ran our own duplicate metric with a second, independent key. The honest figure came back at roughly 18% of the German wave and 14% of the Spanish wave. Nothing in the product had broken. The measurement had.

Below is what the recount found, what we merged, what it cost, and what the learner saw afterwards — which was nothing. The dictionary they open is the same size it was before.

Why a count by headword flatters

Our duplicate report offered one candidate per record. One. So a duplicate whose real match was not the first candidate was not counted as a near miss — it was counted as clean.

report: one candidate per record

  Hersteller  -> candidate #1: (no match)     -> recorded as clean
                 real match:   manufacturer   -> never shown

  Säge        -> candidate #1: (no match)     -> recorded as clean
                 real match:   saw            -> never shown

  servidor    -> candidate #1: (no match)     -> recorded as clean
                 real match:   server         -> never shown

those three, plus 21 more cross-language pairs, were invisible

counted by headword:    understated two to three times
counted with two keys:  de ~18%   es ~14%

The report was not lying. It was answering a narrower question than the one we thought we had asked.

An honest count needs two keys that do not know about each other

One record in our data carries more than one projection: each wave writes an English and a Russian glue label for the same concept. A duplicate is only a duplicate if both projections land on the same node in the shared core. Then a human reads a stratified sample, tier by tier, to check that the machine's answer survives contact with a person.

We also tried the obvious cheap thing first, and it failed:

key A: the English glue label of the record
key B: the Russian glue label of the record
duplicate := A and B resolve to the same core node
             + a hand-read stratified sample per tier

rejected: free-text similarity over example glosses
  -> could not separate a duplicate from a coincidence on our data

Our opinion: a metric that cannot separate a duplicate from a coincidence is not a weak signal, it is a decoration, and we would rather publish the two-key number with its ugly result than a pretty number with nothing behind it.

The same audit surfaced a second figure nobody had measured at all: coverage of the shared core was 46.5% for German and 59.8% for Spanish, against a documented floor of 80%. In our reading that is the cause, and the duplicates were the symptom. The inventory had been built from each target language's own word list instead of from the shared core, so two waves invented two nodes for one concept and neither wave could see the other.

The audit ran as a separate, read-only session, with its own tooling and 16 tests, before a single row was touched. Our opinion: an audit that can edit the thing it is auditing is not an audit, it is a cleanup with an opinion of itself.

The merge, and what the learner saw

                        before  ->  after
surrogate nodes          4,440  ->  3,667     (773 collapsed)
exact core duplicates      129  ->      4
cross-language pairs       301  ->      1

learner-visible dictionary, unchanged:
  es 6,001  ·  en 5,936  ·  ru 5,779  ·  de 5,244
  no sense lost its native side

kept in production: rollback tables + a 748-row merge map

That last pair of lines is the point of the whole exercise. The thing we deduplicated was the plumbing underneath the dictionary, not the dictionary. Nobody woke up to a smaller vocabulary. And the merge map stays in production at 748 rows, because a merge you cannot walk back is a guess you have made permanent.

The half that is not cleanup

Cleaning up is only half of a promise to a learner. The other half is that the next language pair does not arrive with the same duplicates in it. Our two earlier pairs, each built from its own word list, left 295 concept pairs that had to be reconciled by hand months later. So we ordered the next pair the other way round.

blocking gate, measured before any authoring starts:
  catalog coverage     96.9%   floor    80%   pass
  delta                 9.5%   ceiling  20%   pass

build order:
  5,879 concepts taken from the shared catalog
  laid into the new language in 40 batches of 147, each machine-checked
  the source's own word list demoted to a second pass:
    it sets order and shows the difference, and nothing else

prevented up front:  456 cross-language duplicates
the old way:         295 cleaned by hand, months later

The cost, stated plainly: 602 source lemmas were declared as delta and written; another 1,576 were deliberately not declared. We cut them. A word list that disagrees with the shared catalog is a proposal, not an obligation, and we would rather ship a smaller honest overlap than re-open the same duplicate problem one language later.

What one pair actually costs

This is the part that cannot be faked, and it is the answer to "why should we believe you":

one language pair, seven days, 54 commits, phase by phase
  6,664 entries · 7,242 senses · 14,484 examples · 28,968 glosses
  27,431 distractor edges in the link layer
  blind golden set: 120 lemmas, judged by an independent loop
  96.7% pass · 0.0% substantive defects · 3,633 green tests

The build report for that pair carries a section we did not enjoy writing: found in the LIVE pairs — not a regression of this wave, learners see it today. Every new pair audits the ones already shipped, and what it finds goes in the report under its own heading instead of into the next wave's numbers.

FAQ

Why not just publish the bigger, friendlier number?

Because the number on a shelf is a claim about what a learner will actually meet, and a count by display label cannot support that claim. Our opinion: any vocabulary figure should travel with the key it was counted by, ours included.

Did the cleanup take words away from learners?

No. The counts a learner sees came out of the merge identical — es 6,001, en 5,936, ru 5,779, de 5,244 — and no sense lost its native side. What shrank was the surrogate layer underneath: 4,440 nodes down to 3,667.

Can embeddings or fuzzy matching replace the second key?

Not on our data. Text similarity over example glosses could not separate a duplicate from a coincidence, so we dropped it. Two independent projections agreeing on one node gave a usable answer, and a hand-read sample per tier kept it honest.

What would change our mind?

A third key that disagrees with our two. If some independent projection of the same records produced a materially different contamination figure, the 18% and 14% would be wrong and we would publish the correction here, at this length.

How do I know this holds for the pairs shipped after the audit?

By the gate, not by our word: catalog coverage and delta are measured before authoring begins (96.9% against a floor of 80%, 9.5% against a ceiling of 20%), and a pair that misses either does not get written.

The short version, and the dictionary itself

We put the short version of this on camera — the counting mistake, the honest figure, and the line that matters most, that the learner lost nothing: {{anchor_url}}

And the dictionary these numbers describe is open at holt.iloblique.com. Open a language, look at a word, read the senses under it. That is the check we ran on ourselves, and it is the one we would rather you ran too.