A language app is only as good as the week behind each language
One Holt language pair took seven days: 14,484 example sentences and a blind check on 120 dictionary words by a judge that did not write them — 96.7% pass, 0.0% substantive defects. The same report lists what learners in our older languages still see today.
A language app is only as good as the week behind each language. That is the whole answer. The rest of this piece is that week, for one Holt language pair: seven days, 14,484 example sentences, a blind check on 120 dictionary words by a judge that did not write them — and the part release notes usually skip, a report of what learners in our older languages still see today.
What you see, and what it took
When you study a word in Holt, you get a meaning, a translation, an example sentence, and, when the app asks you a question, a few wrong answers that have to be tempting and still be wrong. None of that is filler. Every piece is written and checked. For one pair, the week ended here:
one language pair · seven days
entries 6,664
senses (meanings) 7,242
example sentences 14,484 two per sense
glosses 28,968 four per sense
distractor links 27,431 which wrong answers are plausible
blind golden set 120 dictionary words
judged by a checker that did not write them
pass rate 96.7%
substantive defects 0.0%The last block is the one we care about most. A check run by whoever wrote the material tends to agree with itself. So the 120 words are a fixed set, and the judge had no hand in them. 96.7% of them passed, and nothing the judge flagged was a substantive defect.
The report that makes this not an ad
Here is the part we could have left out. The report that closed the week has a section headed:
found in the LIVE pairs — not a regression of this wave,
learners see it todayIt lists what building the new pair turned up in languages that had already shipped:
6 headwords in one shipped language printing a
grammatically impossible article
9 accent minimal pairs across two shipped languages
outside the grading lists: a learner who answered with
the near-miss word was scored correct, and the wrong
item was scheduled for review
472 of 6,001 senses (7.9%) in one pair where a second
answer is equally right; 528 (7.3%) in anotherThose are real, and they are live. The fix that came out of the week is small and boring: in the new language, accent pairs are now guarded by a test that scans the whole word bank instead of a hand-kept list. The two languages that shipped first are openly owed the same test. It is written down as a debt, not as done.
Our view: adding a language is the cheapest audit of the languages you already have. A new build re-reads every assumption the old ones froze. We would rather publish what it finds than wait for a learner to find it.
How the week is ordered
Seven days of this is not heroics. It is the order of the work.
Every new language arrives with its own word list. The obvious move is to build from that list. Two of our earlier pairs were built that way, and months later 295 concept pairs had to be merged by hand — the same idea, created twice from two sides.
This pair was built the other way round. Start from the shared catalog of concepts, fan them out into the new language, and use the language's own list only to set priority and show what the catalog is missing. Nothing gets written until a gate passes.
before any authoring
catalog coverage 96.9% floor 80%
delta 9.5% ceiling 20%
authoring
5,879 concepts from the shared catalog
40 batches × 147, every batch machine-checked
the language's own word list → second pass: order + delta
result
456 cross-language duplicate concepts prevented up front
vs 295 cleaned up after the fact the previous wayOr, in the words of the build notes: “Ordering in that direction prevented 456 cross-language duplicate concepts up front, against the 295 cleaned up after the fact the previous way.”
What we cut
The inversion has a price, and we chose to pay it in the open.
602 lemmas from the language's own list with no match in
the catalog → declared as delta and written
1,576 lemmas deliberately left outEach of those 1,576 would have produced either a concept that already had a native gloss, or a fresh cross-language duplicate. So they are not in the pair. A bigger number would have been easy. We think a dictionary without double entries is the better one to learn from.
Cleaning up without taking words away
Duplicates already sitting in the shared layer still had to go. The rule for that clean-up was a promise to learners: it may not remove a word anyone can see.
surrogate nodes 4,440 → 3,667 (773 merged)
exact core duplicates 129 → 4
cross-language duplicate
pairs 301 → 1
learner-visible inventory
es 6,001 en 5,936 ru 5,779 de 5,244 unchanged
senses that lost their native side: 0
rollback tables + a 748-row merge map left in production“The learner-visible inventory stayed identical to the row.” That clean-up has its own story, and it gets its own piece in this series.
Why a week, and not an afternoon
Our view, stated plainly: anyone can generate a dictionary in an afternoon. The week goes into what generation does not give you — senses that are actually different, examples that fit them, wrong answers that teach, a judge who did not write the answers, and a report that names what is still broken. That is why adding a language to Holt is slow. We think that is the right speed.
FAQ
Does a week per language mean new languages come slowly? Yes. One pair took seven days of work, and the build is ordered so that coverage gates pass before any writing starts. We would rather add fewer languages than ship one we have not checked.
What does “blind check” mean here? A fixed set of 120 dictionary words, judged by a checker that had no hand in writing the material. On this pair it passed at 96.7%, with 0.0% substantive defects.
The older languages have known issues. Am I affected? Possibly, and that is why the report says so. The listed cases are 6 headwords with an impossible article in one language, 9 accent pairs across two languages, and a class of senses — 7.9% in one pair, 7.3% in another — where a second answer is equally right. The new language has a whole-bank accent test; the first two languages are owed the same one.
Did the clean-up delete words I was studying? No. After 773 nodes were merged, the visible word counts stayed identical — es 6,001, en 5,936, ru 5,779, de 5,244 — and no sense lost its native side.
Try it
Holt is open at holt.iloblique.com. The short version of this week is here: {{anchor_url}}. The full Holt case — what the app measures and why — is at iloblique.com/holt.