Skip to content

A language app is only as good as the week behind each language

One Holt language pair took seven days: 14,484 example sentences and a blind check on 120 dictionary words by a judge that did not write them — 96.7% pass, 0.0% substantive defects. The same report lists what learners in our older languages still see today.

holt · language-learning · data-quality · vocabulary · case

A language app is only as good as the week behind each language. That is the whole answer. The rest of this piece is that week, for one Holt language pair: seven days, 14,484 example sentences, a blind check on 120 dictionary words by a judge that did not write them — and the part release notes usually skip, a report of what learners in our older languages still see today.

What you see, and what it took

When you study a word in Holt, you get a meaning, a translation, an example sentence, and, when the app asks you a question, a few wrong answers that have to be tempting and still be wrong. None of that is filler. Every piece is written and checked. For one pair, the week ended here:

one language pair · seven days

entries                   6,664
senses (meanings)         7,242
example sentences        14,484   two per sense
glosses                  28,968   four per sense
distractor links         27,431   which wrong answers are plausible

blind golden set            120   dictionary words
judged by                         a checker that did not write them
pass rate                 96.7%
substantive defects        0.0%

The last block is the one we care about most. A check run by whoever wrote the material tends to agree with itself. So the 120 words are a fixed set, and the judge had no hand in them. 96.7% of them passed, and nothing the judge flagged was a substantive defect.

The report that makes this not an ad

Here is the part we could have left out. The report that closed the week has a section headed:

found in the LIVE pairs — not a regression of this wave,
learners see it today

It lists what building the new pair turned up in languages that had already shipped:

6     headwords in one shipped language printing a
      grammatically impossible article

9     accent minimal pairs across two shipped languages
      outside the grading lists: a learner who answered with
      the near-miss word was scored correct, and the wrong
      item was scheduled for review

472   of 6,001 senses (7.9%) in one pair where a second
      answer is equally right; 528 (7.3%) in another

Those are real, and they are live. The fix that came out of the week is small and boring: in the new language, accent pairs are now guarded by a test that scans the whole word bank instead of a hand-kept list. The two languages that shipped first are openly owed the same test. It is written down as a debt, not as done.

Our view: adding a language is the cheapest audit of the languages you already have. A new build re-reads every assumption the old ones froze. We would rather publish what it finds than wait for a learner to find it.

How the week is ordered

Seven days of this is not heroics. It is the order of the work.

Every new language arrives with its own word list. The obvious move is to build from that list. Two of our earlier pairs were built that way, and months later 295 concept pairs had to be merged by hand — the same idea, created twice from two sides.

This pair was built the other way round. Start from the shared catalog of concepts, fan them out into the new language, and use the language's own list only to set priority and show what the catalog is missing. Nothing gets written until a gate passes.

before any authoring
  catalog coverage      96.9%    floor    80%
  delta                  9.5%    ceiling  20%

authoring
  5,879 concepts from the shared catalog
  40 batches × 147, every batch machine-checked
  the language's own word list → second pass: order + delta

result
  456 cross-language duplicate concepts prevented up front
  vs 295 cleaned up after the fact the previous way

Or, in the words of the build notes: “Ordering in that direction prevented 456 cross-language duplicate concepts up front, against the 295 cleaned up after the fact the previous way.”

What we cut

The inversion has a price, and we chose to pay it in the open.

602     lemmas from the language's own list with no match in
        the catalog → declared as delta and written

1,576   lemmas deliberately left out

Each of those 1,576 would have produced either a concept that already had a native gloss, or a fresh cross-language duplicate. So they are not in the pair. A bigger number would have been easy. We think a dictionary without double entries is the better one to learn from.

Cleaning up without taking words away

Duplicates already sitting in the shared layer still had to go. The rule for that clean-up was a promise to learners: it may not remove a word anyone can see.

surrogate nodes              4,440  →  3,667   (773 merged)
exact core duplicates          129  →      4
cross-language duplicate
pairs                          301  →      1

learner-visible inventory
  es  6,001   en  5,936   ru  5,779   de  5,244   unchanged
senses that lost their native side:   0
rollback tables + a 748-row merge map left in production

“The learner-visible inventory stayed identical to the row.” That clean-up has its own story, and it gets its own piece in this series.

Why a week, and not an afternoon

Our view, stated plainly: anyone can generate a dictionary in an afternoon. The week goes into what generation does not give you — senses that are actually different, examples that fit them, wrong answers that teach, a judge who did not write the answers, and a report that names what is still broken. That is why adding a language to Holt is slow. We think that is the right speed.

FAQ

Does a week per language mean new languages come slowly? Yes. One pair took seven days of work, and the build is ordered so that coverage gates pass before any writing starts. We would rather add fewer languages than ship one we have not checked.

What does “blind check” mean here? A fixed set of 120 dictionary words, judged by a checker that had no hand in writing the material. On this pair it passed at 96.7%, with 0.0% substantive defects.

The older languages have known issues. Am I affected? Possibly, and that is why the report says so. The listed cases are 6 headwords with an impossible article in one language, 9 accent pairs across two languages, and a class of senses — 7.9% in one pair, 7.3% in another — where a second answer is equally right. The new language has a whole-bank accent test; the first two languages are owed the same one.

Did the clean-up delete words I was studying? No. After 773 nodes were merged, the visible word counts stayed identical — es 6,001, en 5,936, ru 5,779, de 5,244 — and no sense lost its native side.

Try it

Holt is open at holt.iloblique.com. The short version of this week is here: {{anchor_url}}. The full Holt case — what the app measures and why — is at iloblique.com/holt.