Method
How a book’s difficulty is measured
Each book in the library was measured against its own full text — not a sample, not a blurb, not a publisher’s age rating. This page says exactly how, including the places where the method is weaker than the numbers make it look.
What is counted
The text is stripped of front and back matter, then tokenised into words and sentences. From that:
- Total words — running words, so a word repeated fifty times counts fifty times. This is what determines how long the book takes to read.
- Distinct words — how many different word forms appear. Ran and run are two.
- Distinct root forms — the same count after reducing each word to its dictionary form, so ran and run are one. Only the English shelf is lemmatised, so only English books show this.
- Average sentence — words per sentence, across the whole book.
- Sentence length, 99th percentile — one sentence in a hundred is at least this long. Deliberately not the true maximum: in a public-domain text the single longest “sentence” is usually one runaway construction, or a contents list that lost its line breaks, and it tells you nothing about how the book reads.
- Dialogue — the share of the text inside quotation marks. Dialogue tends to use shorter sentences and commoner words, which is why a dialogue-heavy book often reads more easily than its other numbers suggest.
Vocabulary: two different measurements
The English shelf and the other six shelves are not measured the same way, and the difference matters more than any single number on either.
English, against a CEFR-graded wordlist
Every word of an English book is looked up in two vocabulary lists that assign CEFR levels to words. They cover different halves of the scale and are used together:
- CEFR-J Vocabulary Profile 1.5 — A1 to B2, by Tono et al., CC BY-SA.
- Octanove Vocabulary Profile 1.0 — C1 and C2, by Octanove Labs, CC BY-SA.
Together they produce the A1–C2 breakdown, and the “you would recognise this much at B1” figures, which are cumulative: the B1 number includes everything at A1 and A2 as well.
Words the lists do not cover are reported as off-list rather than folded into the hardest band. Most of them are names and places, which a reader absorbs without knowing the word.
The other six languages, against their own corpus
No comparable openly licensed CEFR wordlist exists for French, German, Spanish, Italian, Portuguese or Dutch. Those shelves therefore report something else that is true: what share of the book is built from the two thousand commonest words of that shelf’s own corpus.
They carry no CEFR level, and no A1–C2 breakdown, because assigning one from frequency alone would be a number that looks precise and means nothing. If you see a CEFR band on this site, there is a graded wordlist behind it.
The difficulty score
A number from 0 to 100. Vocabulary dominates it; sentence length is a second term, and a capped one. For English:
(100 − coverage at B1) × 1.9 + off-list share × 1.1
+ syntax
and for the other shelves, where there is no CEFR coverage to use:
(100 − coverage by the commonest 2,000 words) × 2.4 + syntax
where syntax is (average sentence length − 15) × 0.7,
floored at zero and capped at 14 — so sentence length can move a book by
at most fourteen points.
That cap is the most important decision in the formula. Weighting sentence length heavily is what makes readability scores put Ulysses in the middle of the shelf: stream of consciousness is fragmentary, so it scores short sentences while being the hardest book there. It is the same trap the Flesch score falls into with Faulkner, reached by a different road.
Where the book gets harder
Each book is also split into ten equal spans and each span scored on its own. That is what the small chart on a book page shows, and why the hardest span is named: an opening chapter is often not a fair sample of the book behind it.
Reading time
Every “about 11 h at a learner’s pace” on this site is total words ÷ 130 words per minute. 130 is a deliberately slow figure — a realistic pace in a language you are still learning. Native-language reading is roughly double that, and quoting the native figure would understate every book in the library.
What none of this can see
These are measurements of words and sentences. They cannot see how hard the ideas are. Plato is plain-spoken and still Plato; a children’s book about grief is easy to read and not easy. A low difficulty score means the sentences are short and the words are common. It does not promise the book is simple.
Two more limits worth stating. The difficulty score is calibrated within a shelf, so comparing a score across languages is meaningless — a 40-book corpus covers less of any given book than English’s 170-book one, which is a fact about corpus size and not about the language. And the CEFR wordlists disagree with each other at the edges, as all such lists do; a word sitting one band out is normal and does not move a book noticeably.
Reading a book that is above your level
The measurements exist to answer one question — can I read this? — and the honest answer is usually “yes, with help.” Comprehension gets comfortable somewhere around 95–98% coverage, and most learners sit below that on most real books. What makes the gap survivable is not looking up fewer words, it is looking them up without leaving the page.
How many words you need to read a book goes through the research on that threshold.