Official Lexile measures come from MetaMetrics. We are not a licensed Lexile product. We use the same two text features Lexile is built on — word frequency and sentence length — and report an estimated grade band, not an L score. If you already teach with Lexile, the first box is the theory you know. The rest is how this app operationalizes it.
Lexile puts the reader and the text on one developmental scale (typically 0L–2000L; below 0L is coded BR, Beginning Reader). The classroom use you already know is the match: a text near the student’s measure (often cited as about a 75% comprehension target when they are aligned) is instructional; well below is independent; well above is frustration. Stretch bands by grade exist, but Lexile itself is not a grade equivalent.
The text measure is a regression on two proxies. The exact equation is proprietary. The two predictors are public:
Lexile does not score prior knowledge, text structure (how a narrative or informational piece is organized), decoding versus language comprehension, or motivation. Those still matter for a special-ed placement; they just are not in the number.
“The dog ran to the park.” — high-frequency words, mean length 6, low semantic and syntactic demand. “The committee’s decision to postpone the merger was met with considerable skepticism.” — several low-frequency content words and one long sentence, both levers up.
A computer cannot “just read” a story the way a person does. It has to break the page into pieces it can count. That is all this section is.
We only score the meaning-carrying words — nouns, verbs, adjectives, adverbs. Little function words do not count. Names of people are left out too, because a name is not a vocabulary burden.
After NLP has done that sorting, we have the two Lexile-style averages: how common the content words are (semantics), and how many words sit in a typical sentence (syntax). Those two averages are what the next box turns into a grade.
We do not output an L score. We take the same two predictors Lexile uses — mean log word frequency (semantic) and mean sentence length (syntactic) — and map them onto an estimated grade with a linear combination. The coefficients are ours. They are not fitted to a large labeled set of worksheet stories, so treat ordering as trustworthy and the exact decimal as a guide.
Common words raise the “how common” number, and because that term is subtracted, the grade goes down — that is the semantic lever. Longer sentences add to the grade — that is the syntactic lever. NLP finds the tokens, we average them, the formula outputs an estimated grade.
A target of “Grade 4” is a band, 3.5–4.5, analogous in spirit to a Lexile range rather than a single L. Inside the band is a pass. Outside it — too hard or too easy — we try to move semantic demand, syntactic demand, or both.
The formula only moves if we change one of those two averages. So every fix is aimed at rarer or more common words, and at longer or shorter sentences, until the score sits inside the target band.
Words the teacher asked us to keep (a science term the lesson is teaching) stay as they are. We spend the difficulty budget on everything else.
Try it on the Reading Level Corrector. On the Worksheet Generator you will see the first draft’s grade next to the final grade. That gap is this process working.