profile

MLforSEO Newsletter ✨

Write for the chunk, not the page ✨ MLforSEO Newsletter #015


Write for the chunk, not the page
A page that reads beautifully for a human can fail retrieval entirely. Here's how to write so a machine can use you.
MLforSEO MLforSEO Academy

Hi there,

Last week we established the shift: you are writing to be selected, not just ranked. This week, the most practical lever you have — how you actually write.

Here is the sentence Beatrice's content lesson turns on: the page is read by humans, but the chunk is retrieved by machines, and those are not the same thing. A retrieval system does not read your article as a whole. It cuts it into segments of roughly 500 tokens (about 350–400 words), passes 3 to 8 of them to the generator per query, and judges each one on its own. A page produces maybe 8–10 chunks, and only the ones that stand alone win a slot.

Zero carried context — the number that changes everything

A chunk carries zero context about what came before or after it. So when a section says “as noted above,” there is no “above” inside the chunk. A pronoun like “it” has no antecedent. A table arrives stripped of the paragraph that explained it. The generator drops the chunk or misuses it — and your best content quietly loses to a competitor who wrote a weaker point more cleanly.

The page a human reads is not the chunk a machine gets

So the design test is blunt: open your content at random, read any ~400-word window, and you should understand exactly what is being claimed and what it means. If you can't, the machine can't either. This is also why headings matter more than you think — when a system scores a chunk for topic match, the heading is one of its strongest signals. “What does CRM vendor lock-in cost mid-market companies?” is a retrieval target. “A closer look” hands the slot to whoever wrote the specific version.

Three properties every claim needs

The fix is not to write less deeply — it is to make each chunk self-contained. Beatrice's test for a claim is that it should be extract-friendly, verifiable and specific. “Costs $25 per user per month” beats “flexible pricing.” “Founded in March 2018” beats “founded recently.” “Fifty integrations including Slack, Jira and Salesforce” beats “many integrations.” And attribution has to live inside the sentence, because your byline and references list sit outside the chunk and never travel with it. Put the scope, the source and the value in the same sentence as the statistic.

Same facts written two ways — one fails retrieval, one gets cited

Watch how this plays out in a live answer. Your chunk on CRM lock-in says the migration costs €40,000–€120,000 over 6–9 months for mid-market firms — self-contained, scoped. A competitor's chunk leans on “as discussed earlier,” which is unresolvable, so it is dropped. The generator uses yours. Same topic, different structure, one survivor. Vagueness here is not neutral; it is a disqualifier. Beatrice expands on both ideas here: the chunk, not the page and vagueness is a disqualifier.

“As noted above” is a retrieval failure. Every navigational phrase written for a sequential reader becomes an unresolvable broken reference.

Sharp beats comprehensive

There is a second failure hiding inside “good” writing. A well-crafted 200-word paragraph covering a problem, its cause and its fix is unified prose — but to the embedding model it is three topics blended into one vector, and a blended vector matches nothing well. It scores too low for every individual query it competes on. A competitor who split those into three headed sections produces three sharp vectors, and each wins its own query. Covering more in less space is a disadvantage in retrieval, not a virtue.

This matters even more once you remember queries decompose. Your section isn't competing once — it is entered into several independent sub-retrievals whose pools barely overlap. A single sprawling guide can win a sub-question or two and lose the rest to dedicated pages whose headings mirror the sub-question as written. Modularity is how you get to compete in more than one of those contests at all.

Four structural habits cause most of the damage, and they are worth learning to spot: the narrative build-up (each section assumes the one before it), the distributed definition (a term defined once in the intro and used everywhere after), the table without context (data in the table, meaning in the surrounding prose), and the multi-claim paragraph (three topics in one block). Each has the same fix — make the section self-contained, even at the cost of a little repetition. In retrieval that repetition is not sloppy, it is necessary.

✎ Exercise 1 — the 10-minute extraction check

  1. Open your most important page and read a random 400-word window. Understandable on its own? If not, neither can the machine.
  2. Ctrl+F for “as noted above”, “see below”, “as we'll explain”. Drive every instance to zero.
  3. Rewrite one vague heading as the exact question a user would search.
  4. Take one key statistic and put its scope, source and value in the same sentence.

✎ Exercise 2 — rewrite one section proposition-first

  1. Pick a section that matters. Make the heading the question it answers.
  2. Write a first sentence that defines the subject, even if you defined it earlier on the page.
  3. State the claim with scope + source + value in one sentence.
  4. Close with one line on what it means for the reader — no “see next section”. Now that chunk survives on its own.

Resources for this edition

▸ Extraction-Readiness Audit (the full 5-check version)▸ Late-Evaluation Rewrite Checklist▸ Read: The chunk, not the page

Those two checks are the front of the five-check audit above. The full method — the complete four-part proposition-first structure, the four failure patterns and the worked before/after rewrites — is inside the course.

Learn the full method →

Next week: the infrastructure that decides whether an agent can even read you.
— Lazarina

P.S. If the random-window test made you wince at one of your best pages, that is the point — and it is usually a 20-minute fix, not a rewrite.


Course by Beatrice Gamba (Head of Innovation, WordLift) · MLforSEO Academy. Join the free community.

© 2024 - 2025 · MLforSEO and MLforSEO Academy · All rights reserved, property of ML Marketing Consulting, Ltd.


Unsubscribe · Preferences

MLforSEO Newsletter ✨

AI/ML news and concepts, demystified. SEO and digital marketing automations shared regularly, as well as updates from the world of the MLforSEO platform and Academy✨

Share this page