Before an AI engine ever quotes your page, it breaks that page into pieces. It splits your content into passages, stores each one separately, and when a question comes in it pulls the single passage that best matches, then writes its answer from that. The model usually never reads your whole page. It reads a chunk. This is called chunking, and it changes what "good content" means. A page can be accurate, well written, and complete, and still lose the citation because the one sentence that answered the question got split away from the sentence that named the subject.
The unit that has to stand on its own is no longer the page or even the section. It is the chunk, and you do not get to decide where the chunk boundaries fall. What you can do is write content that survives any reasonable split, and audit what you already have for the places it breaks. This guide explains how chunking works, why it quietly costs B2B pages citations, and how to run the audit on your own content in an afternoon.
What chunking is, and why it decides citations
Chunking is the step where a retrieval system cuts a document into smaller passages so it can search them. Each chunk is converted into a numeric representation of its meaning and stored. When someone asks a question, the system compares the question to all those stored chunks and retrieves the closest matches, and those retrieved passages, not the full page, are what the model uses to compose its answer.
That is the part most content teams miss. You are not being evaluated as a page. You are being evaluated as a collection of passages, each of which might be pulled out and read completely alone, with none of the surrounding context that makes it clear on your actual page. A buyer's question maps to one chunk. If that chunk answers the question fully by itself, you can be cited. If it does not, the engine moves on to a competitor's chunk that does.
How systems actually split your content
There is no single way engines chunk, and that is the first thing to understand. Different systems use different methods, and you can neither see nor control which one is applied to your page.
The crude approach is fixed-size chunking: cut the text every so many words or tokens, often with a little overlap between chunks so a sentence on the boundary is not lost. It is simple and common, and it is also the method most likely to slice through the middle of your answer, because it counts length rather than meaning. A better approach splits on natural units: by sentence, by paragraph, or by structural markers like your headings, so each chunk maps to a real section. The most sophisticated approach is semantic chunking, which tries to detect where one idea ends and the next begins and splits there. Chunk sizes vary too, from a single sentence up to a few hundred words, depending on the system.
The practical takeaway is not to guess which method a given engine uses. It is to accept that you cannot know, that it differs from engine to engine, and that it can change without notice. That uncertainty is exactly why the winning move is content that holds together under any of these methods, rather than content tuned to one you happened to guess.
Why a split can quietly cost you the citation
The failure is easy to picture once you see chunking as cutting. Imagine a section that opens by naming your product and then, three sentences later, states the specific fact a buyer is asking about, using "it" to refer back to the product. On the page, that reads fine. But if the chunker splits between those sentences, the passage with the fact begins with "It also supports..." and the model has no idea what "it" is, which product, or which capability. The passage is retrieved and then discarded as unusable, or worse, attached to the wrong subject and misquoted.
The same thing happens when an answer is spread across two paragraphs, when a table is separated from the sentence that explains it, when a list of steps is cut off halfway, or when the real answer sits far below a long wind-up. In every case the information exists on your page, so your team assumes it is working, but the passage that carries it cannot stand alone. This is why a plainer competitor page whose answer sits whole inside one tidy passage can beat a more thorough page whose answer is scattered. Chunking rewards self-contained, not comprehensive.
What a chunk that survives looks like

A passage that survives names its own subject rather than leaning on a pronoun from an earlier sentence, states its answer in full, and does not depend on the paragraph above or below it to make sense. A passage that breaks starts mid-thought, refers to things by "it" or "that" or "this approach" with the antecedent living in a different chunk, or trails off before the answer is complete. The fix is not more words. It is making each passage carry its own context: repeat the subject, put the answer and its qualifier in the same place, and keep each self-contained idea inside a single section rather than stretched across several.
This is the same instinct behind writing extractable content in general, which we cover in how to structure B2B content for LLMs. The difference here is the lens: instead of asking "is this well structured," you are asking "if a machine cut this page into pieces at boundaries I cannot predict, would each piece still hold up." That question is what the audit below answers.
How to audit your content for chunking problems
You can run this on your most important pages, the ones you want cited, in an afternoon. Work section by section.
Start with the removal test, which does most of the work. Take each section under an H2 or H3, copy it out on its own, and read it cold. If it fully answers a question with no help from the rest of the page, it will survive almost any chunker. If you have to scroll back up to understand what it is talking about, it will not, and you have found a problem to fix.
Then run a pronoun and antecedent pass. Scan the opening sentence of every section and every paragraph for "it," "this," "that," "they," and "the tool" where the thing being referred to was named in an earlier paragraph. At the start of a passage, replace the pronoun with the actual subject. It reads slightly more repetitive to a human and far more clearly to a machine that only sees that passage.
Check that each answer is complete in one place. If the response to a likely question is built up across three paragraphs, a chunker may hand the engine only one of them. Pull the core answer into a single self-contained passage near the relevant heading, then elaborate after. Do the same for anything visual or listed: keep a table with the sentence that introduces it, and keep a numbered process together rather than letting it span a boundary.
Finally, front-load the answer under question-shaped headings. A heading that states the question, followed immediately by a complete answer, gives the chunker a clean unit to cut and gives the engine a passage that maps directly to what the buyer asked. Long wind-ups before the answer are where good information goes to get split off from its own question.
A worked example: fixing a section that breaks
Here is what the audit looks like in practice on a typical B2B page. Say a product page has a section headed "Integrations" that reads: "Our platform connects to the tools your team already uses. We built the roadmap around what customers asked for most. It syncs two ways with your CRM, so citation data flows into HubSpot automatically, and there is no extra charge for it on the standard plan."
Run the removal test on the last sentence, the one a buyer's "does it sync with HubSpot" question would pull. On its own it reads: "It syncs two ways with your CRM, so citation data flows into HubSpot automatically, and there is no extra charge for it on the standard plan." What is "it"? The passage never says. A model retrieving this chunk knows something syncs with HubSpot at no extra cost, but not what, so it cannot safely attribute the capability to your product, and the citation goes to a competitor whose passage names itself.
The fix is not longer, it is self-contained. Rewrite the section so the key passage carries its own subject and answer: "Omnibound syncs two ways with HubSpot. Citation data flows into HubSpot automatically on the standard plan, at no extra charge." Now cut that out and read it cold and it still answers the question completely. Notice what changed: the subject is named instead of implied, the filler sentence about the roadmap that added nothing to the answer is gone, and the whole answer lives in one place instead of being set up by a sentence that a chunker might leave behind. That is the entire move, repeated across every section you want cited.
What you control and what you do not
It is worth being clear about the boundary, because a lot of chunking advice overpromises. You do not control where the splits happen, how big the chunks are, whether a given engine respects your headings, or which method it uses this month. Chasing those is a waste of effort, because they are opaque and they change.
What you fully control is whether each passage of your page can stand on its own. That single thing holds up under every chunking method, which is what makes it worth the work: you are not betting on a guess about one engine's pipeline, you are making your content win under all of them. The technical layer underneath this matters too, since a chunk can only be created from content a crawler can actually reach and parse, which is the subject of our B2B guide to technical AEO. But once your content is reachable, self-contained passages are the lever you own.
Frequently asked questions
Is chunking the same as how Google indexes passages?
They are related. Google has long ranked individual passages within a page, and AI retrieval takes the same idea further by splitting your content into chunks, embedding them, and pulling the best match. In both cases the practical response is identical: make each passage answer on its own.
Can I set my own chunk boundaries?
Not for third-party engines. You can influence where natural splits fall by using clear headings, short paragraphs, and self-contained sections, but the engine's pipeline makes the final call, and it varies. That is why writing for any reasonable split beats trying to dictate one.
Does adding more headings fix chunking?
Headings help, because structural chunkers often split on them and they give each section a clear boundary. But headings alone do not save a section whose answer relies on a pronoun pointing three paragraphs up. Structure plus self-containment is what works.
How is this different from just writing clearly?
It overlaps, but the chunking lens catches things clean writing misses, specifically dependencies between passages. A paragraph can be perfectly clear in place and useless once separated from the paragraph it depends on. The removal test is what surfaces those.
What content is most at risk?
Long, flowing thought-leadership pieces where the answer to a question is developed gradually, and pages that lean on a strong intro to set up everything below it. Reference-style pages with self-contained sections tend to chunk well already.
Run the audit on your top pages
Chunking is not a new tactic to adopt, it is a fact about how AI reads that you can either work with or ignore. The move is small and repeatable: take the handful of pages you most want cited, run the removal test on each section, fix the passages that cannot stand alone, and you have made your content resilient to a process you cannot see and cannot control.
The last step is to check whether it changed anything, because chunking fixes show up as citations you start earning, not as anything visible on the page itself. You can see where your pages stand today with our AI Search Visibility Checker, and if you want to watch whether engines start citing the passages you fixed, that is what live model checks are for. Omnibound's AI Search Intelligence shows you which of your pages engines actually pull from, so you can point the audit at the pages where a fixed chunk will earn you the most.
Turn Your Content Into AI-Search Winners
Get cited across ChatGPT, Claude & Perplexity — not just ranked on Google.
- Increase AI citations
- Improve answer visibility
- Track brand mentions in LLMs