Google’s BERT update, rolled out in late 2019, is often summarized as “Google got better at understanding language,” which is true but vague enough to obscure why it mattered so much for how content should actually be written. Understanding the specific mechanism behind BERT explains why keyword-density thinking has aged so poorly since.
What Bidirectional Actually Means Here
Earlier language models generally processed text in one direction — reading left to right, building an understanding of each word based only on what came before it. BERT, short for Bidirectional Encoder Representations from Transformers, processes text in both directions simultaneously, using the surrounding context on both sides of a word to interpret its meaning. A word’s meaning often depends as much on what follows it as what precedes it, and BERT was built specifically to capture that.
Why This Mattered Most for Small, Easily Overlooked Words
The classic example Google used to illustrate BERT’s impact involved a search about a Brazilian traveler needing a visa for the United States. The word “to” in that query carries the entire directional meaning — get it wrong and the search results describe the opposite travel direction entirely. Prepositions, conjunctions, and other small connecting words had historically been treated as near-noise by earlier ranking systems focused heavily on the more prominent nouns and keywords in a query. BERT’s contextual approach gave these small words the interpretive weight they actually carry in natural language.
Why This Directly Undermines Keyword-Density Thinking
Content written primarily to hit a target keyword a certain number of times, often at the expense of natural sentence structure, works against exactly the kind of contextual understanding BERT introduced. A page that repeats an exact-match phrase mechanically, without the natural surrounding language that gives that phrase its actual meaning in context, provides less useful signal to a system built to understand meaning from context than a page written naturally, even if the naturally written page uses the target phrase fewer times.
What This Means for Writing Content on Any Site, Including a Network
Content built to satisfy a search intent thoroughly and naturally, addressing the nuances of a query the way a knowledgeable person would actually explain it, aligns with what BERT and its successors are designed to reward. This applies just as much to content across a network of sites as it does to a single flagship property — content produced at scale that reads as mechanically keyword-stuffed is working against the same contextual understanding mechanisms regardless of which domain it’s published on.
BERT Was a Foundation, Not a One-Time Update
BERT’s contextual language understanding became a foundational component that subsequent language and ranking systems have continued to build on, rather than a single discrete update whose effects faded over time. This means the underlying principle — content should read as genuinely natural, contextual language rather than a mechanically assembled keyword pattern — has only become more relevant as Google’s language understanding capabilities have continued to advance since BERT’s initial rollout.
Applying This to Content Planning Going Forward
Practically, this means content briefs built around hitting a specific keyword density, or forcing an exact-match phrase into a sentence where it doesn’t naturally belong, are optimizing for a model of search that BERT specifically moved away from years ago. Our full breakdown of what BERT actually did goes into more detail on the transformer architecture behind it and what that means for structuring content briefs today.
The practical shift is straightforward even if the underlying technology is complex: write for the person asking the question, in the natural language they’d actually use and expect, and the contextual signals that matter to modern language models tend to follow from that naturally rather than needing to be engineered in separately.
