The Semantic Blueprint · Part I, The Semantic Foundation
The Semantic Ceiling
The first impression given by a large language model (LLM) is astonishing. So astonishing, in fact, that we have since collectively spent approximately $1.5 trillion1 on a future that seems perpetually imminent but somehow never arrives.
We have trained bigger and bigger models, built more datacenters, invented new hardware, dreamt up new tooling, built RAG pipelines, and dug up every scrap of text ever created by humans for training data. Despite all this progress, your LLM-based AI continues to be confidently wrong a worrying percentage of the time. This cannot be solved at the model level.
There are whole classes of problems in AI that won’t be solved with a bigger model, only better data.
The reason an LLM is sometimes confidently wrong is, surprisingly, the same reason it is so often confidently right. Years before the world first saw ChatGPT, Google was working on novel approaches to machine translation2. This had long been a uniquely hard problem. Translation is more than replacing one word with its foreign counterpart. Syntax and structure change. The grammar shifts in ambiguous ways. Then there are interpretation subtleties. Ciao, in Italian, can translate as “hi” or “bye” depending on the context.
This is where semantics come into play. Open a dictionary and you’ll quickly discover that language is vague and overloaded. One word can have many different definitions depending on how it is used. Finding the correct translation for a given word depends more on the concept the word is conveying in each context and less about what the word is.
Semantics is the question of whether we are talking about the same thing.
We humans navigate semantics intuitively. Over time we have learned more than a vocabulary, we have learned the semantic properties of words. We understand that some words imply more than they express. We understand that meaning is contextual. We automatically connect these contexts to language so fluently that we overlook how challenging this can be for machines. We’ve grappled with whether a machine could perform such tasks for over 75 years3.
The transformer architecture was a genuine breakthrough. Ultimately, the team at Google built a highly parallelizable approach to building a model of inferred semantics and différance4 expressed as a high-dimensional vector representation. In essence, meaning is decoupled from the text which enables meaning to be projected into different linguistic and grammatical forms.
At their very core, LLMs begin with a pretrained model that embeds these weightings. Input is tokenized5, translated into underlying vector embeddings, then fed into the transformer model. The output is a probability distribution over many possible output tokens. The model then selects one output token using a non-deterministic selection process to keep model output varied. The output token is added to the stack of input tokens and the process is repeated. In other words, the output becomes part of the input. Each new token attends to all previous tokens simultaneously. Both the original input and the emerging output continue to reshape the response into a statistically plausible completion that is surprisingly accurate in most cases.
The results were a significant improvement over previous approaches to machine translation. In this way the model can programmatically disambiguate the “greeting” and “farewell” forms of the word ciao. It is a mathematical approximation of expressed semantics that works remarkably well.

Although the progress in machine translation was undeniable, research into the broader utility of these models continued. By 2020, researchers at OpenAI began systematically studying what happened as these models scaled7. What we have since discovered is scaling these models results in an ever-more-convincing simulacrum of intelligence. Because the model operates in language itself, conversational interaction with these models can be downright startling. We must not lose sight, however, of the underlying mechanism at play.
The architecture makes a model extraordinarily good at guessing. When the guesses are correct it looks like magic. When the guesses are wrong… well, take your pick from any of the high-profile consequences of an AI implementation that was not ready for prime time.
The only distinction between factual output and hallucination from an LLM is whether the response happens to line up with reality. Hallucinations aren’t an implementation bug; they are the entire mechanism of action for an LLM.
There is no mechanism in the architecture for truth, only probability.
The entire value proposition of artificial intelligence rests on the promise of understanding. Does the model understand the prompt? Does the model understand the facts? Does the operator understand the response? Understanding requires a substrate; understanding requires knowledge.
1.1 An LLM Does Not Know Anything
It is very possible that you and I have just arrived at an impasse. I made the statement that an LLM does not know anything, and there are many who would disagree. Rather than debating who is right or wrong, there is a deeper question that should be answered. What does it mean to know something? What is knowledge?
Intuitively, we might first define it. “Knowledge is the stuff you know.” Then we catch the circular reference and pivot. “It’s things you know… facts” but maybe that does not quite sit comfortably. What if we know things that are wrong? Do we know those things? We know facts and facts are true… but how do we know they are true?
Philosophers have grappled with this question for millennia8. Plato untangled our circular reasoning by defining knowledge on the personal level as justified, true belief. This is known today as the tripartite test.
In other words, to know a proposition, it must be true, you must believe it is true, and you must be able to justify it.
Truth must be justified.
An LLM can be right. An LLM can even be believed. But, based on everything we have just learned about the underlying architecture of LLMs, output cannot be justified. This is where probabilistic inference falls apart and how I justify my statement that an LLM does not know anything.
The implication of this is that an LLM alone cannot reliably perform the most basic artificial intelligence task, that of answering a simple question.
Because the output of an LLM ultimately relies on statistically plausible outputs, truth and fiction are indistinguishable. To trust an LLM response, it must come with justification.
1.2 RAG is Weak Justification
The knowledge gap is most strongly felt in domains and contexts that are under-represented in the training data. Recall our high-level exploration of the transformer architecture in the preceding pages. Namely, input tokens attend to each other to influence output tokens, and how output tokens attend to the input tokens to continuously refine and reshape the response. We quickly learned that if we added context to the prompt, we would yield better results.
The idea was simple: before handing the input prompt to the LLM, we first retrieve additional content from our own collection of knowledge. The LLM might be guessing based on weighted output tokens, but by injecting relevant content into the prompt, we could “put our thumb on the scale” and influence the output in our desired direction. Retrieval-Augmented Generation (RAG) was born.
RAG is a step in the right direction and genuinely valuable. If we presuppose the documents are an authoritative source of truth and an output statement can be linked back to a line in a document, we have the necessary ingredients to justify the truth of a probabilistic ”belief” but how authoritative are those documents? How comprehensively do they capture organizational knowledge? What about the rest of your data?
Most documents capture output, a snapshot of time that immediately begins to decay. Documents do not have schemas that enforce correctness or completeness. Most of the real-time data that is bound to constraints for completeness and accuracy lives elsewhere in the enterprise. This is where the cracks in the architecture begin to show.
A page of prose gives a large language model a rich set of signs and clues to make educated guesses but a JSON document with the property
{
…
"status": 3,
…
}
is completely opaque. Without something explicit to ground the term "status" and 3 the model will revert to its training weights. This interpretation will be wrong almost all the time.
This isn’t a “bug” the next model can “fix.” The real problem stems from a blind spot we all carry.
We think our data means something.
It doesn’t.
That’s the semantic ceiling.
1.3 The Limitations of LLMs are not Uniform
Notably, in my past teaching of the semantic limitations of LLMs, I had to concede that there are use-cases where LLMs perform remarkably well with reasonable consistency, that of coding. In fact, as I write this book new models continue to be introduced and the emphasis is consistently on how well they handle complex coding tasks when compared to previous models or their competitors. There is a reason for this: semantics.
Given the JSON fragment above, what does "status" mean? What about 3? Contrast that with a .cpp file containing the word printf(...). For the former, it is impossible to know based on the given input. The latter is completely unambiguous.
High-level programming languages were constructed as a way of using natural language to identify domain-specific contexts. Each keyword in the language has an explicit meaning—explicit semantics—because the vocabulary was created to resolve ambiguity. The compiler knows exactly how to parse, validate, and interpret language in the narrow context. Moreover, the resulting output is not just a bunch of tokens, it is a computable output. We can do more than evaluate whether an output looks right; we can execute it and verify. This is strong justification.
To break through the semantic ceiling, our data must embody the same properties as source code. Each term has explicit meaning, not hints for a better guess. Responses must be computationally verifiable, not be merely statistically plausible. If you want your pilot or POC to truly deliver in production, output requires strong justification. Otherwise, you’re just shipping expensive guesses that just might make your system the next headline going viral among those who delight in LLM schadenfreude.
The answer, according to industry analyst Gartner, is what they are calling a “Universal Semantic Layer.” The wording of their prediction is absolute:
By 2030, universal semantic layers will be treated as critical infrastructure, alongside data platforms and cybersecurity.
Developing a universal semantic layer is now a must‑do for D&A leaders either leading or supporting AI. It is the only way to improve accuracy, manage costs, substantially cut AI debt, align multiagent systems, and stop costly inconsistencies before they spread. D&A leaders must budget for semantic capabilities as a nonnegotiable foundation.9
Many have interpreted this as a call to industry to “build an enterprise ontology.” However, this intuitive approach can be fundamentally wrong. To build a single enterprise ontology requires standardizing on a single definition of every term across the enterprise. The truth is that there are many contexts within your organization. Rallying around a single canonical definition for each term robs the entire enterprise of necessary nuance.
A better approach is to build your ontology from the bottom up. We do not need global consensus; we need to explicitly express what terms in a context mean. Every term chosen for entities or properties in any enterprise context already carries implicit semantics. The meaning, rules, and relationships already exist in code, foreign keys, and data models. You do not need to build an ontology because it has already been grown and nurtured. Rather, we are simply harvesting it.
The bottom-up approach is, therefore, to identify the terms in each context and express their semantic similarity (or distinctness) in an unambiguous semantic layer so meaning is always part of the data.
This is not as difficult as it may sound. In fact, this is a solved problem… and more within reach than you might expect. The difficulty is not the open standards created in preparation for this moment; the real challenge lies in understanding the mental model. This book is your blueprint for both; preparing you for building these capabilities into your existing enterprise data ecosystem with potentially zero breaking changes.
What you will discover along the way is just how many other problems are secretly semantics problems. This book will not just change the way you think about data and AI, it will change the way you think about integration and architecture universally.
Summary
A large language model (LLM) is an astonishing thing. We speak to it in natural language, and it responds in kind. Ask a question, get an answer. Give it a prompt, it generates something. What it generates can be as simple as a haiku or as complex as a complete, working software platform.
But the results are ruinously inconsistent and model-level fixes are not the answer. LLMs fundamentally operate by guessing tokens. You can make the guesser better; you cannot transform guessing into knowing. You cannot add justification by scaling parameters.
RAG is a step in the right direction by patching the gap with documents, but documents are typically incomplete, stale (or decaying), and schema-less. Our best and most valuable data is our structured data, but the data model is completely opaque to an LLM without being grounded by a semantic layer.
To break through the semantic ceiling, we must build that layer by connecting existing data to strong meaning and formally expressed semantics. But this is not as straightforward as simply “building an ontology” in fact, this is usually the wrong approach. A top-down approach to a world-ontology—or even just an enterprise ontology—quickly collapses under the weight of contextual meaning. The same term might appear in many datasets in your organization, but those terms were chosen by people with an implied context; they plucked one meaning out of the uncountable possibilities and used that. Even within a modestly sized enterprise, those meanings will not be consistent across departments.
The solution is to connect a bounded term to one specific meaning. In other words, the most fundamental thing our semantic layer must communicate is “in this context, when I say X, I mean this specific usage and meaning of X, no more, no less.”
In the next chapter, I will show you how to do exactly that, with the data you already have and the meanings you already have defined.