AI Knowledge
Build the base concepts before moving into advanced workflows.
Start ReadingI've noticed something interesting in my work with entrepreneurs, advisors, and consultants who are starting to use AI.
Many feel they're missing foundational knowledge about how these systems actually work. They can follow tutorials on prompting and use the tools, but there's this nagging uncertainty – like driving a car without understanding what's happening under the hood.
I tell a lot of them that honestly…it doesn’t matter that much! Just learn to use the tools.
BUT understandably people still worry. They lack confidence because of this gap in their knowledge. What if a client asks about how ChatGPT works? What is a customer asks a technical question?
This knowledge gap isn't just academic – it has real business consequences. I've seen consultants struggle to explain AI capabilities to clients, entrepreneurs waste time trying approaches that fundamentally won't work with current AI, and advisors making strategic recommendations based on misconceptions about what AI can and can't do.
This week, we're going to fix that. We'll build a practical foundation of AI knowledge that helps you make better business decisions, communicate more effectively about AI, and cut through the hype to see the real opportunities.
The journey from traditional programming to machine learning to generative AI
Why large language models represent a fundamental shift
How computers evolved from following instructions to "thinking"
The pivotal transition from symbolic AI to neural networks
What this means for your business approach to AI
Last week I was chatting with a developer friend who's been coding since the early 90s. He said something that stuck with me: "In the old days, I had to tell the computer exactly what to do, step by step. Now I just ask it nicely, and it figures things out on its own."
The most profound change is in how we interact with computers. We've moved from:
Programming paradigm: "I will tell you exactly what to do, step by step."
AI paradigm: "I'll tell you what I want to achieve, and you figure out how to do it."
We've gone from painstakingly programming explicit instructions to having conversations with machines that seem to "get it." This fundamental shift isn't just interesting tech trivia – it completely changes how entrepreneurs like us can leverage technology to build and grow our businesses.
Traditional programming is like creating a detailed recipe: "If this happens, then do that." Computers follow these explicit instructions perfectly but have zero flexibility. If you didn't anticipate something in your code, the computer would be completely stuck.
Machine learning introduced a new approach: instead of writing explicit rules, we started showing computers many examples and letting them detect patterns. This was a bit like training a dog – it couldn't explain why it made certain decisions, but it could recognise patterns with remarkable accuracy after seeing enough examples.
What we have now with generative AI and large language models is something entirely different. These systems have ingested vast amounts of human knowledge and can now generate new content, reason through problems, and engage in meaningful dialogue.
It's like the difference between:
A calculator (traditional programming)
A trained animal that can categorise things (machine learning)
A knowledgeable assistant that can understand, create, and explain (generative AI)
AI's evolution hasn't been a straight line – it's been a series of breakthroughs, setbacks, and paradigm shifts spanning over 70 years. It’s not new.
Basically we’ve been trying to build AI since computers were imagined: Alan Turing envisioned how computers could be build and at the same time speculated that they might one day simulate human intelligence. Pretty smart guy!
Here are the pivotal moments:
1950s: The Birth of AI Alan Turing proposed his famous test for machine intelligence, and the field of AI was officially named at the Dartmouth Workshop in 1956. Early researchers were wildly optimistic, believing true machine intelligence was just around the corner.
1970s-1980s: The First AI Winter Early enthusiasm gave way to disappointment as researchers hit hard limitations. Funding dried up as promised breakthroughs failed to materialise.
1990s-2000s: The Rise of Machine Learning Rather than trying to program intelligence directly, researchers focused on statistical approaches where systems could learn from data. This change in direction slowly revitalized the field.
2012: The Deep Learning Revolution A neural network dramatically outperformed traditional methods in the ImageNet competition, triggering massive investment in deep learning approaches.
2022-Present: The Generative AI Era The public release of ChatGPT and similar systems brought sophisticated AI capabilities to the general public, democratizing access to these powerful tools virtually overnight.
For those interested in a deeper dive into AI history, I highly recommend:
This excellent video overview: A Brief History of AI
Michael Wooldridge's book, "The Road to Conscious Machines"
Understanding the shift from older approaches to today's AI requires grasping the fundamental divide between two competing philosophies:
Symbolic AI (also called "Good Old-Fashioned AI" or GOFAI): This approach, dominant in AI's early decades, used explicit symbols and rules to represent knowledge and reasoning.
It's like programming a computer with logical statements: "If X, then Y." These systems were transparent – you could trace their reasoning steps – but they struggled with ambiguity and required manually specifying every rule.
This approach (and its subsequent failure) led to the first AI Winter. Hype was huge but the results were limited. This led to funding being pulled and AI becoming a joke in academic circles.
Neural Networks: These systems learn patterns from data without explicit rules. Modern deep learning is the most successful neural approach, using interconnected layers of artificial "neurons" to identify patterns. Neural systems handle ambiguity well but work as "black boxes" – their internal reasoning isn't easily interpretable.
These systems fell under the “connectionist” approach. They were initially ignored by the symbolic AI practitioners and seen as a weird kooky approach that would come to nothing.
For decades, these approaches competed. Symbolic AI dominated early AI research but ultimately hit walls when trying to handle the messiness of the real world. You simply couldn't program enough rules to cover all possibilities.
Neural approaches, especially deep learning, ultimately surged ahead by letting the computer discover its own patterns from data rather than following human-programmed rules. We basically got out of the damn way.
This transition from symbolic, rules-based AI to neural, pattern-matching AI represents perhaps the most important shift in the field's history. It’s led to where we are today.
This evolution set the stage for today's large language models, which we'll explore in detail in the next Part. These systems don't follow explicitly programmed rules about language; they learn patterns from massive text datasets. That's why they can seem so remarkably flexible compared to earlier systems.
But remember that what we now call AI is just a part of the wider history, trends and approaches that have come before now!
There's often confusion about how different AI terms relate to each other, so let's clear that up right now.
Importantly this is going from widest to most narrow. Each of these fits inside the prior.
Artificial Intelligence (AI): The broadest umbrella term, covering any technique that enables computers to mimic human intelligence.
Machine Learning (ML): A subset of AI where systems learn from data rather than explicit programming. This includes many approaches like decision trees, neural networks, and support vector machines.
Deep Learning: A subset of ML using neural networks with multiple layers (hence "deep"), which excel at finding patterns in unstructured data like images, audio, and text.
Large Language Models (LLMs): A specific type of deep learning model designed to understand and generate human language. They're trained on massive text datasets to predict the next word in a sequence.
Generative Pre-trained Transformers (GPTs): A specific architecture of LLMs developed by OpenAI, using the Transformer architecture to process language more effectively.
ChatGPT: A product built on top of GPT models, specifically designed for conversational interactions, with additional safeguards and optimisations. It’s basically a chatbot app - with lots of bells and whistles.
So when you're using ChatGPT, you're interacting with just one commercial implementation of one type of language model, which is one application of deep learning, which is one approach to machine learning, which is one branch of artificial intelligence. Phew! Complex! But understanding this hierarchy already puts you ahead of 99% of people.
Want to explore this evolution further? Try this prompt with the AI tutor you built last week:
I want to understand how AI development has evolved over time. Can you:
1. Compare and contrast symbolic AI vs neural network approaches
2. Explain why generative AI represents such a significant shift from earlier systems
3. Explain how these changes affect me as an entrepreneur in [your specific industry]
Tomorrow, we'll peel back the curtain on how large language models actually work – without getting lost in technical jargon.
We'll explore concepts like tokens, parameters, and context windows in ways that actually matter for your business decisions.
We'll also examine why this AI revolution is happening now.
Keep Prompting,
Kyle
A common comment I get on my TikTok videos is about the "strawberry problem" – where ChatGPT can’t consistently tell you how many Rs are in the word "strawberry."
“If it’s so smart how many Rs in strawberry?” - “It can’t even count the Rs in strawberry”
It’s a common criticism thrown at ChatGPT and other modern (LLM-based) AIs.
The criticism actually reveals a LOT about people’s misunderstandings about AI.
We're mocking an AI language model for struggling with counting... when we already have perfect tools for counting– they're called computers. Computing is literally what they do. That's like criticising a hammer because it's not good at cutting wood.
Instead LLMs are Large Language Models. The hint is in Language. They aren’t built for maths, but for language.
Understanding this distinction is crucial for entrepreneurs – if you don't understand what LLMs fundamentally are, you'll either expect too much from them or miss their true potential entirely.
In this Part we'll demystify how large language models actually work under the hood – without requiring a PhD in computer science.
Let’s get started:
The basic idea: predicting the next word
Learning from vast amounts of text
The training process and parameters
Turning words into numbers: embeddings
Understanding context: transformers and attention
At their core, large language models do one thing remarkably well: they predict what word should come next in a piece of text.
That's kinda it. Really.
It sounds almost disappointingly simple, but this seemingly basic capability leads to all the "magic" we see.
Here's a concrete example of how this prediction works:
Imagine I start typing: "The Eiffel Tower is located in..."
You, as a human, can probably predict the next word: "Paris."
LLMs do exactly this, but at enormous scale and with extraordinary sophistication. This next-word prediction, when done billions of times with the right training, does something special…
Here's where it gets interesting. Humans intuitively understand that language prediction requires knowledge. To predict that "The Eiffel Tower is located in Paris," you need to know facts about world geography, famous landmarks, French culture etc. etc.
LLMs learn this knowledge through exposure – massive exposure. They're trained on hundreds of billions of words from books, articles, websites, and other texts. It would take a human thousands of years to read this much text.
This process is similar to how a child learns language by being constantly exposed to it, but at a vastly accelerated pace and scale. Through this exposure, the model learns patterns and relationships between words, phrases, and concepts.
Unlike humans who learn through a mix of experiences, explicit instruction, and text, LLMs learn purely through text. This is why they can sound remarkably human-like in some contexts but fail at seemingly simple tasks in others. They haven't experienced the world; they've only read about it. They have a lot of information but not necessarily the context to place it in.
So how does this prediction engine actually work? It's all about pattern recognition at a massive scale.
During training, the LLM learns by repeatedly trying to predict the next word in a given text. Here's how it works:
The model is given a piece of text with the last word hidden
It tries to guess what that word should be
It compares its guess with the actual word
It slightly adjusts its internal settings (called parameters or weights)
This process repeats billions of times
Like a student running questions and then checking the answers and using this to give better answers next time. But at a tremendous scale.
These parameters/weights are values that determine how the model makes predictions – think of it like tuning millions of dials on a complex machine. Modern LLMs have hundreds of billions of these parameters.
This adjustment process uses something called "gradient descent" – a fancy term for a simple idea. The model figures out which direction to turn each of those millions of dials to get a slightly better prediction next time. It's like playing the hot-and-cold game, where the model gets feedback on whether it's getting "warmer" or "colder" and adjusts accordingly.
Human language prediction isn't deterministic – we consider multiple possibilities with different likelihoods. We’d be very boring otherwise!
LLMs do the same thing.
When an LLM predicts the next word, it doesn't just pick one word with certainty. Instead, it assigns a probability to all possible next words in its vocabulary.
For example, after "The Eiffel Tower is located in..." the model might assign:
"Paris" → 95% probability
"the" → 1% probability
"France" → 3% probability
And so on for thousands of other words
When generating text, the model typically doesn't just pick the highest probability word every time (which would make outputs very predictable). Instead, it uses controlled randomness to occasionally select less likely words, making the text more diverse and human-like.
Note: this is also why ChatGPT isn’t “just” autocorrect.
This probabilistic nature explains why you get different responses from ChatGPT when asking the same question multiple times.
Most computer programmes are deterministic. The same input leads to the same output. Not so with LLMs. They are instead probabilistic. The same input leads to different outputs.
This simple fact is both a tremendous strength and weakness of LLMs. It makes them "creative" but also fickle. Like, well, humans.
To make this prediction game work computationally, we need to break down language into manageable pieces and represent them in a way computers can process.
I've been talking about "words" so far, but that's not exactly right. Let’s clear this up a bit and introduce the concept of tokens.
Tokens are the fundamental units in the prediction game. A token can be a whole word, part of a word, a character, or even punctuation. For English text, a token is roughly 3/4 of a word on average.
For example, "Let's understand how AI works" might become: ["Let", "'s", " understand", " how", " AI", " works"]
This tokenisation is crucial because it defines what exactly the model is predicting. It's not always predicting whole words – sometimes it's predicting word fragments or even individual characters.
The next crucial step is creating a "map" of language that captures relationships between words and concepts. This is done through "embeddings."
Imagine a massive multi-dimensional space where every word in the language has a specific location. Words with similar meanings or that are used in similar contexts are positioned near each other in this space. "King" would be near "queen" and "ruler," but far from "bicycle." Continue on in this way for ALL words.
This embedding space is what allows LLMs to understand that:
"Happy" is more similar to "joyful" than to "melancholy"
"Capital" near "country" likely refers to cities, not money
"Bank" near "river" means something different than "bank" near "money"
When the model predicts the next token, it's essentially navigating this semantic space, looking for tokens that make sense in the current context.
Importantly we don’t actually have to teach all of these relationships. That is the fools errand that the symbolic AI practitioners pursued. Imagine trying to formally declare the relationship between “bed” and “Arkansas”. And writing down similar relationships between all pairs of all the words. Yeah…nah…
Our modern LLMs actually go a step further still. Good prediction requires more than just understanding individual words – it requires understanding how all the words in a sentence relate to each other. This is where the breakthrough "attention mechanism" comes in.
Earlier approaches processed text sequentially, one word at a time. The attention mechanism instead lets the model see relationships between all words simultaneously.
It's like the difference between:
Reading a sentence word by word, trying to remember what came before
Reading the whole sentence at once and seeing how all the words relate to each other
When processing "The athlete picked up her trophy because she won the championship," the attention mechanism helps determine that "she" refers to the athlete by measuring the strength of relationships between all pairs of words.
This ability to juggle complex relationships between words is what makes modern LLMs so much better at maintaining coherence and understanding context than earlier models.
This breakthrough only came in 2017 and was fundamental in the rise of our modern LLMs. Here the wikipedia page on the critical paper.
The most fascinating aspect of this prediction system is what happens when we scale it up massively – both in terms of the amount of training data and the size of the model itself.
When we reach hundreds of billions of parameters trained on vast datasets, something remarkable happens. The model doesn't just learn simple word associations – it starts to capture complex patterns related to:
Facts about the world
Reasoning chains
Cultural references
Logical inference
Common sense knowledge
And much more
None of these capabilities were explicitly programmed. This is so important.
They instead emerged naturally from the prediction task when done at sufficient scale. This phenomenon, where new abilities suddenly appear as models grow larger, is called "emergent behaviour."
Prediction, when scaled to this level, begins to simulate many aspects of what we might call understanding or (gulp) intelligence.
Understanding language models as sophisticated prediction engines helps you make better decisions about implementing AI in your business:
1. Play to Their Strengths: LLMs excel at tasks that are fundamentally about pattern recognition in language – content creation, summarisation, translation, etc.
2. Compensate for Weaknesses: For tasks requiring factual precision, mathematics or logical reasoning, supplement LLMs with other tools or human oversight.
3. Design Better Prompts: When you understand that you're guiding a prediction engine, you can craft prompts that lead to better predictions
4. Set Realistic Expectations: Knowing that even the most impressive AI capabilities are built on prediction helps set appropriate boundaries for what these systems can and cannot do.
Perhaps most importantly, understanding LLMs as prediction engines removes some of the mystique and lets you approach them pragmatically – as powerful tools with specific capabilities and limitations.
Want to explore these concepts further? Try this prompt with your AI tutor:
I want to understand how viewing LLMs as prediction engines affects how I should use them in my business. Can you:
1. Identify 3 tasks in [my industry] that would be well-suited for LLMs because they fundamentally involve pattern prediction
2. Identify 3 tasks that would be poorly suited because they go beyond pattern prediction
3. Suggest how I might combine LLMs with other tools to overcome these limitationsIf you're interested in learning more about how LLMs work, these resources provide more depth while remaining accessible:
"What Is ChatGPT Doing... and Why Does It Work?" by Stephen Wolfram (Link): A detailed but accessible explanation of the mechanics behind large language models.
"But what is a neural network?" by 3Blue1Brown (Link): An excellent visual explanation of neural networks with stunning animations. 7 minutes.
"Neural Networks: Zero to Hero" by Andrej Karpathy (Link): A comprehensive 3-hour video course on neural networks by one of the field's leading experts.
Next we'll explore the difference between training and inference in AI systems. We'll unpack why creating these models is so expensive but using them is relatively affordable, and what this means for your AI strategy. We'll also tackle the thorny issue of "hallucinations"—why these models sometimes generate convincing but incorrect information, and how you can safeguard against this in your applications.
Keep Prompting,
Kyle
"It was the best of times, it was the worst of times..." It's the tale of two cities - the city of AI creators and the city of AI users.
Increasingly the world is split between the two. And knowing where we are as business owners is very important.
I was reminded of this divide at a tech event last month when a founder approached me after my talk. "I want to build my own language model from scratch. I don't want to rely on ChatGPT. I want complete control.”
Admirable sentiment. But….wooo…yikes.
This highlighted a crucial misconception that many entrepreneurs have when entering the AI space. Most of us will be citizens of the user city, not the creator city - and that's not just okay, it's smart.
Let’s get started:
The two lives of prediction engines: training vs. inference
Why most entrepreneurs should use, not create, foundation models
The middle ground: fine-tuning and retrieval-augmented generation
Managing prediction failures (hallucinations) in practical applications
Building an AI strategy that leverages the training/inference distinction
Every AI model you interact with exists in two distinct phases, created by two very different groups:
Training: The expensive, time-consuming process where the model learns how to play the prediction game by analysing massive datasets. This is the domain of well-funded AI labs with specialised expertise. Think OpenAI, Anthropic, Google, DeepSeek.
Inference: The relatively quick, affordable process where the model uses what it learned to make new predictions in response to your inputs. This is where entrepreneurs and businesses typically enter the picture.
This distinction creates two separate "cities" in the AI ecosystem:
The Creator City is populated by organisations like OpenAI, Anthropic, Google, and Meta who have the resources to train foundation models from scratch. BIG boys with deep pockets, often funded by entire governments.
The User City is where the vast majority of entrepreneurs and businesses live, building applications and solutions on top of these foundation models
The good news is that living in the User City doesn't limit your ability to create extraordinary value. We can still make amazing things without having to own the land we build on. We’ll discuss why this is so.
We touched on training in the previous Part but let's dive deeper. During training, the model is essentially learning how to play the prediction game through repeated exposure to massive amounts of text.
Here's what makes training so demanding:
1. Data at Scale: Training requires hundreds of billions of words. GPT-4 likely trained on a trillion tokens or more - equivalent to reading millions of books and the entire internet. Remember that Meta got caught using Anna’s Archive to scrape basically ALL the world’s books? This is why - they need data.
2. Computational Resources: Training top-tier models requires thousands of specialised GPUs running for months, costing tens or hundreds of millions of dollars. This is why NVIDIA has rocketed over the last few years - they are the primary manufacturer of these GPUs.
3. Specialised Expertise: Creating these models requires teams of ML researchers with advanced degrees and years of experience. These engineers are increasingly following the best salaries (and stock options!) to Silicon Valley and Hangzhou.
4. Infrastructure Complexity: The technical infrastructure to manage training at this scale is a massive engineering challenge in itself. This is why Elon Musk went ahead and built Colossus and other companies are construction massive data centres and looking into building their own nuclear power stations.
BIG players moving BIG money. Nuclear power station and trade deficit sort of money.
Once a model is trained, it enters the inference phase—when it actually generates responses to your prompts by applying what it learned during training.
Inference is dramatically less resource-intensive than training:
1. One-Way Process: The model is no longer adjusting its billions of parameters; it's simply using them to make predictions. Sure, there’s some training on your responses but that is of a totally different scale to what came before.
2. Single Task Focus: Rather than processing massive datasets, the model only needs to handle the specific text you've provided. It’s working on a tiny tiny subset of its knowledge.
3. Optimised Delivery: Companies have developed highly efficient systems for serving model responses at scale. Making this process fast and (energy) cheap is what gives them an edge over other companies.
The result? While training might cost hundreds of millions our API calls to OpenAI are fractions of pennies and our monthly subscriptions are only around $20/month.
Given the stark contrast between training and inference, there are compelling reasons why most entrepreneurs should focus on using existing models rather than creating their own:
1. Economic Reality: Training foundation models requires capital investments that only the largest companies can justify. The compute costs alone would bankrupt most startups.
2. Expertise Gap: Building these models requires specialised knowledge in machine learning that takes years to develop. OpenAI et al. have been building these skills for years now and have the jump on most of the market. Most businesses need AI solutions now, not after years of skill-building.
3. Opportunity Cost: The time and resources spent trying to create a foundation model could be better invested in building unique applications that solve specific customer problems.
4. Competitive Disadvantage: By the time you could create a foundation model from scratch, existing providers will have released several new generations of more powerful models. They are already rocking and rolling and in production. This is also why we are seeing less competitors entering the market. Instead it is solidifying around a small group of players.
5. Diminishing Returns: For most applications, current foundation models are already good enough that the marginal improvements from creating your own wouldn't justify the cost. Sure, you might “control” the underlying model but how much does that actually matter when it comes to deploying your product or service.
In the startup world, we often talk about focusing on your unique value proposition and outsourcing everything else. Foundation models are the perfect example of something most businesses should outsource rather than build.
We do this all the time with technology. We don’t (generally!) try to reinvent programs that we use. We don’t make our own programming languages. We use what is available and adjust to our needs.
The good news is that there's a middle ground between training your own models from scratch and using pure off-the-shelf solutions. These approaches give you many of the benefits of custom models without the astronomical costs:
Fine-tuning is the process of taking an off the shelf pre-trained model (say, GPT-4o) and adapting it to specific tasks.
Fine-tuning is like sending an experienced professional back to school for a specialised certificate. I might already have an accountancy degree but maybe I need some special training on certain tax issues. Cool - I can get that additional layer of information to sharpen my abilities.
The model already knows how to make predictions about language in general, but fine-tuning helps it make better predictions for your specific domain.
During fine-tuning, you take a pre-trained model and continue training it on a smaller dataset specific to your needs. This might be:
Your company's documentation
Examples of your brand's writing style
Specialised knowledge in your industry
Customer service interactions in your specific domain
Fine-tuning costs a tiny fraction of full training while delivering substantial improvements for specific applications. It's accessible even to small companies and individual developers through services like OpenAI's fine-tuning API or open-source models like Llama.
Retrieval-Augmented Generation is a very fancy way of giving our AI documents. We provide it with a reference library of material connected to the tasks it will be doing. It’s a little different to fine-tuning because we’re not actually running more training - we’re instead providing it supplementary material to refer to during inference.
RAG is sort of like giving an expert a specialised reference book to consult while they work. Instead of expecting the model to have memorised every fact during training, you provide it with relevant information at the time of prediction.
This approach:
Takes your prompt
Retrieves relevant information from your knowledge base
Includes this information as context for the model
Generates a response based on both the prompt and the retrieved context
RAG is particularly powerful because:
It keeps responses grounded in your specific data
It allows the model to access up-to-date information (you can update your data)
It dramatically reduces hallucinations for factual content
It lets you leverage company-specific knowledge that wouldn't be in training data
This all sounds very complex but at its base level it’s sort of like giving an AI access to our company’s Google Drive and all its documents. For most businesses, implementing RAG will deliver far more value than trying to train a custom foundation model.
Even with these customisation approaches, prediction engines sometimes fail. We call these failures "hallucinations" – cases where the model generates convincing but incorrect information.
Here are practical strategies to manage these failures in business applications:
1. Fact-Checking Layers: Implement secondary systems that verify claims made by the AI against trusted databases or search results. Basically a secondary layer of AI that fact checks the outputs of the first. Or, if you’re feeling feisty, multiple layers of checks.
2. Human-in-the-Loop: For critical applications, keep humans in the review process to catch and correct hallucinations before they reach customers. Got your AI generating emails to customers? Have it place the emails in drafts first then a human checks and sends.
3. Domain Constraints: Limit the AI's responses to areas where accuracy can be more easily verified or where errors have lower stakes. This comes from focused prompting, basically restricting the AI’s outputs to what is relevant.
4. Clear Uncertainty: By default AI systems will try to answer your questions even if they have no idea what the correct answer is. They are agreeable. It’s very annoying! So, train your systems to express uncertainty when they don't have sufficient information, rather than generating speculative answers.
These supplementary approaches acknowledge that prediction is inherently probabilistic (as we covered in the last Part!) and build safeguards accordingly. By knowing exactly where AI models are prone to messing up we can better work around these limitations. Which is the whole purpose of this Playbook!
Want to explore how these concepts apply to your business? Try this prompt with the AI tutor you built last week:
I'm considering implementing AI in [specific business function or product]. Help me understand:
1. Whether fine-tuning or RAG would be more appropriate for my specific use case
2. What kind of data I would need to collect or prepare for this approach
3. What verification mechanisms would be most important given the risks in my specific contextIf you want to dive deeper into these concepts, here are some valuable resources:
How AIs like ChatGPT learn (CGP Grey) - an oldie but a goody! CGP Grey knew that ChatGPT was going to be a big deal before the rest of us and made this very handy video on the basic mechanisms!
How Neural Networks Learn (3Blue1Brown): A more technical but highly valuable series on the actual training mechanisms. Much heavier on the maths but excellently explained.
Next up we'll explore the critical role of data in AI systems. It’s…more exciting than it sounds. Promise!
We'll discuss why data quality matters more than quantity, the different types of data used in AI, and how to develop an effective data strategy for your AI initiatives. This will help you make better decisions about what data to collect, how to use it, and how to maintain competitive advantage in the AI era.
Keep Prompting,
Kyle
Boring but important. It’s the lifeblood of AI models and can make or break an implementation.
Lots of people learn a little about machine learning and AI and assume (understandably) that MORE data is always better.
This is (kinda) correct for big foundational models. But as we talked about in the previous Part we’re not trying build our own foundational models for our businesses. That’s not a game we can win.
Instead we’re deploying more strategic smaller plays. And to pull these off we need to talk about what data we need.
"More is better" is actually a misconception in many practical AI applications for businesses. It would be more accurate to say "more relevant, high-quality data is better, up to a point" - but that's not as catchy!!
Let’s get started:
Why data is the new competitive battleground in AI
GIGO: The critical importance of data quality for prediction
The three types of AI data and how they're used
Why tech giants are getting "thirsty" for your data
Data privacy considerations that impact your strategy
Building your data advantage as an entrepreneur
There's an increasingly important shift happening in the AI landscape. While the headlines focus on which company has the "best" model, the real battle is happening elsewhere - over data.
Here's why this matters: foundation models are rapidly commoditising. GPT-4, Claude, Llama, Mistral, and others are becoming more similar in capabilities. What isn't commoditising is proprietary data - the unique information that only your business has access to.
All the foundation models were trained on (basically) the same data at first: the internet. And now that’s insufficient to gain an edge.
Consider these data plays:
The New York Times suing OpenAI over using their content for training then OpenAI striking a deal
Reddit signing a $60 million deal with Google for data access
OpenAI reportedly exploring the creation of a social media network (likely to generate fresh training data)
and countless other examples of AI companies teaming up with content/data providers of all stripes
These aren't random business decisions - they're strategic moves in the new data economy. The companies building foundation models are getting increasingly "thirsty" for high-quality data, and they're willing to pay premium prices to get it.
This creates both challenges and opportunities for entrepreneurs. While you may not be able to compete with OpenAI on model development, you might have access to unique data in your niche that could be extraordinarily valuable.
This is your edge.
For entrepreneurs, this creates a clear opportunity. As we discussed yesterday, you shouldn't build base models from scratch - that's a game for companies with massive resources. Instead, your advantage comes from leveraging your unique data to customise existing models.
Remember the last Part’s message: don't build your own models from scratch. That doesn't mean data doesn't matter - quite the opposite. It means you need to be strategic about how you use data with existing models.
We discussed two primary ways entrepreneurs can leverage data with foundation models. Here’s a quick reminder:
Fine-tuning is like giving the model additional education in your specific domain. You provide examples that teach the model to better understand your industry language, respond in your brand voice, or perform specific tasks relevant to your business.
Retrieval-Augmented Generation (RAG) is like giving the model a custom reference library to consult. Instead of hoping the model already knows about your products, services, or domain, you explicitly provide this information at runtime.
Both approaches let you benefit from the capabilities of foundation models while adding your unique advantage. And both rely entirely on the quality of your data.
Let’s talk about data.
Quality trumps quantity when it comes to AI data. It doesn't matter if you have millions of records if they're inconsistent, inaccurate, or irrelevant.
A constant refrain you must remember is Garbage In, Garbage Out (GIGO). If you flood your AI with crappy data it’s performance will actually get worse. MORE isn’t better.
So what makes data valuable for AI? Relevance is king - data directly related to the problems you're trying to solve is worth far more than massive quantities of tangential information. A thousand examples of exactly what you want to predict are worth more than a million examples of something vaguely related.
If you are building a bot to answer customer service questions give it lots of customer service interactions. Seems obvious but companies make this mistake all the time.
Cleanliness is critical too. Errors, inconsistencies, and noise in your data dramatically reduce its value. I've seen companies spend thousands on sophisticated AI infrastructure only to see it underperform because they skimped on data cleaning. You need to do a few runs to tidy everything up - thankfully AI can help here but don’t be lazy!
Representativeness matters as well. Your data should accurately reflect the real-world conditions where your AI will operate. If your customer service dataset only includes interactions with happy customers, your AI will struggle when it encounters its first angry client.
For entrepreneurs, focusing on these quality factors is far more important than obsessing over quantity. You don't need millions of examples - you need the right examples. Which is great news because as business owners and entrepreneurs we are often sitting on lots of data or (we’ll talk about this momentarily) in a great place to set up collection.
Let's get practical about how to actually implement fine-tuning and RAG systems with your data.
Fine-tuning is becoming increasingly accessible. It sounds scary but here’s a high level run down.
First, prepare your data. For OpenAI's fine-tuning, you'll need to format your data as JSONL files with prompt-completion pairs. It’ll look something like this:
{"prompt": "Customer question: How do I reset my password?", "completion": "To reset your password, click on the 'Forgot Password' link on the login page and follow the instructions sent to your email."}
{"prompt": "Customer question: Where is my order?", "completion": "You can track your order by logging into your account and visiting the 'Order History' section. There you'll find real-time updates on your delivery."}Think of this like model answers. Each prompt-completion pair teaches the model how you want it to respond to similar queries. You'll typically need several hundred to a few thousand high-quality examples for effective fine-tuning.
These can be farmed out on platforms like Mechanical Turk or done in-house. It depends on the specifics of the fine tuning!
Once your data is prepared, you can use platforms like OpenAI's fine-tuning API, Hugging Face, or Replicate to train your custom model. Costs vary but expect to pay hundreds to a few thousand dollars depending on the model size and amount of data. Very low costs compared to building a model from scratch!!
The practical benefit? A fine-tuned model can respond more consistently to queries in your domain, using your preferred tone and following your business policies - without needing to spell these out in each and every prompt.
RAG systems are even more accessible and often more immediately useful for entrepreneurs.
First, gather your knowledge base. This could be product documentation, FAQs, blog posts, internal wikis, or any text relevant to your domain. You'll need to extract the text from various formats (PDFs, websites, databases) using tools like PyPDF, BeautifulSoup, or dedicated services.
Be careful here as there is a tendency to just throw everything and the kitchen sink into our knowledge base. It’s tempting! But remember the GIGO rules above!
Next, chunk your documents into smaller, digestible pieces. Typically, these are paragraphs or sections of 200-1000 tokens. This chunking ensures the model receives relevant context without overwhelming it. You can use tools like LangChain and LlamaIndex for this. Don’t worry it’s not manual!
Then, create “embeddings” for each chunk. Embeddings are numerical representations of text that capture semantic meaning. Sounds hard and complex but again don’t worry tools do this for us. OpenAI's embedding API, HuggingFace models, or services like Pinecone can handle this process. These embeddings allow for semantic search - finding information based on meaning, not just keywords.
Store these embeddings in a vector database like Pinecone or Weaviate. These mean that when a user asks a question, you'll search this database for the most relevant chunks.
Finally, retrieve the most relevant chunks and send them to the model along with the user's query. The model then generates a response using both its general knowledge and your specific information.
Tools like LangChain and LlamaIndex have simplified this entire process with pre-built components for each step. So whilst it’s kinda neat to understand the steps above it’s honestly no longer needed - you choose your files, upload them and let the various tools do the heavy lifting.
The practical benefit? Your AI can reference exactly the information you want it to, staying current with your latest products or policies without retraining. It dramatically reduces hallucinations and keeps responses grounded in your actual business context.
OK that’s all well and good. But how do we do this practically?
Start by identifying your data assets. What information do you have that competitors don't? Customer interactions, domain expertise, proprietary processes - these could all be valuable sources for fine-tuning and/or RAG systems.
Then, design for data collection. Build data capture into your products and processes from the ground up. You decide first what sort of data would be useful and then work out mechanisms to collect that information. Every customer interaction should be seen as a potential opportunity to gather valuable training or RAG data.
AI can help you here! Here’s a prompt to kick off this process:
You are an expert in identifying valuable proprietary data assets in businesses. Help me discover the unique data advantage in my company.
About my business:
- Industry: [your industry]
- Main products/services: [brief description]
- Customer interactions: [how you interact with customers]
- Existing data collection: [what data you already collect]
Based on this information:
1. Identify 5-7 unique data assets my business likely has that competitors may not have access to
2. For each data asset, explain:
- Why it would be valuable for AI applications
- Whether it would be better for fine-tuning or RAG systems
- What specific business problems it could help solve
3. Suggest practical ways to better capture, organize, and utilize this data
4. Identify any potential privacy or ethical considerations
My goal is to leverage these data assets to create AI solutions that provide unique value to my customers.In the final Part we'll conclude our week on AI fundamentals by exploring "The AI Stack"—the various layers of technology that make up modern AI systems.
We'll help you understand build vs. buy decisions, how different components fit together, and how to develop an AI strategy that's both ambitious and pragmatic.
We’ll be pulling everything that we’ve covered so far into a strategy you can deploy in your own business or with others as a consultant.
Keep Prompting,
Kyle
I talk to a lot of business owners who know they ought to do something with AI but have no idea where to start.
The AI landscape feels like chaos to most business leaders - a bewildering array of technologies, terminologies, and techniques that seems to change weekly.
How on earth are they meant to make any lasting AI business decisions when it is all shifting so quickly?
Over the past four days, we've built a solid foundation of knowledge about how AI works. You now understand the evolution of these systems, how they function as prediction engines, the distinction between training and inference, and the critical importance of data.
But the question is still: "OK, but what do I actually DO with all this?"
I'm going to answer that question by giving you a practical implementation framework - a step-by-step roadmap you can use to deploy AI in your own business or help clients as a consultant.
No more paralysis by analysis. No more overwhelm. Just clear, actionable steps to move forward.
Let’s get started:
The three levels of AI implementation maturity
A practical roadmap from simple to sophisticated
Specific tools and technologies to use at each level
Clear indicators for when to level up
Common pitfalls to avoid along the journey
Before we dive into the details, let's look at the big picture. Here’s an outline of a maturity model providing a lear progression from simple beginnings to more advanced implementations:
Level
Description
Key Technologies
When to Use
Examples
Level 1: Getting Started
Simple AI integrations with existing tools and workflows
ChatGPT, Claude, Midjourney, Public APIs
Starting point for all businesses
Content generation, research assistance, basic automation
Level 2: Building Custom Solutions
Implementing RAG systems with proprietary data
Vector databases, embedding APIs, LangChain/LlamaIndex
When you have valuable proprietary data or specific use cases
Customer support bots, internal knowledge bases
Level 3: Advanced Implementation
Fine-tuning models, sophisticated applications
Fine-tuning APIs (OpenAI, Hugging Face), model deployment platforms (Replicate)
When you need deeper customisation and have domain-specific requirements
Industry-specific tools, complex workflows, personalised experiences
We covered RAG and fine tuning in the last Part - now we’re talking about how we actually step up to deploy these.
I know you’ll want to jump to the cool stuff in Level 3 but your company needs to grow at the same time as you are implementing these new tools. Step by step is best!
Level 1 is all about quick wins and low-hanging fruit. You're using existing AI tools and services without any custom development or complex integration. Think of it as "off-the-shelf AI."
This level involves:
Using public AI models through their interfaces (ChatGPT, Claude, Midjourney)
Simple API integrations with existing tools
Learning effective prompting techniques
Establishing basic AI workflows
At this level, you're primarily using tools like :
General AI assistants: ChatGPT Plus, Claude, Bard
Image generation: Midjourney, DALL-E, Stable Diffusion
Simple automation tools: Zapier with OpenAI integration, Make.com
Basic no-code tools: Bubble with AI plugins, Webflow with AI features
AI-enhanced office tools: Microsoft Copilot, Google Workspace AI features
This is non-exhaustive obviously. There are MANY (too many!) tools at this level.
A marketing agency might use ChatGPT to draft initial content outlines, Midjourney to generate concept images, and Zapier to automate posting to social media platforms.
A law firm might use Claude to summarise legal documents, extract key clauses, and generate first drafts of standard correspondence.
A solo entrepreneur might use AI assistants for research, content creation, and email management without any custom development.
This is the bread and butter stuff we cover extensively in Prompt Entrepreneur Playbooks. And it is where all businesses and entrepreneurs need to start to find their footing.
You should consider moving to Level 2 when:
You find yourself constantly feeding the same context into AI tools
Your prompts have become complex multi-page documents
You have valuable proprietary data that could enhance AI outputs
You're spending significant time on repetitive AI interactions
You have a team of people who don’t necessarily have your level of AI sophistication
Watch out for shiny object syndrome (trying every new AI tool rather than mastering a few), unrealistic expectations (expecting perfect outputs without proper prompting), security blindspots (feeding sensitive information into public AI systems), and forgetting human oversight (deploying AI-generated content without proper review).
Once the above seems old hat it’s time to move up a level. Specifically we’re going to put together a RAG system using our company’s “knowledge”.
Level 2 is where you start creating more customised AI solutions that leverage your data. You're no longer just using public tools; you're building simple but tailored applications off the back of internal knowledge.
This level involves:
Implementing Retrieval-Augmented Generation (RAG) systems
Creating specialised AI workflows for specific use cases
Integrating AI capabilities into existing products or services
Some custom development, often using frameworks and libraries
At this level, you're primarily using:
Vector databases: Pinecone, MongoDB, Weaviate, Chroma
RAG frameworks: LangChain, LlamaIndex
Embedding APIs: OpenAI embeddings, Cohere embeddings
No/low-code AI platforms: Retool, Bubble, FlutterFlow with AI components
Document processing: PyPDF, LangChain document loaders
A real estate company might build a RAG system with their property listings, market analyses, and historical transaction data, allowing agents to query this knowledge base for specific client needs.
A software company might create an internal AI assistant that has access to their codebase, documentation, and support tickets to help developers solve problems faster.
An e-commerce business might implement a product recommendation system that combines general AI capabilities with their specific product catalog and customer purchase history.
You should consider moving to Level 3 when:
Your RAG system struggles with complex reasoning about your domain
You need more consistent adherence to specific formats or terminology
You have enough high-quality training data to support fine-tuning
The business value justifies deeper investment in customisation
Be careful of poor data quality (GIGO - Garbage In, Garbage Out), chunking issues (improper document chunking leading to lost context), over-reliance on RAG (sometimes fine-tuning is necessary, see next step), and neglecting the user experience (building technically sound systems that are difficult for end-users).
Let’s now layer in fine-tuning on top of the previous work to really refine our systems.
Level 3 represents a significant step up in sophistication. You're now building deeply customised AI capabilities that are specifically tailored to your domain or use case.
This level involves:
Fine-tuning models for specific use cases
Building sophisticated applications with multiple AI components
More extensive custom development
At this level, you're primarily using:
Fine-tuning services: OpenAI fine-tuning API, Anthropic fine-tuning, Hugging Face
Model deployment platforms: Replicate, Modal
Specialised AI services: OpenAI Assistants API, Azure AI services
Advanced orchestration: Langfuse, LiteLLM
Monitoring tools: Weights & Biases, Helicone
Development frameworks: LangChain (advanced features), DSPy
A healthcare company might fine-tune models on medical documentation to create an AI system that can extract patient information, suggest diagnosis codes, and draft preliminary reports for physician review.
A financial services firm might build a comprehensive system that combines market data analysis, regulatory compliance checking and personalised client recommendations
A manufacturing company might implement an AI quality control system that analyses images and sensor data to detect defects, predict maintenance needs, and optimise production processes.
As your AI journey progresses beyond Level 3, you might eventually consider more advanced implementations when AI has proven so valuable that it needs to be embedded throughout the organisation, or your business strategy relies on AI as a core competitive advantage. But for most businesses, focusing on mastering Levels 1-3 will get them a long way ahead of their competitors.
Watch for over investing in fine-tuning (when RAG might work just as well), data leakage (not properly separating training/test data), fragmented implementation (building isolated systems), scaling too fast (implementing across too many areas simultaneously), and neglecting governance (ie. ignoring privacy and ethical considerations).

We've covered a lot of ground this week:
Part 1: We traced AI's evolution from rule-based systems to the neural networks that power today's generative AI, understanding how this shift fundamentally changes how we interact with computers.
Part 2: We explored how large language models work as sophisticated prediction engines, using patterns in text to generate human-like responses without true understanding.
Part 3: We distinguished between training and inference, understanding why creating models is expensive but using them is relatively affordable, and why we should focus on leveraging existing models rather than building our own.
Part 4: We examined how data fuels AI systems and why quality trumps quantity, especially for entrepreneurs using fine-tuning and RAG to leverage their proprietary data.
Part 5 : We've provided a practical roadmap for AI implementation, from simple beginnings to sophisticated systems, giving you a framework you can use in your own business or as a consultant.
This knowledge gives you a solid foundation for making informed decisions about AI for business.
You understand how these systems work, what they're good at, what they struggle with, and thus how to implement them effectively at various levels of sophistication.
Whether you use this with your own business or with others as a consult is up to you. I will say though that a lot of businesses need help with this right now - so it definitely opens up exciting doorways in advising and consulting once you feel comfortable.
Either way you are now in a much better position to carve through the noise and get to work deploying AI for business.
Keep Prompting,
Kyle
Welcome to the '🧱 AI Fundamentals' playbook, your essential guide to navigating the transformative landscape of artificial intelligence. This playbook equips entrepreneurs, advisors, and consultants with the foundational knowledge necessary to harness AI's potential effectively. You'll gain insights that demystify AI technology, enabling you to confidently discuss its applications and make informed business decisions. Say goodbye to uncertainty and hello to clarity as we explore the history, evolution, and practical implications of AI in your business.
This playbook is designed for entrepreneurs, business advisors, and consultants who recognize the importance of AI but feel overwhelmed by its complexities. If you find yourself struggling to understand how AI works or worried about addressing client inquiries on the topic, this playbook is perfect for you. It addresses your pain points by offering straightforward explanations and actionable steps, empowering you to embrace AI confidently and strategically.
No, this playbook is designed for non-technical professionals and provides clear, accessible explanations of AI concepts.
The playbook is structured to be completed in a week, with manageable sections that allow you to learn at your own pace.
This playbook offers foundational knowledge that will prepare you for future opportunities and enhance your understanding of AI, making you more attractive to potential clients.
Yes, we will cover popular AI tools and provide guidance on how to effectively integrate them into your workflows.
Absolutely! Each section includes practical steps that you can implement right away to start leveraging AI in your business.