← sunnyray.com All books Book 15 min

Free · Read it all here · No signup

The AI Method

Everything I know about building a real business on artificial intelligence.

30,710 words 128 min read 40 chapters By Sunny Ray

Part 01THE GROUND TRUTH


Chapter 01Who this is for

A short orientation before we start, so you can decide quickly whether to keep reading.

This is a field manual for people who intend to build something with AI, not a book about AI. It assumes you are technical enough to be curious and commercial enough to care whether the thing makes money, and it assumes you would rather have a correct picture than an exciting one.

I wrote it because the existing literature splits into two piles and neither is useful to an operator.

The first pile is the research-adjacent explanation of how these systems work, which is accurate and stops precisely where the commercial questions begin. It will tell you what attention is and nothing about what to charge.

The second pile is the money book, and I have read a lot of them. They open with the wealth transfer, they populate themselves with unverifiable case studies about people making improbable sums in improbably short periods, and they end with an invitation to a paid community. The method inside them is often not wrong. It is buried under an amount of hype that makes it impossible to tell which parts the author actually did and which parts they read about.

I have tried to write the version that is left when you remove the second pile's framing and keep its structure. That meant a specific editorial rule which I want to state up front, because it changes what you will and will not find here.

There is not a single invented number in this book. No case studies about a consultant named Sarah earning a figure that cannot be checked. No revenue tables from businesses I did not run. No projections about what you will earn in ninety days. Where I make a quantitative claim it is either arithmetic you can verify yourself, like the relationship between customer lifetime and required acquisition rate, or it is a description of a mechanism rather than a measurement. Where I have an opinion I will say it is an opinion, and where I might be wrong I will say where.

The cost of that rule is that this book is less exciting than its competitors. I think that is the correct trade, because the excitement in that genre is load-bearing: remove the improbable numbers and most of those books have very little left, which is itself the most useful thing to notice about them.

What you will find here, in order. Part One is the ground truth: what actually changed, what did not, and how to think about your own position. Part Two is the machine, which is the technical core, and it is the part I would least want you to skip even if you are not going to write code, because almost every expensive commercial mistake in this field traces back to a wrong mental model of what the system is doing. Part Three is the stack: turning capability into a system that runs without you. Part Four is the four business models and how to choose and price. Part Five is the company: bottlenecks, teams, money and risk. Part Six is the long game and a ninety-day plan to start.

Two things this book is not.

It is not a tool directory. Tools change every quarter and a list of them dates faster than anything else I could write. I name tools where a category needs an example and otherwise describe functions, because the function is the durable thing.

And it is not a promise. I do not know your market, your capacity, your network or your appetite for two hard years, and anyone who claims to forecast your outcome from those unknowns is doing something other than helping you. What I can offer is a method that I believe is correct, an honest account of where it is difficult, and the specific places where I have watched people lose time.

If that is the trade you want, keep going. If you were looking for the other book, it exists in large quantities and it is easy to find.

Chapter 02What the machines actually changed

I spent eight years at Quanser building control systems, mechatronics and haptics. Our products ended up in labs at MIT, Stanford and Georgia Tech. The work taught me a habit I have never lost: before you get excited about a system, you find out what it actually does at the boundaries, and you find out what happens when it fails.

So let me do that with AI before we do anything else, because almost everything written about making money with it starts from a claim that does not survive contact with a boundary.

Here is the claim that does survive. The marginal cost of producing a competent first draft of almost any piece of knowledge work has collapsed. Not to zero, but close enough to zero that the old economics no longer describe the world. A market analysis, a contract review, a set of unit tests, a product description, a research summary, a working prototype of a small application. Things that used to cost a person a day now cost fractions of a dollar and about ninety seconds.

That is the whole event. Everything else in this book is a consequence of that one sentence.

I want to be careful here, because the consequence people jump to is wrong. They hear "production is free" and conclude "output is valuable." The opposite happened. When the cost of producing something falls by two orders of magnitude, the thing itself stops being scarce, and scarcity moves somewhere else. It moved to three places.

It moved to judgment. A model will give you five strategies with equal confidence and no way to tell you which one is right for your situation. Choosing correctly is now the expensive part, and choosing correctly requires knowing the domain well enough to notice when the fluent answer is the wrong one.

It moved to verification. This is the one almost nobody prices properly. A model that is right ninety percent of the time and confidently wrong ten percent of the time is not ninety percent of a worker. In any workflow where a wrong answer causes damage, the cost of checking dominates the cost of producing, and the person who has built a reliable checking process owns something the person with the better prompt does not.

And it moved to distribution. If ten thousand people can now produce the same quality of output that used to take a firm, then the constraint on the business is no longer production capacity. It is whether anyone knows you exist and trusts you enough to hand you a problem. Distribution and trust are the two things AI has not made cheaper, which is exactly why they got more valuable.

The historical pattern is not mysterious. Electricity arrived in factories decades before it showed up as productivity, because the first thing everyone did was replace the steam engine with an electric motor and leave the building exactly as it was. Factories had been designed around a central power shaft: machines were arranged by their distance from the drive belt, not by the logic of the work. The gains only came when someone rebuilt the floor plan around the fact that every machine could now have its own motor. The technology was available for a generation before anyone reorganised around it.

That is precisely where we are. Most organisations have swapped the steam engine for the motor. They have put a chat window next to the existing process and left the process alone. The people who will take the value are the ones who redraw the floor plan, and redrawing the floor plan means asking a harder question than "where can I use AI here." It means asking which parts of this work only existed because production was expensive.

Now, the ownership question, stated honestly rather than dramatically.

The returns from any technology shift accrue to whoever owns the durable pieces: the system, the data it runs on, the distribution that brings work to it, and the customer relationship. If you rent all four, you are a contractor inside someone else's compounding, and that is a perfectly reasonable way to earn a living. It is just not the same thing as building an asset, and a lot of writing about AI blurs the two on purpose.

The version I actually believe, which is less dramatic than the version that sells books: in any period where the cost of production falls sharply, the gap widens between people who convert the saving into more output and people who convert it into a system. More output is a raise. A system is an asset. The raise stops the day you stop working. The asset does not.

011) LABOUR, OUTPUT022) TOOLING, OUTPUT033) SYSTEMS, OUTPUT044) OWNERSHIP05
Four rungs of leverage, and what caps each one.

There is a second thing worth saying about what did not change, because the omission is where most AI businesses die.

Cash still runs out on the same schedule it always did. Enterprises still take four months to sign. Trust is still built at human speed, one conversation at a time, and a language model cannot accelerate the part where someone decides you are safe to rely on. Support still costs money. Churn still kills. Nothing about a transformer architecture repeals working capital.

I have watched a lot of AI businesses in the last few years, from the inside as a director and an investor and from the outside as someone who interviews founders every week for a show. The failures cluster. They almost never fail because the model was not good enough. They fail because the founder built a demo that impressed other builders, priced it against a competitor rather than against the pain, and ran out of money waiting for a market that never had the problem they solved.

That gap, between a demo and a business, is what the rest of this book is about.

Chapter 03The honest ledger

Before any of the strategy, it is worth doing an accounting exercise that almost nobody does, because it prevents a whole class of wasted year.

Take a real process in a real business and work out what AI is actually worth to it. Not what it could theoretically be worth. What it is worth, after everything.

Here is the arithmetic in the shape it usually takes. A task consumes a certain number of hours a week across a certain number of people, each with a loaded cost. That is the gross prize, and it is the number that appears in every pitch deck. Then you subtract.

You subtract the fraction of the task that AI cannot do. Almost no process is fully automatable, and the residue is rarely the boring part. It is usually the exception handling, the judgement call, the conversation with the person who is unhappy. If eighty percent of the volume is routine, you do not capture eighty percent of the cost, because the twenty percent that remains still requires someone who understands the whole process to be available.

You subtract verification. Someone has to check the output, at least on a sample and often on everything, until trust is established, and trust takes months. Verification time is real time and it is charged against your saving. In the early period of a deployment it frequently consumes most of the gain, and teams that did not model this experience the first quarter as a disappointment when it was actually on plan.

You subtract the cost of the failures that get through. Not the visible ones, the ones that produce a downstream consequence weeks later. This is the hardest number to estimate and the one most likely to be zero in a plan and non-zero in reality.

You subtract the change cost. People have to learn a new way of working, they have to be persuaded that it is not a threat, and some of them will use the system incorrectly for a period. This is not friction to be minimised away. It is a line item.

And you subtract the running cost: inference, infrastructure, maintenance, the person who owns quality.

What is left is the honest prize, and in my experience it is a fraction of the gross number, frequently a third or less in the first year and considerably better afterwards as verification cost falls and the exception rate drops.

I want to be clear that I am not making an argument against doing this work. I am making an argument for doing this arithmetic before you build, for three reasons.

The first is that it tells you which processes are actually worth attacking. A process with a large gross prize and a high exception rate can be worth less than a smaller process that is genuinely routine, and the ranking is not obvious from the outside.

The second is that it makes you credible to a serious buyer. Everyone walking into a company right now is quoting the gross number. Someone who walks in and volunteers the subtractions, unprompted, immediately sounds like the only honest person in the room, which is a commercial advantage far larger than it should be.

And the third is that it protects you from the most common way these engagements sour. A project sold on the gross number and delivered at the honest number looks like a failure even when it succeeded. The same project sold on the honest number looks like a win. Nothing changed except which number was said out loud at the beginning.

There is a corollary worth stating. The processes with the best economics are frequently not the most visible ones. The glamorous target is the one everyone can see, the customer-facing thing, the strategic analysis. The economics are often better in the invisible middle: reconciliation, data entry between systems that do not talk, first-pass triage, document extraction, the parts of the operation nobody has ever put on a slide because nobody wants to think about them. Those are high volume, highly routine, and have measurable error rates already, which means the honest prize and the gross prize are close together.

That is where I would look first, and it is where almost nobody is looking, because there is no story in it.

Chapter 04The operator's mindset

I am suspicious of mindset chapters, so let me tell you what this one is not. It is not affirmations. It is not a list of beliefs to repeat until your bank account changes. I have never seen that work and I have watched a lot of people try it.

What I have seen work is a set of operating defaults that make different decisions than the average person makes at the same fork. They are boring, they are learnable, and each one is a specific behaviour you can audit afterwards.

The first default is that you optimise for the asset, not the payment.

An employee optimises for the payment because that is the correct move inside an employment contract: the upside of your improvements belongs to somebody else, so the rational play is to convert effort into salary. An owner optimises for the asset, because the improvement compounds into something they hold. Neither is morally superior. They are just different games with different scoring, and the mistake is playing one while being paid by the other.

The practical test: at the end of a week of work, does something exist that will still be producing value in a year without you touching it? Documentation, a working system, a piece of published thinking, a relationship, a piece of data nobody else has. If the honest answer is no, you had a week of labour, not a week of building, and you should know that rather than tell yourself a story about being busy.

The second default is that you build the loop before you build the output.

This is the engineering habit that transfers most cleanly. A control system is a loop: you measure the current state, compare it to the target, apply a correction, measure again. What destroys a control loop is almost never a bad correction. It is a bad measurement. If the sensor reads the wrong value, every correction after that is confidently wrong, and a more responsive system just drives itself into the wall faster.

Most people building with AI have no sensor at all. They have a prompt, an output, and a feeling about whether the output was good. That is not a loop, it is a vibe, and it will not survive the first time you have to defend a result to someone paying for it. Chapter nine is about how to build the sensor, and it is the most important technical chapter in the book.

The third default is that speed beats polish, but only in the direction of contact with reality.

Speed toward a buyer is always right. Speed toward more features is almost always wrong. The reason the "ship fast" advice so often produces garbage is that people apply it to the wrong axis: they ship a bad product quickly instead of finding out quickly whether anyone wants a good one. What you want to compress is the time between having an idea and having a stranger tell you the truth about it. Everything after that can take as long as it needs.

The fourth default is that you price on the outcome, not the labour.

If your work saves a company an amount of money they can name, your fee is a function of that number, not of the hours it took you. This is not a trick and it is not aggressive. It is the only pricing that stays honest when AI collapses your delivery time, because otherwise every capability improvement you make forces you to charge less for the same result, which is an absurd business to be in.

Now the objections, which are always the same five, and which I will treat as real rather than as obstacles to be affirmed away.

"I do not know enough about AI." You do not need to. The bottleneck in almost every deployment is not model expertise, it is domain expertise, because the person who knows the domain is the only one who can tell a good output from a plausible one. If you have spent a decade in insurance underwriting or logistics or dentistry, you have the scarce half. The AI half is learnable in weeks. The domain half is not.

"What if the technology moves and my work is obsolete." Some of it will be. The way to hold this is to notice which of your assets are model-dependent and which are not. A clever prompt is model-dependent and will be obsolete on a schedule. A relationship, a dataset, an evaluation suite, a reputation, a distribution channel and a documented process are not. Build on the second list and treat the first as consumable.

"I do not have capital." Mostly true and mostly irrelevant at the start, because the input that actually gates you is not money, it is attention and problem access. What money cannot buy you is a real conversation with someone who has the problem. Time can.

"I am too early, or too late." Both of these are the same sentence and neither has ever been answerable. What I can tell you is that the current window has a specific and unusual property: capability is running ahead of adoption by a wide margin. Most organisations are using a fraction of what the models can already do, which means the opportunity is not in waiting for better models. It is in closing the gap on the ones that already shipped.

"I am not qualified to charge for this." Your buyer does not evaluate your credentials. They evaluate whether the outcome arrived. Qualification comes in three layers and you only need the first to begin: enough knowledge to be useful, then evidence you earned by doing the work, then authority which accrues over years and which you cannot manufacture. People try to start at the third layer, and that is why their content sounds hollow.

The last default, and the one that separates people who finish from people who accumulate projects: you have a definition of done.

An enormous amount of AI work stalls at eighty percent, where the demo works and the edge cases do not, because the last twenty percent is unglamorous and involves error handling, logging, retries, and the boring question of what happens at 3am when an API returns a 500. That last twenty percent is the entire difference between something you show people and something people pay for.

Chapter 05From control systems to the convergence

I want to give you my actual background rather than a mythology, because the useful part is not the outcome, it is the transfer.

I trained as an electrical engineer. My first serious professional stretch was eight years at Quanser, where we built control systems, mechatronics and haptic devices. If you have not worked in that world: a haptic device is a machine that lets you feel a simulated force through a handle, and getting it right means closing a loop between a physical actuator and a software model fast enough that a human nervous system is fooled. Our products went into research and teaching labs at places like MIT, Stanford and Georgia Tech.

What that work drills into you is unforgiving. Physical systems do not accept hand-waving. A control loop either converges or it oscillates and the hardware shakes itself apart in front of you. Latency is not an inconvenience, it is a stability parameter. You learn to think in terms of feedback, sample rates, failure modes and margins, and you learn that a system which works in the demo and fails in the lab was never working, it was lucky.

I have carried three things from that decade into everything since. You measure before you conclude. You design for the failure case, not the happy path. And a system's real behaviour lives at the edges, not in the middle of the distribution.

Then I went into bitcoin, in 2011, when explaining it to people got you a specific facial expression I still recognise.

I co-founded India's first bitcoin exchange. We grew it past two and a half million users, which meant building financial infrastructure for a market that had no template, no precedent and no regulatory clarity. Then the central bank effectively cut the industry off from the banking system, which is a polite way of saying they tried to end us. We fought it to the Supreme Court of India, and we won.

That period taught me things that eight years of engineering could not. Regulatory risk is a design constraint, not a footnote. Trust at scale is built through boring operational consistency, not through marketing. And the hardest technical problem in any consumer financial product is not the ledger, it is the support queue at 2am when someone thinks their money is gone.

Today I run Sunny Ray Holding Inc, a family office and venture studio. I sit on the board of Humanoid Global Holdings, which trades on the CSE as ROBO, so I see the robotics side of this convergence from inside a public issuer with real disclosure obligations. I host The Sunny Ray Show, which means I spend a large part of every week in long conversations with founders building at the edge of AI, bitcoin and robotics. I back companies at the point where capital and conviction are both scarce.

The reason I am telling you this is that it explains the shape of the advice in this book. I have shipped physical systems that punish imprecision, and financial infrastructure that punishes carelessness, and I now spend my time evaluating AI companies at the point where someone is asking me for money. Those three vantage points produce a particular bias: I care much more about whether a thing works under load than about whether it demos well.

Here is what I have learned that I would tell a version of myself starting today.

Talk to buyers before you build. This is the most repeated advice in business writing and the most ignored, and I understand why: building is pleasant and talking to strangers about their problems is not. But the failure mode is brutal. You spend six months constructing something technically impressive, and the thing you learn on the day you try to sell it is a thing you could have learned in the first week for the price of ten conversations.

Simplicity wins commercially and complexity wins socially. The elegant, sophisticated version of your product impresses engineers. The embarrassingly simple version gets bought. I have watched this pattern hold in robotics, in financial products and in AI, without exception. If you find yourself explaining why the complexity is necessary, you are usually explaining why you enjoyed building it.

Positioning determines price more than quality does. Two people with identical skill, one describing themselves as an AI consultant and the other as the person who solves a specific expensive problem for a specific industry, will end up an order of magnitude apart on fee. The work is the same. The frame is not.

Authority compounds and cannot be rushed. Everything I have that generates inbound today, the show, the writing, the network, came from years of doing the thing publicly and consistently before it produced anything. There is no shortcut, and people who try to buy the shortcut end up with an audience that does not convert.

And one piece of conventional wisdom I now disagree with. The standard advice is to build multiple revenue streams to reduce risk. In practice, before you have one line that genuinely works, additional lines are not diversification, they are distraction wearing a disguise. Diversify after you have something that compounds, not before. I have made this mistake and it cost me a year.

Where are you stuck?

Fifteen minutes, no deck, no pitch. Tell me what you are building and I will tell you what I would do next.

Book 15 minutes

The last thing worth saying about my own path is that the convergence I now work at, bitcoin plus AI plus robotics, is not three interests stapled together. It is one thesis. Autonomous systems will need to transact with each other without a human in the loop, at frequencies and amounts that no legacy rail was designed for. AI supplies the decision-making, robotics supplies the physical agency, and bitcoin supplies settlement that does not require anyone's permission. I will come back to this in the final part, because it is where I think the largest opportunities of the next decade are, and because it is the part of the landscape that is least crowded.

Chapter 06Positioning, or why nobody hires a generalist

Authority is not a personality. It is an equation with three terms, and most people obsess over the first and wonder why nothing happens.

Expertise plus evidence plus visibility. Expertise alone makes you a well-informed stranger. Expertise plus evidence makes you credible to anyone who happens to find you, which is nobody. Expertise plus visibility without evidence makes you a commentator, which is a fine hobby and a bad business. You need all three, and they are built in that order.

The single highest-leverage decision in this whole book is who you say you are for.

I want to be blunt about why, because "niche down" is advice everyone has heard and almost nobody executes. It is not about market size. It is about the buyer's search behaviour. When someone has an expensive, specific problem, they do not search for a category, they search for their problem. They ask their network who solved this exact thing. A generalist never appears in that conversation, not because they could not do the work, but because there is no sentence anyone can use to recommend them.

Consider the difference between these two descriptions of the same person.

The first: an AI consultant who helps businesses implement AI solutions. That sentence is true, contains no information, and competes against everyone. It cannot be repeated by a third party, which means it cannot travel.

The second: the person who builds automated quality inspection for parts manufacturers. That is repeatable. Someone can say it at a dinner without you in the room, and the listener will either lean in or not, and both outcomes are useful.

The mechanism is that specific positioning makes you referable. Referability is the actual asset. Everything else about niching, the higher prices, the shorter sales cycles, the lower competition, falls out of that one property.

There is a ladder here and you can locate yourself on it honestly.

At the bottom is tool implementation: helping people use AI tools better. This is commodity work by construction. The knowledge is public, the barrier is a weekend, and the price falls every quarter as the tools get easier. Nothing wrong with starting here, but do not build a business on it.

One rung up is process optimisation: taking a specific business function and making it faster and cheaper with AI. Better, because you are now being paid for a result rather than for knowledge. Still limited, because the buyer sees you as a vendor to a department.

Above that is business outcome work: changing how a company competes, using AI as the mechanism rather than the subject. Now you are talking to someone who owns a number, and you are a business advisor who happens to use AI rather than an AI person who happens to know business.

At the top is domain plus AI: you are the only person who can deliver this specific outcome in this specific industry, because you understand the industry's constraints, vocabulary, regulations and politics as well as the technology. This is the position that resists commoditisation for the longest, because the model getting better does not erode your understanding of why a plant manager will reject a system that requires retraining the night shift.

The positioning statement I use is boring and it works. I help a specific kind of buyer solve a specific expensive problem using a specific mechanism, to reach a specific outcome. Every one of those four slots must be filled with something a stranger could verify. If you cannot fill a slot, you have found the thing to go learn.

01AUDIENCE02BOTTOM-RIGHT03TOP-LEFT04TOP-RIGHT0506
Positioning: two axes decide your price.

Once you have the position, evidence is the next build, and evidence is more specific than most people think. A testimonial that says you were great to work with is social decoration. Evidence is a before state, an after state, a mechanism and a number the client would confirm on a phone call. If you cannot yet produce that, the correct move is to do the work at a price that makes the evidence easy to get, and to be explicit with yourself that you are buying a case study rather than discounting out of fear.

Visibility comes last, and there is a right and wrong way to do it in a field this saturated.

The wrong way is opinion. There is an infinite supply of people with takes about AI, the supply grew faster than the demand, and the marginal value of another one is approximately zero. Opinion content builds an audience of people who consume opinion content, which is not a buyer.

The right way is mechanism. Show how something works. Publish the process, the failure, the numbers you are allowed to publish, the thing that surprised you. Mechanism content is harder to produce, which is exactly why it works: it is a costly signal. Somebody reading a detailed account of how you solved a problem learns two things at once, the solution and the fact that you have actually done this.

The pattern I would run if I were starting again: one substantial piece a week that teaches a mechanism, cut into smaller pieces for wherever your buyers actually read, plus one long-form conversation a week with someone in your target industry. The conversations do double duty. They are content, and they are the most natural business development that exists, which is a thing I will come back to in Part Four.

One warning about all of this. Positioning is a decision you should expect to make wrong the first time, and the correction loop is slow, on the order of months rather than days. Do not let that make you avoid the decision. A wrong specific position teaches you something in ninety days. A safe general position teaches you nothing in two years, which is the more expensive error and the one that feels safer.

Get the rest by email

One email when I publish something new. No sequence, no funnel, unsubscribe in one click.

Part 02THE MACHINE


Chapter 07What a language model actually is

You do not need to be able to derive backpropagation to build a business on these systems. You do need an accurate mental model, because almost every expensive mistake I have watched people make traces back to a wrong one.

Here is the accurate version, compressed.

A language model is a very large function that takes a sequence of tokens and returns a probability distribution over the next token. Tokens are fragments of text, roughly three quarters of a word on average in English, and the model has been trained on an enormous corpus to make that next-token prediction as accurate as possible. To generate a response, the system samples a token from the distribution, appends it to the sequence, and runs the function again. That is the entire generation process. Everything you experience as reasoning, style, personality and expertise is what that loop looks like from the outside when the function is good enough.

Three consequences follow, and each one has commercial weight.

The first is that the model has no separate store of facts it consults. Knowledge is not looked up, it is reconstructed from statistical structure. This is why models produce confident, well-formed, wrong answers, and why the phrase "the model lied" is a category error. It did not lie. It generated the most plausible continuation, and plausibility and truth are correlated but not identical. Any workflow that treats fluency as a proxy for accuracy will eventually cost somebody real money.

The second is that everything the model knows about your situation has to be inside the sequence. There is no persistent awareness of you between calls unless something outside the model puts it there. The chat interface hides this by quietly resending your history on every turn. Once you build with the API, the illusion drops away, and you discover that the actual engineering problem is deciding what goes into that sequence and what does not. This is the single most underrated skill in the field and Chapter 7 is entirely about it.

The third is that the model has no privileged channel for instructions. Your system prompt, the user's message, and a document you pasted in all arrive as tokens in the same stream. The model has been trained to treat them differently and mostly does, but "mostly" is doing a lot of work in that sentence. This is the root of prompt injection, which is not an exotic attack, it is the default behaviour of a system that cannot structurally distinguish data from instruction. If your product reads untrusted content, a web page, an email, a customer upload, then that content can attempt to redirect your model, and no amount of please-ignore-instructions in the system prompt fully closes it. Architecture closes it. I will get to that.

01SYSTEM PROMPT +02FAILURE BELOW03FAILURE: COST04FAILURE: NO FACT
What actually happens between your request and the answer.

Two more properties are worth internalising because they show up in every cost conversation.

Generation is sequential and therefore slow in a specific way. The model produces one token at a time, so a long answer takes proportionally longer. Reading your input, by contrast, happens in parallel and is fast. This asymmetry is why a system that reads a hundred-page document and returns a one-line verdict is cheap and quick, while one that reads a line and writes a hundred pages is neither. When you design a workflow, push the volume onto the input side wherever you can.

And the context window, the maximum sequence length, is a hard ceiling but a soft constraint. Modern windows are large enough that most business tasks fit. The real limit arrives earlier: models attend unevenly across a long context, and stuffing everything in degrades performance well before you hit the ceiling. More context is not better context. Relevant context is better context, and the work of selecting it is the work.

Two things people expect models to do that they cannot do reliably, so you should design around them rather than prompt harder.

They cannot count or do exact arithmetic dependably, because both are the wrong shape for next-token prediction. Give them a calculator, a code interpreter, or a database query, and the problem disappears. Every hour spent trying to prompt a model into reliable arithmetic is an hour that a three-line tool call would have solved.

And they cannot reliably tell you what they do not know. Calibrated uncertainty is an active research problem. Asking a model how confident it is produces a fluent number that correlates only loosely with correctness. If your process depends on the model flagging its own uncertainty, your process is built on sand. Build the check outside the model.

Chapter 08The model landscape in 2026, and how not to write it down

I am going to tell you how to think about model selection in a way that will still be correct after the next three release cycles, because the alternative, a table of version numbers, is obsolete before the ink dries. I have watched people build entire business plans around a specific model's specific quirk and then discover the quirk was a bug that got fixed.

Think in tiers of job, not in brands.

There is a frontier tier: the most capable models available, used where the reasoning is genuinely hard, where the output is going in front of a client or a regulator, where a mistake is expensive, or where the task requires holding a lot of interdependent constraints in mind at once. These cost the most per token and are worth it exactly when the alternative is a human hour. In the Claude family this is Opus 5, the most capable model Anthropic currently ships, sitting alongside Sonnet 5 and Fable 5 in the Claude 5 generation. Every serious lab fields something in this tier, and they leapfrog each other on a cadence of months.

There is a workhorse tier: strong general models that handle the overwhelming majority of production traffic at a fraction of frontier cost. This is where most of your volume should live once you have proven the task works. Sonnet 5 occupies this position in the Claude family, and it is the model I would default to for a production pipeline until measurements tell me otherwise.

There is a fast tier: small, quick, inexpensive models for classification, routing, extraction, filtering and any high-volume step where the task is narrow and the correct answer is short. Haiku 4.5 is the current Claude option here. Underusing this tier is one of the most common and most expensive architectural mistakes I see. People run every step of a ten-step pipeline through a frontier model because it is easier to have one API call in the codebase, and then discover their gross margin has evaporated.

And there is an open-weights tier: models you can run on infrastructure you control. Slower to set up, cheaper at scale, and occasionally the only option when a client's data cannot leave their network. The capability gap between the open frontier and the closed frontier has narrowed considerably and continues to. For regulated industries, this tier is not a cost decision, it is a feasibility decision.

The rule that survives every release: route by job. A production system should not have one model in it, it should have three or four, each doing the work it is cheapest at. The classifier that decides which of eight categories an inbound message falls into should never touch a frontier model. The final client-facing analysis probably should.

CLASSIFICATION1234
Route by job, not by brand.

Three practical points about selection that people learn the hard way.

Benchmarks are directionally useful and locally misleading. A model that scores higher on a public evaluation may be worse at your specific task, because your task is not in the benchmark and your data does not look like the benchmark's data. The only benchmark that matters is the one you build from your own examples, which is Chapter 9 again.

Model behaviour changes underneath you. Providers update models, sometimes with a version change and sometimes without one that you notice. A prompt tuned to within an inch of its life against one snapshot can degrade when the snapshot moves. Pin versions where the provider lets you, and keep a regression suite so you find out from a test rather than from a customer.

And do not build a business whose only asset is a wrapper around one provider's API. This is not because wrappers are shameful, plenty of good businesses are thin layers over infrastructure. It is because the defensibility has to live somewhere else: in the workflow, the data, the distribution, the integration depth, the trust. If a competitor can reproduce your product by writing the same prompt, they will, next week.

The last thing on the landscape, and this is the part that changes fastest and matters most. The frontier of usefulness has moved from single responses to sustained work. The interesting question is no longer whether a model can write a function, it is whether a system can hold a goal across many steps, use tools, recover from its own errors, and produce something finished. That shift is the subject of Chapter 10, and it is where the commercial opportunity has moved.

Chapter 09Beyond text: documents, vision and voice

Almost every business conversation about AI is a conversation about text, and that is a hangover from the first two years of the current wave. The models are now genuinely multimodal, and the commercial opportunities in the other modes are less crowded, because fewer people have noticed.

Start with documents, which is the most immediately valuable and the most misunderstood.

Business does not run on clean text. It runs on PDFs of scanned invoices, spreadsheets with merged cells and three header rows, forms with handwritten annotations, contracts with tables inside clauses, and photographs of paperwork taken on a phone at an angle. Every one of those is a place where a person currently sits, retyping.

The old approach to this was optical character recognition, and it worked in the narrow case where the layout was consistent. It broke the moment the layout varied, which meant every new document type was a new engineering project. That is what changed. A model that can see the page understands that this number is in a column headed total, that this stamp is a date, that this handwritten note in the margin modifies the clause above it. It reads the way a person reads, which means layout variation stops being a blocker.

The practical consequence is that document extraction has moved from a specialist engineering problem to something a small team can deploy across a whole category of paperwork in weeks. If your target industry moves documents around, and most of the profitable unglamorous ones do, this is the highest-value place to point a system, and it is a place where the honest prize and the gross prize sit close together because the task is genuinely routine.

Two engineering notes, because this is where people get it wrong. Always extract into a defined schema and validate it with code, so that a missing field is an exception rather than a plausible guess. And always keep the location of what was extracted, so a human reviewing the output can be shown the exact region of the page it came from. That second detail sounds minor and is the entire difference between a review step that takes ten seconds and one that takes two minutes, which at volume is the difference between viable and not.

Then vision on the physical world, which is where my own background makes me biased and, I think, correct.

Inspection, counting, condition assessment, safety compliance, damage assessment, verification that a thing was done. All of these are currently done by a person walking around with a clipboard or a phone, and all of them are now within reach of a system that looks at an image and produces a structured judgement. The economics are strong because the current process is expensive, inconsistent between inspectors, and produces no usable data.

The caution is that physical environments are hostile in ways that demonstrations are not. Lighting varies. Cameras get dirty. The thing being inspected is partially occluded by something nobody anticipated. The gap between a model that works on a curated image set and a system that works in a facility is large, and it is closed by collecting real data from the real environment, including the bad images, and putting them in your evaluation suite. This is exactly the same discipline as Chapter 9, applied to pixels.

Then voice, which has quietly become good enough to matter.

Transcription is effectively solved for most business purposes, and the interesting layer is what sits on top of it. Every phone call, every meeting, every site visit conversation is now a source of structured data: what was decided, what was promised, what the objection was, what needs to happen next. Most organisations are discarding all of it and then complaining that their records are incomplete.

Real-time voice interaction is the newer capability and it is genuinely usable now for constrained tasks. I would be careful here for a specific reason: a synthetic voice on a phone call sits close to a line that customers care about, and the reputational cost of getting it wrong is asymmetric. My position is disclosure, always, and a fast path to a human whenever the caller wants one. The businesses that treat this as an efficiency play without regard for how it feels to be on the receiving end will win in the short term and lose the relationship.

And generation of images and video, which I will treat briefly because it is the most visible and the least commercially interesting for the kinds of business this book is about. It has real value in marketing production and it collapses a cost line that used to be substantial. It is not, on its own, a business, because the capability is universally available and the output is commoditised on arrival.

The general principle across all four modes is the same one from Chapter 7. The model is not the advantage. Access to the pipeline of real inputs is the advantage, and in the non-text modes that access is much harder for a competitor to acquire, because it requires a relationship with someone who has the cameras, the documents or the phone lines. That difficulty is precisely why it is worth pursuing.

Chapter 10Context is the product

If I could get one idea from this book into wide circulation, it would be this one, because it is where the leverage actually is and it is almost never taught.

The model is a commodity. Everyone has access to roughly the same models at roughly the same price. What is not commoditised is what you put in front of it. Two people asking the same model the same question get different answers because one of them assembled a better context, and assembling context is an engineering discipline with its own techniques, failure modes and economics.

Think about what actually goes into a request in production. There is the instruction, which is your system prompt and is stable. There is the user's input, which is variable and untrusted. There is retrieved knowledge, documents or records or prior decisions selected because they are relevant to this specific request. There is state, what happened earlier in this conversation or this workflow. And there are tool definitions, the descriptions of what the model is allowed to reach for.

Every one of those is a design decision with a cost and a quality consequence, and the person who makes them well ships a product that feels like magic while a competitor with the identical model ships one that feels like a toy.

Start with retrieval, because it is the highest-value piece and the most commonly botched.

The naive version, which most tutorials teach, is: chop your documents into fixed-size chunks, embed them into vectors, embed the user's query, return the nearest neighbours, paste them in. It works well enough to demo and badly enough to fail in production, for reasons that are structural.

Fixed chunking destroys meaning. A chunk boundary that lands mid-clause splits the subject from the predicate, and the fragment retrieves poorly because it is not about anything. Chunk on structure instead: sections, clauses, rows, whatever unit your documents actually have. A contract chunked by clause retrieves far better than the same contract chunked every eight hundred characters, and it takes an afternoon to do properly.

Semantic similarity is not relevance. Vector search finds text that is about the same subject, which is not the same as text that answers the question. A query about a refund policy will happily retrieve five paragraphs discussing refunds and miss the one sentence that states the actual rule, because that sentence is short and clinical and does not look like the question. Combining keyword search with vector search fixes most of this, and the combination beats either alone on nearly every real corpus I have seen.

And retrieval quality has a ceiling set by the corpus. If your documents are out of date, contradictory or incomplete, no retrieval sophistication saves you. I have watched teams spend two months tuning an embedding pipeline over a knowledge base that three different departments had been editing in three different directions for a year. The problem was never the vectors.

The second piece is what you leave out, which matters more than what you put in.

The instinct with a large context window is to include everything relevant, on the theory that more information cannot hurt. It can. Models attend unevenly across long inputs, and irrelevant material does not sit there inertly, it competes for attention and pulls the output toward itself. A tight context of four highly relevant paragraphs beats a sprawling one of forty loosely relevant ones, on both quality and cost, and the gap widens as the task gets harder.

Practical discipline: for every element in your context, be able to say what it is for. Anything you cannot justify, remove and measure. You will remove more than you expect.

The third piece is structure. Models follow structure. If you give them a clearly delimited block of retrieved documents, a clearly delimited set of instructions and a clearly delimited user query, they distinguish the three far more reliably than if the whole thing arrives as a wall of prose. Use consistent, obvious section markers. This costs nothing and buys measurable reliability, and it also makes the injection surface smaller, because untrusted content sits inside a labelled fence rather than blending into your instructions.

The fourth piece is memory, and this is where systems become products rather than features.

A useful business system remembers. It remembers what this client decided last quarter and why, which approach failed, what the house style is, which supplier is on hold. None of that lives in the model, so it has to live in a store you control and be selected into context when relevant. This is unglamorous database work, and it is the thing that makes the difference between a tool a user tries once and a tool they cannot leave, because leaving means abandoning accumulated context.

That last point is worth dwelling on commercially. Accumulated context is one of the few genuine moats available to a small company in this field. Model capability is rented and everybody rents from the same landlords. A structured record of a client's decisions, exceptions, vocabulary and preferences, built up over eighteen months of use, is not rented from anyone, and it is the reason switching costs exist.

Chapter 11Prompting as specification

Prompt engineering had a moment as a job title and then stopped being one, which was correct. It was never a profession. It is a subskill, roughly analogous to writing a good bug report or a clear brief, and its importance has fallen as models have improved at inferring intent.

What has not fallen in importance is specification, and I want to draw the distinction clearly. Prompting tricks are the incantations: magic phrases, role assignments, threats and bribes, the folklore that circulates and half of which stops working with each generation. Specification is the discipline of stating precisely what you want, what the constraints are, what the output must look like, and how you will judge it. Tricks decay. Specification does not, because it is just clear thinking written down, and clear thinking is what the model was missing.

The elements of a specification that actually earn their place:

Context, meaning the situation the model needs in order to make a correct judgement rather than a generic one. Not background colour, decision-relevant facts.

The task, stated as a single unambiguous instruction. If your instruction contains an "and" joining two different kinds of work, you probably have two calls, and splitting them will improve both.

The output contract, meaning the exact shape you need back. If you are going to parse the output, ask for a schema and validate against it, and treat a validation failure as a retry condition rather than something to paper over with string manipulation.

Constraints, meaning what the model must not do. These are frequently more useful than positive instruction, because the failure modes are more predictable than the successes. Do not invent citations. Do not answer if the documents do not contain the answer. Do not exceed one page.

And examples, which remain the highest-leverage element and the most neglected. Two or three worked examples of input and correct output communicate more about your standard than a thousand words of description, because they encode the thing you cannot articulate: taste. If you find yourself writing a fourth paragraph explaining what good looks like, stop and write an example instead.

There are a few techniques worth knowing because they change results substantially rather than marginally.

Decomposition. Almost any complex task performs better as a chain of narrow steps than as a single broad request. Extract, then classify, then draft, then check. Each step is easier to specify, easier to test, easier to price into the correct model tier, and easier to debug when it breaks. The instinct to do it all in one call comes from the chat interface and does not survive production.

Letting the model reason before it answers. Asking for the analysis before the conclusion produces better conclusions, because the intermediate tokens are themselves the computation. Newer models increasingly do this on their own, but the principle holds: never ask for a verdict as the first token.

Adversarial passes. Have a second call critique the first, with a specific rubric and permission to be harsh. This catches a meaningful fraction of errors, particularly unsupported claims and missed constraints, and it is cheap because the critique is short. It is not a substitute for real evaluation. It is a filter in front of one.

Constraint-based prompting, which is the one I use most in my own work. Give the model a genuinely restrictive brief: this budget, this timeline, this team, these channels are unavailable. Constraints force specificity and eliminate the generic strategic mush that models default to when the question is open. An unconstrained ask produces an answer that would apply to anybody, which means it applies to nobody.

Two warnings.

Do not build a prompt library as your primary asset. Prompts are consumables. They are tuned to a model generation, they decay, and treating them as intellectual property leads people to guard something that will be worthless in a year while neglecting the evaluation suite that would actually have survived.

And when a prompt is not working, resist the urge to add more words. The usual reason a prompt fails is not insufficient instruction, it is that the task is underspecified in the requester's own head, or that the necessary information is not in the context at all. Adding emphasis to an instruction the model cannot follow because it lacks the facts is a way of feeling productive while changing nothing.

This is the part people get wrong

If you want a second pair of eyes on your version of it, book fifteen minutes and bring the messy version.

Book 15 minutes

Chapter 12Evaluations, or how you know it works

This is the most important chapter in the book and the one most likely to be skipped, so let me put the stake in the ground plainly: if you cannot measure whether your AI system's output is good, you do not have a product. You have a demo with a subscription attached, and you will discover this at the worst possible moment.

The reason this matters more with AI than with ordinary software is the nature of the failure. Conventional software fails loudly. A function throws, a page 500s, a test goes red, and you know. AI systems fail quietly and fluently. The output is well-formed, confident, appropriately toned, and wrong, and it will sail past every check you have unless you built one specifically to catch it. Quality drifts by a few percent a week and nobody notices until a client notices, which is the most expensive way to find out.

So you build a sensor. Here is what that means concretely.

Start with a golden set. Collect real inputs from your actual use case, thirty to fifty to begin with, chosen to include the easy middle and the hard edges. For each one, write down what a correct output looks like, or if there is no single correct output, what the criteria for a good one are. This is manual, unglamorous, takes a day or two, and is the highest-return day of work in the entire build. Almost everyone skips it and almost everyone regrets it.

Then define how you score. There are three mechanisms and you will use all three.

Deterministic checks, where the correct answer is checkable by code. Did it return valid JSON matching the schema. Did every cited document ID exist. Did the total match the sum of the line items. Is it under the length limit. These are cheap, exact and should cover as much of your surface as you can arrange, which is more than you think if you design the output format deliberately.

Model-graded checks, where a separate call scores the output against a written rubric. This works, with caveats that matter. Grader models are biased toward longer answers, toward answers that resemble their own style, and toward agreeing with whatever framing the prompt supplies. Use a different model as the grader than the one that produced the output where you can, give the grader a specific rubric rather than asking whether the output is good, and periodically check the grader against human judgement on a sample. A grader you have never validated is a random number generator with a good vocabulary.

Human review, on a sample, forever. Not because the automation is inadequate but because the automation only measures what you thought to measure, and human review is how you discover the failure modes you did not anticipate. Ten outputs a week, read properly by someone who knows the domain, will surface things no rubric catches.

Then you wire it into the loop, and this is the part that turns a measurement into a system. Every change to a prompt, a model, a retrieval strategy or a chunking rule runs against the golden set before it ships. You compare against the previous run. You look at what got better and, more importantly, what got worse, because AI changes are rarely monotonic: a prompt tweak that improves your five hardest cases will frequently break three easy ones, and without the suite you will ship it and find out from support tickets.

0130-50 REAL020304
The evaluation loop, which is the actual product.

The compounding property is what makes this worth the discipline. Every failure you find in production becomes a new case in the golden set, and once it is in there, that specific failure can never silently return. Over a year, your suite becomes an encoded map of everything that has ever gone wrong in your domain, and it is worth more than any prompt you own. It is also, incidentally, the most convincing thing you can show a serious buyer, because it demonstrates that you know how your system fails, which almost nobody in this market can say.

Some practical measurements that are usually worth having in the suite alongside the quality score.

Cost per unit of work, tracked over time, so you find out that a prompt change tripled your token usage before the invoice does.

Latency at the ninety-fifth percentile, not the average, because the average hides the tail and the tail is what users actually experience as broken.

Refusal and abstention rate, meaning how often the system correctly declines when the answer is not available. A system that never abstains is not confident, it is fabricating, and if that number is zero you have a problem you have not found yet.

And the escalation rate, meaning how often a human has to intervene. This is the number that determines your unit economics in any service business, and it is the number that should be falling month over month if the system is actually learning.

One last thing about evaluation, which is a commercial point rather than a technical one. When you sit across from a serious buyer, the question underneath every other question is: what happens when this is wrong. If you can answer that with a specific, measured account, you separate yourself from every competitor in the room. Not because your model is better. Because you can prove you know what your system does at the edges, and everyone in that room has been burned by someone who could not.

Chapter 13Agents, tools, and the loop

The step from a model that answers to a system that acts is the largest change in practical capability of the last few years, and it is where most of the remaining commercial upside sits. It is also where the failure modes get sharp, so it deserves a careful treatment rather than an enthusiastic one.

The mechanism is simpler than the vocabulary suggests. You give the model a set of tools, which are just functions with descriptions: search this database, send this email, run this query, read this file, call this API. Instead of only producing text, the model can produce a request to call one of those tools. Your code executes the call, returns the result into the context, and the model continues with the result now in hand. Loop until the goal is met or a limit is reached.

That loop is the whole idea. Everything called an agent is some arrangement of it.

What this unlocks is real. The model can now fetch the fact instead of recalling it, which removes an entire class of fabrication. It can do exact arithmetic by writing code rather than by predicting digits. It can take multiple steps toward a goal, notice that step two failed, and try a different approach, which is the difference between a tool and a worker. And it can operate on live state rather than on a frozen snapshot of the world.

What this breaks is equally real, and the breakages are systematic.

Errors compound across steps. If each step in a chain is ninety-five percent reliable, a ten-step chain is around sixty percent reliable end to end, and a twenty-step chain is a coin flip. This arithmetic is unforgiving and it is the single largest reason ambitious agent demos do not become products. The response is not to hope for better models, it is to shorten the chain, checkpoint aggressively, and make each step verifiable independently.

Cost becomes unpredictable. A conversational call has a bounded price. An agentic loop can decide it needs eleven more tool calls, and each of those carries its accumulated context. Loops need hard budget caps, on tokens, on steps and on wall-clock time, enforced by your code rather than by an instruction in the prompt. An instruction is a request. A cap is a guarantee.

Actions are not idempotent. If a step sends an email and the loop retries it, the email goes twice. This is a distributed systems problem, not an AI problem, and it has known solutions: idempotency keys, an action log, a distinction between operations that are safe to repeat and operations that are not. Any system with side effects needs this from day one, and most prototypes do not have it, which is why the first production incident is usually a duplicate.

And the injection surface widens dramatically. The moment your agent reads content from outside your trust boundary, that content is talking to your model. A web page can contain instructions. A customer email can contain instructions. A PDF can contain instructions in text the human reader will never see. If your agent has a tool that sends email or moves money, and it also reads untrusted content, you have built a system where a stranger can influence what it does.

The architectural answers are known and are worth stating explicitly, because I see them omitted constantly.

Separate the reading context from the acting context. The component that processes untrusted content should not be the component that holds the dangerous tools. Have the reader extract structured data, validate that structure with code, and pass only the validated structure onward. Untrusted prose never reaches the actor.

Give every tool the narrowest scope that works. Not a database connection, a specific parameterised query. Not an email client, a function that sends one template to one address from an approved list. The blast radius of a confused model is exactly the size of the permissions you handed it.

Put a human in the loop wherever the action is irreversible, and be honest about which actions those are. Sending an external message, moving money, deleting data, publishing anything. The right design is usually that the agent prepares the action completely and a human approves it in one click, which preserves nearly all the time saving while removing nearly all the risk.

And log everything, in a form a human can read. Every prompt, every tool call, every result, every decision. When something goes wrong in an agentic system, and it will, the only way to understand it is to read the trace. Teams that skip this spend days on incidents that a good log resolves in ten minutes.

The design principle I would put above all the others: build the smallest agent that does the job. Multi-agent architectures with specialised roles are intellectually appealing and mostly a way of multiplying failure surface. If a single well-specified loop with four tools solves the problem, that is the correct architecture, and the fact that it is unimpressive to describe at a conference is not a defect.

Want this as you go?

Drop your email and I will send the new chapters and the tools as they land.

Part 03THE STACK


Chapter 14Building an arsenal without becoming a tool collector

There is a specific disease in this field and I have had it. You subscribe to everything. You spend your Sundays testing the launch of the week. You have forty tabs of AI products open and a monthly bill that would fund a part-time contractor, and at the end of the quarter you have learned a great deal of interface trivia and shipped nothing.

The cure is to notice that tools are not the asset. Your workflow is the asset, and tools are interchangeable parts inside it.

I evaluate anything new against three questions now, and the order matters.

Does it remove a step from a process I actually run? Not a process I could imagine running. One that consumed my hours last month. If I cannot name the step, the answer is no, regardless of how impressive the demo was.

Will the capability still exist in eighteen months, and will it be mine? A lot of what gets sold as a product in this market is a feature that the model providers will absorb into their platforms, and when they do, the company disappears and takes your workflow with it. Building a critical dependency on something that is one platform release away from being free is a real risk and you should price it.

Does it deepen something or widen something? Widening is adding a new capability you will use occasionally. Deepening is making something you already do faster or more reliable. Deepening compounds. Widening mostly does not, and the pleasure of widening is exactly why people accumulate subscriptions.

Now, the layers that a working setup actually needs, described by function rather than by brand, because the brands rotate.

You need direct access to the models, not just their chat products. This is the layer people skip longest and it is the one that separates a user from a builder. The chat interface is a demonstration of what the model can do. The API is where you build a system. The moment you can make a scripted call in a loop over a thousand rows, your capability changes category.

You need somewhere to run code. Not necessarily as a professional developer. But a large fraction of the leverage in this field is available only to someone who can write forty lines that call an API, transform the result and write it somewhere, and the models themselves have made this dramatically more accessible. If you have avoided this because you are not technical, understand what you are trading away: it is the difference between doing a task with AI once and having AI do the task for you a thousand times.

You need an automation layer for the things that should happen without you. There is a spectrum here from visual builders that anyone can use through to self-hosted workflow engines with full programmability. Start at the simple end. Move down the spectrum only when you hit a specific wall, and expect to hit it later than you think.

You need a system of record. A database, a CRM, a structured store of some kind, that holds the state your workflows read and write. This is the least exciting item on the list and the one whose absence causes the most chaos, because without it every automation is stateless and every mistake is invisible.

You need version control and a place to keep prompts, evaluations and configuration as files. Prompts kept in a chat history or a document are not managed, they are lost. Prompts kept in a repository with a change history are an asset you can reason about, roll back and hand to someone else.

And you need observability: logs, cost tracking, and a dashboard that tells you what your systems did while you were not watching. In practice this is the difference between a set of automations and an operation.

That is six layers, none of them glamorous, and a setup with all six beats a setup with thirty subscriptions and none of them, every time.

On the question of mastery versus breadth, my honest position after doing both: pick a small number of tools and go deep enough that you know their failure modes. Knowing what a tool cannot do, and what it does badly under load, is worth more than a passing familiarity with ten alternatives. The person who has run one automation platform hard for a year knows things that no amount of comparison shopping teaches.

One caveat on the discipline, because I do not want to talk you into stagnation. The field is moving genuinely fast, faster than any technology cycle I have worked in, and a rigid stack becomes a liability. The way I hold both is to schedule the exploration rather than doing it continuously. A fixed window, once a quarter, where I look at what has changed, run two or three candidates against a real task, and decide. Outside that window, the stack is frozen and I do the work.

Chapter 15From the chat window to a system

Almost everyone's AI journey has the same shape. You use a chat product and it is genuinely useful. You start copying things into it and copying results back out. Then you notice you are doing that fifteen times a day and think, this should be automated. And that is the moment the difficulty steps up by an order of magnitude, because the chat window has been hiding all of the hard parts from you.

Here is what it was hiding.

It was managing your context. Every message you sent silently included the whole conversation. Once you build with the API, you own that decision, and it turns out to be most of the design work.

It was retrying for you. When a call failed, the interface quietly tried again or showed you a spinner. In your own system, a failed call is an exception in the middle of a loop over four thousand records, and if you did not plan for it, you now have a partially completed job and no record of where it stopped.

It was rate-limiting you invisibly. Providers cap requests per minute and tokens per minute. A script that works perfectly on ten rows will start returning errors at two hundred, and the fix is queuing and backoff, not a faster loop.

It was giving you a human in the loop by default. You read every output before doing anything with it. In an automated pipeline nobody reads anything, which means the quality bar has to be enforced by code, and anything you were catching by eye is now shipping.

So the work of turning a prompt into a system is mostly plumbing, and the plumbing is where the reliability lives. Here is the shape that works.

Ingest, with validation. Something arrives: a form submission, a row, an email, a file. Validate its structure before anything else touches it. Reject or quarantine what does not conform. A surprising share of pipeline failures are malformed inputs that were allowed in and blew up four steps later where the cause was unrecognisable.

Enrich, cheaply. Add whatever context the decision will need, from your database, from an external source, from a prior record. Do this before the model call, not inside it, because deterministic lookups are cheaper, faster and more reliable than asking a model to remember.

Decide, with the cheapest model that can. Classification and routing steps belong on the fast tier. Get the branching done before you spend frontier tokens.

Generate, on the appropriate tier, with a strict output contract. Ask for structured output, validate it against a schema, and treat validation failure as a retry with a repair instruction rather than as something to handle with regular expressions.

Verify, with the deterministic checks you defined in your evaluation work. This is where the golden set pays rent: the same checks you run in testing run in production on every item, and anything that fails gets routed to review instead of shipped.

Act, idempotently, with a log. Every side effect gets an idempotency key and an entry in an action log, so a retry does not double-send and so you can answer the question "what did the system do on Tuesday" without guessing.

Review, on a sample and on all exceptions. Someone competent looks at a fixed number of items every week, and at everything the verification step flagged. This is not a temporary measure until the system is good. It is permanent, and its output is new evaluation cases.

01VALIDATE02FAILURE EXIT03NOTE CHEAPER04NOTE BRANCH
The seven stages of a production pipeline, and where each one fails.

A few hard-won notes on the plumbing.

Design for partial failure from the beginning. Long-running jobs will be interrupted. Write progress to durable storage as you go, make jobs resumable, and never assume a batch will complete in one run. This costs an hour at the start and saves days later.

Separate the expensive step from the cheap step in your retry logic. If a job does an expensive generation and then a cheap write, and the write fails, do not retry the generation. Cache the intermediate result. This is obvious in the abstract and routinely missed in code.

Version everything that affects output. The prompt, the model identifier, the retrieval configuration, the schema. Store the versions alongside the result. When a client asks why this output differs from one they got in March, you want to be able to answer that with a lookup, not a shrug.

And build the alarm before you need it. A daily summary of volume, cost, error rate and escalation rate, delivered somewhere you actually look. The failure mode of automation is not that it breaks loudly, it is that it degrades quietly and produces plausible garbage for three weeks.

Chapter 16The economics of inference

This chapter exists because I have watched several otherwise sound AI businesses discover, well after launch, that their gross margin was structurally negative and nobody had done the arithmetic.

The arithmetic is not hard. It is just usually not done, because during development you are running dozens of calls a day and the cost rounds to nothing, and then you go to production and run hundreds of thousands.

Start from the unit. Define one unit of work in your product: one document processed, one report generated, one conversation handled, one lead qualified. Then count what that unit actually costs in tokens, remembering that you pay for input as well as output, that input is usually much larger than people estimate once retrieval and history are included, and that in an agentic loop the accumulated context is resent on every iteration. That last one is the killer, and it is why a ten-step loop can cost far more than ten times a single call.

Once you have cost per unit, put it next to price per unit, and you have your gross margin before support, before infrastructure, before the human review you are certainly going to need. If that number is not comfortable, you do not have a pricing problem, you have an architecture problem, and there are five levers.

Route down. The largest single saving available to most systems, usually by a wide margin. Audit every model call in your pipeline and ask what the cheapest tier that passes your evaluation suite is. In most systems I have looked at, the majority of calls are running on a model far more capable than the task requires, purely because that is what was used during development.

Cache. Two kinds. Prompt caching, offered by the providers, reuses computation on a repeated prefix, which is enormously effective when you have a long stable system prompt and short variable inputs. Design your prompts so the stable part comes first, which is free and is frequently not done. Then application-level caching: if the same question has been answered before and the underlying data has not changed, do not pay to answer it again. In domains with repetitive queries this alone can remove a large share of your traffic.

Batch. Providers offer discounted asynchronous processing for work that does not need to happen this second. A large fraction of business workloads are not interactive, and moving them to batch is a straight price reduction for a latency you do not need.

Shorten. Every unnecessary paragraph in the system prompt is paid for on every single call forever. Every irrelevant retrieved document is paid for twice, once in tokens and once in degraded output quality. Trimming context is the rare optimisation that improves cost and quality simultaneously.

Precompute. If you can do the expensive work once, offline, and serve it many times, do. Summarising a document on ingest rather than on every query is a common example and a large saving.

SYSTEM PROMPT8862412714
Where inference cost actually goes, and what each lever removes.

Two structural pricing points that follow from all of this.

Never price a variable-cost service at a flat rate without a cap, unless you have measured the tail of your usage distribution and can survive it. The pathological customer is real, they exist in every cohort, and in a usage-based cost structure a single one can consume the margin of dozens. Either meter, or cap, or price with the tail included.

And be careful about the seductive logic of pricing against the human cost you replace. If a task costs a client four hundred dollars of labour and costs you four dollars of tokens, the temptation is to price near the client's number and celebrate the margin. That works until a competitor with the same models does the arithmetic and prices at forty. The defensible version is to price on the outcome and make the moat something other than the inference, which is the subject of the next chapter.

One more note that is easy to miss. Token prices for a given level of capability have fallen consistently and steeply, and there is no sign of that stopping. This has a strategic implication: a business model that is marginal today at current prices may be comfortable in eighteen months without you changing anything. It also has the inverse implication, which is that any advantage you hold purely because you are willing to pay for expensive inference is temporary, because the price of that capability is falling for everyone.

Chapter 17Data is the thing you actually own

I said in Chapter 14 that proprietary data is one of the few real moats available, and I want to spend a chapter on it, because most people agree with the statement and then do nothing structural about it.

Here is the distinction that matters. There is data you can buy, which is not a moat because your competitor can buy it too. There is data you scraped, which is not a moat and is increasingly a legal liability. And there is data that exists because your system was used, which nobody else can obtain at any price, because it did not exist before you built the thing that produced it.

That third category is what to design for, and designing for it is a decision you make at the start or do not get.

What does it look like in practice? Every time your system produces an output and a human corrects it, that correction is a labelled example. Every time a user chooses one of three options, that is a preference signal. Every time an extraction fails and someone fixes it, that is a hard case. Every time a client tells you their exception rule, that is domain knowledge that exists in no public corpus.

The failure is that in most systems all of this is discarded. The correction is made in a text field, the output is overwritten, the fix is applied in a spreadsheet, and the information evaporates. A year later the team has learned an enormous amount and none of it is in a form that can be used.

So the first piece of advice is structural and simple: capture the corrections. Not as free text. As a record with the input, the system output, the corrected output, the person, the timestamp and, where you can get it, the reason. That record costs almost nothing to store and it is the raw material for everything else.

What it buys you, in ascending order of value.

It buys you a growing evaluation set, which is the immediate return and reason enough on its own. Every correction is a case where the system was wrong, which makes it exactly the case you want in your golden set. A team that captures corrections has a suite that improves automatically as a byproduct of operating.

It buys you a map of where you are weak. Cluster the corrections and the pattern appears: this document type, this edge case, this customer segment. That map tells you what to fix, in priority order, based on frequency rather than on whoever complained loudest.

It buys you examples, which are the highest-leverage element of a specification. Corrections are, by construction, examples of the difference between what the system produced and what good looks like. Feeding a well-chosen handful back in as few-shot examples is the cheapest quality improvement available and most teams never do it because they did not keep the corrections.

And eventually it buys you the option of training. I want to be measured about this, because fine-tuning is oversold. For most businesses, most of the time, better context beats a fine-tuned model, and the correct order of attempts is: fix the context, then fix the prompt and examples, then consider fine-tuning. Fine-tuning earns its place in a narrow set of cases, mainly when you need a consistent format or style at very high volume, when you want to run a smaller and cheaper model at the quality of a larger one, or when the task involves a domain vocabulary that general models handle badly. In each of those, the input is a large set of high-quality examples, which is exactly what a year of captured corrections is.

A word on the ethics and the contracts, because this cannot be an afterthought. If you are going to learn from data your clients generate, say so, in the agreement, in plain language, before you do it. Say what you use it for, whether it is aggregated, whether it ever leaves your environment, and what happens to it when they leave. Some clients will say no and that is fine. What is not fine is doing it quietly and having the question asked in a security review two years later, because at that point it is not a data conversation, it is a trust conversation, and you will lose it.

There is a related asset that people undervalue even more than data, which is the accumulated record of decisions.

Why did this client want the output formatted that way. Why did we stop using that supplier. What was the reason for the exception on invoices from this region. In most companies the answers live in a few people's heads and leave when they do. In a company that writes them down in a structured, retrievable form, they become context that gets loaded into every relevant decision from then on, which is a compounding advantage that has nothing to do with model quality and that no competitor can shortcut.

That is the real answer to the moat question, and it is unglamorous. The durable thing is not the technology. It is the accumulated, structured record of what happened and what was learned, held in a form a machine can retrieve. Build the capture mechanism early, even when there is nothing to capture, because the version of you in two years cannot go back and collect it.

Chapter 18Build, buy, or wrap

The question I get asked most by founders on the show is some version of: is this defensible, or will the model providers just do it.

The honest answer is that it depends entirely on where your value sits, and there is a clean way to think about it.

Ask what a competitor with the same models, the same budget and no access to your specific situation would need in order to reproduce what you have. Whatever is on that list is your moat. If the list is "write a similar prompt," you have no moat, and that is worth knowing early rather than late.

Things that reliably are not moats, in this market, right now.

A prompt, however sophisticated. It is a text file. It can be reconstructed by anyone who reads your outputs carefully, and it decays with each model generation anyway.

A choice of model. Everyone has the same menu.

A user interface, on its own. Interfaces are copied in weeks, and in this field they are copied faster because the underlying capability is public.

Being first. First matters when it converts into something durable, distribution, data, contracts. On its own it is a head start measured in months.

Things that reliably are moats.

Proprietary data, specifically data that is generated as a byproduct of your product being used. Not a dataset you bought, which someone else can buy. A dataset that exists because your customers used your system and that grows with usage. This is the strongest position available and it is worth designing your product to create it deliberately.

Workflow depth. If you are the system of record for a process, if the work happens inside you rather than passing through you, replacing you means migrating a live operation. Companies do not do that casually.

Integration surface. Every connection to a customer's existing systems is a switching cost. Ten integrations is a project to unwind, and projects to unwind do not get prioritised.

Regulatory and compliance position. In regulated industries, the certification, the audit trail and the documented process are a barrier that has nothing to do with model quality and that takes competitors years, not weeks.

Distribution and trust. The unfashionable answer and often the correct one. If you are the person the industry calls, that is not reproducible by writing better code.

And the evaluation suite, which I keep coming back to because it is genuinely underrated. A year of encoded knowledge about how this specific task fails in this specific domain is not something a competitor can prompt their way to. It has to be earned by shipping and failing.

On the build-versus-buy question specifically, the rule I use: buy the layers everyone needs and build the layer that is yours. Nobody should be building their own vector database, their own authentication, their own queueing. Everybody should be building their own evaluation, their own domain logic, and their own context assembly, because those are the parts that encode what you know.

The wrapper question deserves a direct answer too. Thin wrappers over model APIs are not shameful and some of them are good businesses. What determines the outcome is whether the wrapper is a starting position or a final position. If you launch as a wrapper and use the early revenue and usage to build data, integrations and workflow depth, you are running a legitimate strategy. If you launch as a wrapper and stay a wrapper, you are in a price war with people whose costs are identical to yours, and the model providers are shipping your feature set as a checkbox.

Want me to look at yours?

Bring the thing you are least sure about. That is the part worth the fifteen minutes.

Book 15 minutes

There is one more thing I would say to anyone weighing this, from the seat I occupy when founders pitch me. I discount claimed technical advantage almost entirely, because in this field it has a half-life of months. What I do not discount is evidence that a company has found a specific expensive problem, has proof that their system solves it, and has a reason customers cannot easily leave. Those three things, in that order, are what a durable AI business looks like from the outside.

Part 04THE BUSINESS MODELS


Chapter 19Consulting, priced on the outcome

Consulting is the fastest route from capability to revenue, and it is the one I would tell almost everyone to start with, for a reason that has nothing to do with the money.

It puts you in the room. Every week you are inside a real organisation, watching the actual process, hearing the actual objection, discovering that the thing you assumed was the problem was a symptom of a political fight between two departments. That information is unavailable from the outside, and it is the raw material for everything else in this book. The product ideas that work almost always come from consulting engagements, because that is where you find out what people will actually pay to have removed from their lives.

The revenue is a side effect of being educated at the client's expense, and the fact that it is good revenue is a bonus.

Now, how to do it without ending up in the commodity trap.

Your price is set by the frame, not by the work. Two people delivering an identical result can be an order of magnitude apart on fee, and the difference is which question the buyer thinks they are answering. If the buyer is answering "how much does an AI person cost per hour," you are being compared to a labour market and you will lose. If the buyer is answering "what is it worth to stop losing this specific amount every quarter," you are being compared to the problem, and the problem is expensive.

Getting into the second frame is a sequencing move, not a persuasion move. It happens in discovery, before any number is discussed, and it consists of getting the client to quantify their own pain out loud. Not you asserting it. Them saying it. How many hours does this take now, across how many people, at what loaded cost. What does it cost when it goes wrong, and how often does it go wrong. What has not happened because this consumes the team. When they have said those numbers, your fee is a fraction of a number they produced, and the conversation is different.

If they cannot quantify it, that is important information. Either the problem is not actually expensive, in which case you are about to sell a project nobody will value, or nobody has measured it, in which case measuring it is your first engagement and it is an easy one to sell.

On structure, the shape that works reliably has three tiers and each one leads to the next.

An assessment. A short, fixed-scope, fixed-fee piece of work that produces a decision-grade document: here is where AI helps in your operation, here is what it is worth, here is the sequence, here is what it will cost, here is what will go wrong. This is easy to buy because the price is small relative to the decision it informs, and it qualifies the client ruthlessly. A client who will not pay for the assessment will not pay for the implementation.

An implementation. The actual build, scoped from the assessment, with defined deliverables and a defined end. This is where the real money is, and the discipline that matters is that it has an end. Open-ended implementations become resource drains that neither side is happy with.

An ongoing arrangement. After delivery, systems need monitoring, tuning and extension, and the models underneath keep changing. A monthly fee for that is legitimate, valuable and predictable, and it is how a consulting practice stops living quarter to quarter.

On terms, I will be direct about how I think this should be done, because there is a lot of bad practice in this market.

Flat fees. A defined scope for a defined price, agreed before the work starts. Not a percentage of results, not equity in place of payment, not a share of savings, all of which create misalignment and disputes about attribution that poison relationships. If you cannot scope it, scope a smaller piece of it.

No guarantees of outcome, and no predicted numbers in the proposal. You are supplying technology and expertise. You do not control the client's staffing, their data quality, their politics, their market or their willingness to change how they work, and any of those can sink a good system. Promising a percentage improvement is writing a cheque against variables you do not hold. Describe the mechanism, describe what you will deliver, describe what you have done before, and let the client draw the inference.

And the client owns every relationship the work touches. If your engagement involves reaching their market, their customers, their prospects or their investors, those relationships are theirs, in their systems, under their control, from the first day. You are supplying the machinery and the operating expertise. You are not standing between a company and the people it does business with. This is both the right way to do it and the only version that survives the day the engagement ends.

On scope creep, which kills more consulting practices than pricing does: the fix is not firmer language in the contract, it is a change-order process that is genuinely easy to use. Clients ask for extra work because they need it, not because they are exploiting you. Give them a frictionless way to buy it and the problem converts from a source of resentment into a source of revenue.

One last thing about consulting that people miss. The deliverable is not the report. The deliverable is that a person inside the client organisation can now do something they could not do before, and can keep doing it after you leave. Systems that only work while you are there do not get renewed, do not get referred and do not become case studies. Build for the day you are gone and every other metric improves.

Chapter 20Productised services, and the agency that scales

The agency model is where consulting goes when it grows up, and it is the model I think is most underrated right now.

The idea is straightforward. Instead of custom work for every client, you find the piece of the work that repeats, standardise it, and sell the standardised version many times. You keep the customisation at the edges, where clients feel it, and you industrialise the middle, where they do not. AI makes this dramatically more viable than it used to be, because the parts that used to require a person now require a well-evaluated pipeline with a person checking it.

The economics are better than consulting because delivery is not linear in headcount, and better than software because you can charge for outcomes rather than for seats. The trade is that you are running an operation, with all the management that implies.

Getting to a productised service is a specific process and it takes two or three clients of custom work to find it.

Do the custom work first. Deliberately. Take three engagements in a narrow domain, do them properly, and pay attention to what you did the same way each time. That repeated core is your product. You cannot find it by planning, only by delivering.

Then draw the line between standard and custom, and be aggressive about where you put it. Everything on the standard side is a documented process running on a measured pipeline. Everything on the custom side is scoped and priced separately. The temptation is always to move the line to accommodate a client, and every time you do, you are converting a scalable business back into a consulting practice one exception at a time.

Then build the onboarding, because that is where these businesses actually break. The first thirty days determine whether a client renews, and the failure pattern is consistent: a closer sells the work, hands it to delivery, and for three weeks nobody owns whether the client has actually got value yet. The fix is structural. Someone who is not the person who sold it owns day one to day thirty, there is a defined proof point that must be hit inside that window, and it is checked, not assumed.

On retention, which is the whole game in a recurring model, one uncomfortable truth. Retention is decided long before the cancellation. A client who leaves at month four decided at week two, when the thing they were promised did not visibly happen and nobody addressed it. Monitoring your churn rate is monitoring an outcome. Monitoring whether every client hit their proof point in the first month is monitoring the cause.

The pricing structure that works for these businesses is a setup fee plus a monthly fee, both flat. The setup fee covers the real work of onboarding and, more importantly, filters out clients who are not committed, because a client who has paid nothing to start will not do the work required to succeed. The monthly fee is priced against the value of the outcome and does not vary with volume unless volume genuinely drives your costs.

The same terms discipline from the previous chapter applies here and applies harder, because agencies operate closer to the client's revenue. Flat fees. No guaranteed results, no promised numbers, no income projections in the pitch. You are supplying technology and a service, and the client owns their relationships, their data, their accounts and their pipeline. If your service touches the client's market, everything you generate belongs to them in their systems, and it stays theirs whether they renew or not. Any agency that structures it otherwise is holding a client hostage rather than serving one, and that shows up in retention eventually.

On the operational side, three things separate agencies that scale from agencies that plateau at the founder's capacity.

Documented process, to the level where a competent new hire can execute it in week two rather than month three. Most agencies think they have this and actually have tribal knowledge in the founder's head.

Quality control that does not depend on the founder reviewing everything. This is where the evaluation work from Part Two becomes a business asset rather than a technical one: an automated check that catches the same errors the founder would catch is what makes delegation safe.

And client communication that runs on a schedule regardless of whether there is news. The single most common cause of a surprise cancellation is silence. A client who hears from you every week without fail will tell you they are unhappy while it can still be fixed. A client who hears from you when there is something to report will tell you by cancelling.

Chapter 21Software, and why the road is longer than it looks

Software is the model with the highest ceiling and the longest, most brutal road to get there, and the AI wave has made people underestimate the road badly.

The reason is that building the first version got easy. Genuinely easy. A capable person with modern tools can produce a working application in days that would have taken a small team a quarter three years ago. And so the natural conclusion is that the hard part has been removed.

The hard part was never building the first version. The hard part is finding the specific problem that enough people have, that they know they have, that they are already spending money on, and that they will change their behaviour to solve. That has not become one bit easier, and the collapse in build cost has made it harder in one respect, because your competitors can now build as fast as you can, so the software itself is not the differentiator it once was.

If you are going to do this anyway, and it is a legitimate ambition, here is the sequence I would actually follow.

Sell it before you build it. Not a landing page with an email capture, which measures curiosity. An actual conversation where you describe the thing and ask someone to commit, and where the commitment costs them something. A deposit, a signed pilot, a scheduled implementation date. What you learn from ten of those conversations is worth more than six months of building, and if you cannot get any of them, you have just saved six months.

Deliver it manually first. Take the first customers and do the work by hand, or half by hand. This feels wrong and is right. You learn the actual shape of the problem, including all the exceptions that would have destroyed your architecture, and you get paid while you learn. Then you automate the parts that repeat, in the order of how much time they consume. The product that emerges from this process is one that fits the work, rather than one the work has to be bent to fit.

Then build the narrow version and refuse to widen it. The gravitational pull in early software is toward more features, because every prospect who says no gives you a reason, and every reason sounds like a feature. Most of those noes are not feature problems, they are positioning problems or budget problems in disguise, and building for them produces a product that does eleven things adequately and nothing compellingly.

On pricing, the specific advice that matters for AI products: your costs are variable in a way that traditional software's were not, so per-seat pricing can be actively dangerous. Ten seats at a company where two people use it heavily and eight never log in is fine. Ten seats where all ten run large jobs is a different business. Either price with a usage component, or set limits that are generous enough not to annoy and firm enough to protect you, and measure the distribution of usage in your customer base before you commit to a model.

On the metrics, only a few actually matter early and the rest are decoration.

Whether users come back without being prompted. This is the only early signal that is hard to fake, and it is the one that predicts everything downstream.

Time to first value, meaning how long between signup and the moment the user gets something they wanted. Every day in that gap costs you a large fraction of your cohort, and it is almost always compressible with work that nobody wants to do because it is not building features.

Net revenue retention, once you have enough customers for it to be meaningful. Whether existing customers are worth more this year than last is the number that determines whether you have a business or a leaky bucket.

And the one specific to this field, which most SaaS advice will not tell you: your cost of goods per active customer, tracked monthly. Because it can move underneath you when a customer changes how they use the product, and because a healthy-looking revenue chart can sit on top of a margin that is quietly going the wrong way.

I want to be honest about the odds here, since nobody else in this genre is. Most software attempts do not reach meaningful revenue, and the ones that do usually take longer than the founder projected by a factor of two or three. That is not a reason to avoid it. It is a reason to fund it with something else, which is exactly why I recommend consulting first, and why the sequence consulting to productised service to software is a more reliable path than jumping straight to the end.

Chapter 22Content as distribution, not as a product

I want to reframe content entirely, because the standard treatment of it in books like this one is wrong in a way that wastes years.

Content is not a business model. Content is a distribution mechanism, and the mistake is treating it as a thing to monetise directly rather than as the thing that makes everything else cheaper to sell.

Here is the mechanism, stated plainly. Every business has an acquisition cost. You either pay it in money, through advertising, or in time, through outbound, or you build an asset that lowers it permanently. Content is that asset. Its value is not the audience number, it is that when someone finally does have the problem you solve, you are already in their head as the person who solves it, and the sales conversation starts three steps further along.

Which changes what you should make. If content is a product, you optimise for reach, and you end up making the kind of thing that gets attention from people who consume content. If content is distribution, you optimise for the right hundred people, and reach becomes almost irrelevant.

I would rather have five hundred readers who own the problem than fifty thousand who are curious about AI. This is not false modesty about scale, it is arithmetic about conversion, and it means the content that works for a business is often the content that performs worst by public metrics: specific, technical, unglamorous, addressed to a narrow group.

On what to make, the rule that has held for me: publish mechanism, not opinion. Opinion about AI is in infinite supply and its marginal value is zero. Mechanism is scarce because producing it requires having actually done the thing. When you write about how you solved a specific problem, including what failed, you are simultaneously teaching something useful and proving you were there. Nobody can fake the second part convincingly for long.

The format that has produced the most for me is the long conversation. I host a show, and I want to be precise about why it works commercially, because it is not the reason people assume.

It is not the audience. The audience is a nice byproduct. The mechanism is that an invitation to a substantial conversation is the most welcome message a busy senior person receives, because it is the only one in their inbox that is offering rather than asking. It reverses the entire posture of outreach. Nobody accepts a sales call from a stranger. A lot of people accept an hour to talk about their work with someone who has clearly prepared.

And the preparation is the actual asset. To interview someone properly, you research their company, their market, their thesis and their problems. By the time you sit down, you understand their situation better than most people who have been trying to sell to them for a year. Whether or not anything commercial ever follows, you now have a real relationship with someone senior in your target market, built on a foundation of you having given them something.

One conversation also produces a great deal of downstream material: the full piece, the clips, the written version, the notes, the specific insights that seed other work. That is a genuine multiplication and AI helps enormously with the mechanics of it. But the multiplication is not the point. The relationship is the point, and the content is the residue.

A SPECIFIC
The compounding loop that content actually drives.

The honest part about content that the genre never says: the loop above is slow. The first six months produce almost nothing measurable, and this is where nearly everyone quits, usually while telling themselves the market is too crowded. It is not a growth channel in any short-term sense. It is a fixed investment in lowering a recurring cost, and like any such investment, the return is invisible right up until it is obvious.

The counsel I would give is to size it accordingly. Do not build a content operation before you have a business. Do build a consistent, narrow, mechanism-heavy publishing habit alongside whatever else you are doing, at a cadence you can hold for two years without heroics, because two years is the timescale on which this actually pays.

Chapter 23How this work actually gets bought

I have been on both sides of this table. I have sold technical work and I now sit on the receiving side, evaluating people who want money from a company I am responsible to. The buying process for AI work has a specific shape and it is different from ordinary software or consulting, so it is worth mapping.

The first thing to understand is who is actually in the room. There are usually three parties and they want different things.

There is the person with the problem, who is operational, who feels the pain daily, and who is your champion if you have one. There is the person with the budget, who does not feel the pain and who is evaluating this against four other things. And there is the person with the veto, which in AI deals is almost always security, legal, compliance or IT, and who is not evaluating the upside at all. They are evaluating whether this creates a problem for them.

Most deals in this field die at the third party, and they die because the seller spent all their energy on the first. You can have an enthusiastic champion, a signed-off budget and a dead deal, because nobody prepared for a data processing question that was always going to be asked.

So prepare for it. Have a written answer to the standard set before the first meeting: where does the data go, who processes it, is it retained, is it used for training, where geographically, what is your subprocessor list, what happens on termination, what is your incident process. This is a two-page document and having it ready moves deals by weeks. Not having it is the single most common self-inflicted delay I see.

The second thing is that the objections in AI deals are consistent, and they are reasonable, and they should be answered rather than handled.

"How do we know it is right." This is the real question underneath most of the others, and it is why Chapter 9 is a sales chapter as much as a technical one. The answer is not reassurance. It is a specific account of how you measure quality, what your error modes are, and what happens when it is wrong. Nobody expects perfection. They expect you to know your failure rate, and almost nobody in this market can state one.

"What happens when the technology changes." Answer honestly: parts of what we build will be superseded and here is which parts, and here is why the parts that matter to you, the process, the integration, the data, the measurement, are not model-dependent. Pretending your work is permanent invites the follow-up you cannot answer.

"We are thinking of building this ourselves." Also reasonable, and the answer is not to disparage their team. It is to be specific about what the build actually includes: not the demo, which they can absolutely build, but the evaluation suite, the edge cases, the maintenance as models change, the review workflow, and the opportunity cost of their team not doing the thing only their team can do. Some of them should build it themselves and telling them so is worth more than the deal you would have lost anyway.

"Our data is a mess." True, always, everywhere. The right response is that this is normal, that it changes the sequencing rather than the feasibility, and that the first engagement can be scoped to find out how much of a problem it is. A seller who is surprised by messy data has not done this before and the buyer will notice.

The third thing is about the demo, and this is where I would push against common practice.

A polished demo on curated data is now a negative signal to a sophisticated buyer, because everyone has one and everyone knows how they are made. The thing that lands is running your system on their data, live, including the failures. Take a handful of their real documents, or their real tickets, or their real records, and process them in front of them. Show what worked and show what did not and explain why. You will lose a few deals doing this that a polished demo would have won, and the ones you lose were going to fail in delivery anyway.

The fourth thing is the pilot, which is where most AI sales cycles now go and which is a trap if you structure it badly.

A pilot with no defined success criterion is a way for a buyer to feel productive without deciding, and it can consume months of your delivery capacity for nothing. Before agreeing to one, get three things in writing: what specifically will be measured, what number constitutes success, and what happens if it is met. That last one is the important one. If the answer to "and then what" is vague, the pilot is not a step toward a purchase, it is a substitute for one.

And price the pilot. A paid pilot with a defined criterion converts at a completely different rate than a free one, because payment is the cheapest available test of whether the buyer is real.

The fifth thing is timing, and it is the least controllable. Most AI purchases in established companies happen because of a trigger: a person left and their work has to be absorbed, a volume increase broke the existing process, an audit found something, a competitor did something visible, or a budget appeared. Your job is largely to be the obvious call when the trigger fires, which is a positioning and content problem rather than a sales problem, and which is why the publishing habit from Chapter 4 is a sales asset.

The last thing I will say is about how you talk about what you do. The strongest position, and it took me a long time to learn this, is to be the person in the conversation who is most willing to say what the technology cannot do. In a market saturated with people claiming everything, calibrated honesty is a differentiator that is nearly free to produce and almost impossible for an over-claiming competitor to match, because they have already spent their credibility.

Chapter 24Pricing, concretely

Pricing is the highest-leverage lever in any of these business models and the one people spend the least time on, so let me be concrete rather than philosophical about it.

Start with the observation that makes AI pricing genuinely different. Your cost to deliver is falling, sometimes fast, and it is falling for your competitors at the same rate. If you price on cost plus a margin, you have signed up to reduce your own price every year in exchange for nothing. That is not a hypothetical risk, it is the default outcome, and it is why cost-based pricing is disqualified before we start.

That leaves two coherent approaches and one hybrid.

Value-based pricing sets the fee as a fraction of the quantified benefit. It is the right answer when the benefit is genuinely quantifiable and the client will say the number out loud. Its weakness is that quantification is often disputed, and that it requires a discovery process good enough to establish the number before the price is discussed. Do not attempt this backwards. If you name a price and then try to justify it with value, you are arguing. If the client establishes the value and you then name a fraction of it, you are calculating.

Market-based pricing sets the fee by reference to what comparable alternatives cost: the salary of the person who does this now, the tool they currently pay for, the agency they use. It is weaker than value pricing and much easier to execute, and for most people starting out it is the realistic option. Its strength is that it is easy for a buyer to sanity check, which reduces friction.

The hybrid, which is what I would actually do in most cases, is to anchor on the market comparison and structure so that the upside is captured through expansion rather than through the initial price. Price the first engagement where it is easy to say yes, deliver something measurable, and let the second engagement be priced against the evidence the first one produced. This is slower and it converts far better, and the second number can be a multiple of the first without any argument, because there is now proof.

Some specific structures and when each is right.

Fixed fee for a fixed scope is the default and should be your default. It is easy to buy, it makes your revenue predictable, it rewards you for getting faster, and it aligns you with the outcome rather than with elapsed time. Its only real requirement is that you can scope, and if you cannot scope, the fix is a smaller first engagement rather than an hourly rate.

A monthly retainer for ongoing operation is right when the system needs continuous attention, which most do. Price it against what it would cost the client to have the capability internally, not against your hours, and be explicit about what it includes so that it does not silently become unlimited support.

Setup fee plus monthly is the structure I would use for anything recurring. The setup fee funds the most expensive part of your delivery, which is onboarding, and it filters out clients who will not do the work. A recurring business with no setup fee is subsidising the acquisition of customers who are least likely to stay.

Usage-based pricing is right when your costs genuinely scale with usage and when the unit is something the client understands and can predict. It is wrong when the unit is opaque, because a price the buyer cannot forecast is a price they cannot approve, and finance departments reject unforecastable spend regardless of how good the value is.

Structures I would avoid, and why. Hourly billing, because it penalises you for every efficiency gain you make, which in this field is most of them. Percentage of results, because attribution becomes a dispute and disputes end relationships. Equity in lieu of fees from a company you have no information rights in, because you are taking the risk of an investor without the position of one. And anything with a guaranteed outcome attached, because you do not control the client's implementation, staffing, data quality or market, and a guarantee against variables you do not hold is a liability disguised as a sales tool.

Now some practical mechanics that matter more than the theory.

Present three options rather than one. Not as a manipulation, but because a single price invites a yes or no decision, while three invite a which decision. Make them genuinely different in scope, not the same thing at three price points, and make the middle one the one you actually want to sell.

Put the price in writing on the same day as the conversation. The energy in a deal decays fast and a proposal that arrives four days later arrives into a different context.

Never discount for nothing. If you reduce the price, take something out or get something back: a shorter scope, a longer term, a case study, a reference call, faster payment. A discount given freely teaches the buyer that the original number was theatre, and that lesson is permanent.

Raise prices on new customers only, and do it before you feel ready. The signal that you are underpriced is not complaints, it is the absence of them. If nobody has flinched at your number in six months, it is too low, and the cost of that is not just the margin, it is the customer profile you are attracting.

And the last one, which is the hardest and the most valuable. Be able to walk away, out loud, when the fit is wrong. Not as a technique. The reason it matters is that a business which cannot decline work will accept work it cannot deliver well, and badly delivered work costs more than the revenue it brought in, in support burden, in reputation and in the opportunity it displaced. The ability to say no is downstream of having enough pipeline that no single deal is decisive, which makes pricing power, in the end, a function of demand generation rather than of negotiation skill.

Chapter 25Choosing among the four

These four models are not a menu you pick from once. They are a sequence, and the sequence has a logic.

Consulting first, because it requires no capital, generates revenue immediately, and buys you the market knowledge that everything else depends on. Its ceiling is your time and it does not compound, which is exactly why it is a starting point rather than a destination.

Productised service second, because you build it out of the pattern that consulting revealed. It converts your knowledge into an operation with margin that is not linear in your hours. It is the model with the best ratio of achievability to economics, and I think most people should stop here and be very happy.

Software third, if the pattern you found is broad enough to serve without customisation and if you can fund the road. Highest ceiling, longest odds, and the correct time to attempt it is when you have already proven the demand with a service business.

Content underneath all of them, permanently, because it lowers the acquisition cost of whichever one you are running and it is the only asset that keeps working while you sleep in a way that does not require maintenance.

01CONSULTING02PRODUCTISED03SOFTWARE MODERATE04CONTENT LOW BUT05PRODUCTISED06SOFTWARE SIX
The four models compared on the dimensions that actually decide.

The failure mode I see most often is choosing the model by aesthetics rather than by fit. Software is the prestigious answer, so people who should be running an excellent productised service spend three years building an application nobody asked for. Consulting is the unglamorous answer, so people avoid the thing that would have taught them what to build.

Pick by what you are actually optimising for. If you want your time back, consulting is the wrong end state. If you want a sellable asset, a productised service with documented process and real retention is a legitimately valuable company and is far more achievable than software. If you want the largest possible outcome and you can survive a long stretch of no revenue, software. If you want optionality while you figure it out, run consulting and publish.

The shortcut is a conversation

I do a handful of these a week. No charge, no obligation, and you leave with a next step.

Book 15 minutes

And a note on doing more than one at once, since I flagged earlier that I disagree with the standard advice. Running two models before either is working is not diversification, it is splitting a scarce resource, and the scarce resource is not money, it is your attention on one specific market. Get one line to the point where it produces revenue without your daily intervention. Then add. The sequencing is the whole discipline, and the reason people ignore it is that adding a new line feels like progress while fixing the existing one feels like admitting something is broken.

Keep the thread

I write up what I learn from these conversations. Leave an email if you want it.

Part 05THE COMPANY


Chapter 26The bottleneck is always you

Every business that stalls has a bottleneck, and in a business under about ten people the bottleneck is almost always the founder. Not because the founder is inadequate. Because the founder is the only node in the system that every process routes through, and a system with a single shared node has a throughput ceiling that no amount of effort raises.

The symptoms are recognisable and worth naming honestly, because the internal experience of being the bottleneck feels like being important rather than like being a constraint.

Work waits for you. Not urgent work, all work, and the waiting is invisible because the people waiting fill the time with something else. Revenue tracks your working hours with a lag of about a month. You cannot take a week off without either working through it or accepting that the month after is worse. Growth requires more of your hours, so growth and your life are in direct conflict. And you are doing several things every week that somebody less experienced could do adequately, which you know, and which you keep doing because explaining is slower than doing.

That last one is the trap, and it is a maths error. Explaining is slower than doing once. It is faster than doing on the fifth repetition, and after the twentieth it is not close. People stay in the trap because the cost of explaining is paid today and the saving arrives later, which is the same cognitive bias that makes people not exercise.

The way out has four stages and they must be done in order, because each one depends on the last.

Document. You cannot delegate what you cannot describe, and the act of describing reveals that half of what you do is undocumented judgement rather than process. Write the process down as you do it, not from memory afterwards, because memory smooths over exactly the exception handling that makes the process work. AI genuinely helps here in a way it does not help elsewhere: narrate what you are doing, have it transcribed and structured, then correct the result. The correction is fast and the blank page is what was stopping you.

Systematise. A documented process is a description. A system is a description with inputs, outputs, a quality standard and a way of knowing it failed. The addition that turns one into the other is the check. If you hand someone a process with no defined standard of done, you have not delegated the work, you have delegated the work and kept the anxiety.

Delegate, in ascending order of scope. First tasks, where you specify exactly what to do and check the output. Then processes, where you specify the outcome and the standard and check periodically. Then functions, where someone owns a whole area and reports on results. Then strategy, where someone owns a goal and decides the approach. Most founders try to jump from the first to the third and then conclude that delegation does not work, when what actually happened is that they skipped the rung where the person learned the judgement.

Lead. The final stage is that other people make decisions you would have made, and some of them are decisions you would have made differently, and the business is better because a decision was made in a day rather than waiting three weeks for you. Getting comfortable with a decision that is eighty percent as good as yours and arrives ten times faster is the actual transition, and it is emotional rather than procedural.

There is a specific version of this problem in AI businesses that is worth calling out, because it catches technical founders hard.

If you are the person who understands the model behaviour, the prompts, the evaluation and the failure modes, you become an unusually severe bottleneck, because the work genuinely does require judgement that nobody else has. The instinct is to conclude that this part cannot be delegated. It can, but the delegation runs through the evaluation suite rather than through explanation. Once quality is measured rather than judged, someone else can make changes safely, because the suite tells them whether they broke something. Without the suite, every change has to go through you forever, and you have designed yourself a life sentence.

That is the strongest business argument for the eval work in Part Two, and it is worth restating: evaluation is not a quality practice, it is a delegation mechanism.

Chapter 27Automation that survives contact with reality

I have built and watched a lot of automation, and the difference between the ones that last and the ones that get quietly abandoned comes down to a few principles that are learned expensively.

The first principle is that you never automate a broken process. Automating a bad process produces a bad process that runs faster and is now harder to change, because the badness is encoded in software rather than living in someone's habits. Before you automate, fix. Frequently the fix eliminates the need for the automation, which is the best possible outcome and one people resist because they wanted to build the automation.

The second is that you automate in order of pain multiplied by frequency, and you ignore everything else. A task that takes an hour once a quarter is not worth automating, no matter how annoying it is, and people automate it anyway because it is annoying. A task that takes four minutes forty times a day is worth an enormous amount of engineering and nobody notices it because each instance is small.

01IMPLEMENTATION0203040506
What to automate, and what to leave alone.

The third principle is that maintenance is the real cost and nobody budgets it. Every automation is a small piece of software with dependencies that change underneath it. An API version deprecates. A form field is renamed. A model updates. Each of those breaks something, silently, and the breakage is discovered when someone notices a downstream effect weeks later. Before you build, ask who will notice when this breaks and how. If the answer is nobody, you have built a future incident.

The fourth is that every automation needs a heartbeat. Not an alert when it fails, because a system that dies rarely announces it. A regular signal that it is alive and processing at the expected volume. Silence is the failure mode, and the only way to detect silence is to expect noise.

The fifth is about the human in the loop, and it is the principle I would defend hardest. For any action that is irreversible or externally visible, the right design is almost never full automation. It is preparation plus approval. The system does ninety-five percent of the work and presents a finished thing for a human to approve in one action. You keep nearly all of the time saving and you eliminate nearly all of the tail risk, and the tail risk in an automated system that sends things to the outside world is much larger than people model.

I hold this position from experience on both sides. Fully automated outbound communication is a category of system that can do real damage in a short time, because the failure is not one bad message, it is four thousand bad messages before anyone looks. Approval gates cost seconds and prevent that entirely.

The sixth is that automation should reduce the number of systems, not increase it. A common pattern is that automation gets added between existing tools to paper over the fact that they do not integrate, and after two years there are thirty small workflows nobody fully understands, several of which duplicate each other and at least one of which is doing something wrong that has been absorbed into normal operations. Periodically, audit. Turn things off and see who complains. Nobody complaining is your answer.

And the seventh, which is a scoping principle: automate the middle, not the ends. The beginning of most processes involves judgement about what to work on and the end involves judgement about whether it is good. The middle is transformation, and transformation is what machines are for. Systems that try to automate the judgement at either end tend to be the ones that produce confident nonsense, and systems that automate the middle tend to be the ones people keep.

Chapter 28The team that AI actually needs

The received wisdom is that AI lets you build a large business with a tiny team. That is partly true and the part that is false is expensive.

What is true: the ratio of output to headcount has genuinely shifted. Work that required a department can be done by a small group with good systems, and the shift is largest in production work, drafting, first-pass analysis, research, code, documentation, anything where a competent first version was previously the expensive part.

What is false: that the small team can be composed of the same roles, just fewer of them. The composition changes, and the roles that matter now are not the ones that mattered five years ago.

Here is what a small AI-native team actually needs.

Someone who owns quality. Not a QA function in the traditional sense. Someone whose job is the evaluation suite, who reviews samples, who investigates when a number moves, who says no to a change that regressed the hard cases. In a business built on probabilistic systems, this is the most important role and it is the one almost nobody staffs deliberately. It usually falls to whoever is least busy, which is the wrong selection criterion.

Someone with deep domain knowledge, who is not a technologist. The person who has spent fifteen years in the industry you serve and who can look at an output and say, immediately, that a real practitioner would never write that. This person catches things no metric catches, and their judgement is the thing you are eventually encoding into your evaluation rubrics.

Someone who can build. Not necessarily a career engineer, but someone who can wire systems together, debug an integration, read a log and write the glue. One capable builder with modern tools covers a surprising amount of ground.

And someone who owns the customer, whose entire job is that clients get value and say so. In a recurring revenue business this role pays for itself several times over and it is always hired too late, usually after the first cluster of cancellations makes the case that nobody could make in advance.

Notice what is not on the list. A large production team, because production is what got cheap. A prompt specialist, because that is a subskill not a job. A data scientist, unless you are genuinely training models, which most businesses in this space are not and should not be.

Two structural observations about managing this kind of team.

The first is that the review burden shifts and grows. When production was expensive, review was a small fraction of total effort. When production is nearly free, review becomes the dominant cost, and a team that has not reorganised around that fact will be drowning in unreviewed output while feeling extremely productive. Budget review time explicitly. It is now the work.

The second is about hiring for judgement over speed. The traditional signal in a hiring process is how quickly and well someone produces. That signal has decayed, because everyone produces quickly now. The signal that has appreciated is whether someone can look at a plausible output and identify what is wrong with it. In interviews, I would rather watch a candidate critique a flawed piece of work than watch them create one, and the difference in what you learn is stark.

One more note, on the ethics of this, because it is not an abstract question if you are actually running a company. AI changes what roles are needed, and a founder who pretends otherwise while quietly not backfilling is being dishonest with their team. The version I think is right is to be explicit: this is what the tools now do, this is what we need people for, here is the support to move from one to the other. People handle a stated change far better than an unstated one, and the unstated one destroys trust in a way that outlasts the transition.

Chapter 29Money, margin, and the numbers that decide

I have looked at a lot of AI companies from the board seat and from the investor seat, and I can tell you which numbers I look at first and why, because it is not the ones on the front page of most decks.

Revenue growth is the headline and it is the least informative number on the page in isolation, because it does not distinguish between a business acquiring customers who stay and a business acquiring customers who leave. Two companies with identical growth curves can be in completely different situations, and the difference shows up in the second number.

Retention is the number that decides. Specifically, how long a customer stays and whether they are worth more over time. This determines everything: how much you can afford to spend acquiring, whether growth compounds or merely replaces, and whether the business has any terminal value.

The arithmetic is worth doing explicitly because it is so unforgiving. If you want a given level of recurring revenue and your customers stay a short time, you need to acquire a large number of new customers every month, forever, just to stand still. If they stay a long time, you need a small fraction of that. The required acquisition rate is inversely proportional to the lifetime, which means a change in retention has a leveraged effect on how hard your business is to run. Doubling retention does not make acquisition twice as easy, it makes the entire business a different business.

REQUIREMENT HALVES2461224
Why retention decides everything about how hard the business is.

The reason I emphasise this is that when a recurring business is struggling, the instinct is always to fix the top of the funnel, because that is the visible symptom. More leads, more meetings, more pipeline. And it almost never works, because if customers leave at the rate they are leaving, the funnel has to run faster and faster to stay in place, and the funnel has a maximum speed. Retention is the leveraged variable and it is the one people address last, because fixing it means confronting the possibility that the delivery is not good enough, which is a harder conversation than a marketing conversation.

The third number is gross margin per unit of work, tracked over time, which I covered in Part Three and which deserves repeating here because of a specific failure. In AI businesses, margin can degrade without any decision being made. A customer changes how they use the product, a prompt gets longer, a retrieval step gets more generous, and the cost per unit drifts upward while revenue stays flat. Nobody notices until the quarter closes. Put it on the dashboard.

The fourth is cash, and I will be short about it because it is not specific to this field. Know the date. Not the runway in months, which is a soft number that people round in their favour. The date on which the account reaches zero on current behaviour. Every founder should be able to say it without looking, and most cannot, and the ones who cannot make different decisions than the ones who can.

A few pricing notes that come out of all of this.

Charge for onboarding. It filters, it funds the most expensive part of your delivery, and it materially improves retention because a client who has invested at the start behaves differently than one who has not.

Raise prices on new customers before you think you should. The information that tells you your price is too low arrives late and quietly, in the form of everyone saying yes immediately. If nobody is pushing back on price, you are leaving margin on the table and, worse, you are attracting a customer segment that shops on price and churns on price.

Do not discount to close. Discount to change the deal, meaning give something up in return: a shorter scope, a longer commitment, a case study, a reference. A discount given for nothing teaches the client that your number was never real, and that lesson persists into every future conversation.

And meter what is variable. If your costs move with usage, your price must have a component that moves with usage, or you have written someone an option at your expense.

Still reading?

Then this is probably live for you right now. Fifteen minutes usually settles it.

Book 15 minutes

Chapter 30The operating cadence

A short chapter on rhythm, because strategy without a cadence is a document, and I have written enough documents to know the difference.

Every business I have seen run well has a small number of recurring reviews that happen whether or not anyone feels like it, and each one looks at a different time horizon. The specific frequencies matter less than the fact that they are fixed and that they do not get cancelled when things are busy, which is precisely when they matter most.

Daily, look at whether the machines are alive. Volume processed, errors, escalations, cost. Five minutes. The point is not analysis, it is detecting silence, because a system that has quietly stopped is the failure mode you will otherwise find out about from a client.

Weekly, look at the work. What went out, what got corrected, what a sample of outputs actually looked like when a competent person read them properly. This is where quality drift is caught, and it needs a human eye on real output, not a dashboard. An hour, with someone who knows the domain.

Weekly, also look at the pipeline and the commitments. What did we promise, to whom, by when. Most damage to client relationships is not caused by failure, it is caused by silence on something that was promised and then quietly not done.

Monthly, look at the numbers that decide: retention, gross margin per unit, escalation rate, cash date. Four numbers, tracked as a series rather than a snapshot, because the direction carries the information and the level does not.

Quarterly, look at the position. Is the buyer segment still right, is the price still right, what should we stop doing. This is also the window for the technology review from the previous chapter, and putting them together is deliberate: the question of what the field can now do is only useful next to the question of what our buyers need.

Annually, look at the asset. What exists now that did not exist a year ago and that would still have value if you stopped working tomorrow. This is the only review that measures the thing this book is actually about, and it is the one nobody does, because the answer is sometimes uncomfortable.

Two notes on making this stick.

Write down what you decided, not just what you observed. A review that produces observations produces nothing. A review that produces a written decision with an owner and a date produces change, and the writing takes four minutes.

And keep the cadence when it is boring. The reviews are valuable precisely in the periods when nothing appears to be happening, because that is when drift accumulates. The temptation to skip them during a busy quarter is the same temptation that causes the problem the review would have caught.

That is the whole system. It is not sophisticated and it does not need to be. The thing that separates operators from people with good ideas is almost never the quality of the plan. It is whether anything checks, on a schedule, that the plan is being executed and still makes sense.

Chapter 31Risk, and the things that actually end companies

I want to close this part on risk, and I want to be specific rather than gestural, because the risks in AI businesses are concrete and most of them are manageable if you look at them before they arrive.

Data handling is the first and the one that ends client relationships. When you process a client's data through a third-party model provider, you have made a decision on their behalf, and you need to be able to explain it: which provider, whether the data is retained, whether it is used for training, where it is processed geographically, and what your contract with them says. Most providers offer terms that are acceptable for business use, and most builders have never read them. Read them. Then write down your own data handling policy, in plain language, before a client's procurement department asks, because being asked and not having one is a bad position that is entirely avoidable.

Prompt injection is the second, and it is the one that is systematically underestimated because it does not look like a traditional vulnerability. The rule is simple to state: any content from outside your trust boundary that enters a model's context is potentially an instruction. If that model also holds tools with real power, you have built a path from a stranger to your systems. The mitigations are architectural and I covered them in Chapter 10. What I will add here is the governance version: keep an explicit inventory of every place untrusted content enters your systems and every tool that can cause an external effect, and make sure no path connects the two without validation and, for anything irreversible, a human.

Third is dependency. Your business probably rests on one or two model providers, and that is an acceptable risk that you should nonetheless price. What happens if pricing changes materially, if a model you rely on is deprecated, if a rate limit is imposed, if service is unavailable for a day. The mitigation is not to avoid dependency, it is to keep your model layer swappable, to have run your evaluation suite against at least one alternative so you know what would break, and to have a graceful degradation path rather than a hard failure.

Fourth is the output itself. If your system produces something a client acts on, you have a liability question, and the answer should be in your contract rather than discovered afterwards. Describe what the system does, describe its limitations, state clearly where human review is required, and do not let a salesperson describe it as more autonomous than it is. Most disputes in this field are not about a system failing, they are about a gap between what was promised and what was delivered, and that gap is created in the sales conversation.

Fifth is intellectual property, which is genuinely unsettled and which you should therefore handle conservatively. Ownership of AI-generated output varies by jurisdiction and is evolving. Training data provenance is contested. If your business depends on owning something a model produced, get proper advice rather than an opinion from the internet, and structure so that your value sits in things whose ownership is not in question: your data, your process, your integrations, your relationships.

Sixth is regulation, which is arriving unevenly and faster in some sectors than others. If you operate in healthcare, financial services, employment, education or anything touching consumer credit, assume rules apply and find out which. The cost of compliance discovered early is a design constraint. Discovered late, it is a rebuild.

And seventh, the one nobody puts on a risk register: quiet quality drift. This is not a dramatic failure, it is the slow erosion that happens when a system degrades a few percent at a time and everybody is too busy to look. It is the most likely thing to actually damage your business, and the entire defence is the measurement discipline from Chapter 9. A company that measures finds out in a week. A company that does not finds out from a client, and by then it is not a quality conversation, it is a trust conversation.

Chapter 32When it goes wrong

It will go wrong. Not might. Will. Probabilistic systems in contact with a messy world produce bad outputs, and if you operate long enough one of them will reach a customer and cause a real consequence. How you handle that hour determines whether you keep the account.

I want to give this its own chapter because most people building in this field have never thought about it until it happens, and the difference between a prepared response and an improvised one is enormous.

Prepare three things in advance.

The first is the ability to answer what happened. This requires the logging from Chapter 12: every prompt, every retrieved document, every tool call, every version of every component, stored and queryable. When a client says the system told them the wrong thing on Tuesday, you want to be able to pull the exact trace within minutes and see the actual cause, rather than speculating. Teams without this spend days reconstructing and their client watches them fail to answer a simple question, which is worse than the original error.

The second is a kill switch. A way to stop the system, or to route everything to human review, without a deployment. It should be a configuration change that takes seconds and that more than one person can perform. The moment you need it, you need it immediately, and if the only way to stop the system is a code change reviewed by someone who is asleep, you will spend hours watching it make things worse.

The third is a written incident process, which does not need to be elaborate. Who is notified. Who owns communicating with the client. What gets checked. What gets written down. Having it on paper means that during the incident people execute rather than debate.

Then there is the handling itself, and here I would borrow entirely from operational engineering practice because it is correct and hard-won.

Tell the client before they tell you, wherever it is possible. The reputational difference between discovering a problem yourself and having it reported to you is very large and entirely disproportionate to the technical difference, which is none. A client who hears from you first concludes that you are watching. A client who has to report it concludes that you are not.

Say what happened plainly, including the part that reflects badly on you. Every instinct pulls toward minimising, and every attempt at minimising is detected. What buys back trust is a specific, unflattering, accurate account, because it demonstrates that you understand the failure, and understanding the failure is the only credible basis for a claim that it will not recur.

Separate the immediate fix from the systemic fix and communicate both. The immediate fix stops the bleeding. The systemic fix is the change that prevents the class of error, and it is almost always a new case in the evaluation suite plus a check that would have caught it. Naming the systemic fix is what turns an incident from a reason to leave into evidence that you operate seriously.

And do not blame the model. It is technically accurate and it lands as an excuse, because from the client's position you chose the model, you designed the system, and you decided what checks it ran through. The failure is yours. That posture is uncomfortable and it is the one that preserves the relationship.

There is a broader point here about how quality actually gets built in this field. Every mature engineering culture I have worked in treats failures as inputs rather than as embarrassments, and runs the review without hunting for someone to blame. This matters more with AI systems than with ordinary software, because the failures are more frequent, more subtle, and more likely to be produced by an interaction between components rather than by anyone's mistake. A team that punishes people for surfacing bad outputs will simply stop hearing about bad outputs, and will then discover them at scale.

The metric I would watch, over a year, is not the error rate. It is whether errors are being found internally or externally. A system where most problems are caught by your own checks is healthy even at a moderate error rate. A system where most problems are reported by customers is unhealthy at any error rate, because it means the sensor is not working, and everything downstream of a broken sensor is guesswork.

Part 06THE LONG GAME


Chapter 33The convergence

I want to spend a chapter on where I think this goes, and I want to be careful about the difference between a thesis and a prediction. What follows is a thesis: a claim about architecture that I think is structurally sound, that I have organised my working life around, and that could still be wrong on timing by a decade.

Three curves are moving at once, and almost everyone is watching only one of them.

The first is machine intelligence, which is the curve everybody is watching. Capability per dollar has been improving at a rate that has no recent precedent, and the improvement is not only in the models but in the scaffolding around them: tools, memory, evaluation, orchestration. The practical effect is that the set of tasks a machine can complete without a human in the loop keeps widening.

The second is embodiment. Robotics has been a slow field for a long time, and I say that as someone who spent eight years inside it. The reason was never mechanical. It was that the perception and control problems required an amount of hand-engineering per task that made general-purpose machines uneconomic. What has changed is that the same architectural advances driving language models are being applied to perception and control, which means the per-task engineering burden is falling. I sit on the board of a public company in this sector, so I will be careful and general: the direction of travel is toward machines that are programmed by demonstration and instruction rather than by explicit specification, and that changes the economics of the entire category.

The third curve is settlement, and this is the one almost nobody connects to the other two.

Here is the argument. If machines are going to act autonomously in the world, at some point they need to pay for things. An autonomous system that needs an inference call, a data feed, a compute allocation, a charging slot or a physical service has to acquire that resource, and acquiring resources means transacting. Today, every transaction a machine makes is actually a transaction made by a company on that machine's behalf, settled through rails designed for humans in the twentieth century: account-based, permissioned, reversible, with minimums that make small payments uneconomic, and with settlement times measured in days.

Those rails do not fit. Not because of ideology, because of specification. A machine economy would generate payments that are very small, very frequent, machine-to-machine, and continuous. That is a different set of requirements than a card network was designed against, and the mismatch is structural rather than a matter of upgrading.

Bitcoin is the settlement layer I think this eventually runs on, and my reasoning is engineering reasoning rather than enthusiasm. It is permissionless, which matters because a machine cannot open a bank account. It is final, which matters because reversibility requires a human dispute process. It is neutral, which matters because no participant can be de-platformed by a counterparty. And it has payment layers built on top of it capable of small, fast transfers, which is the specification the machine case actually demands.

I hold that view having spent years building financial infrastructure in a market where the permissioned rails were withdrawn from an entire industry by a single decision from a central bank. I have seen what it costs when your settlement depends on permission. That experience is not theoretical for me, and it is the reason I weight neutrality more heavily than most people do.

01S MARKED02030405
Three curves, one architecture.

What does this mean for someone building a business now, which is the only question that matters practically.

It means the least crowded opportunities sit at the intersections rather than in the middle of any one field. Pure AI application companies are the most crowded market I have ever seen. AI applied to physical operations is meaningfully less crowded, because it requires you to understand two hard domains. Machine-native payments and identity is barely populated at all, because it requires three.

It means proprietary data from the physical world is going to be disproportionately valuable, because it is the input that cannot be scraped. Text is exhausted. Sensor data, operational data, and data from real processes in real facilities is not, and whoever holds it holds something that the next model generation cannot make redundant.

And it means the boring infrastructure layers will matter more than they appear to now. Identity for non-human actors. Authorisation and spending limits for autonomous systems. Audit trails for machine decisions. Insurance and liability frameworks. All of these are unglamorous, all of them are required, and most of them do not have a credible provider yet.

Let me also state the ways this thesis could be wrong, because a thesis you cannot falsify is not a thesis.

Timing could be much slower than I expect. Physical deployment cycles are measured in years, industrial buyers are conservative for good reasons, and regulation of autonomous physical systems will be slower still. The intelligence curve can run far ahead of the embodiment curve for a long time.

The settlement layer could end up being something else. Central bank digital currencies, incumbent networks building machine-appropriate rails, or a stablecoin layer that captures the volume for reasons of convenience. I do not think that is the durable answer for the reasons above, but I would be arguing from principle against a large amount of installed distribution.

And the economics of general-purpose machines could remain worse than special-purpose ones for longer than expected, which has been the historical pattern in robotics for forty years and which I would be foolish to dismiss.

What I am confident about is narrower and, I think, more useful. Machines will do more autonomous work each year. Autonomous work requires acquiring resources. Acquiring resources requires transacting. That chain does not depend on any particular timeline, and it means the intersection is worth understanding now even if the mass market arrives later than anyone wants.

Chapter 34What survives the next release

The most common anxiety I hear from people building in this field is a version of: what if the next model release makes what I built pointless.

It is a reasonable fear and it has a precise answer, which is that you can sort everything you own into two piles and the sorting is not difficult.

The first pile is things that get better when the models get better. The second is things that get worse. The entire strategic question is which pile your effort is going into.

In the pile that gets worse: any capability that exists because the current models are limited. Elaborate workarounds for context limits. Multi-step chains that decompose a task the next model does in one call. Tuning that compensates for a weakness that is about to be fixed. Products whose entire value is making a model easier to prompt. These are all real work and they all have a shelf life, and the tell is that a model improvement is bad news for you.

In the pile that gets better: your evaluation suite, which becomes more valuable as models change because it is how you know whether the change helped. Your proprietary data, which a better model extracts more value from. Your integrations, which are unaffected by model quality and which get more useful as the thing they connect gets smarter. Your distribution and your reputation, which are entirely orthogonal. And your domain judgement, which is the ability to tell a good answer from a plausible one, and which becomes more valuable as the plausible answers get more convincing.

Notice that the second pile is almost entirely things that are not the model, and that most of them are boring. That is the whole lesson. In a field where the exciting layer is rented from three companies and improves for everyone simultaneously, the durable advantage is necessarily somewhere unfashionable.

There is one more item for the second pile that deserves its own treatment, because it is the least discussed and possibly the most important.

Taste.

I mean something specific by that. Taste is the ability to look at four competent outputs and know which one is right for this situation, and to know why, and to be able to say what is missing from the other three. It is not aesthetic preference. It is compressed judgement, built from having seen a great many attempts and their consequences.

Taste has become dramatically more valuable and the reason is arithmetic. When producing an option was expensive, you had two options and choosing was easy. Now you have forty options in ninety seconds, all of them fluent, and the bottleneck has moved entirely into selection. The person who can choose well is now the constraint on quality, and there is no way to acquire that except by doing the work over a long period and paying attention to outcomes.

This is also, incidentally, the honest answer to the question of what humans are for in this. Not production. Not the first draft. Judgement about what matters, taste about what is right, and responsibility for the consequences, which is a thing a model cannot hold because it cannot be accountable.

The practical version of all this is a filter I apply to my own work now. Before spending a week on something, I ask which pile it goes into. If a better model would make this week's work irrelevant, I want a very good reason to be doing it, and usually the reason is that it unblocks something in the other pile. That filter has saved me more time than any tool I have adopted.

One conversation beats ten chapters

Book the time, bring the specifics, and we will work out what actually moves for you.

Book 15 minutes

Chapter 35Working with machines, not beside them

A short chapter on personal practice, because how you actually operate day to day is where most of the compounding happens and it rarely gets written down.

The shift that matters is from using AI as a better search engine to running it as a team you delegate to. The difference is not the tool, it is the posture, and there are four habits that mark the transition.

The first is that you brief instead of ask. A question gets a generic answer, because a generic answer is the correct response to an underspecified question. A brief gets useful work: here is the situation, here are the constraints, here is what I have tried, here is what the output needs to look like, here is how I will judge it. If you would not send it to a competent contractor and expect a usable result, do not send it to a model and expect one.

The second is that you keep context rather than rebuilding it. Most people start every interaction from nothing, which means they spend a large share of their effort re-explaining their business. The fix is to maintain durable context: files describing your business, your standards, your customers, your voice, your constraints, kept as documents you can attach or that your tooling loads automatically. This is a few hours of one-time work and it changes the quality of everything downstream, because the model is now working with the situation instead of the average of all situations.

The third is that you verify structurally rather than by feel. Decide, before you look at the output, what would make it wrong. Then check that. Reading an output and thinking it seems good is not verification, it is fluency detection, and fluency is the one thing guaranteed to be present.

The fourth is that you run things in parallel. This is the habit that changes throughput most and it is the least natural, because it fights the sequential way we are used to working. If you have four independent pieces of research, start all four rather than doing them in order. If you want three approaches to a problem, generate all three and compare rather than iterating on one. The constraint on your day is no longer how fast work gets done, it is how many things you have set in motion, and most people are running one thing at a time out of habit.

A note on where I would spend the tokens, since attention and budget are both finite. Use the cheap tier liberally for anything you will check anyway. Use the expensive tier for the things that go out under your name and for problems where you genuinely do not know the answer, which is where the reasoning difference actually shows. And accept some waste. The cost of generating three approaches and discarding two is trivial compared to the cost of committing to the first idea, and optimising token spend at the expense of exploring options is a false economy that engineers are particularly prone to.

The habit I would leave you with is a weekly one. Once a week, look at what consumed your hours and ask which of it was production and which was judgement. Production is the part to systematise. Judgement is the part to protect. Businesses go wrong when the ratio drifts the wrong way and nobody is measuring it, and the drift is always toward production, because production feels like progress.

Chapter 36The first ninety days

Here is a plan. It assumes you are starting from a professional background and some domain knowledge, and no existing AI business. It does not promise a revenue figure, because I do not know your market, your network or your capacity, and anyone who gives you a number for those three unknowns is selling you something.

What it does promise is that at the end of ninety days you will have either a paying client and a documented process, or a clear and specific reason why not, which is worth roughly as much and which most people never obtain because they never run the experiment properly.

01CHOOSE ONEBUYER02HOLD TWENTY03BUILD ONEWORKING04PUBLISH TWOPIECES05DELIVER ATLEAST
Ninety days, and the gate at the end of each month.

Let me put content on that skeleton.

Days one to thirty are about position and proof, and the sequence matters.

Pick the buyer first, not the technology. One segment, narrow enough that you can name the job title and the industry, and chosen because you have some genuine advantage there: you worked in it, you know people in it, you understand its vocabulary. Advantage beats market size at this stage by a wide margin.

Then find the expensive problem by asking, not by assuming. Twenty conversations. Not sales calls. Conversations where you ask what takes their team the most time, what goes wrong most often, what they have tried, what it costs when it fails. Do not pitch. If you pitch, they become polite and you learn nothing. The goal is that five or more of them describe the same problem without prompting, which is your signal.

While those conversations are happening, build one thing. Narrow, working, solving a small version of the problem you are hearing about. Not a product, a demonstration that the mechanism works. Building it teaches you what is actually hard, which is information you cannot get any other way, and it means that when the right conversation happens you have something to show rather than something to describe.

And start publishing, one piece a week, about mechanism. This will feel pointless for the entire ninety days. Do it anyway. The return arrives in month seven and it does not arrive at all if you did not start in month one.

Days thirty-one to sixty are about the first transaction, and the most important thing here is that the first sale is small and fixed.

Go back to the warmest of your twenty conversations and offer a bounded, paid piece of work. An assessment, an audit, a pilot on a narrow slice. Priced low enough to be an easy yes and high enough that it is a real commercial decision, because a free engagement teaches you nothing about whether anyone will pay. Money changing hands is the only reliable signal in the whole process.

Then deliver it properly, and while delivering, do two things people skip. Build your evaluation suite out of the real examples this engagement gives you, because now you have actual data rather than imagined data. And document the process as you go, in enough detail that someone else could follow it, because you will never reconstruct it accurately afterwards.

Days sixty-one to ninety are about repetition, which is the least exciting phase and the one that determines whether you have a business or a job.

Deliver two more engagements using the documented process. Measure how many hours each takes. If the third takes as long as the first, your process is not real and you should find out now. If it takes meaningfully less, you have found the repeating core, which is the thing that eventually becomes a productised service.

Raise your price on the third one. Not enormously. Enough to test where the ceiling is, because the only way to find it is to hit it, and the information from a lost deal at a higher price is more valuable than the revenue from a won deal at a low one.

A few notes on how this actually goes, because the plan is clean and reality is not.

The twenty conversations will be the hardest part and the part you will want to shorten. Do not. Every hour you cut here you pay back tenfold later, building something for a problem that turned out not to be expensive.

You will be tempted to widen your position when the first three conversations do not land. Resist for at least fifteen conversations. Three is not a sample, it is an anecdote, and widening in response to early friction is the reflex that produces the generalists we discussed in Chapter 4.

Something will take three times longer than planned. It will probably be an integration with a client system, because it always is. Build that assumption into your scoping from the first engagement.

And you may reach day ninety with no client. If that happens, the useful question is which gate you failed. No shared problem across twenty conversations means the segment was wrong and you should change it, which is a cheap lesson. A shared problem but nobody will pay means the problem is annoying rather than expensive, which is the most common finding and the most valuable one. Payment but delivery that did not repeat means you have a consulting practice rather than a business, which is fine, and the fix is in Chapter 16.

Failing a specific gate is a result. Vague discouragement is not, and the entire reason for structuring the ninety days this way is to convert one into the other.

Chapter 37After ninety days

The plan above gets you to a repeatable engagement. What happens next is less scripted, but the sequence that works has a shape and I will describe it briefly, because knowing the shape prevents a few common wrong turns.

Months four to six are for the operation. You have a process that repeats, and now you make it not depend on you: quality checks that run without your eye, a delivery person who is not you, onboarding that has a defined proof point in the first thirty days. This period feels like a step backward because revenue may flatten while you build infrastructure. It is the highest-return period in the whole arc and the one most people skip in favour of chasing more clients, which is how a delivery business becomes a trap.

Months seven to twelve are for the position. By now you have real evidence, and evidence changes what you can charge and who will talk to you. This is when the publishing you have been doing since day one starts to produce inbound, and when the correct move is to raise prices, narrow the client profile, and start turning down work that does not fit. Turning down work is the specific behaviour that marks the transition from a practice that survives to one that compounds.

Year two is for the asset. Whatever it is: the productised service with documented process and real retention, the software built from a validated pattern, the data set that only exists because your clients used your system. The thing that has value independent of your continued attention.

And the discipline that runs through all of it, which I will say one final time because it is the thing I would most want to have understood earlier: measure what you are building, not what you are doing. The two feel identical from the inside and they diverge completely over five years.

Chapter 38Staying current without drowning

The field moves fast enough that keeping up is a real cost, and most people manage it badly in one of two directions. Either they consume constantly and build nothing, or they tune out entirely and wake up eighteen months behind on something that mattered.

Here is the allocation I use, and I offer it as one working solution rather than the answer.

Most of the time, ignore the news. The overwhelming majority of what is published about AI on any given day is either an announcement that will not affect you for a year, a benchmark result that will be superseded, or someone with an opinion. Consuming it feels like staying informed and functions as procrastination with a good alibi.

Instead, run a fixed review on a schedule. Once a quarter, spend a day. Read the actual technical documentation of the providers you depend on, not the coverage of it, because the documentation contains the specifics that change what you can build and the coverage contains adjectives. Run your evaluation suite against whatever is new. Try two or three tools that people you respect have actually adopted rather than merely mentioned. Then decide what changes and freeze it again.

Between those days, keep three narrow inputs open.

Changelogs and deprecation notices for the services you depend on, which are the only announcements that can break your business and are the ones nobody reads.

A small number of people who build rather than comment. The signal-to-noise difference between someone shipping systems and someone summarising other people's work is enormous, and you can identify the first group by whether they ever describe something that did not work.

And direct conversation with people in your target industry, which is the highest-value input by a wide margin and the one that appears on no reading list. What you actually need to know is not what the frontier can do, it is what your buyers are trying to do and failing at, and there is no publication that carries that.

On depth, one recommendation. Understanding the mechanism of these systems at the level of Chapter 5, tokens, next-token prediction, context, sampling, is worth acquiring properly once. It does not go stale, because it is the architecture rather than the release, and it inoculates you against a large fraction of the nonsense. Beyond that level, additional theoretical depth has sharply diminishing returns unless you are training models, and most people should get their next increment of understanding by building something rather than by reading.

On reading more broadly, my honest position is that the durable books in this space are not about AI. They are about how technologies diffuse, how businesses actually work, and how markets adopt things. The strategy classics on disruption, on crossing from early adopters to a mainstream market, on building a business through iteration rather than planning, all still apply, and they apply with more force now because the technical bottleneck has moved and the commercial bottlenecks are the ones they describe. Reading a book about how a previous general-purpose technology took thirty years to reorganise industry will tell you more useful things about the next five years than another survey of model capabilities.

And a note on conferences and communities, since the source of most people's information is social. The value of a community is almost entirely in whether its members are building things you could not build, and almost none of it is in the volume of discussion. A small group of people doing serious work, where you can ask a specific question and get a specific answer from someone who has hit that exact wall, is worth more than any number of large forums. Those groups are usually private, they are joined by contributing something first, and the entry price is having done work worth discussing.

Which brings the whole thing back around. The way to stay current, in the end, is to be building, because a person with a live system and real users encounters the meaningful changes as problems rather than as headlines, and problems are how anyone has ever actually learned anything.

Chapter 39What to do Monday

I have given you a lot. Here is the compression.

The event that happened is that producing a competent first draft of knowledge work became nearly free. Everything else follows. Scarcity moved to judgement, to verification, and to distribution, and those are the three things worth building.

The model is a commodity and everybody rents it from the same few places. The context you put in front of it, the evaluation that tells you whether it worked, the data your customers generate, the integrations that make you hard to remove, and the trust that makes people call you first: those are yours. Build there.

Position specifically enough that a stranger can repeat your sentence without you in the room, because referability is the actual asset and everything else about niching falls out of it.

Price on the outcome, in flat fees, with no promises about results you do not control. Let the client own every relationship the work touches. That is both the honest structure and the one that survives.

Measure. If you take one technical thing from this book, make it the golden set and the evaluation loop, because it is simultaneously your quality system, your delegation mechanism, your sales differentiator and the asset that appreciates when the models improve.

Sequence: consulting to learn the market, productised service to escape your hours, software only if the pattern generalises and you can fund the road. Publish mechanism underneath all of it, every week, for years.

And on Monday, specifically, do this. Pick one buyer segment you have a genuine advantage in. Write down the one sentence. Then send five messages to people in that segment asking for twenty minutes to hear about their work, offering nothing and selling nothing.

That is the whole first step. Not a tool. Not a course. Not more reading, including this. Five messages, and then the conversations that follow, and then a small thing built for a real problem you actually heard someone describe.

The people who build something durable in this period will not be the ones who understood the technology first. They will be the ones who found a real problem, proved they could solve it, and could show a stranger exactly how their system behaves when it is wrong. That has always been the job. The machines just made the production part cheap.

Where are you stuck?

Fifteen minutes, no deck, no pitch. Tell me what you are building and I will tell you what I would do next.

Book 15 minutes

Chapter 40About this book, and why it is free

A short closing note, and then I will get out of the way.

I am Sunny Ray. I am an electrical engineer by training. I spent eight years at Quanser building control systems, mechatronics and haptic devices, with products used in research and teaching labs including MIT, Stanford and Georgia Tech. I have been in bitcoin since 2011. I co-founded India's first bitcoin exchange, which grew past two and a half million users, and when the central bank cut the industry off from the banking system we fought it to the Supreme Court of India and won.

Today I run Sunny Ray Holding Inc, a family office and venture studio, from Toronto. I am a director of Humanoid Global Holdings, which trades on the CSE under ROBO. I host The Sunny Ray Show, where I spend most weeks in long conversations with people building at the intersection of bitcoin, AI and robotics. I back founders early, where capital and conviction are both scarce.

That is the whole biography and it is relevant only because it explains the bias in this book. I have shipped physical systems that fail loudly, and financial infrastructure that fails expensively, and I now evaluate AI companies from a seat where someone is asking me for money. All three of those experiences push in the same direction: toward caring about what a system does under load and at the edges, and away from caring about how it demonstrates.

This book is free and there is nothing to buy at the end of it. There is no course, no upsell, no gated bonus, no community with a monthly fee. I have deliberately removed every number I could not personally stand behind, which meant cutting a great deal of material that would have made the book more exciting and less true. If a claim in here has a number attached, it is either arithmetic you can check yourself or something I have actually seen.

The only thing I will ask for is a conversation, and only if it is useful to you. If you are building something at this intersection, or trying to work out whether the thing in front of you is worth the next year of your life, book time with me. I am genuinely more interested in the second conversation than the first, which is to say I would rather hear what you built after reading this than talk about the book.

One last thing, which is the sentence I would keep if I could keep only one.

The machines made production cheap. They did not make judgement cheap, they did not make trust cheap, and they did not make it any easier to find a problem worth solving. Those three things were always the job. They are just now, finally, the only job, and that is better news than it sounds like.

This is the part people get wrong

If you want a second pair of eyes on your version of it, book fifteen minutes and bring the messy version.

Book 15 minutes
That is the whole book

If any of it landed, the fastest next step is a conversation. Fifteen minutes, bring the specifics.

Book 15 minutes