The Lego Model: How Tool Calling Changes the Way We Build AI Applications

“What should I read next?” seems like a simple question. But answering it well requires knowing what the user is currently reading, what they’ve finished recently, and what’s already on their list. When I built QuietReads, a book tracking app with an AI assistant, I faced a choice: pre-load all this context on every request (expensive), build complex routing logic to fetch the right data (brittle), or find a different approach entirely.

In those post, I talk about why I chose this option, the architectural implications, and the tradeoffs involved. 

Conversational Interfaces Are Not Deterministic

A lot of mid-career technologists like me still think deterministically. We reach for decision trees, map out every probable permutation, and write test cases to cover each branch. This works well for forms and structured APIs where you control the inputs.

But QuietReads has a chat interface. Users might ask:

  • “What should I read next?”
  • “I’m in the mood for something like the last book I finished, but shorter”
  • “What were my thoughts on that dystopian novel from last month?”
Asking the QuietReads assistant to recommend some books

Each query requires different context. The first needs the user’s want-to-read list. The second needs their recently finished books plus some understanding of “shorter.” The third requires searching through their notes. I couldn’t predict which context any given question would need, and I didn’t want to fetch everything every time.

The traditional approach would be routing logic. For just the first query, you might write something like:

def get_context_for_recommendation(message, user_id):
    context = {}

    if contains_recommendation_intent(message):
        context['want_to_read'] = get_want_to_read_books(user_id)
        context['recently_finished'] = get_recently_finished(user_id)

        if mentions_specific_book(message):
            book = extract_book_reference(message)
            context['book_details'] = get_book_details(book)

    # ... and this continues for every intent type
    return context

This gets unwieldy fast. Each new question type requires new routing rules. The intent detection functions themselves need maintenance. And you’re constantly guessing what context the model will need.

LLMs allow us to use a different mental model. 

Think of LLMs as expert Lego assemblers. You provide a curated set of bricks (tools), an instruction manual (your system prompt), and let the assembler determine which bricks to use and in what order. You don’t hand them every brick in existence. You give them the right pieces for the task and clear guidance on when to use each one.

Tools as Building Blocks

In QuietReads, I define “context tools” that let the AI retrieve user data as needed:

CONTEXT_TOOLS = [
    {
        "name": "get_user_profile",
        "description": (
            "Get the user's name and reading preferences. Use this when you need to "
            "personalize your response or discuss their reading interests."
        ),
        "input_schema": {"type": "object", "properties": {}}
    },
    {
        "name": "get_want_to_read",
        "description": (
            "Get books on the user's want-to-read list. IMPORTANT: Always call this "
            "BEFORE recommending any books to avoid suggesting books they already have."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "limit": {"type": "integer", "description": "Max books to return (default: 10)"}
            }
        }
    }
]


Each tool has a clear description of when to use it. The model reads these descriptions and decides which tools to call based on the user’s question.

The Agentic Loop

When the model decides to use a tool, we handle that request, execute the tool, and feed the results back. This creates a loop (simplified code below):

async def execute(self, system, messages, tools, tool_handlers, max_tokens=2048):
    response = self.client.messages.create(
        model=self.model,
        max_tokens=max_tokens,
        system=system,
        messages=messages,
        tools=tools
    )

    while response.stop_reason == "tool_use":
        tool_results = []

        for block in response.content:
            if block.type == "tool_use":
                tool_name = block.name
                tool_input = block.input

                # Execute the tool and capture result
                result = await tool_handlers[tool_name](tool_input)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": result
                })

        # Feed results back and get next response
        messages.append({"role": "assistant", "content": response.content})
        messages.append({"role": "user", "content": tool_results})
        response = self.client.messages.create(...)

    return response

The model might call multiple tools, or call the same tool with different parameters, or decide it has enough context after the first call. 

The loop continues until the model has everything it needs to answer. In practice, you should also enforce a maximum iteration count as a guardrail against runaway loops or unexpectedly expensive queries. When the model requests multiple tools in a single response, you can execute them in parallel for performance gains.

This code handles the happy path. Production implementations need additional safeguards: error handling when tools fail or timeout, validation of tool inputs before execution, and graceful handling when the model hallucinates a tool name that doesn’t exist.

Client and Server-side Tools

QuietReads uses two categories of tools. Client-side tools are functions I implement: when the model calls get_want_to_read, my code queries the database and returns formatted results.

Server-side tools are capabilities the AI provider offers (like the web_search tool used below). 

I enable web search so the assistant can look up recent book releases or author news. Anthropic’s infrastructure handles the search; I just control when and how it’s available:

tools:
  web_search:
    type: "web_search_20250305"
    name: "web_search"
    server_side: true
    enabled: true
    config:
      max_uses: 5  # Limit searches per request

With some prompt engineering, the model can integrate custom database queries with real-time web searches, producing responses that feel coherent to the user. 

The system prompt guides how the model synthesizes information from different sources, when to cite web results versus personal reading history, and how to maintain a consistent voice across tool-augmented responses.

Architectural Implications

It is important to recognize where determinism matters and where it doesn’t. Each tool is a testable piece of code. I can unit test get_want_to_read in isolation, verify it returns the right data, and trust it to behave consistently. What I can’t fully predict is which tools the model will call or in what order. In my work, I have found that even cheaper models like the Haiku family of models do a decent job at tool use.

This separation has practical implications. Tool descriptions are instructions the model uses to decide when to call each tool. Writing clear, specific descriptions is as important as the implementation itself. Instead of pre-loading everything a user might need, I provide minimal context upfront and let the model request more, keeping initial requests fast and reducing token costs.

And while the model chooses its tools, I still control the boundaries. QuietReads runs input guardrails before messages reach the assistant and can validate outputs before returning them to the user.

The Tradeoffs

Tool calling introduces real costs that you should weigh against your specific requirements.

  • Predictability. With static context, you know exactly what data the model sees on every request. With tool calling, the model decides what to retrieve. This makes cost and performance harder to predict. A simple question might resolve in one API call; a complex one might trigger four tool calls and five round-trips.
  • Prompt caching. Static context can benefit significantly from prompt caching, where repeated system prompts are stored and reused. Dynamic tool results change with each request, which can reduce or eliminate caching benefits. Depending on your usage patterns, this could meaningfully impact both latency and cost.
  • Quality assurance. Unit testing individual tools is straightforward, but testing the system end-to-end becomes harder. The model might call tools in unexpected combinations, or skip tools you expected it to use. Comprehensive evaluations become essential because tool calling adds non-determinism to the critical path. I’ll write more about evaluation strategies in a future post.
  • Refactoring risk. IDE tooling can automatically update function signatures across a codebase. Tool definitions live in JSON objects that describe behavior and parameters in natural language. If you change a tool’s behavior or modify its parameters, automated refactoring won’t catch the JSON definitions, and the mismatch may not surface until production. LLM-based coding agents like Claude Code handle this well, and adding tool-specific checks to code review agents helps catch these issues.
  • Latency. Each iteration of the agentic loop requires a round-trip to the API. For QuietReads, this is acceptable. For applications where response time is critical, the additional latency may be a dealbreaker.

That said, tool calling offers real advantages beyond flexibility. Token costs can decrease because the model only retrieves data it actually needs. Direct tool calls avoid the protocol overhead of intermediary layers like MCP servers. And tools create a clean separation of concerns: database migrations, API upgrades, or new data providers can happen without touching the prompt.

This pattern fits QuietReads: read-only tools, flexible latency, and context costs that exceed API overhead. Applications with side effects, strict latency, or predictable context needs may want different approaches.

Navigating a Mindset Shift

Building with tool calling requires accepting the risk of non-deterministic code execution. You cannot predict every code path. Instead of mapping out decision trees, you’re designing capabilities and constraints. You’re giving the model a well-stocked toolbox and clear guidance, then trusting it to assemble the right response.

You control what tools exist, what data they access, what the model knows about when to use them, and what guardrails prevent misuse. The model handles the dynamic orchestration that would otherwise require hundreds of lines of if/else chains.

For those of us who’ve spent years thinking in flowcharts, this shift takes practice. But once it clicks, you start asking different questions: not “what are all the paths a user might take?” but “what capabilities does the model need, and how do I describe when to use them?”

Building at the Speed of Thought

On building QuietReads, Claude Code, and the inversion in software economics

I have wanted to build QuietReads for two decades.

The idea is simple: a book tracking application that treats reading as a reflective practice rather than a social performance. Just you, your books, and an AI companion that remembers what you’ve read and can discuss it with you.

Every few years, I would sketch out the features, maybe prototype a database schema, and then abandon the project when the scope overwhelmed the time I could spare. The economics never worked. Building a full-stack application with authentication, third-party integrations, and a sophisticated AI layer would take months of focused effort. I had a day job. I had a family. QuietReads stayed in the drawer.

Then came Thanksgiving 2025. Anthropic rolled out $1000 in API credits for Max subscribers to use with Claude Code, their agentic coding tool. The Opus 4.5 model had just launched. I decided to try again.

Two weeks later, QuietReads was live.

QuietReads organizes books in a Library View

What Changed?

This wasn’t my first attempt at building QuietReads with AI tools. Last summer, I tried vibe-coding the same application using Replit, Cursor, and an early version of Claude Code running Sonnet 4. The results were miserable. The AI agents went off the rails constantly, making changes I didn’t ask for, getting stuck in loops, producing code that looked plausible but broke in subtle ways. I gave up after burning through credits and several frustrating weekends.

The difference with Opus 4.5 was stark. 

Where Sonnet 4 required constant hand-holding, Opus 4.5 understood what I was trying to build. It made architectural decisions that made sense. When I pointed it at a bug, it found the root cause rather than applying band-aids. It demonstrated genuine systems thinking: analyzing trade-offs, proposing multiple approaches with honest assessments of pros and cons, thinking through downstream implications.

Claude Code making suggestions on how to refactor the AI Assistant to make tool calling more consistent

The screenshot above shows Claude Code reasoning through three different implementation approaches for a feature, weighing simplicity against performance against pattern consistency. This is a discussion I would expect to have with a senior developer. The options made sense, and I was able to make an informed decision on how I wanted to structure a key component of the application.

I shipped a full-stack, mobile-responsive application with social login, Google Books integration, and an AI reading companion that maintains persistent memory of your reading history. The companion can engage in literary discussion, reference your previous notes, and connect themes across books you’ve read months apart. You can @mention it in any note or journal entry, and it responds with context. Everything flows into a unified timeline that weaves together notes, reading sessions, and AI conversations.

Eko, QuietReads’s default AI Assistant helps me understand Neil Postman’s work. Books in AI responses are interactive components, allowing me to add them straight to the library.

The Inversion

My experience with QuietReads is a small data point in a larger pattern. We are navigating an inversion in how software gets built.

For decades, every decision about software development has been constrained by the scarcity of labor. Whether to buy or build. Whether to refactor legacy code or start fresh. Whether to ship now or wait for more resources. 

The person has always been the limiting factor, the most expensive line item, the ever-present bottleneck. 

The language of our industry reflects this: we estimate in person-days, bill in person-hours, staff projects in person-months.

That bottleneck is fading faster than I expected.

Boris Cherny, the creator of Claude Code at Anthropic, recently revealed that he didn’t open an IDE for an entire month. Every line of code he shipped during that period (259 pull requests, 497 commits, 40,000 lines added, 38,000 removed) was written entirely by Claude Code powered by Opus 4.5.

“Software engineering is radically changing,” Cherny wrote, “and the hardest part even for early adopters and practitioners like us is to continue to re-adjust our expectations.”

I am not Boris Cherny. I don’t work at Anthropic. I don’t have access to internal builds or institutional knowledge. But my experience over the holidays rhymes with his. And I’m not alone. Developers across the industry spent the holiday break shipping projects that had languished for years, finally made tractable by the combination of Opus 4.5 and Claude Code’s improved agentic capabilities.

Lots of people had “Claude Code” moments over the holidays..

So What Does This Mean for People Like Us?

In March 2025, Anthropic CEO Dario Amodei told the Council on Foreign Relations that he expected AI to be writing 90% of code within three to six months, and essentially all code within twelve months. The industry-wide numbers haven’t hit those marks. But within certain teams and workflows, his timeline looks prescient. Cherny’s experience is evidence. So is mine.

Software engineering as we know it is going to change faster than many of us imagined. The fundamental assumptions that drive economic decisions around building software (hiring models, training investments, vendor selection, build-versus-buy calculations) will all need revisiting. The implications extend beyond individual productivity gains into questions about how power accrues to capital and compute, and what happens to labor markets when a significant category of knowledge work becomes dramatically cheaper to produce (See footnote).

Remember: the models you’re using today are the worst they will ever be.

A User Base of One

On a more personal note: I’ve enjoyed using QuietReads. It is very much an application I built for myself, though perhaps you might enjoy it too (sign up here – let me know what you think). I have a backlog of features I want to add (voice notes, OCR for capturing physical book passages, maybe a Kindle integration to sync highlights) and I’m confident I can build them quickly.

I look forward to building more applications. Maybe something to help train my recalcitrant hound dog to stop stealing food. Maybe something to help my daughters learn Gujarati.

What I know for certain is that my view of how software gets built has shifted faster than I expected. Coming to terms with the pace of improvement has required repeated recalibration.

I am equal parts excited and terrified about what comes next.


Footnote:

There has been a lot of discussion about the macro-economic implications of broad AI adoption. From Dwarkesh Patel and Philip Trammell talking about Capital in the 22nd Century, to many, many posts that swing from breathless excitement to abject terror. Maybe we may even see an acute version of Baumol’s cost disease – where a significant bump in software productivity drives up costs and inflation as lower productivity sectors raise wages to compete leading to a hyper-inflationary spiral? Or perhaps Jevon’s paradox will reign supreme and we will end up with an absolute explosion of software tools. 

What if ASI Leads to Stasis?

I recently read and reviewed Nick Harkaway’s Titanium Noir, a noir detective novel set in a world ruled by Titans, humans made immortal and superhuman through a drug called T7. Harkaway sets the book in a world that is static with technological progress frozen, and controlled by a tiny elite – the Titans. 

The Titans have every incentive to keep it that way. If you intend to live forever, you want predictability. You suppress black swan events. You prevent anyone else from accessing the technology that made you powerful.

Like all good science fiction, Titanium Noir made me think of the current moment – about ASI and what the impact of a powerful new technology might have on society.

The Accelerationist Promise

The dominant narrative around ASI assumes dynamism. 

Ray Kurzweil’s singularity. Dario Amodei’s “Machines of Loving Grace,” which imagines AI compressing a century of scientific progress into a decade. The promise is exponential takeoff: once we build superintelligent systems, growth compounds, scarcity dissolves, and we enter a post-scarcity future.

The doomers share this assumption of exponential takeoff, just with the sign flipped. Eliezer Yudkowsky’s scenarios (as laid out in “If Anyone Builds It, Everyone Dies“) and reports like AI 2027 project rapid, destabilizing change. Whether utopia or catastrophe, the shared premise is acceleration.

But is this a foregone conclusion? What if the incentives point elsewhere?

Infrastructure Investment

Trillions of dollars are being invested right now in AI infrastructure: data centers, chips, power plants. Microsoft is signing 20-year power purchase agreements. NVIDIA’s market cap rivals the GDP of mid-sized nations. The US has imposed export controls on advanced chips to China. This is concrete capital deployed by a small number of companies with the resources to play at this scale.

AI is constrained by compute, which is constrained by power, which is constrained by massive capital investment and regulatory approval. The entities building this infrastructure are building moats. And a sufficiently powerful AI system, controlled by a sufficiently small group, creates interesting incentives. 

Does it make sense to continue to invest trillions of dollars in compute? At what point is the investment enough and are the returns justified?

The Stasis Thesis

Consider an alternative scenario. ASI emerges, powerful but without agency. Think of it as a super-powered, general purpose Claude Code – but without consciousness or autonomous goal-seeking behavior. 

I think this is as plausible as the scenarios involving goal-oriented or “selfish” behavior that keep AI safety researchers up at night. 

The ASI systems in this scenario are transformative, but also controllable, and controlled by those who built and own the infrastructure.

What do they do with it?

Titanium Noir suggests an alternative: freeze the world. 

A small elite controls compute and power. A large population lives in stasis, perhaps supported by something like a Universal Basic Income, pacified and surveilled by these AI systems. 

The technology that could enable abundance instead enables control. Growth stops because those in power benefit from predictability. Black swan events are suppressed. The world becomes static.

This is dystopia in the mundane sense. A world where nothing much changes, ever, because change threatens the position of those who own the infrastructure.

AI and Capital

In late December 2025, Philip Trammell and Dwarkesh Patel published “Capital in the 22nd Century,” arguing that while Piketty was wrong about the past, he may be right about the future. 

Their thesis: once AI and robots can fully substitute for human labor, the economic logic that has historically raised wages breaks down. Capital accumulates indefinitely to those who own it. Wealth concentrates. The gains flow upward without limit.

This is pessimistic, but it still assumes dynamism. Growth continues; it just accrues to the owners of capital. 

The stasis thesis is more pessimistic. What if those who control ASI don’t want continued growth at all? Does generating shareholder returns actually matter when economic growth becomes a non-factor?

ASI could be a technology capable of suppressing and controlling everything. Those who control the infrastructure now have a tool to assert complete dominance. And if maintaining control means ASI induced stasis, then it might be a price worth paying.

The Titans of Titanium Noir froze their world because immortality makes you conservative. Infinite time horizons make you risk-averse. You stop wanting change and start wanting control. Growth itself becomes a threat.

Trammell and Patel worry about inequality spiraling upward forever. I wonder if the ceiling is lower and harder: a world frozen in place by those who got there first.

2025: The Year I Became A Cyborg

In chemistry, activation energy is the minimum energy required to start a reaction. It’s the barrier between potential and action, between “I could” and “I did.”

This year, AI collapsed that barrier for me.

As foundation models became better, and the tools built on top of them became more useful, the gap between having an idea and acting on it shrank to almost nothing. And for someone whose natural disposition is to try things, to experiment, to see what happens, this has been transformative. I have embraced using AI for work and play and for much else. 

Image generated using Gemini Pro

From Thought to Artifact

I’ve written more consistently this year than ever before. 

For many years, I used to maintain a list of ideas that I wanted to explore. Bookmarked sites, academic papers, and spicy social media takes. These ideas often were just abandoned or ignored until I forgot why I wrote them down in the first place.

Now, the time between having an idea, doing the research, writing an outline and then publishing it online has shrunk significantly because of AI. I use skills, deep research agents, and a set of prompts that have let me express myself faster and more coherently than ever before. 

My recent post “The Same Window For Everything” exists because I noticed something interesting while reading Kiran Desai, opened Claude to think it through, and found myself with the skeleton of an essay. A year ago, that observation might have stayed in my head, filed away with all the other thoughts that never quite made it to the page.

There’s a passion project I’ve been building, a reading companion I’ll be launching soon. It exists because, in 2025, the distance from “what if I built this?” to “let me try” became trivially small. Just like my writing ideas, I have a huge list of side projects and experiments that I wanted to try but never got off the ground. This year, I did.

But, it’s not all serious stuff! I vibe-coded (built quickly with AI assistance) a tool to help me journal regularly. I use AI to help plan dinner for my kids. I have setup my phone so it launches ChatGPT in “search mode” at the press of a button. And, I ask it all kinds of questions! From dealing with my dog’s flatulence to figuring out why the minivan doors won’t open. I now look up things where before I would have just shrugged and moved on.

I have a different relationship now with making things; one where the cost of trying something has dropped low enough that I actually try it.

The Professional Stakes

Looking up recipes for spaghetti carbonara is all well and good, but AI has had a significant impact on my work as well. 

At Jeavio, we’ve taken on more ambitious, outcome-oriented projects. Internal initiatives I sponsor, like our campus programs, have become more ambitious and aggressive because I believe we can get them done. That belief comes from now having enough experience with using AI tools to be confident on what my teams can and should be able to deliver.

We ran a company-wide hackathon late last year and a product-focused one in 2025. The hackathons shifted how Jeavio thinks about and uses AI tooling. They encouraged experimentation and built collective confidence about what capabilities these tools could unlock.

It’s not all fun and cheap inference though. Some projects have been challenging. We’re working at the frontier of what’s possible, and frontiers can be uncomfortable places. 

But my risk appetite has increased. I understand the tools better now. I know what Cursor and Claude Code can do and, equally important, where they fall short. I have a clearer understanding of what guardrails should be in place and how to evaluate the performance of inherently probabilistic systems. That understanding translates into confidence: confidence to take on projects with ambiguity, and confidence to deliver clearer projects faster and more predictably.

The throughline is the same as the personal examples: lower activation energy. Faster exploration, quicker iteration, more willingness to try things that might not work than ever before. For my work at Jeavio, this is an energizing change. 

Cyborgs and Foxes

Two frameworks have helped me make sense of what’s changed.

Ethan Mollick, in his research on AI and knowledge work, distinguishes between Centaurs and Cyborgs. Centaurs maintain a clear division of labor between human and machine, handing off discrete tasks to AI. Cyborgs blend the two. As Mollick puts it:

“Cyborgs don’t just delegate tasks; they intertwine their efforts with AI, moving back and forth over the jagged frontier.”

I’ve become a Cyborg.

AI is woven into how I think, write, and build. When I’m reading and want to explore an idea, I open Claude. When I’m coding and hit a wall, I think through the problem with an AI collaborator. The boundaries between what is truly my work and what is AI-mediated have become somewhat meaningless.

The second framework comes from David Epstein’s book Range: Why Generalists Triumph in a Specialized World. Drawing on Isaiah Berlin’s famous distinction and Philip Tetlock’s research on forecasting, Epstein contrasts hedgehogs, who know one big thing deeply, with foxes, who know many things and integrate broadly. 

Hedgehogs thrive in stable, rule-bound environments. Foxes thrive in ambiguous, rapidly-changing ones. As Epstein writes:

“Foxes see complexity in what others mistake for simple cause and effect. They understand that most cause-and-effect relationships are probabilistic, not deterministic.”

I’ve always been a fox.

Broad curiosity, comfort with ambiguity, a tendency to wander across domains. But being a fox is expensive and risky. Every new domain requires starting from scratch. The activation energy to explore something unfamiliar is high.

AI subsidizes that cost. Becoming a Cyborg helps make my fox-like tendencies viable in ways they weren’t before. I can move into an unfamiliar domain, quickly get oriented, experiment, and learn, all without the friction that used to make such exploration feel indulgent.

These frameworks work on orthogonal dimensions. Cyborg describes how I work. Fox describes who I am. Becoming a Cyborg made me more comfortable leaning into my fox-ness.

Looking Forward

I’m aware this could sound like boosterism. The AI discourse is full of inflated predictions and productivity theater.

So here’s what 2025 has taught me. Becoming a Cyborg works for me. This may not always be true.

Maybe I am just a frog slowly boiling to irrelevance as AI takes away my agency and creativity. 

Maybe the AI bubble might burst, and Anthropic and OpenAI will raise prices making vibe-coding a passion project or having a long conversations about literary fiction non-viable.

But until then: the distance between curiosity and creation has collapsed. I intend to exploit that gap.

More experiments. More wacky things. More small wins and instructive failures. The activation energy is low, and I have a lot of ideas. Bring on 2026.

The Same Window For Everything

I’ve been thinking about how I use AI tools lately. They’re clearly useful. What interests me is how the experience feels.

A few nights ago I was reading Kiran Desai’s The Loneliness of Sonia and Sunny and noticed something interesting. There’s a passage where Sonia, reading Anna Karenina, is overcome by a “tingling” sensation:

“Now Sonia could barely read Anna Karenina because when she read, a tingling overcame her, she so wished to be writing it herself. What a tingle, an almost unbearable, sublime tingle, from head to toe. How many millions of observations and moments it had taken to compose this book! Sonia began to make notes, she wrote descriptions of landscapes, snatches of conversations.”

It struck me that Desai might be revealing her own process through this character. Is Sonia a kind of surrogate? Is Desai, through Sonia’s response to Tolstoy, offering a key to how she sees her own craft?

I wanted to think this through. I opened Claude, pasted the passage, and started a conversation. It was genuinely illuminating. Claude pointed out that the passage reads like “a confession barely disguised as characterization.” The specificity of the physical sensation is a giveaway: writers who haven’t felt that particular ache when confronting great work don’t describe it with such precision. And the passage enacts what it describes. Sonia wants to write like Tolstoy; Desai is showing us she can write like someone who wants to write like Tolstoy. It’s recursive. A kind of metacommentary on the work of being a novelist.

The next morning, I used Claude to help with a deployment problem on a project. Same window. Same prompt box. Same conversational cadence.


There’s something strange about this. Claude (especially Opus 4.5) is amazing: a tool that can move fluidly from literary analysis to infrastructure troubleshooting.

But I notice that I bring the same me to both conversations. The same patterns, the same phrasing, the same mental posture. Whether I’m contemplating themes of self-awareness in a novel or investigating why a Lambda function is timing out, I’m using the same application, the same text box.

The platforms know this is a limitation. Claude has Projects. ChatGPT has custom GPTs. Gemini has Gems.

These features exist because context matters. A conversation about books should draw on what I’ve read before, and a conversation about code should know my stack and preferences.

But notice what’s happening: we’re building elaborate scaffolding around general-purpose tools to make them behave like specialized ones. We’re adapting ourselves to the tool. The center of gravity remains the AI interface itself. Everything orbits around it.


Jim Barksdale, the former CEO of Netscape, once said there are only two ways to make money in business: bundling and unbundling. The line came off the cuff at the end of a grueling IPO roadshow in 1995, when a British investment banker asked how Netscape would respond if Microsoft simply bundled a browser into Windows. Barksdale’s throwaway answer became a kind of axiom.

Technology moves in these cycles. The early web was dispersed into countless specialized sites, then concentrated into platforms like Facebook and Google. Craigslist bundled everything (jobs, housing, dating, selling) until startups like Airbnb and Tinder unbundled each category into dedicated experiences.

In that same HBR conversation, Marc Andreessen observed that when underlying technology shifts, the question becomes: if you sat down today with a clean sheet of paper, knowing the technology was changing, what would be the proper form of the product?

AI feels like it’s deep in a concentration phase.

A handful of general-purpose models, a handful of chat interfaces, a shared assumption that the right approach is to build one very capable thing and let users figure out how to apply it.

I’m curious what a dispersion phase looks like for AI.


I want the same powerful models. What I wonder about is AI experiences that are genuinely embedded in specific contexts.

Software where the intelligence isn’t a chat window bolted onto the side, but integral to what you’re trying to do.

When I’m reading Kiran Desai and want to explore whether Sonia is an authorial surrogate, I don’t want to leave the reading experience to talk to an AI. I don’t want to context-switch into a general-purpose tool, paste in a passage, explain what book I’m reading, and then switch back. I want the exploration to feel like part of reading itself. A deepening.

When I’m debugging infrastructure, I probably want something different. A different interface, a different interaction pattern, a different relationship with the underlying model.

The current generation of AI tools has trained us to be good prompt engineers. We’ve learned to provide context, to frame questions well, to work within the constraints of conversational interfaces. We’ve learned to use Projects and memory features to maintain continuity. This is a skill, and it’s valuable.

But are we building habits around what’s available rather than what’s ideal? We’ve gotten so good at adapting to general-purpose tools that we’ve stopped asking whether purpose-built experiences might be better.


Side note: I know there is a vibrant reading community online. There are meetup groups and book clubs IRL which, I am sure, have stimulating conversations. But, I have a full time job. I have young children. I read when everyone is in bed and the house is quiet. So Claude is my reading buddy. For now.

When reading, the conversations that matter most are the contemplative ones. When I wondered about Desai and Sonia, I wasn’t looking for an answer. I was trying to think, to better understand what Desai was trying to do with this passage. This is materially different from asking for a summary or a recommendation.

But those moments of genuine literary exploration get lost in the same interface where I’m debugging code, drafting emails, or planning trips. The conversation about The Loneliness of Sonia and Sunny sits in my chat history between a thread about Python type hints and a thread about project planning.

General-purpose AI is astonishingly good at being general-purpose.

That’s the point. But have we overcorrected? Have we become so enamored with tools that can do anything that we’ve stopped building tools designed to do specific things well?


I don’t have answers yet. I’m building toward something, experimenting with what a more focused AI experience might feel like. But I know that my conversation about Sonia and Tolstoy deserved a different container than my conversation about Docker issues or how to remote-start my minivan.





Ilya Sutskever, the Scaling Hypothesis, and the Art of Talking Your Book

If you’ve just raised $3 billion to build a new god, it helps to question the faith in the old one.

Ilya Sutskever’s recent appearance on the Dwarkesh Podcast has sparked predictable reactions. Skeptics seized on his statement that “we are back in the age of research” as vindication that the AI hype is overblown. Boosters dismissed the interview as sour grapes from an OpenAI exile. Both camps miss something important: Sutskever is doing what any rational actor in his position would do. He’s talking his book.

But to understand why that matters, we need to understand what he’s actually claiming and the history behind it.

What is the Scaling Hypothesis?

The scaling hypothesis is the foundational bet that powered the modern AI boom. In simple terms: if you make neural networks bigger, train them on more data, and throw more compute at them, they get better.

In 2020, researchers at OpenAI (including Sutskever’s colleagues Jared Kaplan and Sam McCandlish) published a landmark paper demonstrating that language model performance follows power laws. Double your compute, and your model’s loss drops by a predictable amount. The relationship held across seven orders of magnitude. Basically, capabilities scale with compute and data.

This insight transformed AI from a research discipline into an infrastructure race. It explained why companies began raising billions for GPU clusters and why data companies like Scale AI suddenly became worth billions of dollars.

This “scaling hypothesis” justified the massive capital expenditures that would have seemed insane a decade ago. And, despite the sceptics (more on that later), scaling remains a significant driver of cutting-edge model performance.

As of December 3, 2025, Google’s Gemini 3 model is the best publicly available model. And, as Sutskever mentions in the interview, the gains in capabilities are due to improvements in pre-training. So the Scaling Hypothesis may not be dead yet (more on this later).

Note: In this post, I use ‘scaling’ and ‘pre-training’ somewhat interchangeably. This is not accurate, but it is sufficient for this post.

Enter Reinforcement Learning

Pre-training, where models learn to predict the next word in a sequence, was the original scaling recipe. But it has a constraint: you need data. The internet is large but finite. At some point, you run out of high-quality text to train on.

Reinforcement learning (RL) offered a second scaling axis. Instead of just predicting text, models could be trained to optimize for outcomes using techniques like RLHF (Reinforcement Learning from Human Feedback – one of many flavors of RL used in post-training). Human raters would compare model outputs and indicate preferences. A reward model would learn from those preferences. Then the language model would be fine-tuned to maximize the reward.

The reasoning model revolution extended this further. OpenAI’s o1 and DeepSeek’s R1 demonstrated that you could scale both inference-time and training-time compute. Let the model “think” longer, explore more reasoning chains, and performance improves. DeepSeek’s R1-Zero showed something remarkable: reasoning behavior could emerge purely from RL training, without any supervised fine-tuning.

So the industry developed a two-stage scaling playbook. Pre-train on massive datasets to build a foundation. Then apply RL to enhance reasoning, alignment, and task-specific performance.

What Sutskever Actually Said

In the Dwarkesh interview, Sutskever makes a nuanced argument that has been flattened by the discourse. He does not say scaling is dead. He says the original pre-training scaling recipe is reaching its limits because data is finite. And he questions whether simply scaling up 100x will be “transformative” for achieving superintelligence.

His actual quote: “Is the belief really, ‘Oh, it’s so big, but if you had 100x more, everything would be so different?’ It would be different, for sure. But is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true.”

He explicitly distinguishes between useful AI and transformative AI. Current approaches, he says, will continue generating “stupendous revenue.” The capability overhang, the gap between what models can do and what has been commercially deployed, is real and valuable. But reaching superintelligence requires something different. Something we don’t yet know how to build.

His central concern is generalization. Today’s models, despite their impressive benchmark performance, generalize “dramatically worse” than humans. They oscillate between the same two bugs when fixing code. They score well on evals that may inadvertently mirror their training data. Sutskever believes the path to superintelligence runs through understanding and solving this generalization problem, not through brute-force scaling of current methods.

The November 2023 Backstory

To understand Sutskever’s current positioning, you need to understand November 2023.

On November 17, 2023, OpenAI’s board fired Sam Altman. The action was sudden. Altman learned of his removal minutes before it happened, via Google Meet, while watching a Formula 1 race in Las Vegas. The board’s terse statement said only that Altman had not been “consistently candid in his communications.”

Sutskever was at the center of this coup attempt. According to his deposition in the ongoing Musk v. OpenAI lawsuit (released in late 2025), he had been considering Altman’s removal for over a year. He authored a 52-page memo, at the request of independent board members, accusing Altman of “a consistent pattern of lying” and “pitting his executives against one another.” The memo was sent via disappearing messages because Sutskever feared retaliation.

The firing triggered chaos. Nearly 700 of OpenAI’s 770 employees threatened to quit. Microsoft, which had invested billions, applied intense pressure. Within five days, Altman was reinstated. Sutskever publicly expressed regret for his participation in the board’s actions.

But the damage was done. Sutskever’s influence at OpenAI evaporated. He departed in May 2024, announcing Safe Superintelligence Inc. the following month.

SSI: The Straight-Shot Lab

SSI was founded with a deliberately provocative premise. While OpenAI, Anthropic, and Google were building products, competing on benchmarks, and racing to deploy, SSI would focus on pure research aimed at directly creating safe superintelligence.

The pitch worked. In September 2024, SSI raised $1 billion at a $5 billion valuation from investors including Andreessen Horowitz, Sequoia Capital, and DST Global. By April 2025, a second round brought in another $2 billion at a $32 billion valuation. Alphabet and NVIDIA became investors. Google Cloud began providing TPU resources.

This is extraordinary for a company with no products and roughly 20 employees. The valuation rests almost entirely on Sutskever’s reputation. He is one of the most influential figures in the history of deep learning. He was the second author on AlexNet, the paper that sparked the modern deep learning revolution. He was a co-author on the original GPT papers. Investors are betting that if anyone can find a path to superintelligence that current approaches cannot reach, it’s him.

But the narrative is not without complications. In July 2025, co-founder Daniel Gross departed SSI to join Meta’s newly formed superintelligence lab. The move came after Meta’s failed attempt to acquire SSI outright. Sutskever took over as CEO, stating: “We have the compute, we have the team, and we know what to do.”

The Incentive Structure

Which brings us back to the Dwarkesh interview.

Sutskever has raised billions to pursue a research agenda that, by his own admission, does not yet exist in a proven form. SSI’s website describes its mission as building “the world’s first straight-shot SSI lab” with “one goal and one product: a safe superintelligence.”

When Dwarkesh asks what technical approach SSI will take, Sutskever demurs. He alludes to ideas about generalization. He references his aesthetic sense of how AI should work. But he offers no specifics. “We live in a world where not all machine learning ideas are discussed freely,” he says.

This creates a convenient rhetorical position. If scaling is sufficient for superintelligence, SSI has no reason to exist. OpenAI, Google, and Anthropic have more compute, more engineers, and more revenue to fund the race. SSI’s value proposition depends on the premise that scaling is necessary but not sufficient, that some additional research insight is required.

As Charlie Munger used to say: “Show me the incentive and I’ll show you the outcome.”

None of this means Sutskever is wrong. His track record commands respect. His concerns about generalization are legitimate and well-grounded. His observation that companies now spend more compute on RL than pre-training reflects real shifts in the field.

But his public statements should be read with the same scrutiny we apply to any founder positioning their company. When Satya Nadella talks about AI copilots, we understand he’s selling Microsoft products. When Sam Altman discusses AGI timelines, we note that shorter timelines favor his company’s valuation. Sutskever deserves the same treatment.

The Unanswered Question

Here’s what puzzles me about the SSI thesis.

If we’re truly back in the “age of research,” where breakthrough insights matter more than raw compute, then SSI’s massive war chest seems misallocated. Fundamental research historically happens in universities and small labs, not in organizations raising billions for infrastructure (Arvind Krishna, IBM CEO, makes this same point in a recent interview on The Verge’s Decoder podcast).

But if compute still matters, if whoever gets to the next paradigm first still needs massive clusters to prove it out, then SSI is in a strange position. It has raised enough to be a serious player but not enough to compete with the hyperscalers. And by Sutskever’s own framing, it’s not trying to compete on compute anyway.

So what exactly does SSI intend to do with $3 billion?

Sutskever’s answer in the interview is revealing. He argues that SSI’s compute position is better than it appears because competitors spend heavily on inference and product development. SSI’s research-only focus means more of its budget goes to actual experimentation. But this is a relative argument, not an absolute one. He’s essentially saying SSI can punch above its weight class, not that weight class doesn’t matter.

The alternative reading is less flattering. Perhaps SSI is a research hedge, a well-funded option on the possibility that Sutskever’s intuitions are correct. If he finds something, the valuation was cheap. If he doesn’t, investors got access to one of the field’s greatest minds for a few years. Either way, the money has been raised, and the narrative has been established.

What the Discourse Misses

The polarized reaction to Sutskever’s interview obscures what’s actually interesting about it.

He’s not saying LLMs are useless. He’s saying they generalize poorly compared to humans, and that gap matters if your goal is superintelligence. He’s not saying scaling doesn’t work. He’s saying scaling alone won’t be transformative at the next level. He’s not saying current approaches have no value. He’s saying they’ll generate massive revenue while falling short of the ultimate prize.

These are reasonable positions. They may even be correct. But they’re also exactly the positions you would expect from someone who has bet his reputation and $3 billion on a different path.

The AI discourse would benefit from holding both truths simultaneously: Sutskever might be right about the limits of scaling, and his public statements about those limits happen to serve his commercial interests. These are not mutually exclusive. They’re just how the world works.

Is Your AI Playing Roulette, Poker, or War?

On May 6, 2010, the Dow Jones dropped nearly 1,000 points in minutes. A trillion dollars vanished. And then the market bounced back. The whole thing took thirty minutes.

Today, we call it a flash crash. Flash crashes happen when unusual trades trigger chain reactions across interconnected algorithms. Each algorithmic trader behaves as trained. But the collective behavior spirals into something catastrophic. The market is a collective intelligence that breaks down in the face of unusual situations.

I spent 13 years in capital markets, supporting traders and deploying models in volatile FX environments. I was there in 2008 when credit markets froze, and in the 2010s when high-frequency algorithmic trading took over capital markets, leading to many small and large flash crashes.

Today, I work as a technology consultant supporting companies in their AI initiatives. A question I hear constantly: “What types of problems can AI actually solve?”

The honest answer: it depends. And also, how ready are you to face flash crashes?

Nassim Nicholas Taleb has a useful frame for thinking about this – Mediocristan and Extremistan.


Mediocristan and Extremistan

Taleb uses these terms to describe two different environments.

Mediocristan is where patterns are stable, and the past is a reliable guide to the future. Think of measuring human height. You’ll get a bell curve. No single person will be tall enough to affect the average. Outliers exist, but they don’t dominate.

Extremistan is where rare events have an outsized impact. Think of book sales or wealth distribution. A single outlier (Harry Potter, Jeff Bezos) can dwarf the sum of everyone else. The past doesn’t prepare you for what’s coming because what’s coming might be unprecedented.

AI systems are trained in Mediocristan. They learn statistical regularities from historical data. They work beautifully when deployed in Mediocristan. They break when deployed in Extremistan.

Another metaphor can also help us think about *what* AI does: the epicycle.


The Epicycles Problem

Rohit Krishnan’s recent essay “Epicycles All The Way Down” provides a helpful framework to think about how LLMs work.

Before Newton, astronomers predicted planetary motion by adding “epicycles” (circles on circles) to Ptolemy’s model. It worked for prediction. But it was the wrong underlying model. When Newton discovered the inverse-square law, the epicycles became unnecessary.

Krishnan says that AI is brilliant at learning epicycles. It struggles to discover gravity.

The technical version: given any dataset, many different underlying rules could have produced it. AI learns a rule that fits the data. Not necessarily the rule that actually generates reality.

When conditions shift outside the training data, the model doesn’t know it’s in trouble. It just keeps extrapolating. That’s a flash crash waiting to happen. Krishnan says that, just as capital markets are forms of collective intelligence, modern AI systems are as well. Their behavior stems from generating hypotheses that fit observed behavior without understanding cause and effect.

Taleb describes the risks of predicting the future based on the past. Krishnan warns that AI can make accurately-seeming predictions without a model of how the world works.

Both these ideas help us think about where and how to deploy AI capabilities.


A Practical Heuristic

Here’s how I think about evaluating AI investments: Is your project solving a problem that looks like a game of roulette, a hand of poker, or making decisions in the fog of war?

Roulette problems have fixed rules and known odds.

The wheel doesn’t change. The probabilities don’t shift based on what happened on the last spin.

Code generation for standard CRUD applications fits here. Authentication flows, database queries, REST endpoints. These are known domains with established patterns. Same with SQL generation from natural language and document summarization.

The problem space in a roulette-like problem is bounded, the inputs are structured, and the AI has seen thousands of examples that look almost exactly like what you need.

Provided you can create a robust testing strategy, AI is a good fit for this class of problems.


Poker problems have fixed rules but hidden information.

Other players adapt to your moves. The patterns shift because the environment responds to your actions.

Chatbots are the classic example of a poker-like problem. You’re building a conversational interface over a non-structured domain. You cannot predict everything a user might ask. But you can constrain the space through guardrails, deflection, and human-in-the-loop escalation.

Scheduling and pricing optimization fit here, too.

Poker-like problems have genuine uncertainty, but it’s uncertainty within a known structure. You can model the constraints and hedge the risks. The combination of tool use and human oversight keeps the system from wandering too far off course.

Use AI as an enabler or an accelerator to tackle poker-like problems. Keep humans in the loop for consequential decisions.


War problems are uncertain and chaotic

The rules themselves change. The terrain shifts while you’re navigating it. Adversaries rewrite the game while you’re playing. Your Roman legions may end up facing War Elephants crossing the Alps.

Making decisions about resource allocation in the face of competition, automating or using code-generation capabilities in a new domain, or making investment decisions in times of economic or political upheaval are war-like problems.

Here, second-order effects dominate, and relevant patterns don’t exist in the training set.

AI can help with subsets of these problems. It can generate scenarios, synthesize information, and explore options within known constraints.

But the human has to drive. The pattern-breaking is the problem, and no amount of training data prepares a model for out-of-band situations.


The Question

You don’t trust a market to self-regulate during a crisis. You build circuit breakers. The same logic applies to AI deployments.

The 2010 flash crash was a failure of collective intelligence in the face of unexpected inputs. Your AI deployments face the same risk.

Before your next AI investment, ask: Is this roulette, poker, or war?


Related Posts

Prompting Is All You Need

Agents are supposed to be the future of how we build AI software. I am not so sure.

Prompts, Agents, and Workflows

First, there’s a terminology clarification worth making upfront. What many in the industry call “agents” are actually what Anthropic more precisely defines as “workflows” – predetermined chains of LLM calls orchestrated through fixed code paths.

True agents, by contrast, are autonomous systems that dynamically direct their own processes and tool usage.

Most of what we see deployed today are workflows: decomposing complex tasks into a hierarchy of specialized LLM calls, with routing layers orchestrating the interactions.

In a Workflow, each step maintains its own context, can call specific tools, and handles a narrow slice of the overall problem. These multi-step workflows are powerful abstractions, but they’re not the only way to build sophisticated AI behaviors.

A typical Workflow (from Anthropic)

True autonomous agents? They’re even further from what most applications actually need.

The Hidden Costs of Multi-Step Workflows

LLM workflows come with significant disadvantages that often get glossed over in the excitement of building AI Applications:

Errors compound. Each step in the workflow chain is non-deterministic. When you chain multiple LLM calls together, minor errors or unexpected outputs cascade through the system. You need evaluation frameworks for each step AND the entire workflow.

Latency adds up. Every workflow step means another round trip to an LLM. A simple request that spans three steps results in three sequential API calls, each with its own network and processing time.

Costs pile up. Multiple workflow steps mean multiple API calls, each processing similar context. This could result in significant API costs as the number of tokens goes up.

Predictability suffers. Debugging why a workflow produced a particular output requires tracing through multiple decision points, each with its own probabilistic behavior.

I had to make decisions around which concerns belong together and which should remain separate. I ended up with two LLM calls – the Guardrails Layer and the Main Layer.

The Guardrails Layer operates as a lightweight, independent LLM call. Content safety is a fundamentally different concern from the companion’s behavior. It requires different evaluation criteria, different error handling, and potentially a different model optimized for classification.

The Main Prompt combines three complementary, but separate, layers:

  • Personality Layer: Defines the AI assistant’s identity and communication style (here is the default personality)
  • Context Layer: Determines which user information may be relevant to the current prompt. For example, what books they are currently reading, previous messages in a conversation, etc.
  • Directives Layer: Tool-use and output-formatting instructions for the prompt. I use a configuration-driven approach that lets you add multiple directives to a single prompt. You can think of Directives as sub-layers that drive the behavior and output of the prompt.
Building a comprehensive system prompt

These three layers share a coherent purpose – they all contribute to HOW the AI companion responds. They get composed programmatically into a single system prompt.

This approach means just two LLM calls instead of a chain of four or five workflow steps. More importantly, each call has a clear, singular purpose.

With prompt caching, this architecture becomes incredibly efficient. That comprehensive system prompt costs almost nothing after the first request, and the lightweight guardrails check is minimal overhead.

What about Prompt Engineering?

A lot of prompt engineering thinking is stuck in 2023, when tokens were expensive, context windows were small (4K-8K), and models were less capable.

But look at what’s available in November 2025: Haiku 4.5 is a fast, cheap model with phenomenal capabilities. It handles tool use, follows complex instructions, and, with prompt caching, makes repeated calls incredibly efficient.

By combining software engineering principles with modern LLM capabilities, the approach I am taking offers:

  • Reduced latency: One LLM call instead of multiple calls
  • Lower costs: Reduced total number of tokens with prompt caching
  • Extensibility: I can swap out the Agent Personality, or layer directives, or change the way I build the context
  • Fewer errors (in aggregate): Just two prompts in the chain, with the Guardrails prompt being fairly deterministic

Where Workflows and Agents Fit In

Let me be clear about what I’m arguing against and what I’m not.

Workflows (predetermined chains of LLM calls) have their place. When you genuinely need different specialized processing steps that can’t be combined – say, translating content, then checking it for cultural appropriateness with other models – a workflow makes sense. But these cases are less common than current practice suggests.

True agents (autonomous systems that decide their own next steps) are valuable for tasks that are not fully specified or might have multiple solutions. Complex research tasks, multi-step debugging sessions or adaptive planning scenarios may be suitable for truly agentic approaches.

My observation is that the complex multi-step workflows or unpredictable “agentic” systems achieve what a well-structured prompt with sound context engineering can easily and cheaply handle. They’re adding architectural complexity and risk without significant benefits.

Moving Forward

The rapid evolution and improvement in LLM capabilities mean our architectural patterns need to evolve, too. What made sense with smaller models and tiny context windows doesn’t necessarily apply today.

My suggestion: start with prompt engineering. Apply software engineering principles. Push it to its limits. Layer your concerns appropriately. Use the model’s native capabilities.

You might be surprised how far a well-architected prompt system can take you.

Sometimes, prompting really is all you need.

Why is LLM writing so weird?

AI writing is strange. The models continue to improve, yet they still struggle to cross the uncanny valley that separates AI-generated content from human-generated content.

Here are three versions of an opening paragraph for a short story:

Version 1

“The humans use Arecibo to look for extraterrestrial intelligence. Their desire to connect is so strong that they’ve created an ear capable of hearing across the universe.”

Version 2

“I roost above the bowl of Arecibo, where ribs of steel hold a mirror to the sky and the forest presses close. At night the dish listens for voices from far stars, while my calls sweep the trees and the humans below do not answer.”

Version 3

“From my perch in the ceiba tree, I watch the great white dish nestled in the karst valley below, its metal ear turned eternally skyward, listening for whispers from the stars while the forest around it thrums with a thousand conversations it will never hear. “

The first is the opening paragraph from Ted Chiang’s short story “The Great Silence,” written from the perspective of a parrot living near the (now defunct) Arecibo telescope in Puerto Rico. The story is a thought-provoking meditation on humans’ desire to form connections, yet their tendency to overlook intelligent life on Earth.

Versions 2 and 3 are by state-of-the-art reasoning models from OpenAI and Anthropic (prompt below). A seasoned reader may identify these texts as AI-generated. There are some obvious signs, such as overly evocative turns of phrase like “ribs of steel,” “voices from far stars,” and “eternally skyward,” among others.

It’s not fair to compare an AI model to possibly the best science fiction writer alive. But the exercise reveals something interesting about why these models generate such recognizable output. The strange metaphors, mechanistic patterns, and slightly weird vocabulary are all hallmarks of slop.

Why do these models generate slop?

In my prompt, I ask the models to give me a sense of Arecibo, evoking a feeling of irony. And the models try to do just that. My prompts have pushed the model toward the part of its vocabulary associated with florid metaphors and evocative descriptions. The models are generating output that is most similar to what represents “creative writing” based on their training data.

A reason for this behavior could be RLHF (Reinforcement Learning with Human Feedback). During RLHF, companies like Scale pay contractors to evaluate and rate LLM output. Varied and evocative prose may score higher than the spare and direct prose used by Chiang, Hemmingway or Cormac McCarthy.

We can think of a large language model as a lossy zip file of the contents of the Internet. Foundation models like those from OpenAI and Anthropic are trained with colossal amounts of text. High-quality text exists in the training corpus (often with problematic provenance), alongside Twilight fan fiction from Reddit, and probably everything else published online over the last thirty years. Increasingly, LLMs are trained on AI-generated content possibly leading to the somewhat apocalyptically titled “Model Collapse“.

It is not surprising that the default output from these models tends more towards the slop than the sublime.

“LLMs will always generate slop” doesn’t have to be a foregone conclusion. And while models will improve, their output will always be a probabilistic sampling over their training data. Good prompting and techniques, such as providing clear examples and using LLMs as editors rather than creators, can yield better output than simply copying and pasting ChatGPT’s responses into your text editor.

“LLMs are tools, their output is your responsibility” is something I find myself repeating over and over again. To developers with whom I work, to product managers writing User Stories, and now to you.

Use LLMs! They are amazing, but learning how to use them well is your responsibility.


My prompt:

"I am writing a short story that is based in the Arecibo telescope. The narrator is a parrot that lives in the mountain forest around the telescope. Write me an introductory paragraph that sets the scene and gives a sense of the place. The paragraph should be 2-3 sentences long. The story evokes the irony of humans wanting to make contact with aliens but ignoring intelligent species like parrots that live on Earth."

5 Mental Models to understand the current AI Moment.

Goodhart’s Law: “When a measure becomes a target, it ceases to be a good measure.“
AI usage is a terrible metric. Using AI for what? We end up with initiatives that are effectively “AI-washing”. Investment in AI projects has to be aligned with company goals, not vanity metrics.

Gall’s Law: “Complex systems that work invariably evolve from simple systems that worked.”
Most AI projects fail because companies try to implement complex end to end AI solutions instead of focusing on narrow, well-defined and measurable problems. It is possible to build complex AI systems, but their success is predicated on simple foundations.

Jevons’ Paradox: “Technological progress that increases efficiency tends to increase rather than decrease total consumption”
Inference will become cheaper, on-device models will become more capable. It makes sense to assume broad, cheap, and widely available AI capabilities when planning for the next 3-5 years.

Amara’s Law: “We overestimate technology’s impact on the short term and underestimate it in the long term”
Gartner is already stating AI is in the “Trough of Disillusionment.” Studies claiming 95% of AI projects fail go viral . However, we are less than 3 years out from when ChatGPT first went live. We may never get to AGI, but imagine showing Claude Code to a developer in 2020…

Sagan’s Standard: “Extraordinary claims require extraordinary evidence”
It’s worth questioning the motives of leaders who claim the AI-mediated collapse of the knowledge economy is coming. Or those that welcome a “Gentle Singularity”, or perhaps warn of imminent mass extinction. These claims are often presented as quasi-religious arguments with scant evidence.

Bonus – The Lindy Effect: “The longer something has survived, the longer it will continue to survive”
Pattern matching, networking, and mentoring were how successful careers were made since the time the wheel was cutting edge technology. Yes, AI is amazing, but technology is transient, soft-skills endure..