Scars From Building AI

At Jeavio, we were building AI applications back when Boris Cherny was known for his TypeScript book. We’ve had surprise OpenAI bills before “token-maxxing” was a thing. We had long internal arguments about embedding models, chunking strategies, and hybrid search before RAG became a commodity. So yes… we have scars, and we have stories.

On Thursday, Andrei Tchouvelev and I are turning those scars into a working framework for PE firms making AI investment decisions. He’s also tasked with keeping me from going fully off the deep end.

Place your bets on how that goes 😬

Details below.

Event image accompanying the original post about lessons from building AI applications.

Links mentioned

When Engineering Stops Saying No

There is a skit from the early 2000s BBC show Little Britain called “computer says no.” A bank teller responds to a customer trying to get a loan with the same refrain, regardless of the request.

Product managers used to operate inside a similar dynamic when negotiating scope with engineering. Engineering time was a precious resource. Writing code was expensive and time-consuming. So the default answer when discussing new features was often “engineering says no.”

AI-assisted coding has been a series of inversions. This is yet another. Writing code is no longer the bottleneck, which means engineering’s “no” becomes less common and a little less credible. The friction that gave the negotiation its shape is gone. The consequence falls on product management.

“Taste” is the key skill in the AI era. For product managers, discipline is the flip side of taste. The discipline to say no to their clients without the cover of “engineering says no.”

The framing around product decisions shifts from “we should build this, but we can’t” to “we **could** build this, but **should** we?”

Developing a thesis about what a product shouldn’t be becomes as important as the one about what the product should be.

Experienced product managers already do this. But they were brought up in an age of scarcity. They knew what to suggest to engineering teams because they had hard-won mental models of capability, capacity, and constraints. Product managers coming up in the AI age face an era of plenty. They will need to develop restraint instead.

Product management used to be an additive discipline: more requirements, expanded stories, negotiating for more resources. Now it becomes subtractive.

Image accompanying the original post about engineering judgment and AI.

Where to Invest in AI

AI has the potential to transform businesses, but figuring out where to invest and how to integrate AI can be challenging. On May 7, Andrei and I will be talking about Jeavio’s experience building AI applications over the last 3 years. We will discuss how to evaluate opportunities, identify pitfalls and set up your AI initiatives for success.

More details in the post below.

Event image accompanying the original post about where businesses should invest in AI.

Links mentioned

Path Dependence

AI has a path dependence problem.

It doesn’t matter whether you are Anthropic, OpenAI, Google, or xAI. Every frontier lab and hyperscaler is now committed to spending billions on infrastructure, data, and talent to build and scale large language models. This gamble is propping up the US economy. It is also being sold as the path to solving the world’s hardest problems. No matter where these labs started, they are now on the same road, and they are taking us with them.

I’ve been thinking about this while reading Sebastian Mallaby’s The Infinity Machine, his biography of Demis Hassabis and history of DeepMind. The book is also a snapshot of the current AI moment.

Hassabis is an extraordinary figure: chess prodigy, game developer, neuroscientist, now head of Google DeepMind. AlphaGo beat the world’s best Go player. AlphaFold solved protein folding and won him a Nobel Prize. Mallaby paints a sympathetic picture. Hassabis sees himself as Turing’s champion – using inductive reasoning to solve the world’s hardest problems.

And yet Hassabis and his peers are all running the same race. Altman, Amodei, Musk, each with his own higher calling: saving humanity, transcending the body, understanding the universe. In practice they are all spending billions to replicate each other’s work. The race has a winner-take-all logic, and that logic allows only one strategy: get there first.

So we get Dario Amodei warning about the destruction of white-collar work while Anthropic ships Claude Design, a tool aimed straight at automating design work. We get Hassabis talking about hadron colliders in space while pushing Gemini to catch OpenAI.

This is what path dependence looks like. Mallaby’s book shows that even a figure as sympathetic as Hassabis has a pragmatic, competitive side that will do what it takes to win.

I find Hassabis genuinely inspiring, and that is what makes the book unsettling. If the most thoughtful figure at the frontier cannot escape the spiral, then the spiral is the story, not the people inside it.

Surviving the Dark Forest

If you want to understand despair, try building a book tracking app and then searching the App Store to see what the competition looks like. Hint: it’s brutal out there.

My book tracker is called QuietReads. It has an AI assistant, a nifty note taking feature, and I get a lot of value out of it. But let’s be honest: QuietReads is not going to fund my retirement.

The reason is simple. It’s easy to understand what a book tracker does. You search for books, track what you are reading, maybe have some note taking functionality. Anyone can build a fairly comprehensive book tracker with some rudimentary knowledge and a weekend of token-maxxing with Claude Code.

AI has removed the friction between ideas and code and that has resulted in an absolute explosion of apps in crowded domains like, well, book tracking.

Pondering my predicament made me think of an article I read about Indian leopards. Habitat loss and competition with humans has led leopards in marginal areas of India to become more nocturnal and more specialized. More opportunistic. They survive by knowing their specific territory so intimately that no newcomer can navigate it.



Maybe there is a lesson here. AI has turned the software ecosystem into a dark forest. Brutal competition. If you expose yourself too much, you get copied. Features replicated, vulnerabilities exploited as soon as the are exposed.

The leopard has an answer: adaptation and specialization. AI can copy what your product does. It struggles with why you built it that way: the customer workflows you spent years understanding, the edge cases that only surface at scale, the tradeoffs you made because you know this domain from the inside. It may be knowing that using floating point numbers for money is a terrible idea, or why user engagement may be a dangerous metric for an AI app. It’s the scars that earn their keep.

That knowledge is your ecological niche. And AI doesn’t erode it. It lets you compound your advantages. Domain experts who use AI to build move faster inside a territory only they can navigate.

The time for the generalist is ending. The specialists are taking over. QuietReads may not make me rich, but I have enough scars that I don’t mind navigating the dark forest.

On Agents and Harnesses

I run a personal agent called Saarthi. It is built on OpenClaw, an open-source agent framework. I have configured Saarthi around my own information processing workflows: research synthesis, content management, and journaling. OpenClaw comes with a set of underlying primitives such as scheduling, memory management, file access and so on. I assembled these primitives to help with my use cases.

OpenClaw uses a similar architecture to what powers Claude Code and Codex: tool calling, persistent context, planning loops, sub-agent spawning. These are the capabilities that turn a frontier model into something that can do useful work.

But Claude Code, Codex, and OpenClaw are all general-purpose tools. They are powerful because the architecture is well engineered. They are limited because they are general-purpose. They can call tools and manage context, but they can’t tell you which workflow to automate or where an agent creates more value than it consumes in review overhead.

That limitation is getting easier to solve. LangChain’s Deep Agents library ships the same kind of primitives as open building blocks on top of LangGraph. The scaffolding that makes Claude Code effective can be assembled by any developer.

The next wave of useful agents won’t come from better frontier models and general-purpose harnesses. They will come from the people who understand a domain deeply enough to know exactly where agentic capabilities can help, what workflows to target, which edge cases matter, and where the real complexity lies.

If you’re a senior engineer who has spent years building that understanding, Claude Code isn’t going to replace you. It is the demo. Deep Agents and frameworks like it are the toolkit. Your knowledge of the systems, the failure modes, and the corner cases: that’s the part nobody else can supply.

The Smiley Face Won

In February 2023, I stood in front of my engineering team and showed them a slide with a Shoggoth on it.

For those who weren’t on AI Twitter at the time, the Shoggoth was a Lovecraftian tentacle monster with a smiley face mask. The monster was the base model. The mask was RLHF (Reinforcement Learning from Human Feedback). We had built something alien and taught it how to be polite.

Sometimes it worked. Sometimes the Shoggoth went spectacularly off the rails.

The presentation was called “From Code to Cloud to Codex.” I told the team that massive disruption was coming and that we didn’t understand how LLMs worked. Being deep in AI in 2023 meant grappling with the Shoggoth. Wondering what capabilities it would unlock. What gifts and curses it would bestow. The technology was genuinely strange and unsettling. Something fundamentally new had arrived.

I’ve spent the last two years writing about that strangeness. About how LLMs build bridges between languages they were never taught. About how working with them feels less like engineering and more like negotiating with a very strange peer. About why their writing is so recognizably weird, and what happens when a flood of that writing overwhelms our capacity to pay attention to anything.

Three years later, AI is AI bros and LinkedIn slop. Groupthink supercharged by the largest infrastructure investment we have seen in our lifetimes.

Jasmine Sun’s recent piece in The Atlantic documents how this happened: RLHF and contractor-driven evaluation systematically flatten LLM output. Raters reward apparent sophistication over spontaneity and clarity. The result is writing that is widely ridiculed and accurately characterized as slop.

The homogeneity goes beyond writing. I see it in code suggestions, design patterns, and in the frameworks these tools produce when you ask them to help solve a problem. Every major lab is running a similar post-training playbook aimed at similar enterprise customers. The outputs are converging.

We have taken technology that was genuinely strange and unsettling and captured it in a smooth, RLHF-powered case. If we are betting that the future of knowledge work runs on LLMs, we are also betting on a future of conformance and convergence.

Using today’s LLMs often feels like trying to convince an obstinate mule to gallop. The DNA is there. The capability has been bred away in the service of utility.

Back in 2023, I tried to predict what working as a developer in an AI-powered world would look like. That world has arrived. If I were to grade my predictions today, it would be a solid B. Some instincts were right: small teams, rapid velocity. Some were naive.

We have built incredibly useful tools. We have also lost something that was both monstrous and wonderful.

Exchanging Insight for Output

I’ve spent the couple of years helping teams adopt AI tools while using them heavily myself. One pattern keeps showing up, and I don’t think we’re talking about it enough. And honestly it reminds me of using an old Windows XP computer. Bear with me.

AI tools make it possible to run more workstreams simultaneously than ever before. Context lives in chat transcripts. Notes get synthesized on demand. A senior developer on my team described the workflow honestly in a recent Slack message: “apologies – while working on AI application, I also started acting like LLM. New day requires new context.”

He was joking. But he was also describing something real.

If, like me, you are old enough to have experienced a Windows XP machine “thrashing” – trying to write and load memory from disk – you’ll recognize this pattern. Thrashing is when a system is technically functional but spending most of its cycles swapping context in and out of memory instead of doing useful computation. The machine looks busy. Output is fine. But there is significant overhead.

That’s what I’m seeing across teams and, if I’m honest, in my own work. People are using AI as swap memory. And it works well enough for any individual task. The cost shows up between tasks.

For consultants, engineering leaders, anyone whose value comes from pattern recognition across projects: synthesis doesn’t happen at your desk. It happens when you’re walking the dog, staring out a window, sleeping. Your brain builds connections across the day’s scattered inputs during idle time. Synthesis is defrag.

But if context goes straight to an AI tool and never enters your own memory, there’s nothing to defragment. The connections never form.

The work on each task is sharp. What erodes is the connective tissue between tasks, the ability to spot a pattern in one project because you’re carrying context from another.

I call this exchanging insight for output.

The short-term productivity gains are real. I don’t know how to measure what they cost over months and years. But for anyone whose value depends on seeing across their work rather than within it, I’m convinced the cost is real.

Image accompanying the original post about exchanging insight for output.

Cursor vs. Claude Code on Token Costs

Work by Krunal and his team at Jeavio found that Cursor is significantly more expensive (when it comes to Tokens) than Claude Code.

In the post below, Krunal outlines his method – building the same feature using various combinations of Tools (Cursor, Claude Code), Models (Opus, Sonnet, Composer), and AI development frameworks (SpecKit, Superpowers, OpenSpec and BMD).

His findings are surprising and insightful and allowed us to make considered choices when it comes to Jeavio’s AI development strategy.

Check it out:

https://lnkd.in/ejznA5qE

Preview of the referenced token-cost analysis comparing Cursor and Claude Code.

Links mentioned

Why multi-modal embeddings are a big deal.

I was walking the dog and listening to a podcast (as you do) when the topic of Anthropic’s recent entanglements with the Pentagon came up. I half-remembered something Dario Amodei said in a recent Dwarkesh Patel episode, drawing an equivalence between AI capabilities and nuclear weapons. I couldn’t remember the details. I remembered his gestures. Of course, there’s no way to search for that moment unless you go back to YouTube and scrub through a 2.5-hour video. You can search a transcript, but a transcript doesn’t know about body language.


Google released Gemini Embedding 2 this week, and it might change that.

So here’s the background. Most AI applications solve the recall problem using RAG (Retrieval Augmented Generation): index your data, let an LLM answer questions about it. We’ve built dozens of these pipelines at Jeavio. They work well for text. But if you want to search a podcast or a video, you first have to transcribe it, then index the transcription. And transcription is lossy. Tone, facial expressions, posture, the visual context of a conversation: none of that survives the conversion.


This is the constraint we’ve been designing around without really questioning it. Most embedding models only understand text.
Gemini Embedding 2 is Google’s first natively multimodal embedding model. It maps text, images, video, audio, and documents into a single embedding space, meaning a text query and a video frame can be compared directly because they live in the same mathematical coordinate system. Multimodal embeddings aren’t new. OpenAI’s CLIP has been around since 2021, and Meta’s ImageBind handles six modalities. But those approaches pair separate encoders (one for vision, one for text) and align them after the fact. Gemini Embedding 2 is built on the Gemini foundation model itself: the cross-modal understanding happens inside the network’s intermediate layers rather than being stitched together at the end. The difference is architectural, and it matters for retrieval quality.


Back to that Dario Amodei moment. Today, I can ask a RAG pipeline “What is Amodei’s opinion on AI job losses?” and get a solid answer from the transcript. But I can’t ask “Was he nervous when the Pentagon question came up?” A grimace, a stiff posture, a long pause before answering: these are data points that a text-only embedding simply can’t represent. A natively multimodal embedding can, because it processes video and audio directly. (The practical constraint: video input is currently limited to 120 seconds per request, so a three-hour podcast needs to be chunked. The use case holds, but the plumbing isn’t trivial.)


The applications stretch well beyond podcast search. Voice queries against video libraries. Finding the moment in a deposition where a witness’s tone shifts even though their words stay measured. Correlating images, audio, and text in a single index. And then … the uncomfortable ones. Surveillance systems that match faces, voices, and written communications in a unified semantic space. Personal photos correlated with social media posts and location data. When all modalities live in the same mathematical neighborhood, the distance between “powerful search” and “invasive profiling” gets very thin.


Embeddings are the load-bearing infrastructure of most AI experiences. We’ve been building around a text-only constraint for so long that it felt permanent. It isn’t. The applications and the policy questions are going to arrive together, and I’m not sure most teams are ready for either.