Building a Truth-Seeking AI is a Sisyphean Endeavor

Elon Musk’s goal for xAI’s Grok model is to be “maximally truth-seeking.” When Grok generated responses that were not aligned with Musk’s ideas of Truth, he promised to “fix” Grok, which appears to involve tweaking its system prompt. The results included Grok calling itself MechaHitler after being made less ‘politically correct’. Problematic.

But is it even possible to build a truth-seeking AI?

LLMs are probabilistic machines. They predict the next token based on patterns from a massive corpus of Internet text.

When xAI added “don’t shy away from politically incorrect claims” to Grok’s prompt, they weren’t accessing Truth but adjusting probability distributions and nudging the bot’s behavior into problematic spaces.

Training Grok4 reportedly cost almost half a billion dollars. It was so expensive because model capabilities grow with the size (parameters) and the amount and diversity of the data used to train the model. LLM capabilities are an emergent behavior driven by the amount of data used to train the model.

LLMs are “grown, not crafted”. Trying to ensure that an LLM becomes “maximally truth-seeking” is a Sisyphean task.

You could train an LLM only on data that is politically acceptable – oh sorry – certified to be True. Musk is, of course, building Grokipedia – guaranteed to be free of bias and presumably used as a corpus for training “the son of Grok”.

Good luck with the benchmarks!

Elon Musk has a phenomenal track record, but he will fail to build a maximally truth-seeking AI. LLMs operate in a probabilistic world. They are phenomenally capable black boxes for which we have no coherent theoretical framework to explain their behaviors.

Tweaking the system prompt, or tweeting angrily, may nudge LLM behavior, but with unpredictable and potentially undesirable outcomes. Instead of engaging in an endless culture war, it might be more prudent to use engineering resources to develop better guardrails on LLM behavior and take a realistic assessment of current AI capabilities.

Review: If Anyone Builds It, Everyone Dies

by Elizier Yudkowsky and Nate Soares (2025)

We are at an interesting moment in artificial intelligence. Massive investment continues to pour into AI infrastructure, with McKinsey estimating $5.2 trillion in capital expenditures by 2030 for data centers alone. We’re seeing the first documented cases of what’s being called “ChatGPT-induced psychosis,” where users spiral into severe mental health crises after becoming obsessed with AI chatbots. We’re watching significant job displacement begin to unfold, with Anthropic CEO Dario Amodei warning of a “white-collar bloodbath,” predicting that AI could eliminate half of entry-level white-collar jobs and push unemployment to 20% within five years. And yet, despite all this disruption, there’s still no clear path to artificial general intelligence.

My P(doom), the probability I assign to AI causing human extinction, is low. There are significant risks associated with the widespread adoption of poorly understood technology. However, I don’t believe current foundation models represent a viable path to ASI (Artificial Super Intelligence, also sometimes referred to as AGI – Artificial General Intelligence). We’re more likely to experience a dot-com-style correction than achieve exponential growth toward superintelligence.

It’s in this context that “If Anyone Builds It, Everyone Dies” arrives. The authors, Eliezer Yudkowsky and Nate Soares, run the Machine Intelligence Research Institute (MIRI), where they’ve spent decades working on AI alignment and safety. Their new book makes an extreme claim: if anyone builds artificial superintelligence, humanity will go extinct. Not “maybe possibly,” but inevitably.

I found the book compelling in parts, incomplete in others. It succeeds at making alignment challenges accessible to a general audience. It fails to grapple with where we actually are today with AI: massive investments, uncertain results, and significant challenges and risks to the broad adoption of the technology.

The Book’s Structure

Yudkowsky and Soares have written the book for a general audience, using parables, stories, and examples to explain how AI is “grown, not crafted.” This approach makes understanding how modern AI systems work surprisingly accessible, even to readers without a technical background.

The authors divide the book into three main sections. First, an introduction to machine learning and AI concepts that grounds readers in the fundamentals. Second, a fictional scenario where a misaligned AI releases a bioengineered plague to facilitate its takeover of human society. Third, a call for a nuclear non-proliferation-style moratorium on AI development, enforced by military action if necessary.

The title leaves nothing to interpretation. It is a strident warning about humanity’s impending doom if we don’t stop the march towards ASI.

Their core thesis rests on several interconnected arguments. ASI will inevitably lead to extinction because we cannot understand current AI architectures; the interpretability problem remains unsolved. We cannot predict emergent behaviors, which they illustrate through evolution’s production of the peacock’s elaborate tail. And crucially, we cannot guarantee alignment with human welfare when the systems are “grown, not crafted.”

The fictional section follows Sable, a near-future AI platform created by a company that serves as a transparent stand-in for OpenAI or Anthropic. Sable releases a bioengineered plague to facilitate its takeover of human society. The AI’s motives remain deliberately inscrutable; that’s the author’s point. We won’t understand what drives a superintelligence any more than an ant understands human motivations.

In the final section, Yudkowsky and Soares draw parallels with the Chernobyl disaster, arguing that perverse incentives will always lead someone to take catastrophic risks. Their solution: a treaty that makes it illegal to conduct AI research that could lead to the development of ASI. Military action undertaken by the signatories will enforce this treaty. This connects to Yudkowsky’s 2023 TIME magazine piece where he called for airstrikes on rogue data centers training unauthorized AI systems.

Where the book falls short..

The fictional scenario is the book’s weakest element. Any casual science fiction fan has encountered this scenario before, from the Reapers in Mass Effect harvesting civilizations for inscrutable reasons to the Matrix’s machines farming humans for energy. An AI going rogue and taking over the solar system doesn’t offer fresh insight when we’ve seen these narratives unfold across books, movies, and video games for decades.

More critically, the book contains a glaring omission: no discussion of timelines or pathways to ASI. The authors just sort of wave their hands and assume that it will happen at some point.

Are large language models even the right approach? What if they’re a dead end? What if we never solve hallucinations?

The authors offer no guidance for our current moment, where we have invested trillions of dollars in AI infrastructure. It’s unlikely we’ll just let those investments go to waste. Will we?

The book illuminates the bind decision-makers are already in. Even a small probability of AGI makes development rational from a game-theoretic perspective; it could be a winner-takes-all scenario. Companies pursuing ASI despite risks aren’t delusional. They are responding to competitive pressure; FOMO on steroids. The race dynamics are rational, even if the outcome might be catastrophic. The book uses this bind to show that we risk triggering an uncontrollable, recursively self-improving, non-aligned ASI by simply doing what seems rational in the moment.

The authors acknowledge this dynamic and address it in the book. Here’s a passage from the book:

“Imagine that every competing AI company is climbing a ladder in the dark. At every rung but the top one, they get five times as much money: 10 billion, 50 billion, 250 billion, 1.25 trillion dollars. But if anyone reaches the top rung, the ladder explodes and kills everyone. Also, nobody knows where the ladder ends.”

But why the despair?

Yudkowsky and Soares argue that believing we can solve the alignment problem, as OpenAI and Anthropic claim, represents pure hubris. Reading their work, one senses an almost religious veneration of ASI.

It brings to mind Aquinas:

“This is the ultimate in human knowledge of God: to know that we do not know Him.”

Their thesis remains that it’s better to stop the creation of an unfathomable, unexplainable power than to try bargaining with it.

But Yudkowsky and Soares ignore the trillions already invested. There’s enough AI overhang that resources could shift to deployment and inference optimization rather than capability development. The book offers no practical path forward from where we are, only where we shouldn’t go.

Broader Implications

On a personal note, my father worked at the OPCW for many years, serving as the enforcement arm of the Chemical Weapons Convention. It’s proof that we can coordinate at a global scale and agree that some technologies are best banned and not developed further.

But just as there will always be a North Korea or Syria that ignores conventions and develops chemical weapons anyway, enforcement of an AI moratorium would be extraordinarily challenging. Moreover, a Pyongyang-aligned AGI is a far worse scenario than a localized sarin gas attack. So what can be done?

We’re not dealing with hypothetical future risks but immediate present concerns: the economic disruption and social impact of current AI systems. These challenges require attention now, not after we’ve solved the alignment problem for hypothetical superintelligences.

A More Grounded Alternative

For a more coherent treatment of the current moment, I found Arvind Narayanan and Sayash Kapoor’s “AI as Normal Technology” more compelling than either Yudkowsky and Soares’s urgent doomerism or the e/acc posturing of AI evangelists. While Yudkowsky warns of extinction, leaders like Dario Amodei paint utopian visions, and Sam Altman promises a “gentle singularity,” Narayanan and Kapoor treat AI as a transformative but manageable technology, similar to electricity or the internet before it. Narayanan and Kapoor are writing a book based on the paper, which I look forward to.

Conclusion

You should read “If Anyone Builds It, Everyone Dies.” For a layperson, it effectively lays out the risks and alignment challenges of unchecked AI acceleration in accessible terms. The book succeeds as a provocation and a warning.

But it’s maximalist doomerism that ignores incentive structures and our current technological reality. While serious, it’s not a sufficient treatment of the topic. It fixates on one particular scenario, which the authors consider inevitable, while ignoring where we are today.

The book succeeds at making alignment challenges vivid and accessible. It fails at providing actionable guidance for a world that has already invested trillions in AI infrastructure. We need frameworks for managing the AI we have, not just warnings about the AI we might build.

My P(doom) remains low because I don’t think current foundation models lead to AGI. I suspect we’re in for a significant correction as massive AI infrastructure investments fail to bear fruit. A dot-com-bust style pullback is more likely than runaway exponential growth to ASI.

Moreover, even before we confront ASI, we must deal with the economic and social impact of current AI systems, something Yudkowsky and Soares don’t seem particularly interested in addressing. The apocalypse may not be coming, but the disruption has already begun.

We’re DDoS-ing ourselves with AI Slop.

A DDoS (Distributed Denial of Service) attack overwhelms a scarce resource with a flood of traffic, making it unavailable to its intended users.

We’re DDoS-ing ourselves with AI Slop.

I came across a post on Hacker News that captures this moment. Someone filed what seemed like a comprehensive vulnerability report about cURL, a widely used command-line utility. The entire report was AI-generated and made no sense.

When called out, the reporter published a polished apology that was also clearly AI-generated. (See screenshot).

The entire exchange is surreal, and it wasted the time of someone maintaining tools we all depend on.

My LinkedIn feed is filled with the kind of AI slop that is now easy to detect. Glib prose that says nothing in paragraph after paragraph of polished text. Emails are getting longer, Confluence pages are more verbose, and PRs arrive with hundreds of lines of changes with little explanation.

On forums like Hacker News and technology subreddits, there are posts from leads and managers in absolute despair as they try to cope with this flood.

What we are DDoS-ing is attention.

When attention is not given to reviewing code and providing thoughtful feedback on documentation, the entire ecosystem that is nurtured by attention degrades. Is poisoned.

There is a flip side to this problem. When so many things bear the hallmarks of AI slop, it becomes easy to bring a jaundiced eye to everything we encounter. An em-dash? Slop. It’s not just pervasive, it’s annoying.

I read my posts from a few weeks ago and wonder when I became a slop-peddler.

AI tightens up prose and fixes typos, but it also applies a uniform, flat AI-slop-primer to all output. And it’s not just prose. AI-generated code reads the same. AI-designed websites have the same blue-neon styling.

I am no Luddite. I love using AI, I write about it, and I work on projects that focus on building AI capabilities. AI is a valuable tool that has significantly improved my life.

However, unlike an IDE, there is very little friction in using ChatGPT or Claude. You can write a half-baked, two-sentence prompt, and the AI will enthusiastically go about writing a post, building a website, or submitting a vulnerability report.

As leaders, we need to think carefully about building a culture that encourages both the open-ended exploration of these tools and their disciplined use in day-to-day work.

Otherwise, we are going to DDoS ourselves into a quagmire of AI slop.


Related Posts

How do LLMs understand Gujarati?

One of my favorite ways of testing LLM-powered apps at Jeavio is to ask them questions in transliterated Hindi or Gujarati. I ask questions in Latin script and see how the application responds.

When building chat apps, we are often given instructions by our clients that the bot should only support English. This is an interesting test case on the type of guardrails that our engineers have built into the app.

The more interesting point is why models behave this way 🤔.

Take the screenshot below. Here I ask ChatGPT a question in phonetic Gujarati about Horza, a character in Iain M. Banks’s “Consider Phlebas.” The model understood and responded in phonetic Gujarati. The style was very formal and not quite like how most people speak, but it was recognizable as Gujarati.

Intuitively, you would assume that models are trained on multilingual data and can respond to questions in multiple languages. Gujarati training data in -> Gujarati output out.

However, it is unlikely that a significant amount of Gujarati language analysis of an Iain M. Banks book is available.

There is some interesting research in this space (citations below):

Shared semantic spaces across languages
It appears that LLMs learn a shared semantic space, allowing them to take content from a high-resource language, such as English (with lots of nerdy sci-fi commentary), and respond in a lower-resource language, like Gujarati. The model somehow maps concepts across languages even when direct translations don’t exist in the training data.

Transliteration without explicit training
Gujarati has its own script, of course, so how do models understand transliterated languages? While tokenizers are typically biased toward their training distribution, models appear to learn mappings between transliterated tokens and semantic concepts, despite not being explicitly designed for this purpose. The model figures out that “Horza” in Latin-script Gujarati refers to the same entity as “Horza” in English.

Cross-lingual knowledge alignment
Research also shows that the internal knowledge representations seem to align across languages. This enables translation between language pairs that lack a shared vocabulary. The model builds bridges where none existed before.

Emergent, not designed
This behavior is all emergent! Models weren’t explicitly trained on transliteration pairs or given instructions to handle Latin-script versions of non-Latin languages. They figured it out on their own! Smart models 🧐.

So LLMs are weird.

While we learn more about them, some of their behaviors are still emergent and unpredictable. This makes evaluations extremely important when building with LLMs – something that my team at Jeavio is learning very quickly.

Meanwhile, my patient and generous QA teams continue to tolerate my weird edge cases involving transliterated languages and decades-old science fiction.


Citations (sourced via ChatGPT’s Deep Research mode)

  1. Language Models are Unsupervised Multitask Learners (OpenAI – 2019)
  2. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models (Google Research – 2022)
  3. Crosslingual Generalization through Multitask Finetuning (2023)

Here’s the output from DeepResearch summarizing the results:


Modern multilingual LLMs don’t keep a single “language-free knowledge graph,” but they do learn a shared semantic space that lets them read transliterated inputs (e.g., Latin-script Gujarati), retrieve facts learned mainly from high-resource languages like English, and answer back in the user’s language. This works because their tokenization and modeling are language-agnostic: subword/byte tokenizers (e.g., SentencePiece) and byte-level models (e.g., ByT5) reliably parse mixed or nonstandard text and are robust to spelling/romanization noise.  During pretraining on many languages, the model’s internal representations align across languages—even without a shared vocabulary—so knowledge learned in one language can be accessed from another.  Evidence from multilingual machine translation shows similar “interlingua-like” behavior via zero-shot translation between unseen language pairs.  Probing studies further show that models can recall factual knowledge across languages, though performance varies by language and prompt.  Public resources also provide training/evaluation signal for romanized/code-mixed text (e.g., the Dakshina dataset for Indic languages), which reinforces these abilities.  The caveat: cross-lingual answers are not perfectly consistent—recent analyses find significant variability in factual consistency across language pairs—so quality can be uneven, especially for low-resource languages. 

Why do LLMs hallucinate?

Your palms are sweaty, knees are weak, arms are heavy – it’s not a rap battle but a test for a class which you may or may not have slept through for the entire semester. The test is multiple choice – do you guess or leave the answer blank? There’s no negative marking, so you guess. There’s a one-in-five chance that the answer is correct, and you take it.

I didn’t mean to lose myself in the traumas of my misspent youth 😰, but recent research by OpenAI on why models hallucinate made me take this rather unwelcome trip back to those days.

So, why do models hallucinate?

Let’s take a step back to think about what is happening under the hood.

The first step is called pre-training, where models are taught to predict what word should follow by training on extremely large amounts of text. The problem starts here: some data is very rare or doesn’t exist in the training corpus. Take the birthday of one of the paper’s authors – the model confidently spits out wrong dates because this fact barely appears in training.

The next step usually involves some sort of reinforcement learning (RL). Here, the model is given labelled data and further trained to become more accurate.

OpenAI claims that this training for accuracy is a factor leading to hallucinations. When models are trained to be more accurate, it makes more sense to guess an answer than to say “🤷🏾‍♂️- I don’t know.” After all, a slight chance of being correct is better than a zero chance of being correct.

So, let’s bring this together: LLMs are first trained to predict plausible answers and then further trained to optimize for accuracy on rare facts where they have limited training data. So we end up with behaviors where a model will confidently BS instead of saying “I don’t know.”

OpenAI suggests that we could fix this by changing how we evaluate models, giving them explicit confidence targets in prompts like “Only answer if you’re more than X% confident” and scoring uncertainty appropriately. The article (and associated paper) is worth a read!

MIT Study – 95% of Generative AI Investments Fail (or do they?)

MIT released a study showing that 95% of organizations are getting zero return from their GenAI investments.

While some may claim this proves AI is all hype, a closer reading suggests the findings aren’t a death knell for the technology. Instead, they reveal critical truths about how to succeed.

🎯 It’s Goodhart’s Law writ large
The report shows GenAI adoption is driven by areas like sales and marketing, where success is easier to measure. Pilots are optimized for visible, top-line metrics. However, the study suggests the most dramatic cost savings come from the back office – reducing BPO contracts and agency spend, where the ROI is clear but less flashy.

💡Knowledge and memory are sensitive to each organization
General-purpose AI tools will never work perfectly because each company has its own ontology, its own ways of making sense. Building tools sensitive to this is critical. But there is a contradiction – the study finds that these highly-contextual internal projects fail twice as often as those led by external partners. This is the gap where a strategic partner can help bridge deep internal context with external expertise. (🙋🏾‍♂️ – Jeavio)

🤨 There is a productivity paradox at the heart of GenAI adoption
Workers from over 90% of companies surveyed reported regular use of personal AI tools. If individuals are seeing productivity gains, why does it fail at the aggregate? The report suggests the reason is simple: the most successful AI adoption is bottom-up, not top-down. Successful organizations source initiatives from “frontline managers” and “power users,” not central labs.

At Jeavio, we live this principle. We host hackathons and sponsor open-ended projects to explore how AI can address real-world problems. The ADAPT platform, our flagship AI initiative, began as an internship project in 2023.

Roadtrippin’

Fifteen hours alone in a minivan will take your mind to strange places.
Last weekend, as I drove from the Gulf Shore back home, mine wandered from gas station hot dogs to the future of AI.

📎 It feels like we’re already in a “paperclip maximization” loop.
Each new model is just good enough to justify ongoing jaw-dropping investments in data, compute, and talent. Data center construction now seems to be propping up a flagging US economy. But the benefits of AI don’t yet show up in the numbers.
Is the point of AI simply… to build more AI?

🙏🏾 AI research has often been overtly religious undertones. Kurzweil imagined post-singularity AI as an omnipotent God — the Old Testament kind: awesome, inscrutable, alien.
But maybe we don’t get that.
With so many teams building frontier models, maybe we get something closer to the Hindu pantheon — a whole cast of deities, each with their own agendas. Some awe-inspiring. Others… a little kooky.

🎭 Calling a startup a “ChatGPT wrapper” used to be an insult.
Now I think we’ve all become AI wrappers — sometimes just the meat-interface for LLMs.

🐉 On vacation, my kids and I made up stories:
Pink glitter dragons.
Mean unicorns.
Friendly witches.
Fearsome fairies.
I tried asking ChatGPT for stories, but even the most expensive model couldn’t match my three-year-old’s chaotic creativity. That made me hopeful — because what are humans, if not storytellers?

💀 There’s probably a billion-dollar business in a “dead man’s switch” for chatbots.
An app that erases your entire chat history when you die.
Because I’d rather not be remembered as the guy who once asked ChatGPT why the minivan’s doors wouldn’t close.
(It was a switch. Of course it was.)

Do agentic coding tools make us less productive?

Are agentic coding tools like Cursor making us less productive?

An interesting study from METR (HT The Pragmatic Engineer newsletter) found that use of LLM-assisted coding tools like Cursor made some experienced developers *less* productive than before. The catch? Most of these developers had used Cursor for less than 50 hours.

I’ve been vibe coding a webapp as my summer project using Claude Code and Cursor, and this finding resonates with my experience.

Here’s what I think is happening:

➡️ IDEs are incredibly complex tools with steep learning curves
Fifty hours isn’t nearly enough time to develop the muscle memory needed for effective coding with any new IDE. I’ve stuck with IntelliJ for years precisely because of this learning curve investment, and even now when I’m not evaluating Cursor, I use Claude Code with IntelliJ.

➡️ AI tools are black boxes in ways that aren’t obvious
You might assume Cursor simply passes what you type to the underlying LLM, but these tools use code search, context compaction, and LLM routing behind the scenes. What you type gets significantly modified before reaching the AI, making it nearly impossible to develop a mental model of what’s happening. Cursor updates far more frequently than any IDE I’ve used, and those “optimizations” can change behavior without warning.

➡️ Agentic workflows break the flow state
It’s easy to get distracted when your AI agent is off deciding what file to change next. By the time the agent is done, I’ve often forgotten what I was working on. I honestly don’t understand how people running multiple instances of Claude Code stay productive.

📝 My experience so far
The magic moments when Claude Code one-shots a complex solution are incredible. But like I discovered while switching from Replit’s authentication to Supabase, it can be a grind. The auth switch seemed to work perfectly at first, but broke other parts of my app in subtle ways. Later, migrating the database to Supabase was a disaster – I forgot to update my CLAUDE.md file with the new libraries, so Claude Code got confused and burned through tokens trying to figure out my setup. Technically user error, but exactly the kind of mistake you only learn to avoid through experience.

I expect my productivity will improve significantly as I develop better workflows around these tools. Until then, vibe coding feels like being a new parent – periods of genuine frustration interspersed with moments of sheer amazement.

Agency and LLMs as Mentors

Cate Hall, writing on Substack, suggested that a way to do hard things is to ask “What would someone 10x better do?” and then do it. Reading the post made me realize that agency – the belief that we can ‘just do it’ – is what holds most of us back from success.

I saw this play out again and again growing up in India in the 1980s. You were expected to do as you were told, to not question authority, to focus on getting credentials – degrees, membership of networks, family connections – to get ahead in life. I was lucky to have mentors who encouraged me to read and indulged my curiosity. But millions of kids weren’t so lucky. Who didn’t have anyone to talk to about books, art, movies, and… well… dinosaurs.

What’s forgotten in debates about AI replacing human creativity, taking jobs, or generating slop is that these models can act as expert mentors.

Don’t have anyone to help you with a presentation? Ask AI Simon Sinek to give you feedback. Thinking about how to live your life? Get AI Marcus Aurelius to talk to you about Stoicism. Never written a love letter? Maybe AI Jane Austen could help..

This democratization extends beyond learning into creation itself. Those who sneer at people creating Miyazaki-style images are practicing the same gatekeeping that once kept mentorship exclusive.

Yes, those profile images may seem trite, but they’re also gateways – introducing millions to the worlds of ‘My Neighbor Totoro’ and ‘Spirited Away.’ More importantly, they make it possible for anyone to create, explore, express…

People today have access to world-class mentors via AIs on their phones, regardless of whether they live in Boston or Baroda. These mentors may just be the catalyst to inspire a generation of young people to say, “I can do it.” To embrace agency and set themselves up to do impossible things.

Without gatekeepers, without credentials, without someone telling them what they can or cannot do.

PS – This image is an AI rendering of one of my favorite portraits of my daughter and our dog. Photography was also once looked down upon as “not really art”. Today, it remains one of my favorite hobbies and ways of creative expression.

What is a vibe coder really?

20+ years ago, in my first job (junior business analyst – first class), I developed two critical skills: the ability to type fast and the ability to click around an application until something breaks.

As part of my ongoing exploration into vibe coding, I’ve come to realize that these two skills are more important than anything else I’ve learned since then 🤔.

Touch typing has become my emergency brake for when Claude Code or the replit Agent go bananas, deciding to rewrite 400 lines of perfectly good code when all I wanted was to change a variable name. Being able to quickly jump in and course-correct these enthusiastic agents before they decide to rewrite the entire app and burn a ton of credits has been critical.

The functional QA instincts have been even more crucial. AI coding agents looove to refactor code with the enthusiasm of a junior developer who has just discovered a new frontend framework (why are there so many?? 😭). They’ll elegantly refactor the entire authentication system and accidentally break the login button. Thanks, Claude 👍🏾👍🏾👍🏾!

Test automation assumes some stability in your interfaces, but coding agents treat every function signature like a creative writing exercise. So I find myself writing detailed bug reports as GitHub issues for Claude Code to pick up and fix: “Steps to reproduce: 1. Ask agent to add error handling 2. Agent rewrites entire error handling system 3. Original happy path now throws exceptions.”

My current vibe coding workflow has become suspiciously familiar:
1. Write detailed specs with explicit guardrails
2. Hand over to the developer, sorry – coding agent, and 🙏🏾
3. Do manual testing
4. File detailed bug reports for the agent to fix

Wait – this was my first job! I have become an old-school business analyst again 😬.

Just like baggy jeans and flip phones, late 90s ways of building software are back, baby!