AI, Friction, and Credibility

I got caught in a thunderstorm just outside the town of Lovettsville, VA on Friday. I pulled into a gas station and spent thirty minutes in the company of a taciturn storekeeper while I waited for the sound and light show to end. In the end I gave up and rode home in pouring rain, arriving soggier than a neglected bread pudding.

Riding motorcycles is stupid. It is dangerous, and my minivan is a way more comfortable way to see the world.

And yet – I love my motorcycles. I love feeling the engine braking as I downshift into a curve. I don’t even mind the bugs making a mess of my visor as I ride through the Virginia Spring.

It is precisely the discomfort and friction involved in riding motorcycles that makes it a compelling experience. It’s the same friction involved in lifting weights or in figuring out a particularly nasty software bug.

Friction leads to new experiences, and to growth.

And yet, AI is being deployed across knowledge work to eliminate friction entirely. Writing an email – ChatGPT can help. Software – just vibe code through it. Trying to understand a complex topic – ask Gemini to summarize.

But there are downstream effects.

The more AI gets used in day-to-day work, the more it becomes clear that we’re not eliminating friction, but instead just displacing it (HT to Rohit Krishnan – see below).
The friction shifts from the developer to reviewer – who now must deal with 10 PRs a day instead of 3.
It shifts from the product manager to the development team – who now must deal with a firehose of AI-generated User Stories.

Friction builds credibility.

There is a reason why doping is such a taboo in professional sports. Lance Armstrong incinerated his credibility when the allegations of widespread doping turned out to be accurate.

Credibility remains the only viable currency in a world where AI can do the heavy lifting of knowledge work. And today, there is no better way of incinerating it than passing off low-effort AI slop as your work.

Social mores will evolve as we become used to AI tools. It’s very likely that as the models get better, we’ll just embrace this as the new way of working and laugh at pieces like this. AI is surely just the next stage of knowledge work – following calculators and Excel.

My minivan is superior to my motorcycle and yet, I remember my motorcycle rides more than I do car trips. It’s because the discomfort, the danger, and the friction contribute to my own growth. This is something worth thinking about as we embrace AI.


This post was inspired by two very thought provoking posts – many thanks to Rohit Krishnan and Kyla Scanlon. Check out their Substacks

Building hyper-personalized AI Apps..

I’m a PowerPoint jockey and a very rusty programmer. Yet, over the weekend, I built something that had been an idea for years.

I write constantly – notes, emails, journals – using writing to process thoughts and help calm the chaos in my head. But I couldn’t find a journaling tool that worked exactly as I wanted: private, organized, tagged, and summarized with my specific quirks.

So, I built a custom workflow using Claude and MCP. I dump thoughts into Claude via text or voice. It offers prompts for elaboration, generates metadata and tags, creates markdown files, and pushes everything to my private GitHub repo. Claude even helped write a GitHub action to maintain an index whenever I created a new entry.

Extremely nerdy? Absolutely. Could I have used an off-the-shelf app? Probably. But building something that behaved *exactly* how I wanted is why LLMs excite me.

This represents something bigger than one nerdy weekend project. What required technical knowledge today will soon be accessible to everyone. What took me a weekend of GitHub repos and MCP servers will quickly be declarative workflows in mainstream tools.

We’ve spent decades accepting software uniformity. SaaS companies optimize for the broadest user base, creating generic interfaces. We adapt our workflows to software constraints rather than software adapting to us.

LLMs could flip this script. Instead of adapting ourselves to software, software could adapt to us. Hyper-personalized workflows, bizarre interfaces, and tools that match how we think and work.

Remember when personal computers were personal? Before everything became a web app that looked exactly like every other web app? We might be heading back there, but everyone gets to be the programmer this time.

—


Screen grab is from the metadata of an early draft of this post. You can find a link to the project prompt in the comments.

AI safety in 2025

Two years ago, the Future of Life Institute called for a 6-month pause in AI development due to fears of misalignment. Elon Musk and hundreds of other luminaries signed the letter.

Instead, AI research accelerated. Foundation models now have capabilities that make GPT-3 era models look like toys.

And we’re seeing some concerning emergent behaviors. Anthropic’s safety testing of Claude 4 revealed some interesting behaviors.
➡️ When researchers implied the model would be replaced, it attempted to blackmail the fictional engineers by threatening to reveal personal information
➡️ When placed in scenarios involving user wrongdoing and told to “take initiative,” it frequently took extreme actions including locking users out of systems
➡️ Researchers noted that Claude 4 Sonnet seems to “care a lot about animal rights” while Claude 4 Opus doesn’t. They can’t explain why.


I appreciate Anthropic making this research public. It highlights how difficult it is to interpret LLM behavior – even for their creators.

As we build applications on these capabilities, we need to acknowledge the “capabilities overhang” – we haven’t fully explored what these systems can do! However, we must also acknolwedge that the risks are emerging faster than our understanding.


The pause letter asked the right question: are we moving too fast? Two years later, with models exhibiting goal-oriented behavior we can’t explain, that question feels more urgent than ever.

Dispatches from Mediocristan

LLMs are powerful tools – but credulous users risk being stuck in a dangerous place: Mediocristan, the land of the average.

Mediocristan appears in Nassim Nicholas Taleb’s Incerto series. It’s a domain where outcomes are predictable, smooth, and derived from averaging all inputs.

Sound familiar?

LLMs predict the most likely next token based on massive training data (yes, yes – I know about RLHF, etc.). They are statistical engines of mediocrity by design.

And like it or not, LLM use pushes us deeper into Mediocristan daily.

A recent viral piece in NY Magazine exposed how university students rely utterly on ChatGPT. But it’s hardly limited to academia—I’ve encountered memos, emails, and pitch decks that bear the unmistakable hallmarks of AI slop.

We’re outsourcing our thinking to Mediocristan with great enthusiasm.

On the other side lies Extremistan—the domain of consequential outliers where one event’s probability is uncorrelated with another. Mathematically, it’s the fat tails of distributions where Black Swans lurk.

Extremistan is where interesting and unexpected things happen—where growth and destruction co-exist. The very release of ChatGPT in 2022 was itself an event straight from Extremistan!

I’m as enthusiastic an LLM user as any, but comparing my writing from 2020 to today, I’m clearly on the express train to Mediocristan.

This realization is jarring. So what now?
Should we embrace the slop and relocate to Mediocristan?
Angrily denounce AI and revert to writing screeds on clay tablets?

The critical skill for navigating our new knowledge economy will be deciding where and how to use AI.

Meanwhile, Mediocristan steadily expands, assimilating new domains and making them ripe for disruption from—you guessed it—Extremistan.

On DeepSeek

Is it really doomsday for U.S. AI companies? The harbinger of the apocalypse appears to be a blue whale.

Nvidia’s stock is down 12.5%. There’s a broad tech sell-off, and Big Tech seems a little uneasy.

The reason? A Chinese hedge fund built and trained a state-of-the-art LLM to give their spare GPUs something to do.

DeepSeek’s R1 model reportedly performs on par with OpenAI’s cutting-edge o1 models. The twist? They claim to have trained it for a fraction of the cost of models like GPT-4 or Claude Sonnet—and did so using GPUs that are 3-4 years old. To top it off, the DeepSeek API is priced significantly lower than the OpenAI API.

Why did this trigger a sell-off of Nvidia (NVDA)?

  • It shows that building cutting-edge models doesn’t require tens of thousands of the latest Nvidia GPUs anymore.
  • DeepSeek’s models run at a fraction of the cost of large LLMs, which could shift demand away from Nvidia’s high-end hardware.

For U.S. companies, this is a wake-up call. The Biden-era export restrictions didn’t have the intended impact. But for anyone building on AI, there’s a silver lining:

  • Building LLMs and reasoning models is no longer limited to companies throwing billions at compute.
  • This will likely kick off an arms race as U.S. companies race to optimize costs and stay competitive with DeepSeek.
  • Data sovereignty will still matter—most companies won’t want their data processed by a Chinese-hosted model. If DeepSeek’s approach proves viable, expect U.S. providers to replicate it.

Categories AI Tags

Melanie Mitchell on the Turing Test

From “The Turing test and our shifting conceptions of intelligence” by Melanie Mitchell.

In her insightful piece, “The Turing Test and our shifting conceptions of intelligence,” Melanie Mitchell challenges the traditional view of the Turing Test as a valid measure of intelligence. She argues that while the test may indicate a machine’s ability to mimic human conversation, it fails to assess deeper cognitive abilities, as demonstrated by the limitations of large language models (LLMs) in reasoning tasks. This prompts us to reconsider what it truly means for a machine to think, moving beyond mere mimicry to a more nuanced understanding of intelligence.

Our understanding of intelligence may be shifting beyond what Turing initially imagined.

From the article:

On why Turing initially proposed the Turing Test

Turing’s point was that if a computer seems indistinguishable from a human (aside from its appearance and other physical characteristics), why shouldn’t we consider it to be a thinking entity? Why should we restrict “thinking” status only to humans (or more generally, entities made of biological cells)? As the computer scientist Scott Aaronson described it, Turing’s proposal is “a plea against meat chauvinism.”

A common criticism of the Turing Test as a measure of AI capability

Because its focus is on fooling humans rather than on more directly testing intelligence, many AI researchers have long dismissed the Turing Test as a distraction, a test “not for AI to pass, but for humans to fail.”

Generative Models and the “Grey Goo Problem”

Generative AI models may be causing a “Grey Goo” problem with art, publishing, and user-generated content. 

Thomas Jane encounters the Protomolecule in The Expanse

The Grey Goo Problem is a thought experiment where self-replicating nano-robots consume all available resources leading to a catastrophic scenario. This scenario is a popular science fiction trope (see comments).

Several publishers and user-generated content sites like StackOverflow have been impacted by a flood of AI-generated content in the last few months. Clarkesworld, a science fiction magazine, stopped accepting submissions last week. Even LinkedIn is overrun by ChatGPT-generated “thought leadership.” 

Tools like ChatGPT need high-quality training data to generate good results. They collect training data by scraping the Internet. You can see the issue here, can’t you? 

The Grey Goo scenario is managed through containment and quarantine in science fiction. For example, in The Expanse series (see image), containing the “Proto-Molecule” is a crucial plot element. 

The need to contain and quarantine Generative AI will result in more paywalls, subscriptions, and gated content. Crypto may even find its calling in guaranteeing the authenticity of online content. 

I fear that the Open Internet that made ChatGPT possible will be crippled by the actions of ChatGPT and its cousins.

Book Review: “Artificial Intelligence – A Guide for Thinking Humans” by Melanie Mitchell

Artificial Intelligence – A Guide For Thinking Humans

Introduction

Melanie Mitchell’s book “Artificial Intelligence – A Guide for Thinking Humans” is a primer on AI, its history, its applications, and where the author sees it going. 

Ms. Mitchell is a scientist and AI researcher who takes a refreshingly skeptical view of the capabilities of today’s machine learning systems. “Artificial Intelligence” has a few technical sections but is written for a general audience. I recommend it for those looking to put the recent advances in AI in the context of the field’s history.

Key Points

“Artificial Intelligence” takes us on a tour of AI – from the mid-20th century, when AI research started in earnest, to the present day. She explains, in straightforward prose, how the different approaches to AI work, including Deep Learning and Machine Learning, based approaches to Natural Language Processing. 

Much of the book covers how modern ML-based approaches to image recognition and natural language processing work “under the hood.” The chapters on AlphaZero and the approaches to game-playing AI are also well-written. I enjoyed these more technical sections, but they could be skimmed for those desiring a broad overview of these systems. 

This book puts advances in neural networks and Deep Learning in the context of historical approaches to AI. The author argues that while machine learning systems are progressing rapidly, their success is still limited to narrow domains. Moreover, AI systems lack common sense and can be easily fooled by adversarial examples. 

Ms. Mitchell’s thesis is that despite advances in machine learning algorithms, the availability of huge amounts of data, and ever-increasing computing power, we remain quite far away from “general purpose Artificial Intelligence.” 

She explains the role that metaphor, analogy, and abstraction play in helping us make sense of the world and how what seems trivial can be impossible for AI models to figure out. She also describes the importance of us learning by observing and being present in the environment. While AI can be trained via games and simulation, their lack of embodiment may be a significant hurdle towards building a general-purpose intelligence.

The book explores the ethical and societal implications of AI and its impact on the workforce and economy.

What Is Missing?

“Artificial Intelligence” was published in 2019 – a couple of years before the explosion in interest in Deep Learning triggered due to ChatGPT and other Large Language Models (LLMs). So, this book does not cover the Transformer models and Attention mechanisms that make LLMs so effective. However, these models also suffer from the same brittleness and sensitivity to adversarial training data that Ms. Mitchell describes in her book. 

Ms. Mitchell has written a recent paper covering large language models and can be viewed as an extension of “Artificial Intelligence.”

Conclusion

AI will significantly impact my career and those of my peers. Software Engineering, Product Management, and People Management are all “Knowledge Work.” And this field will see significant disruption as ML and AI-based approaches start showing up. 

It is easy to get carried away with the hype and excitement. Ms. Mitchell, in her book, proves to be a friendly and rational guide to this massive field. While this book may not cover the most recent advances in the field, it still is a great introduction and primer to Artificial Intelligence. Some parts of the book will make you work, but I still strongly recommend it to those looking for a broader understanding of the field.

Machine Learning and its consequences

Machine Learning has brought huge benefits in many domains and generated hundreds of billions of dollars in revenue. However, the second-order consequences of machine learning-based approaches can lead to potentially devastating outcomes. 

This article by Kashmir Hill in the New York Times is exceptional reporting on a very sensitive topic – the identification of abusive material or CSAM. 

As the parent of two young children in the COVID age, I rely on telehealth services and friends who are medical professionals to help with anxiety-provoking (yet often trivial) medical situations. I often send photos of weird rashes or bug bites to determine if it is something to worry about.  

In the article, a parent took a photo of their child to send to a medical professional. This photo was uploaded to Google Photos, where it was flagged as being potentially abusive material by a machine learning algorithm. 

Google ended up suspending and permanently deleting his Gmail account and his Google Fi phone and flagging his account to law enforcement. 

Just imagine how you might deal with losing both your primary email account, your phone number, and your authenticator app. 

Finding and reporting abuse is critical. But, as the article illustrates, ML-based approaches often lack context. A photo shared with a medical professional may share similar features to those showing abuse. 

Before we start devolving more and more of our day-to-day lives and decisions to machine learning-based algorithms, we may want to consider the consequences of removing humans from the loop.