The world is too loud. Read what matters.

This Week in Startups

90% of AI prototypes die at POC: what's missing isn't the model, it's the harness

No matter how strong the model is, it can't survive a mid-run crash — the process dies halfway through issuing a refund and the customer gets paid twice. Pulling state out of the agent and handing it to a harness is the line between the laptop and production.

AI agentsdurable executionproductionizationdistributed systemsenterprise security

The video won't play here. Listen to the audio instead:

19 minutes to explain an overlooked engineering truth: the bottleneck for putting agents into production isn't model capability, it's durable execution. For teams pushing a demo toward production.

The argument · tap a timestamp to hear it

2:05

90% of AI prototypes die after POC

Samar says anyone can now turn an idea into an app with a coding agent, but 90% of ideas die after POC and never see the light of day. The cause of death isn't that the model isn't smart enough — it's that they're "unstable, brittle, unreproducible." Jason says he sees this play out in organizations every day: someone excitedly demos something, he asks "can it go to production," and the answer is always "I don't know, I feel like no, but maybe." That gap between POC and production-grade software is what this episode takes apart.

— Samar Abbas
3:08

The more valuable the agent, the more it looks like a distributed system

Samar's observation: once an AI application goes to production, it runs longer and longer, gets more and more asynchronous, takes more and more actions in the real world, and creates more value. But the moment you get there, you run into infrastructure-level challenges — reliability, scalability, durability — exactly like the cloud platform migration era. His judgment: these problems are fundamentally distributed, not brand-new AI problems. Once an agent graduates from the laptop into a genuinely distributed environment, the nature of the problem changes.

— Samar Abbas
5:21

The refund crashes halfway through — do you send a second one?

Samar's picture is concrete: an agent is processing a refund, has just finished the refund but hasn't yet notified the customer, and the machine restarts or the process crashes. Now you have only two options — either rerun it, or send another refund. Jason chimes in: that could get expensive. Samar says handling this class of failure is exactly where every production system gets stuck, because AI is starting to take actions in the real world. What Temporal provides is called durable execution: when a failure happens during code execution, the platform remembers all the state for you, the developer doesn't write a single line of code, and the application still moves forward through all kinds of failures.

— Samar Abbas
7:30

Show the agent's HUD to a human

After looking at Temporal's interface, Jason says the top left is the workflow code, the right is the runtime event timeline where you can watch the agent execute tools, call the LLM, get responses back, and move forward step by step in real time, and the bottom is the output the application actually produces. He says he's never seen anyone expose the behind-the-scenes process, and he wishes every consumer product had a HUD like this so he'd have more confidence about "where did I get to last time." Samar adds: Temporal isn't just a transaction engine, it also gives you full visibility into agentic applications — an agent is just picking an LLM, giving it a prompt, and configuring a set of tools, with the LLM deciding how to proceed on its own.

— Jason Calacanis
11:52

Code mode keeps security teams up at night

Samar raises an underrated risk: there's now a code mode where the agent generates more code on the fly while solving a problem and executes it. Imagine a large enterprise handing a big business process to this kind of code mode agent — you're effectively running code the LLM spits out at runtime straight into the business environment, which from a security standpoint is deeply unsettling. Temporal addresses this with two constructs, workflows and activities: workflows are responsible for generating commands, and the platform can intercept those commands and add the necessary guardrails before they actually enter the sandbox or runtime.

— Samar Abbas
13:02

The harness is the step from MS-DOS to the cloud

Samar uses an analogy: we're now transitioning from the MS-DOS era of agents to a genuine cloud environment. The way most people run agents today is to install a coding agent on a laptop and run it. But people have already started building loops, and terms like loop engineering and graph engineering have even appeared, meaning agents will run more independently and for longer, and will necessarily move off the laptop into a distributed environment. At that point the harness becomes the core component — it's the "brain" that gets pulled out and placed outside the agentic loop, so that when the agent goes down the wrong path it doesn't crash the whole process.

— Samar Abbas
16:01

Enterprises will never allow agents to run on a laptop

Jason asks: how do you decide it's time to move the agent you're playing with on your desktop into the production loop? Samar's answer is hard: enterprises simply cannot adopt these agent architectures unless they can run business processes under the required guardrails, so no enterprise will allow agents to run on a laptop — for an enterprise this is table stakes. He offers a more general graduation signal: when you go from one person solving a problem to a team solving a complex business problem, the only viable path is to run the agent in a distributed environment. At that point a forward deployed engineer has to be embedded in the business unit, building it right with the right tools so it isn't brittle.

— Samar Abbas

In their own words · checked verbatim

But 90% of those ideas die after a POC, essentially. They never see light of the day.

Samar Abbas2:05

these problems start to look more and more similar to like distributed in nature rather than a completely new AI problem

Samar Abbas3:08

Imagine an agent which is kind of building, processing a refund and literally, immediately after processing a refund, a failure happens.

Samar Abbas5:21

Where during an execution of a code, if a failure happens, we remember all of that state. Without you as a software developer writing a single line of code for it.

Samar Abbas6:26

i feel right now we are transitioning from an ms dos era of agents to a real cloud environment essentially

Samar Abbas13:02

which means like none of the enterprises will ever allow these agents to run on a laptop, essentially. So for them, it's core table stakes, in my opinion.

Samar Abbas16:01

I think at this point, running agents on your laptop is no longer even an option.

Samar Abbas17:03

Figures

Share of AI ideas that die after POC90%2:05

Glossary

harness
The "brain" pulled out of the agent loop, responsible for coordination, durability and fault tolerance.
durable execution
When a failure happens during code execution, all state is saved automatically so the application keeps moving forward.
code mode
A way of working where the agent generates and executes new code on the fly while solving a problem.
forward deployed engineer
An engineer embedded in a business unit who builds the agent system to production standards with the right tools.

How to listen

Who it's for

Engineering leads pushing an AI demo toward production, platform architects, and CTOs who need to put guardrails on agents.

Skip

0:00-1:00 opening intro and guest background, skippable.