You feed the exact same data and the exact same skill to two different internal AI agents. Agent A tells the client, “Your hotel occupancy last week was 76.3%, up 2.1% month-over-month.” Agent B tells the same client, “Recent performance is good, above market average.”
The client screenshots the contradiction, drops it in the group chat, and asks why your system is fighting itself. The bosses, completely unaware of the engineering nuances, demand absolute data consistency. There is no room for debate. You take the blame.
If you are building AI products, this nightmare is probably already on your calendar. But to fix it, we have to kill the biggest distraction in the industry today.
“Emergence” is just a fig leaf for missing engineering. It’s not a feature; it’s a failure of context management.
Over the last two years, the AI space has violently swung like a pendulum. First, we chased the multi-agent hype. CrewAI and LangGraph made us feel smart. We split our architectures into query agents, diagnostic agents, and action agents. It felt like classic software engineering—modular, clean, familiar.
But then cross-agent communication started dropping semantics. State synchronization broke. Debugging a single case meant parsing logs across three different agents. Google Research found that interaction overhead scales at n^1.724: three agents cost six times the overhead of one. We spent months bleeding on this architecture.
So the pendulum swung back. “Single Agent + Skills” became the new gospel. OpenAI and others proved that one beefy agent with a pile of tools could handle most tasks. The community instantly flipped its narrative.
But swapping the flag didn’t fix the output. You still get sushi from one chef and a stew from another using the exact same knife. Why?
Because everyone is arguing about the wrong thing. The number of agents doesn’t matter. The real watershed is the quality of your Harness engineering.
Agent = Model + Harness. The model provides intelligence; the Harness provides control. When DeepSeek open-sourced their Harness, it sent a massive signal: the model layer is being commoditized. The control layer is the future value high-ground.
You aren’t offering a deterministic API; you are offering a component embedded in a system of uncertainty. If you don’t control the context, you don’t control the product.
Stop thinking of an agent as just a brain with tools. The Harness is a three-layer system. At the bottom, the Connection Layer handles tool calling and API adapters. In the middle, the Capability Layer manages context assembly, memory retrieval, and prompt compression. At the top, the Orchestration Layer handles task breakdown, error recovery, and loop control.
When Agent A gave precise numbers and Agent B gave vague feelings, it wasn’t the Skill’s fault. The Skill is just a part in the Connection Layer. The difference was in the upper layers. Agent A’s team injected user preferences and historical data into the Capability Layer before calling the Skill. Agent B fed the model a bare query. Agent A had strict output templates and validation in the Orchestration Layer. Agent B let the model “freely generate.”
When engineers tell you to “let the model decide,” they are actually saying they have no context management strategy, no output constraints, and no validation mechanisms. You aren’t giving the model freedom; you are giving it nothing.
This completely redefines the Product Manager’s role. In the AI Native era, PMs aren’t obsolete. You must manage deeper. You no longer define UI flows; you define verifiable runtime contracts.
At the Connection Layer, you must dictate the exact tool names, descriptions, and schemas. Change one word in a description, and the model’s call accuracy drops from 90% to 60%. At the Capability Layer, you must define the context contract: what must be included, what should be recommended, and what is strictly forbidden. At the Orchestration Layer, you must define prerequisites and fallback strategies.
But contracts are useless if they aren’t enforced. Traditional PRDs told engineers what to build. AI evaluation sets tell everyone what “done” means.
If your boss asks you to evaluate two competing agents in a hackathon, don’t just test a few cases. Build a structured framework. Test the Connection Layer for tool accuracy. Test the Capability Layer to see if the agent asks clarifying questions when context is missing, or if it just hallucinates. Test the Orchestration Layer with 3-to-5-step tasks to see if it can recover from a broken step.
Unconstrained freedom is chaos. Constrained freedom is intelligence.
The best AI products—Claude’s Computer Use, Cursor’s code generation, Perplexity’s search—don’t just chase tech narratives. They do the exact same thing: they give the model maximum freedom within strictly defined boundaries.
Architectures will change. Models will iterate. Organizations will restructure. But the user’s demand for consistent, trustworthy output will never change. Stop chasing the tech pendulum. Find the unchanging baseline, define the contract, and hold the line.
FAQ
Q: Isn't 'emergence' just how AI naturally works?
A: No, it's an excuse. In deterministic product scenarios, 'emergence' is a fig leaf for missing engineering. If you can't predict the output, you haven't built the context management or validation layers.
Q: What's the practical implication for Product Managers?
A: Your job is shifting from designing UI flows to defining verifiable runtime contracts. You must dictate exactly what context the model sees, what tools it calls, and how the output is formatted. Evaluation sets are the new PRD.
Q: What's the contrarian take?
A: Model companies open-sourcing their Harness scaffolding is a signal that the model layer is being commoditized. The real value and future moat isn't in the AI model itself, but in the control layer surrounding it.