The Agent Harness Revolution: what open vs. closed AI means for your organisation

The Agent Harness Revolution: What Open vs. Closed AI Means for Your Organisation

The Engine Isn't the Car

Most of the AI conversation in boardrooms today is about the model — which frontier model to subscribe to, which API to call, whose weights to worry about. That's understandable, but it's also, increasingly, a distraction.

Think of a car. A Ferrari engine bolted onto a wooden cart doesn't give you a Ferrari. The engine matters, but what actually determines whether you arrive anywhere useful is the whole vehicle: the chassis, the steering, the gearbox, the driver. In AI, the model is the engine, and the thing that wraps it into something that can actually do work is the harness — the software that gives the model memory, tools, files, a terminal, a way to plan, act, and correct itself.

This is the agent harness revolution. The battleground for AI is no longer (only) who has the smartest engine. It's who controls the vehicle that turns raw intelligence into finished, dependable work.

The 2025-2026 Milestone: Agents Stop Being Demos

Before 2024, "AI agents" were largely research curiosities — impressive on a benchmark slide, unreliable in production. In 2025 and 2026 that changed. Agentic systems moved from proof-of-concept to real spending. Let's put numbers on it, because the shift isn't vibes; it's measurable.

These are not small numbers hiding in a niche. This is infrastructure money. And infrastructure money flows toward whoever owns the harness.

The Open Players

If you are of the "open is the real revolution" camp, the good news is that 2025-2026 has been a historic stretch for open-weights and open-source agents. A few anchors:

Hermes Agent, by Nous Research. The single cleanest example of a fully open harness. It is an agent framework that runs on top of open-weights models (Hermes, DeepSeek, Qwen, Llama — whatever you point it at), and it is the vehicle used to write many of the posts on this site. The repo lives at github.com/NousResearch/hermes-agent. Open harness, open models, no gatekeeper. This is the model-vs-harness argument made concrete: Nous does excellent open weights, and Hermes Agent is the chassis that turns them into work.

DeepSeek's open-weights models. DeepSeek-V3 is the reference point for why "you need billions to train frontier models" is no longer a complete sentence. From the DeepSeek-V3 technical report: it is a Mixture-of-Experts model with 671 billion total parameters (37B activated per token), and the full training run (pre-training, context extension, fine-tuning) required only 2.788M H800 GPU hours. At commonly-cited $2/GPU-hour pricing, that works out to roughly US$5.6 million for the entire training run — compare that to the hundreds of millions often associated with frontier-scale training. Two caveats that the reporting usually leaves out and you shouldn't: (1) the figure only covers the GPU compute cost, not researchers, data curation, the R&D that produced the architecture, or the RL stages; and (2) it was called misleading by some outside observers for exactly that reason. But even with those caveats, the number reshaped the entire cost-of-frontier conversation in January 2025, and it remains the single most-cited figure on open-weights economics.

The open ecosystem. It is not just one champion. The stack is rich and maturing:

The point isn't to name every repo. The point is that the open side now has every layer of the stack — models, harness, tools, memory, graph, local inference — maintained by real, identifiable teams, most of it Apache/MIT licensed and auditable. Nothing is a black box you have to take on faith.

The Closed Players

The closed side is dominated by the two coding agents that defined the category and by the harnesses around the frontier APIs.

Claude Code, by Anthropic (github.com/anthropics/claude-code, ~143k stars). The most famous harness in the world right now. It sits on top of Anthropic's frontier models, is deeply integrated with the terminal and with code, and its marketing has been aggressive and effective. People pay real money for Claude Code subscriptions because the harness does real work.

OpenAI Codex (github.com/openai/codex, ~117k stars). OpenAI's answer, wrapped around GPT-series frontier models, also terminal-based, also capable of shipping code. OpenAI's strategy is to make the harness feel like the natural thing you reach for when you already pay OpenAI.

The pattern is identical on both sides: closed harness = one vendor owns model + harness + the tool integration + the data. You are renting the whole car from the dealership, and the dealership can change the engine, the steering, or the price at any time.

Open vs. Closed: The Head-to-Head

This is the decision your organisation will actually make. Here are the eight axes that matter, side by side.

  1. Cost. Open: pay for compute + people; the model runs on your hardware, marginal cost per run approaches zero once you own the boxes. Closed: per-token/per-seat pricing you rent forever; as usage scales, so does your bill, and the vendor sets the rate.
  1. Data privacy. Open: data stays on your hardware; you can run air-gapped. Closed: data transits the vendor's cloud; whatever you let the model see is, operationally, in the vendor's datacentre.
  1. Portability. Open: you can move models, swap engines, keep the harness. You are not locked in. Closed: harness is tuned to one model; switching models means switching harnesses and rebuilding the tooling.
  1. Control / auditability. Open: you can read the code, patch it, self-host it, and know exactly what it does — important for regulated industries. Closed: black box; you depend on the vendor's promises and roadmap.
  1. Speed of innovation. Closed wins this axis. Anthropic and OpenAI ship polished features weekly, backed by enormous engineering budgets. The open ecosystem is catching up fast but it's a community-maintained marathon, not a sprint.
  1. Support. Closed wins here too: a vendor, SLAs, a support channel, a roadmap you can hold them to. Open: community forums, the repo, and your own people — great if you have them, a problem if you don't.
  1. Capability ceiling. As of late 2025, the frontier closed models still hold a real edge at the very top of raw intelligence, especially in hard reasoning and long-horizon tasks. The gap narrows every quarter, and for many ordinary tasks it's already academic.
  1. Risk profile. Open = technology risk (the harness can drift), but vendor risk is low. Closed = you've traded technology risk for vendor dependence: a price hike, an API change, or a model you don't like is a business problem, not just a technical one.

There is no universal winner. The right answer is "it depends on the axis your organisation most fears."

Deployment Strategies

You don't have to bet the house on one side. The pragmatic organisations use these four patterns:

The Outlook

A few things worth betting on.

The Mark Cuban integrator thesis. Cuban has said, essentially, that the value isn't in the model — it's in knowing how to apply the technology to a business problem. The big wins will go to integrators and domain people who take off-the-shelf engines and harnesses and point them at specific business problems, not to whoever owns the next weight-update. If the harness is the car, the integrator is the mechanic and the driver.

The $500-5,000 digital employee. The done-for-you market matures: agents sold as products, with setup, like consultants, but then they stay and do the work. This creates a huge skills re-organisation in-house: fewer people to do routine work, more people to define, supervise, and correct agents.

Hybrid stacks become the default. Few organisations will be all-open or all-closed. The sensible default is a hybrid: open models and harnesses for control and cost, closed frontier models where the marginal intelligence is worth the money.

Agent metrics become the discipline. Once agents make 15% of decisions autonomously, governance is everything: audit trails, evaluation harnesses, fallbacks, guardrails. The organisations that succeed will be the ones that treat the harness as infrastructure to be managed, not as a shiny thing to be demoed.

Why Every Employee Must Understand This

This is not a topic only for engineers. Your frontline people are the ones who will be asked to hand work to an agent, and to judge whether the agent's output is trustworthy. They need a mental model of what the agent can and cannot do, and they need to know the difference between renting your car from a vendor and owning it. When an employee says "the AI made a mistake," the first question should be: was that a model failure or a harness failure? They're different failures with different fixes, and only people who understand both can tell them apart. Literacy here is not a luxury; it's a risk-control skill.

Why It Matters for the Public Sector

The public sector runs on procurement, compliance, and data protection. For a public agency that needs reliable, auditable, and legally defensible automated processes, the open-vs-closed choice is not neutral. Open weights and open harnesses let you inspect the system, keep citizen data on your own hardware, and audit the decision pipeline. Closed systems push the data and the decision-making into a vendor’s cloud, which creates legal and trust complications that a public body cannot hand-wave away. If you are going to put a model in a state-mandated, regulated process, you need to be able to say what it does and why. That is far easier with an open, auditable harness than with a black box.

The Bottom Line

The engine wars got all the attention; the harness war decides who owns the road. Organisations that treat AI as "subscribe to the smartest model" will find themselves renting their entire vehicle from a vendor that can change the terms any month. Organisations that build or buy their own harness — and treat models as pluggable engines — get to keep the road. The open revolution was never really about the model. It was always about who drives.


Sources