The Agent Harness Revolution: what open vs. closed AI means for your organisation
The Agent Harness Revolution: What Open vs. Closed AI Means for Your Organisation
The Engine Isn't the Car
Most of the AI conversation in boardrooms today is about the model — which frontier model to subscribe to, which API to call, whose weights to worry about. That's understandable, but it's also, increasingly, a distraction.
Think of a car. A Ferrari engine bolted onto a wooden cart doesn't give you a Ferrari. The engine matters, but what actually determines whether you arrive anywhere useful is the whole vehicle: the chassis, the steering, the gearbox, the driver. In AI, the model is the engine, and the thing that wraps it into something that can actually do work is the harness — the software that gives the model memory, tools, files, a terminal, a way to plan, act, and correct itself.
This is the agent harness revolution. The battleground for AI is no longer (only) who has the smartest engine. It's who controls the vehicle that turns raw intelligence into finished, dependable work.
The 2025-2026 Milestone: Agents Stop Being Demos
Before 2024, "AI agents" were largely research curiosities — impressive on a benchmark slide, unreliable in production. In 2025 and 2026 that changed. Agentic systems moved from proof-of-concept to real spending. Let's put numbers on it, because the shift isn't vibes; it's measurable.
- The AI agents market was worth roughly US$7.92 billion in 2025 and is forecast to climb to about US$294.66 billion by 2035, a compound annual growth rate of 43.57% from 2026 to 2035 (Precedence Research, AI Agents Market, 2026).
- The narrower agentic AI market was valued at about US$7.55 billion in 2025, up from US$5.25 billion in 2024, and is projected to reach US$199.05 billion by 2034, a CAGR of roughly 43.84% (Precedence Research, Agentic AI Market, 2025).
- Gartner now forecasts that AI will autonomously make 15% of day-to-day work decisions by 2028 — a threshold that has direct, practical consequences for how work gets staffed, governed, and audited in any organisation that isn't in the AI-skeptic minority.
- On the spend side, Gartner projects worldwide end-user spending on AI platforms and models to reach US$64 billion in 2026, up 63.4% from about US$39 billion in 2025 — a sign that enterprises are shifting from experimentation to scaled, budgeted deployment.
These are not small numbers hiding in a niche. This is infrastructure money. And infrastructure money flows toward whoever owns the harness.
The Open Players
If you are of the "open is the real revolution" camp, the good news is that 2025-2026 has been a historic stretch for open-weights and open-source agents. A few anchors:
Hermes Agent, by Nous Research. The single cleanest example of a fully open harness. It is an agent framework that runs on top of open-weights models (Hermes, DeepSeek, Qwen, Llama — whatever you point it at), and it is the vehicle used to write many of the posts on this site. The repo lives at github.com/NousResearch/hermes-agent. Open harness, open models, no gatekeeper. This is the model-vs-harness argument made concrete: Nous does excellent open weights, and Hermes Agent is the chassis that turns them into work.
DeepSeek's open-weights models. DeepSeek-V3 is the reference point for why "you need billions to train frontier models" is no longer a complete sentence. From the DeepSeek-V3 technical report: it is a Mixture-of-Experts model with 671 billion total parameters (37B activated per token), and the full training run (pre-training, context extension, fine-tuning) required only 2.788M H800 GPU hours. At commonly-cited $2/GPU-hour pricing, that works out to roughly US$5.6 million for the entire training run — compare that to the hundreds of millions often associated with frontier-scale training. Two caveats that the reporting usually leaves out and you shouldn't: (1) the figure only covers the GPU compute cost, not researchers, data curation, the R&D that produced the architecture, or the RL stages; and (2) it was called misleading by some outside observers for exactly that reason. But even with those caveats, the number reshaped the entire cost-of-frontier conversation in January 2025, and it remains the single most-cited figure on open-weights economics.
The open ecosystem. It is not just one champion. The stack is rich and maturing:
- GraphRAG (Microsoft, github.com/microsoft/graphrag) — graph-structured retrieval for RAG, so a harness can ground answers in real relationships, not just keyword hits.
- Neo4j — the property-graph database underpinning graph-based retrieval.
- Text2CypherRetriever — lets the model generate Cypher queries against the graph, turning a language model into a query engine.
- Ollama (github.com/ollama/ollama, ~179k stars) — the easiest way to run open weights locally; one command, a local OpenAI-compatible endpoint, zero cloud.
- LM Studio — a polished desktop runtime for running local models offline.
- Mem0 (github.com/mem0ai/mem0, ~64k stars) and Honcho (github.com/plastic-labs/honcho) — open memory layers, giving agents long-term persistence rather than a blank slate each turn.
- Composio (github.com/ComposioHQ/composio, ~30k stars) — an open tool/action layer connecting agents to hundreds of external apps.
- OpenCode (github.com/anomalyco/opencode, ~200k stars) — an open, terminal-based coding agent that runs against any model, including local ones.
The point isn't to name every repo. The point is that the open side now has every layer of the stack — models, harness, tools, memory, graph, local inference — maintained by real, identifiable teams, most of it Apache/MIT licensed and auditable. Nothing is a black box you have to take on faith.
The Closed Players
The closed side is dominated by the two coding agents that defined the category and by the harnesses around the frontier APIs.
Claude Code, by Anthropic (github.com/anthropics/claude-code, ~143k stars). The most famous harness in the world right now. It sits on top of Anthropic's frontier models, is deeply integrated with the terminal and with code, and its marketing has been aggressive and effective. People pay real money for Claude Code subscriptions because the harness does real work.
OpenAI Codex (github.com/openai/codex, ~117k stars). OpenAI's answer, wrapped around GPT-series frontier models, also terminal-based, also capable of shipping code. OpenAI's strategy is to make the harness feel like the natural thing you reach for when you already pay OpenAI.
The pattern is identical on both sides: closed harness = one vendor owns model + harness + the tool integration + the data. You are renting the whole car from the dealership, and the dealership can change the engine, the steering, or the price at any time.
Open vs. Closed: The Head-to-Head
This is the decision your organisation will actually make. Here are the eight axes that matter, side by side.
- Cost. Open: pay for compute + people; the model runs on your hardware, marginal cost per run approaches zero once you own the boxes. Closed: per-token/per-seat pricing you rent forever; as usage scales, so does your bill, and the vendor sets the rate.
- Data privacy. Open: data stays on your hardware; you can run air-gapped. Closed: data transits the vendor's cloud; whatever you let the model see is, operationally, in the vendor's datacentre.
- Portability. Open: you can move models, swap engines, keep the harness. You are not locked in. Closed: harness is tuned to one model; switching models means switching harnesses and rebuilding the tooling.
- Control / auditability. Open: you can read the code, patch it, self-host it, and know exactly what it does — important for regulated industries. Closed: black box; you depend on the vendor's promises and roadmap.
- Speed of innovation. Closed wins this axis. Anthropic and OpenAI ship polished features weekly, backed by enormous engineering budgets. The open ecosystem is catching up fast but it's a community-maintained marathon, not a sprint.
- Support. Closed wins here too: a vendor, SLAs, a support channel, a roadmap you can hold them to. Open: community forums, the repo, and your own people — great if you have them, a problem if you don't.
- Capability ceiling. As of late 2025, the frontier closed models still hold a real edge at the very top of raw intelligence, especially in hard reasoning and long-horizon tasks. The gap narrows every quarter, and for many ordinary tasks it's already academic.
- Risk profile. Open = technology risk (the harness can drift), but vendor risk is low. Closed = you've traded technology risk for vendor dependence: a price hike, an API change, or a model you don't like is a business problem, not just a technical one.
There is no universal winner. The right answer is "it depends on the axis your organisation most fears."
Deployment Strategies
You don't have to bet the house on one side. The pragmatic organisations use these four patterns:
- Hybrid stacks. Use closed frontier models where you need the last 5% of intelligence, but run open models and open harnesses for everything else — with open models doing the high-volume, privacy-sensitive, or cost-sensitive work. This is now the default for serious shops.
- Agent platforms. Buy or build the harness as a platform, and treat models as pluggable engines. This is the biggest strategic unlock: the harness becomes your own, models come and go beneath it.
- Local inference with Ollama / LM Studio. For privacy or air-gapped environments, run open weights locally. The data never leaves, and the cost structure is predictable.
- The "done-for-you" digital employee. This is where the revolution touches everyone, not just engineers. A finished agent — a "digital employee" — can now be set up by a competent integrator for a one-off cost in the low thousands of dollars: $500 to $5,000 setup per agent, then it works ongoing. At that price point, a business's blocking question is no longer "can we afford automation?" It's "which work do we actually hand over?"
The Outlook
A few things worth betting on.
The Mark Cuban integrator thesis. Cuban has said, essentially, that the value isn't in the model — it's in knowing how to apply the technology to a business problem. The big wins will go to integrators and domain people who take off-the-shelf engines and harnesses and point them at specific business problems, not to whoever owns the next weight-update. If the harness is the car, the integrator is the mechanic and the driver.
The $500-5,000 digital employee. The done-for-you market matures: agents sold as products, with setup, like consultants, but then they stay and do the work. This creates a huge skills re-organisation in-house: fewer people to do routine work, more people to define, supervise, and correct agents.
Hybrid stacks become the default. Few organisations will be all-open or all-closed. The sensible default is a hybrid: open models and harnesses for control and cost, closed frontier models where the marginal intelligence is worth the money.
Agent metrics become the discipline. Once agents make 15% of decisions autonomously, governance is everything: audit trails, evaluation harnesses, fallbacks, guardrails. The organisations that succeed will be the ones that treat the harness as infrastructure to be managed, not as a shiny thing to be demoed.
Why Every Employee Must Understand This
This is not a topic only for engineers. Your frontline people are the ones who will be asked to hand work to an agent, and to judge whether the agent's output is trustworthy. They need a mental model of what the agent can and cannot do, and they need to know the difference between renting your car from a vendor and owning it. When an employee says "the AI made a mistake," the first question should be: was that a model failure or a harness failure? They're different failures with different fixes, and only people who understand both can tell them apart. Literacy here is not a luxury; it's a risk-control skill.
Why It Matters for the Public Sector
The public sector runs on procurement, compliance, and data protection. For a public agency that needs reliable, auditable, and legally defensible automated processes, the open-vs-closed choice is not neutral. Open weights and open harnesses let you inspect the system, keep citizen data on your own hardware, and audit the decision pipeline. Closed systems push the data and the decision-making into a vendor’s cloud, which creates legal and trust complications that a public body cannot hand-wave away. If you are going to put a model in a state-mandated, regulated process, you need to be able to say what it does and why. That is far easier with an open, auditable harness than with a black box.
The Bottom Line
The engine wars got all the attention; the harness war decides who owns the road. Organisations that treat AI as "subscribe to the smartest model" will find themselves renting their entire vehicle from a vendor that can change the terms any month. Organisations that build or buy their own harness — and treat models as pluggable engines — get to keep the road. The open revolution was never really about the model. It was always about who drives.
Sources
- Precedence Research, "AI Agents Market Size, Share and Trends 2026 to 2035" — AI agents market US$7.92B (2025) → US$294.66B (2035), CAGR 43.57%. https://www.precedenceresearch.com/ai-agents-market
- Precedence Research, "Agentic AI Market Size, Share and Trends 2025 to 2034" — agentic AI US$5.25B (2024), US$7.55B (2025) → US$199.05B (2034), CAGR 43.84%. https://www.precedenceresearch.com/agentic-ai-market
- Gartner (via The Economic Times / CIO ET reporting), "Gartner forecasts AI will make 15% of work decisions autonomously by 2028" — https://cio.economictimes.indiatimes.com/news/enterprise-services-and-ai/from-hype-to-horizon-preparing-the-workplace-for-the-data-ai-and-quantum-convergence/... ; corroborated in Gartner press-news coverage "Gartner Forecasts AI Will Autonomously Complete 15% of Day-to-Day Work Decisions by 2028" (March 2025).
- Gartner (via press-news coverage "Global AI platforms and models market to grow 63% to $64 billion in 2026"), worldwide end-user spend on AI platforms & models to US$64B in 2026, +63.4% from US$39B in 2025. https://www.thehindubusinessline.com/info-tech/global-ai-platforms-and-models-market-to-grow-63-to-64-billion-in-2026-gartner/article71244644.ece
- DeepSeek-AI, DeepSeek-V3 Technical Report, arXiv:2412.19437 — 671B params MoE (37B activated); 2.788M H800 GPU hours for full training. https://arxiv.org/abs/2412.19437
- Wikipedia, DeepSeek-V3 — training cost table: total 2,788K GPU hours ≈ US$5.576M, with the "cost called misleading" caveat. https://en.wikipedia.org/wiki/DeepSeek-V3
- European Commission, Regulatory framework on AI (AI Act) — entered into force 1 Aug 2024; prohibitions 1-8 from Feb 2025; GPAI rules effective Aug 2025; transparency rules effective Aug 2026; high-risk obligations; human oversight requirements. https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
- European Parliament, "EU AI Act: first regulation on artificial intelligence" — ban on unacceptable risk applied 2 Feb 2025; GPAI transparency rules 12 months after entry into force; high-risk obligations 36 months after entry into force (Aug 2027). https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence
- Wikipedia, Artificial Intelligence Act — entry into force 1 Aug 2024 (OJ L 2024/1689, 12.7.2024); fines up to €35M or 7% of worldwide annual turnover. https://en.wikipedia.org/wiki/Artificial_Intelligence_Act
- GitHub — Nous Research Hermes Agent (open harness). https://github.com/nousresearch/hermes-agent
- GitHub — Anthropic Claude Code (harness, ~143k stars). https://github.com/anthropics/claude-code
- GitHub — OpenAI Codex (harness, ~117k stars). https://github.com/openai/codex
- GitHub — OpenCode (open coding agent, ~200k stars). https://github.com/anomalyco/opencode
- GitHub — Ollama (local model runtime, ~179k stars). https://github.com/ollama/ollama
- LM Studio — offline desktop runtime for open models. https://lmstudio.ai
- GitHub — Microsoft GraphRAG (graph-based retrieval, ~36k stars). https://github.com/microsoft/graphrag
- GitHub — Mem0 (open memory layer, ~64k stars). https://github.com/mem0ai/mem0
- GitHub — Plastic Labs Honcho (open memory/personas). https://github.com/plastic-labs/honcho
- GitHub — Composio (open tool/action library, ~30k stars). https://github.com/ComposioHQ/composio
- AI Act Explorer (Future of Life Institute) — AI Act structure incl. Article 5 prohibitions, Article 6 high-risk classification, Article 8+. https://artificialintelligenceact.eu/ai-act-explorer/