---
title: "Which Model is All the Way Down?"
date: 2026-09-29
type: essay
summary: "Multi-model workflows, specialized models like Jev, and enterprises owning their intelligence all cut against everything collapsing into one lab's OS. So where does AI value live?"
url: https://dazuck.com/writing/which-model-is-all-the-way-down
---

As recently as early August I wrote a group chat, "I've assumed for a while that all functionality will collapse into the harness, and harnesses will collapse into the OS."

This was met with general consensus. It seemed models were eating everything, everyone was playing to expand their stack, and more access to more OS was going to be the primary differentiator. But that was a whole month ago.

Three things have happened, or accelerated, to force a rethink.

**First, the benefits of multi-model workflows persist**, and as we look to AI to do more they matter more. Any model family shares an architecture, training approach, and data set - so blind spots and weaknesses. Combining model families in execute/verification or other agent swarm approaches reduces gaps and combines strengths. This is obvious in my own work, and research backs it: [Sakana's Fugu](https://arxiv.org/html/2606.21228v1) beat the best single model on all four benchmarks it tested. Enterprises are moving the same way, with [81% now using 3+ model families](https://a16z.com/leaders-gainers-and-unexpected-winners-in-the-enterprise-ai-arms-race/), up from 68% a year earlier. If mixing models is essential, can one OS (dominated by one company and its own model) hold everything?

**Second, [Jev arrived](https://typesafe.ai/blog/introducing-system-one-models-and-jev).** It excels at efficient classification, probabilities, routing, judging, and many other use cases. [Tom Tunguz found](https://tomtunguz.com/ai-comes-for-the-if-statement/) it beat a frontier model 80% to 47% on real email threads, at 82x lower cost. And it hints that after a period of large LLMs eating everything, we may head into a period of differentiation by specialization, and speciation (aka more model types). Maybe Anthropic, OAI, Google, etc. can replicate Jev quickly ([Kev](https://github.com/jaredpalmer/kev) already did), but then won't that type commoditize, and every other type besides? Either way model types multiply, and a harness (consumer) or platform (enterprise) that can orchestrate independent model architectures has a growing case for independence.

**Third, [Pat Grady's great video](https://x.com/gradypb/status/2103536289438130670) makes the case for enterprises wanting to own their intelligence.** Lab APIs offer a handful of points on the cost/performance curve, and real workflows fall between them. Owning and tuning your own models is how you get on the Pareto frontier for every workflow (and role in a multi-model workflow, presumably) you run. This would have seemed crazy in January, except to the most advanced organizations. Most large enterprises were just learning how to use this stuff, and the conventional wisdom was that models are commoditizing, so why train your own?

The labs still capture most of the dollars today; open models were [under 4% of spend](https://trendingtopics.eu/open-weight-models-from-china-are-capturing-a-growing-share-of-ai-usage/) on Vercel's gateway even at nearly 30% of tokens. But as we shift rapidly from tokenmaxing to a consensus that inference spend is going to increasingly dominate costs, maximum efficiency and performance is a natural business imperative. That suggests a collapse into the dominant models and OS's is far less likely, on the enterprise side at least.

On the surface, consumer looks like a different story, with a recent onslaught of do-everything agents in a good interface. [Instinct](https://techcrunch.com/2026/09/28/viral-ai-agent-instinct-raises-1b-series-c-at-a-10b-valuation/) is a $10B company growing 10% per day; Muse [hit #1 on the App Store](https://www.cnbc.com/2026/09/21/meta-muse-personal-ai-agent-downloads.html) ten days after launch. The narrative today is *most users just want something easy and simple*, not Claude Code or OpenClaw to battle with.

But that mirrors the enterprise/developer narrative nine months ago, at the "one model for everything" stage. Given the pace and cycles of AI, it seems likely consumer is just at an earlier stage, and differentiation comes there too. Many people, e.g. [Dennison Bertram](https://x.com/DennisonBertram/status/2104712154800632108), think so.

Which all leads back to the original question: **does everything collapse into the OS, and what is the OS for all these agents, whether one model for all or many used together?**

Software-defined harnesses were all the rage just a few short months ago, but as apps build in agents and platforms build in agents and labs build in agents, will they be squeezed out? Or will enterprises insist on owning their harness as the control plane of intelligence?

Or will the OS settle closest to the hardware - Android/iOS for consumer, and the data center itself for enterprise? That buys maximum access, scale and raw intelligence. And it seems an implicit hypothesis for Google as they prioritize Cloud over Gemini, and for Nvidia as they [buy Hugging Face](https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000078/nvda-20260902.htm) and add to their developer platform.

Or can vertical players carve off segments from the hyperscalers? Thomson Reuters shipped a model it says is competitive with the strongest frontier models, and Harvey post-trained its own on an open-weight base. Vivek Ramaswami and Sabrina Albert call these ["neo neo labs"](https://aspiringforintelligence.substack.com/p/the-rise-of-the-neo-neo-lab), ready to serve their segments with more custom intelligence.

![models-all-the-way-down.png](https://dazuck.com/images/models-all-the-way-down.png)

In most of these scenarios, we get the proverbial "models all the way down," with a model trained to help you manage, observe, and continuously tune the other models as needed. **Which model will live at the bottom of that stack?**

---
*This is the source markdown for https://dazuck.com/writing/which-model-is-all-the-way-down.*  
*Full writing + projects index: https://dazuck.com/llms-full.txt*
