Amazon and Cloudflare Ship Jev Clones in 48 Hours, Turning AI Agent Routing Into Its Own Hardware Layer
Introduction
On September 15, 2026, TypeSafe AI released Jev, the first model in a category it called "System One" models. The pitch was deceptively simple. Instead of answering a question with prose, Jev accepts an application state and a set of typed questions, and returns bounded structured values: a choice from a list, a score, a yes-or-no probability. No sentences, no paragraphs, no reasoning chain. For months the model had no serious open competition, and TypeSafe priced it accordingly at $0.042 per million input tokens with no output-token fee.
Within a matter of days the category detonated. Amazon's Strands Labs published Strands Decider 2B, a roughly two-billion-parameter model under Apache 2.0 that strips out text generation entirely. Cloudflare, almost simultaneously, open-sourced Clef and Clef-flash and launched a reinforcement-learning fine-tuning service built on its own edge infrastructure. Two of the largest infrastructure companies on earth arrived at the same architectural conclusion within about 48 hours of each other, and both framed their releases as direct answers to Jev.
That is the story worth telling. Not that a benchmark got beaten, but that agent routing has now separated cleanly out of the large-language-model layer and into a category with its own economics, its own hardware target, and its own competitive battlefield.
What Decision Models Actually Do
The term is new but the idea is not. Classifiers have existed for decades. What TypeSafe introduced with Jev was packaging that idea into an API shape that agent frameworks can call in the hot path, which is to say synchronously, while an agent is mid-turn and waiting.
Consider a concrete case. A customer writes in to a support system saying that checkout has been failing for every customer for the last hour. A decision model can be handed that state and asked three typed questions: is this urgent, yes or no; which team should handle it, with criteria for billing, technical, and sales; and how severe is the impact, scored from "no impact" through "critical." The model returns probabilities against each option. The calling code then routes the ticket, triggers an escalation, or hands off to a human. No human needs to read the message before the decision is made, and no large model burns tokens writing an explanation of why the checkout is broken.
Cloudflare described exactly this shape in the technical writeup accompanying its Clef release, and the request body in their documentation looks like a schema declaration rather than a prompt. The distinction matters because a schema is enforced. A model that must emit one of three enumerated team names cannot hallucinate a fourth.
The framing that made the category legible is that large language models are non-deterministic but open-ended, suited to reasoning and generating text and tool calls. Decision models are bounded and cheap, and they are not a replacement for the reasoning model. They are the thing that decides, quickly, whether reasoning is even required, and which of a fixed set of paths to take.
The Architecture That Makes It Fast
The performance claims in this category are entirely architectural, and the reasoning is straightforward enough to follow without reading a paper.
A conventional language model generates a token, reads it, generates the next token, and repeats. Every answer costs dozens or hundreds of sequential steps. Cloudflare's Clef sidesteps this. It runs the base model in a prefill-only pass, then scores the valid schema choices in parallel. As Cloudflare put it, the decision step is non-autoregressive, so there is no intermediate text being generated token by token. Rather than generating reasoning and then parsing it into structure, Clef derives schema choices directly from the backbone's internal representations.
Amazon arrived at something structurally similar from a different direction. According to VentureBeat's reporting, Strands Decider 2B begins with a pretrained Qwen3.5-2B base and its developers removed the component that predicts the next word, replacing it with a small pointer component that scores the supplied answer options. A rank-16 low-rank adapter plus that scoring head add roughly a million trainable parameters, under an internal architecture name AWS calls Hobson.
Two unaffiliated engineers arriving at the same shape, one from a cloud edge provider and one from a consumer-scale lab, is meaningful evidence that this is a genuine architectural convergence rather than one company's peculiar implementation.
Cloudflare also detailed its training recipe: the base model is frozen, and a routing head is optimized alongside rank-256 low-rank adapters, using label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to calibrate probability output. They also developed a reinforcement learning objective they call Reinforcement Learning for Calibrated Decisions, which grants partial credit to adjacent ordinal choices and rewards fully precise records, with a reference penalty to prevent distribution shift.
The Numbers Are Real, With Caveats
Cloudflare published a benchmark table comparing Clef, Clef-flash, Jev, and three other decision models across ten benchmarks. The headline figures are striking. On BANKING77 macro-F1, Clef scores 94.20 against Jev's 79.74. On API-Bank accuracy, Clef-flash reaches 93.11 against 88.19. On BFCL case-exact, Clef-flash scores 98.76 versus Jev's 95.75.
Latency is where the category's real argument sits. Cloudflare reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, against 524.1 milliseconds for Jev. At the 95th percentile, Clef-flash measures 122.4 milliseconds against Jev's 536.0.
Amazon's numbers come from different hardware, so they are not directly comparable, but they point the same way. AWS claims a median of 106 milliseconds and a 95th percentile of 296 milliseconds for Strands Decider 2B on an Nvidia RTX 3090, roughly 150 milliseconds median on an M3 MacBook for small tasks, and about 72 percent accuracy on the JevBench v19 public set.
Two honest caveats belong here. First, all of these are vendor-published figures. Cloudflare's Clef results come from Cloudflare's own evaluation harness, and no independent replication was available at the time of writing. Second, quality is genuinely mixed rather than uniformly dominant. Jev beats Clef on two of the ten benchmarks Cloudflare listed, taking When2Call accuracy at 80.97 against Clef's 72.37, and BRIGHT nDCG@10 at 47.52 against 45.91. On Cloudflare's four-workflow suite, Jev led agent trace observability at 71.6. The honest reading is that Cloudflare's models win the majority of the table and the latency comparison decisively, but they do not sweep it.
Amazon's framing was more cautious still. VentureBeat's Carl Franzen declined to call it a win, arguing that AWS's distinguishing proposition is less a demonstrated Jev-killer than an open, reproducible decision layer tied to an existing agent framework. That is a fair characterization, and the same piece flagged that adversarial text in an agent's input can influence the model's verdict, a risk that a bounded-output model does not automatically eliminate.
The Licensing Divide Is the Actual Fault Line
What makes this week's releases consequential is not the benchmark table. It is that Cloudflare and Amazon both chose Apache 2.0.
TypeSafe AI's Jev is a hosted, closed, paid service. There is no weight to download and nothing to inspect. It launched on September 15, 2026 and is available in early access, with reported end-to-end latency in the 70 to 500 millisecond range. Both of its fast-moving competitors answered not only with better latency but with weights that anyone can download, run locally, and fine-tune.
That reframes the economics for anyone building agents at volume. A decision model sits in the hot path, which means it is called constantly, and a per-token API bill is the wrong cost structure for a component that emits roughly one structured value per call. Cloudflare makes the infrastructure argument explicitly: because Clef is hosted on Workers AI running on Cloudflare's own GPUs at the edge, a decision can be made with low network latency and combined with one of Cloudflare's own language models on the same platform.
Amazon makes the locality argument: 106 milliseconds median on a consumer RTX 3090 means a developer can self-host the weights on hardware they already own and handle synchronous agentic decisions without a cloud round trip at all.
For enterprise buyers, the fine-tuning path is arguably the more interesting part of both announcements. Cloudflare described internal use cases including evaluating Trust and Safety submissions, triaging support requests, and deciding whether an inbound crawler is a good bot or a bad bot, noting that the company has more than fifteen years of labelled network data across domains. Its new service will start with hands-on work from forward-deployed engineers and then become a self-serve platform stitching together AI Gateway for capturing request data, Workers AI for generating rollouts, Containers as a reinforcement-learning sandbox, a new trainer for updating weights, and Bring Your Own Model deployment. Amazon published its training recipe alongside the weights for the same reason: fine-tuning on proprietary workflow data without a vendor dependency.
The category's centre of gravity is still hosted. The direction of travel is not.
What It Means for Agent Architecture
The near-term consequence is architectural. Teams building agent systems have been using the large language model for everything, including the small classification decisions that gate each step, because there was no alternative that was fast enough and structured enough to call synchronously. That is changing within a week.
The trade-off developers are actually making is now explicit and quantifiable. A routing decision that takes 500 milliseconds and costs per-token pricing is a bad fit for a step that fires on every turn of an agent loop. The same decision at 38.8 milliseconds, local, open-licensed, and fine-tunable on your own data is a component you can put in the critical path and leave there.
There is a broader signal too. Both announcements arrived as part of a wave that has also seen Anthropic open Claude Code's internals to small TypeScript mods that can intercept events, rewrite prompts, block or retry tool calls, and redact secrets mid-session, and separately a reinforcement-learning fine-tuning service. The common thread is that the infrastructure providers are moving up the stack into the runtime layer, competing on the programmable substrate that agent frameworks run on rather than only on model weights.
For anyone watching the agent stack, the interesting question is no longer whether large models get better. It is what happens to the layer underneath them as the decisions that used to be free inside a model's context window get pulled out and become their own purchasable, benchmarked, contestable components.
Conclusion
Forty-eight hours is a short window in which to reorganize an architecture, but that is roughly what this week produced. Two companies independently concluded that the decisions an agent makes before it reasons are distinct enough from reasoning to deserve their own model class, their own latency budget, and their own price. Both shipped weights under Apache 2.0. Both targeted the same incumbent, and both reported latency wins that were large enough to matter for synchronous agent loops.
Cloudflare's numbers are the more complete, and the more vendor-produced, with Clef winning most of a ten-benchmark table and losing two of them. Amazon's are the more modest and the more locally deployable. Neither has been independently replicated yet.
What is already clear is the shape of the layer. Decisions get scored, not generated. Latency drops by an order of magnitude when you remove autoregressive text from the answer path. And open weights plus a published fine-tuning recipe changes who gets to build on top of it.
Images
![]()
Illustrative: a generic data-center aisle, not Cloudflare or AWS infrastructure. Photo: Wikimedia Commons.
![]()
An opened processor die under microscope illumination. Illustrative of the inference hardware behind these models, not one of the models discussed. Photo: Wikimedia Commons.
References
- Cloudflare: Introducing Clef, our open-source decision models, and new RL fine-tuning platform
- TypeSafe AI: Introducing System One Models & Jev
- VentureBeat: Amazon unveils a free, fast, open source Jev killer: Strands Decider 2B
- The New Stack: Anthropic's mods let you change Claude Code's look and behavior
- AI Weekly: AI news for Friday, October 2, 2026
- Related coverage on artificial intelligence and cloud and edge computing