Jev Engineering: what a decision layer changes about the agent loop, and what it doesn't
On 18 September 2026, three days after TypeSafe AI opened early access to Jev, a developer posting as codila published an X article titled "Jev Engineering: Full 10-Step Roadmap to Set Up and Use a New Brain for AI." Two more guides with the same name followed within a day, along with a small ecosystem of harnesses, skills and curated lists. All compress the idea into one line: an LLM writes, Jev decides, code acts.
That line is the popular definition, and it merges two different things. One is a routing pattern: stop spending a chat model on decisions that are effectively typed switch statements. The other is a falsifiable architectural argument that TypeSafe's founder Diogo Almeida laid out in design notes for a coding harness built around Jev — independently compiled into a document titled Jev Engineering for Coding Agents, which states plainly that it is not affiliated with or endorsed by the company. The first is cheap to adopt. The second contains the interesting engineering, and one piece of arithmetic that cuts against the intuition the whole category is sold on.

What the substrate actually is
Jev is not a smaller chat model. A request is a block of state — a string, a JSON object, or an array of text — plus one or more typed questions. The model evaluates every question against that state in a single parallel pass and returns structured answers. There are three question types:
| Primitive | What it does | What comes back |
|---|---|---|
Choice |
Picks one option from a set you define | The chosen option, per-option probabilities, a confidence value |
Score |
Places the state on an ordered rubric | A score, per-level probabilities, confidence |
Noul |
Answers a yes/no statement | A probability between 0 and 1 |
Because the answer space is declared up front, the model cannot return a value outside your schema — a claim TypeSafe calls mathematically guaranteed, and which is easy to believe because the sampler is not generating tokens that your code then has to parse and validate. Output tokens are not generated at all in the usual sense, which is why output is priced at zero.
The published numbers for jev-1.13.0 are: $0.042 per million input tokens ($42 per billion), output free; end-to-end response of 70–500 ms, with most queries around 100 ms; a 64k context per request, of which 32k covers the state plus the longest single question; text input only. The endpoint is POST /v1/systemone, and jev-latest is an alias that moves when a new release ships — the response reports the versioned ID that answered, so if you tuned thresholds against one version, pin the version. Rate limits are 250,000 tokens per second and 1,200 requests per minute, and the docs say these are adjusting dynamically while the company lands GPU capacity.

The three-way split, and the test for each step
Jev Engineering's central habit is sorting every step of an agent loop into one of three kinds and giving it to the part built for it:
| If the step | It goes to | Example |
|---|---|---|
| Creates text | An LLM | Draft the briefing, summarise a paper |
| Picks, scores, or answers yes/no | Jev | Which worker acts next, is this command safe |
| Follows an exact rule | Code | Stop after ten actions, never publish without approval |
The test is deliberately mechanical, which is what makes it useful in review. A step that produces prose is an LLM step. A step whose output is one of a short list of options is a decision. A step that a regex, a parser, or a comparison operator can settle is code, and asking a model to do it is a defect — TypeSafe's own failure-mode documentation lists "asking the model something code can compute exactly" as the first thing to avoid.
State and questions: where the engineering time actually goes
The client call is a few lines. Everything hard happens in two places: the state you send and the threshold you act on.
State is evidence, not summary. Jev decides on what you put in state and nothing else. "The researcher finished" tells it far less than the sources, what they found, and what is still missing. Accuracy also falls as state grows with material unrelated to the question — the docs call it context rot, and it is listed as a distinct failure mode. Retrieve and filter first; send the fields the question needs.
Question IDs never reach the model. A field named safe_to_publish is a name, not an instruction. The requirement goes into the instruction, and each option needs criteria saying what belongs in it, what belongs in the neighbouring option instead, and a representative example each. Use the same field names across options so the model can compare them directly.
Atomic over broad. "Should we escalate?" hides several judgments behind one answer. Decomposed, it becomes severity, reproducibility, and blast radius as separate questions, combined by rules your code owns. TypeSafe calls this the most important concept in its build guide, and it is what makes a decision auditable: when the wrong branch is taken, you can see which question was wrong instead of arguing with a paragraph.
Thresholds scale with risk, and confidence is not accuracy. Every Choice and Score answer carries probabilities plus a confidence value derived from the shape of that distribution — concentrated means confident, flat means the options are ambiguous. The docs are explicit that this is a convenient default rather than a principled measure. A read-only action and a destructive one should not share a threshold; the risk tolerance belongs in code, set from your own labelled examples.
The harness argument: take the KV cache away
The founder's design notes open with a question that is more useful than it first looks: how would you design a coding agent if language models had no KV cache?
The answer is that almost every design choice in today's agents is downstream of one economic fact — reusing a cached prefix is cheap, and changing anything early in the context forces the model to reprocess everything after it. That is why agents are append-only transcripts. Remove the cache from the picture and six familiar behaviours stop looking like properties of the model and start looking like properties of the harness: routing fails because context is reprocessed, tool schemas crowd the window because they are declared up front, compaction compresses before the question is known, passing state makes sub-agents rare, restarts discard good state along with the bad, and the batteries debate exists because every built-in consumes context forever.
The proposal is one move: make state explicit and typed, and let Jev answer the per-turn questions the transcript currently answers by default — which chunks are visible for this query, whether to reuse the cached prefix or rebuild, whether a subtask can leave the frontier model, which tool fits the intent, whether a command may run, how sensitive the files are. Each is a Choice, Score, or Noul, asked thousands of times per session.
Two sub-proposals are worth calling out because they are unusually concrete. The first is a visibility ladder: a 2,400-line grep result can be twelve relevant hits for one question and invisible for the next, without ever being deleted from state. That is query-aware compression, and it keeps compaction's benefit while removing its timing flaw. The second is tiered tool disclosure: one-line snippets for hundreds of capabilities, full schemas only for the few the model selects, documentation only for a one-off query — which, if it works, dissolves the batteries debate rather than settling it.
The routing arithmetic that argues against naive routing
This is the most useful piece of the notes, because it is arithmetic rather than philosophy, and it says the obvious optimisation is often a loss.
Take list prices as the notes did — $5 per million input and $25 per million output for the frontier model, $3 and $15 for the mid-tier one. Let X be context tokens, Y generated tokens, and Z additional tokens read during the work (command output, file reads).
| Path | Cost |
|---|---|
| Pure frontier | 25Y + 5Z |
| Frontier → mid-tier → frontier | 3X + 20Y + 8Z |
The routed path carries a context load on the way down (3X) and a reprocessing pass on the way back (8Z is where the frontier model reloads what changed), and those together outweigh the per-token discount. Plugging in a plausible session shape — X = 0.65, Y = 0.12, Z = 0.23 — gives 4.15 for pure frontier against 6.19 for the routed path. Staying on the expensive model costs roughly two thirds of what the route intended to save.

The conclusion is not that routing is wrong. It is that routing priced per token instead of per context rebuild is wrong, and it becomes viable only when the harness can hand the cheap model a small, purpose-built context and merge the result back as a chunk rather than a transcript the frontier model must reread. That is a much stronger claim than "cheap model good," and it is testable: if your harness can assemble a small context for a subtask, routing and sub-agents both get cheaper; if it cannot, they do not.
The same notes put a number on where the budget goes in a typical CLI coding session: reading files 30–40%, searching 10–18%, command output 10–20%, system prompt and schemas 5–12%, reasoning 5–15%, editing 4–10%, explaining 2–5%. Writing code is among the smallest line items; retrieval dominates. Microsoft's fastcontext figures for GPT-5.4 trajectories — 56.2% of tool-use turns and 46.5% of main-agent tokens — point the same way: the largest available saving in a coding agent is smarter and shared retrieval, not a better model.
Fan-out: where the numbers are hard to argue with
The one performance claim that survives scrutiny without a benchmark harness of your own is structural. Questions in a single request are evaluated in parallel and cannot see each other's answers, so a speculative extra question typically adds no latency and only the input tokens you send. TypeSafe's own test found 13 questions in one call were 11.5× cheaper and 9.6× faster than 13 separate calls.
That changes the shape of a decision tree. Instead of asking for a ticket's category and then its severity in a second round trip, you ask both at once and let code ignore the severity answer when the category turns out to be a feature request. The cost of being wrong about which questions matter collapses to the token price of asking.
The constraint to design around: because questions are independent, one that depends on another's output cannot be in the same call — search first, then decide.
What is verified, and what is not
TypeSafe is unusually candid about the shape of its own evidence.
The 193.6× faster and 444.6× cheaper figures come from the company's workflow evals, where four workflows were designed by its own model-capabilities team and scored against reference probabilities built by averaging two frontier models. The launch post concedes that some bias could exist because its own team built them, that a frontier consensus reference biases the comparison toward those vendors, and that "we expect that these are on the higher end of real world gains." A community compilation of Jev builds notes that the 200×/400× figures quoted in the guides trace back to a founder's talk rather than to a measured workload. TypeSafe also cannot prove its pricing is not subsidised, and says so.
What third parties report is smaller and more interesting. A developer classified 1,018 AI research papers for $0.08 at a 256 ms median latency per paper, after spending $3.99 on summarisation with a different model to feed the classifier. Another classified 500 emails for 3.5 cents. Neither is a controlled benchmark, and both costs are dominated by choices made upstream of Jev — which is the point the pattern is making.
Where jev-1.13 breaks
The failure-mode page is unusually specific:
- No counting. Characters in a word, occurrences of a term, items in a long list — the model recognises the shape of an answer instead of tallying, and the error grows with the size of the thing being counted. Iterate in code and ask one question per item.
- No arithmetic, no dates. Math belongs in code;
Scorelevels are weak at reconstructing a magnitude by interpolation. Dates are read as text, not ordered quantities. Extracting a date component is a judgement and can be aChoiceover twelve months or thirty-one days; comparing and ordering belongs in code. - Literal reading. "It answers the question you wrote, not the one you meant." When you find yourself explaining what you really meant by an instruction, that explanation is the missing half of the instruction.
- No guaranteed structural invariants. The sharpest example in the docs, worth reading twice: asked whether a customer is requesting a refund, the same ticket yields
noul= 0.22 while a yes/noChoiceanswersnowith 0.99 probability and 0.97 confidence. A question and its negation as two separateNouls sum to 1.19. Do not carry a threshold tuned on aNoulover to aChoice, and do not hold the model to arithmetic identities across questions. - Context rot and adversarial state. Unrelated material in
statecosts accuracy. Injected instructions, or text that argues for its own classification, can move the answer — state is data, and it is not treated as hostile. - English first. English is the primary training language and where accuracy is best; other languages, including CJK, are handled but not equally well, and the docs say to test on your own content before relying on it.

And the one distinction that matters most for adoption: type-safe is not the same as correct. The model cannot return a value outside the schema you declared, and it can still pick the wrong option. "Zero type errors" is a real guarantee. "Zero wrong answers" is not, and the confidence score exists because of that gap.
Architecturally, the company has not published a paper, weights, or its architecture. It describes Jev as transformer-based, trained on synthetic data, with Reinforcement Learning for Calibrated Decisions (RLCD) optimising probabilities against outcomes rather than against human rater preference. Outside observers have suggested it may be built on an open-weight base model. That is not disqualifying — RLCD is the claim, not the lineage — but it does mean the most load-bearing claim in the stack is currently unverifiable from outside.
What it does not replace
TypeSafe's own coding-agent page is blunt: Jev is not a drop-in replacement for the LLM behind Claude Code, Cursor, or similar tools, and no model: "jev-latest" setting turns a coding agent into a Jev-powered one. The recommended uses are inside software you are building — routing to a fixed set of destinations with a confidence you can threshold, scoring on a rubric, checking a statement against a record before acting, and replacing a prompt that asks an LLM to "return JSON" with a call that returns typed values by construction.
The most defensible adoption path is the boring one the guides converge on: log one run of your agent, mark every model call as text, decision, or rule, pick the decision that runs most often, write its state and options carefully, try it in the playground before writing any code, then run old and new paths against the same labelled cases and compare accuracy, latency, and cost per finished task. That last metric keeps the pattern honest: a cheap decision that sends a worker down the wrong branch costs more than the expensive call it replaced.
Where this still hurts
- The headline numbers are self-reported. Treat 200× and 400× as a marketing ceiling; the honest planning assumption is the structural one — parallel sampling is faster, output is free, and your savings depend on how much of the loop you actually move.
- The 64k context is a real ceiling once state, questions, and the longest question share one budget, and the fix — retrieve and filter before calling — is work you have to build. Rate limits also move without notice while capacity is being added, which is awkward for anything latency-critical.
- Confidence needs its own calibration. It is derived from the distribution shape by a definition TypeSafe calls a solid default but not a principled one, and thresholds set from someone else's examples will be wrong for your traffic.
- The harness blueprint is a design document, not a product. Its routing arithmetic uses list prices the notes themselves flag as changeable, and the token-share table is described as an illustrative estimate rather than a measurement.
- Single-vendor dependency, with no fine-tuning on your data: domain behaviour is shaped entirely through the request, so your state and criteria cannot be handed to a different decision model without rewriting.
References
- Introducing System One Models & Jev — TypeSafe AI
- System One — TypeSafe AI docs
- How to build with TypeSafe — TypeSafe AI docs
- Models and pricing — TypeSafe AI docs
- Confidence — TypeSafe AI docs
- Jev 1.13 jaggedness — TypeSafe AI docs
- Jev with coding agents — TypeSafe AI docs
- Jev Engineering for Coding Agents — compiled from design notes by Diogo Almeida
- Building a harness with Jev — LangChain