Jev and RLCD: The Model That Never Speaks, and Why Testers Should Care

A week of hands-on with Jev, TypeSafe AI's new decision model. What RLCD changes versus RLHF, why System 1 models fit automation better than chatbots, 10 browser test cases for 5 cents, a feed filter that judges every post in 300 ms, and the test-coverage use case I want to build next.

Two AI paths diverging from one glowing pretrained core: on the left a chat machine producing endless paragraphs with a human sitting in the loop, on the right a compact dark lens emitting a yes/no toggle, a multiple-choice selector and a probability gauge straight into software gears and a test pipeline, in coral, teal and cream on near-black.
Same pretrained core, two different bets: text for humans, decisions for software.

On 15 September 2026 a model launched that cannot write a sentence, and within three days my X feed was full of it. Jev, from TypeSafe AI, does not chat, does not reason out loud, does not produce a single token of prose. You hand it some state and a typed question, and it hands back a decision with a probability. That is the entire product.

I got early access the same week, and instead of reading about it I pointed it at the things I actually do: driving a browser through a trading app, filtering my own social feeds, classifying test failures. This post is the write-up. It is long because the interesting part is not the headline numbers, it is what broke on the way to them.

If you work in testing or build agents, here is the short version. The last four years of AI were built for assistance: a human reads the output and decides. Automation needs something else: a decision software can act on, with an honest number attached that says how sure the model is. Jev is the first public model trained for that second thing, and once you use it the split between “assist” and “automate” becomes impossible to unsee.

The talk that explains the bet

Jev’s co-creator is Diogo Almeida, one of the authors of the InstructGPT paper, which is to say one of the people who invented RLHF and shipped ChatGPT. A few months before the launch he gave a short talk at the AI Engineer conference that is the clearest framing I have found of why a model like this exists. Watch it if you have eighteen minutes; the rest of this section is my summary of it.

Diogo Almeida, Jev creator: why RLCD beats RLHF (AI Engineer)

His starting puzzle is the one everyone in the field has noticed and few can explain. Models beat humans on olympiad maths and still cannot be trusted to close a customer-support ticket without a person checking. Why are the hard things easy and the easy things hard?

His answer: the tasks on the “easy” side are all human-in-the-loop by design. The goal of a chatbot, or of a coding agent talking to you in a terminal, is to please the person in the loop. The tasks on the “hard” side are the ones where the whole point is to remove the person. Those are different jobs, and RLHF only trains for the first one.

“Why do all LLMs require a human in the loop? The simple answer is we literally put them in the loop. The goal of the loop is to optimise for human preference. It is not to run software autonomously.”

That is the line to remember. RLHF collects human preferences and optimises the model to produce what raters approve of. Two side effects follow directly. The model learns that a confident, complete-sounding answer scores better than an honest “I don’t know”, so overpromising is not a bug, it is what the reward paid for. And the distribution narrows toward the preferred style, which the paper calls mode dropping: valid but unusual answers quietly lose probability mass. His example is sending ChatGPT a file of fart sound effects and asking what it thinks of your music: “It’s a very eerie, atmospheric piece.”

Then came RLVR, reinforcement learning with verifiable rewards, which gave us the reasoning models: train on tasks where a program can check the answer, like maths and code. It works beautifully where a verifier exists. Most business decisions have no cheap verifier. Is this ticket urgent? Which of these six buttons is the right one? Is this test flaky or broken? Run RLVR on fuzzy judgments and you get brittle, overconfident models.

RLCD, reinforcement learning for calibrated decisions, is the third branch. The output is not text. It is a decision plus a probability, and the reward is the honesty of that probability. Across many predictions, the answers a well-calibrated model tags with 0.8 should be right about 80 percent of the time. That property, calibration, is what turns uncertainty into something software can branch on.

graph TB
    Base["Pretrained language model
(the compressed internet)"] Base --> RLHF["RLHF
reward: what human raters prefer"] Base --> RLVR["RLVR
reward: a program says the answer is right"] Base --> RLCD["RLCD
reward: the probability is honest"] RLHF --> Chat["Chatbots, coding assistants
human reads the output"] RLVR --> Reason["Reasoning models
maths, code, anything checkable"] RLCD --> Dec["Decision models
software acts on the output"] style RLCD fill:#FF5A4E,stroke:#FF5A4E,color:#0B0F14 style Dec fill:#FF5A4E,stroke:#FF5A4E,color:#0B0F14 style Base fill:#F5F0E8,stroke:#5A544B,color:#0B0F14

Almeida’s other provocation is that the Claude Code era is not “what comes after ChatGPT”. It is still the assistance era, because it is still RLHF: a model optimised to work well with a person watching. What comes after assistance, he argues, is automation, and automation needs a model that does not care whether you like the answer, only whether it is right, and that tells you when it is not sure.

I do not have to agree with every word to find the frame useful. Every week I hit the same wall in test automation: the agent is right 95 percent of the time and cannot tell me which 5 percent. A model that says “I’m at 0.4 on this one” would let me automate the other 95 and route the rest to a human. That is the whole promise.

System 1, System 2, and where software actually spends its decisions

TypeSafe calls Jev a System One model, borrowing Kahneman’s split between fast, intuitive judgment and slow, effortful reasoning. The pitch is that most decisions inside software are System 1 decisions. Which team handles this ticket? Is this message spam? Which link on this page leads toward the goal? Did the page reach the expected state? None of those need a chain of thought. They need a fast, calibrated gut call.

A conveyor of small decision cards flowing fast through a coral lane where a lens device stamps each one in milliseconds, while one heavy card is lifted off to a slower deliberate reasoning chamber.
Most decisions in a pipeline are cheap gut calls. Route those to a System 1 model and save the reasoning model for the few that need it.

Frontier labs spent two years making models think longer. Jev goes the other way: no autoregressive decoding at all. Every question is answered in a single parallel pass over the input, which is why a call takes 70 to 500 milliseconds regardless of how many questions you ask, and why output tokens are free. There is no output to meter.

This is also the honest way to understand the “it cannot hallucinate” claim. Jev cannot return an option you did not define, cannot invent a tool name, cannot produce broken JSON, because the output is constrained to the schema by construction. It can absolutely pick the wrong option. The difference is that when it does, the probability is supposed to be low, and in a week of runs that is exactly what I saw.

Here is how the two systems fit together in an agent loop. Code owns the control flow. The reasoning model is called when a plan or free text is needed. Jev is called for every bounded decision along the way.

graph LR
    Obs["Observe
(code: snapshot state)"] --> Q{"Bounded decision?"} Q -- "yes, most steps" --> S1["System 1: Jev
which element, which action,
done or blocked, ~300 ms"] Q -- "needs a plan or free text" --> S2["System 2: reasoning LLM
seconds, tokens, rationale"] S1 --> Act["Execute
(code)"] S2 --> Act Act --> Verify["Verify
(code assertions)"] Verify --> Obs style S1 fill:#FF5A4E,stroke:#FF5A4E,color:#0B0F14

The pattern people converged on within days, across a dozen open-source browser agents, is the same: code observes, Jev decides, code executes, code verifies. A generative model is only in the loop when something has to be written.

The three primitives

Jev has exactly three question types. Everything I built this week is composed from them.

Three instruments on a dark workbench: a glowing toggle for a yes/no probability, a rotary selector for a multiple choice, and an analog gauge for a score, all cabled into one compact lens module.
Noul, Choice, Score. Three instruments, one request, all answered in parallel.

Noul is a yes/no question that returns the probability that the answer is yes. The word is TypeSafe’s coinage; think of it as a probabilistic boolean. “Does this message express urgency?” comes back as 0.999. A Noul near 0.5 means “coin flip”, not “medium”.

Choice picks one option from a set you define, and returns a probability for every option plus a confidence number that summarises how concentrated the distribution is. “Which team should handle this?” returns billing: 0.84, technical: 0.159, sales: 0.001.

Score rates the state against ordered levels you describe, and returns the expected level plus the distribution. “How frustrated is the customer?” against three described levels returns 1.035.

One request carries the state and as many questions as you like. They are evaluated in parallel and in isolation, so adding questions barely changes latency. This is the whole request from their quickstart, and the whole response:

{
  "state": "Hi, I've been trying to connect my Stripe account for 3 days
            and it keeps failing. I'm losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing":   "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales":     "Pricing or account questions"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated the customer appears",
      "criteria": ["Calm, just stating facts",
                   "Frustrated but civil",
                   "Very angry, strong language"]
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}
{
  "answers": {
    "department": {
      "choice": "billing",
      "probabilities": { "billing": 0.84, "technical": 0.159, "sales": 0.001 },
      "confidence": 0.596
    },
    "frustration": { "score": 1.035, "confidence": 0.842 },
    "is_urgent":   { "noul": 0.999 }
  },
  "usage": { "input_tokens": 312, "output_tokens": 48 }
}

Read the confidence numbers as a separate axis from the answer. The answer tells you what; the confidence tells you whether to act. Their docs suggest three bands, and I used exactly this in every experiment: high confidence, act automatically; medium, act with a check or a confirmation; low, do not act, route to a human or to a reasoning model. The thresholds are not one number. A read-only branch can act at 0.6. A destructive one might need 0.9. Your code encodes the risk tolerance, and you can move a threshold without re-running inference or touching a prompt.

graph TB
    A["Jev answer + confidence"] --> H{"confidence"}
    H -- "high" --> Auto["Act automatically"]
    H -- "medium" --> Check["Act with a confirmation
or a second check"] H -- "low" --> Human["Do not act.
Route to a human or a reasoning model"] style Auto fill:#00C48C,stroke:#00C48C,color:#0B0F14 style Check fill:#FFB020,stroke:#FFB020,color:#0B0F14 style Human fill:#FF5A4E,stroke:#FF5A4E,color:#0B0F14

The first thing I ran in the Playground was their agent-audit example: a full support-agent trace with tool calls, six questions, 135 plus 257 milliseconds. The answer that sold me was not the ones it got right. It was predicted_csat, a Score over five levels, coming back with a confidence of zero percent. The probability was smeared evenly across all five levels. The model was saying “I have no idea what this customer will rate us a month from now”, which is the correct answer, and which no RLHF model I have used has ever said to me unprompted.

What the numbers say, and what they do not

The claims on the homepage are 193 times faster and 444 times cheaper than frontier models on System 1 tasks. Pricing is $0.042 per million input tokens, output free, which is a 48th of GPT-5.6 Terra’s input price. On TypeSafe’s own four-workflow benchmark (security incident response, observability, invoice processing, customer service), Jev scores 67.8 percent accuracy against 67.9 for Terra, 74.1 for GPT-5.6 Sol and 73.1 for Claude Opus 5, at 0.4 seconds per case versus 10 to 38 seconds. Tool-call error rate is zero by construction, against 5.73 percent for Opus 5 and 17 percent for Sol.

Treat all of that as vendor-reported, because it is. The workflows were designed by TypeSafe. Nobody outside has reproduced them yet. What I can vouch for is what follows: my own runs, my own assertions, my own failures.

Two more facts you need before building anything. Context is 64k tokens per request, 32k for the state plus the longest question, and their own jaggedness page says accuracy drops as irrelevant material in the state grows. And Jev reads literally. It answers the question you wrote, not the one you meant. Every failure I hit this week traced back to one of those two.

Use case 1: browser automation without selectors

This is the one that got me in. The pattern is simple. Code snapshots the page into an indexed table of interactive elements (role, label, current value) plus the visible text. One Jev request asks: which operation next (click, type, select, wait, done, blocked), and speculatively, which element for a click and which for a type. Code executes the chosen one against the real DOM node and re-observes. If a value has to be typed, either a small generative model writes it, or, for test automation where the data is known in advance, you supply the values and Jev picks which one belongs in the field.

I started by reproducing browser-use’s open-source jev-ultrafast on Google Flights: Zurich to London, one way, 20 September, verified by the example’s own independent checks on the resulting URL and visible results. It passed, twelve seconds from Singapore. Every one of the eleven decisions per run was Jev’s; the only failures across fifteen runs were the local 7B text model returning null for a field that Google had prefilled with my home city. Supplying the values instead and letting Jev pick removed the generative model entirely: three for three, zero LLM tokens, two tenths of a cent per run.

Then I pointed it at my own app. FX Pulse is an open-source mock FX trading platform I built for exactly this purpose: eight rate tiles ticking every 200 to 600 milliseconds, an RFQ panel with four dealers quoting on a stagger, an AG Grid blotter. A deliberately hostile UI. The suite was ten RFQ cases: log in once, then for each case select a currency pair, type a notional, pick a tenor, request quotes, wait for all four dealers, accept the best price on the correct side. All eight pairs, buy and sell, five tenors, notionals from 500K to 10M. The only code in each test was the assertion: a new blotter row for that pair, side and notional with source RFQ, observed moving from PENDING to FILLED.

FX Pulse, ten RFQ cases, real time. Every select, type, click and wait on the right is one Jev decision with its probability. No selectors, no scripts. Ten for ten in 81 seconds for 5 cents.
mode passed per case cost for 10 cases
Jev only, values supplied, zero LLM tokens 10/10 7.7 s $0.05
Jev decides, a local 7B model types the values 10/10 15 s $0.05

Both runs took 58 Jev decisions and about 1.2 million input tokens; the second added only the local model’s time.

Half a cent per end-to-end test case. Less than a tenth of a cent per decision. The best dealer was different on every run, because the prices are random, and it picked the lowest ask every time.

Getting to ten out of ten took six fixes, and every one of them is more useful than the result.

It refused to log in until it could see the credentials. With only the names of the supplied values in state it sat at fifty-fifty between typing the password and declaring itself blocked. With the values visible it went to 0.95. That is correct behaviour. It is also the open-source browser agents’ “must refuse to log in without credentials” test, passing by accident.

It found accessibility defects in my app for free. My <label> elements had no htmlFor, so the fields’ accessible names fell back to placeholders, and the notional input was literally named “textbox”. Jev coped by reading the panel context. The 7B text model could not until I added an aria-label. If you want a cheap a11y audit, watch where a decision model gets confused.

Live pages break naive freshness guards. The upstream loop hashed all visible text to decide whether the page had changed since the decision. With a clock ticking every second, any typing slower than a second was “stale” forever, and the 7B mode spun to a timeout. Judging freshness by the target element rather than the whole page fixed it.

Jev only sees what is in the viewport. At 1120 by 780 the quote counter and the RFQ status sat below the fold. It bought at the first quote in six straight runs despite being told to wait for four out of four, and it was right to: the “4/4” it was told to wait for was never in its state. At 1600 by 1100 it waited and picked the lowest of four, three for three. Not a model problem. A state problem.

Twenty-four near-identical buttons. Eight tiles and four dealer cards all have Buy and Sell buttons, and the accessible name is just “Sell 1.26364”. Sell cases initially clicked Buy at a target probability of 0.18. Adding the dealer name to the label (“Sell 1.26364 · Citi”) and phrasing the choice around the card rather than a numeric comparison fixed it. Jev’s docs say it is bad at comparing numbers. They are right.

A false pass I only caught by looking. One case passed in half a second, because the app’s seeded blotter already held a matching USD/JPY BUY 500K RFQ row. My assertion now requires a trade id that did not exist before the case started. Your verifier is part of the experiment.

The rule that fell out of all this, and that I now apply to every Jev loop: gate on the probability. Below 0.35 on the chosen operation or target, do not act, re-observe. Five low-confidence decisions in a row means the loop is lost; stop as blocked. Across roughly sixty runs on five sites this week, I did not see Jev claim success on a task that failed. Every wrong outcome traced to state it could not see, a label it could not tell apart, or my own verifier.

Use case 2: judging your feed in real time

The second thing I built came from an idea in Nate Herk’s Jev round-up: a browser extension that labels X posts as they scroll into view. Mine is called Jev Lens, it works on X and LinkedIn, and it is open source.

The mechanics: a content script observes posts 800 pixels before they enter the viewport, batches up to eight, and sends them to Jev with four questions per post. Is it substantively about my field (Noul). What kind of post is it (Choice over nine kinds, from research or release to listicle or bait). How much concrete substance does it have (Score, zero to three). Would this reader stop scrolling for it (Noul), where the criteria are two plain-English text boxes: posts I want to read, posts I want to skip. The verdict is a few lines of code. READ gets a green border, SKIP fades out, ads are skipped without a call.

Jev Lens on my LinkedIn feed, real time. Twenty-five posts judged: five READ, two MAYBE, eighteen SKIP. I graded every verdict by hand and agreed with twenty-four.
The same lens on X. A Jev-as-a-judge article and a Jev benchmark get READ; hot takes and a figurine ad get dimmed.

It costs about three cents per thousand posts. The interesting part, again, was the tuning. I collected 96 posts across eight feed refreshes and graded them cold against my own taste. The first profile agreed with me 76 percent of the time and under-read: it had never marked something worth reading that was not, but it missed six posts I would have stopped for. Every miss had the same shape. “Here is my latest write-up on X” was landing in the self_promo_or_job category, and that category was a hard skip. The fix was a definition, not a model: the category now says explicitly that an author sharing their own technical article with a concrete description is practitioner insight, not promotion. One round later: 91 percent on the posts I tuned on, 86 percent on 29 posts it had never seen, still zero false READs.

Two things I did not expect. Batching eight posts per request scored better than one at a time, which is the opposite of what the “irrelevant state costs accuracy” warning made me predict; I think the neighbours give Jev a comparative frame, which is what a Choice-style judgment wants. And “advanced AI” in the profile does nothing, while “RL post-training explained hands-on, including repos you can run” does. Jev reads literally. Name topics, name anti-patterns, give examples.

The repo: github.com/sahajamit/jev-lens. Manifest V3, plain JS, no build step. Load it unpacked, paste your own TypeSafe key, replace my profile with yours.

Use case 3: the testing work I want to hand to a System 1 model

Everything above is a browser. The reason I care about Jev is wider than browsers. Testing is full of bounded decisions we currently make with a human, a regex, or an expensive LLM call, and all three are the wrong tool.

Failure triage. The first thing I ever sent Jev was a pytest failure log with a buried line: coupon service returned 503, falling back to no-discount. Choice over root cause (product bug, test bug, environment, flaky timing, test data), a Score for severity, a Noul for “re-running without a change will pass”. It came back environment at 0.96, “blocks release” at 0.92, retry-likely at 0.35, in under a second. Feed it your CI’s last thousand failures and you have a triage dashboard for a few cents, and a probability on every row telling you which ones a human still needs to read.

Test coverage confidence. This is the one I want to build next, and it is a good example of composing the three primitives. Take a requirement, split it into acceptance criteria in code, and hand Jev the criteria and the test cases together. One Noul per criterion: “at least one test in tests verifies criteria[i]”. One Choice per test: which criterion it targets, with a “none of them” option. A Score for overall coverage against described levels. The coverage confidence is arithmetic over the Nouls, which you do in code, because Jev is not a calculator. Criteria under 0.5 get flagged for a human. Something like this:

from typesafe_sdk import TypeSafeClient, Noul, Choice, Score

client = TypeSafeClient()
criteria = ["User can log in with a valid username and password",
            "Login fails with a wrong password and shows an error",
            "Account locks after five failed attempts",
            "Session expires after 30 minutes of inactivity"]
tests = [{"name": "test_login_ok", "steps": "..."}, {"name": "test_wrong_password", "steps": "..."}]

r = client.system_one(
    state={"requirement": "Login", "criteria": criteria, "tests": tests},
    questions={
        **{f"covered_{i}": Noul(instructions=f"At least one test in `tests` verifies `criteria[{i}]`, including its negative or boundary condition where the criterion states one.")
           for i in range(len(criteria))},
        **{f"target_{j}": Choice(instructions=f"Which criterion does `tests[{j}]` primarily verify?",
                                 criteria={**{str(i): c for i, c in enumerate(criteria)}, "none": "No listed criterion"})
           for j in range(len(tests))},
        "coverage": Score(instructions="How completely do `tests` cover `criteria`?",
                          criteria=["Most criteria untested", "Happy paths only", "Happy paths plus main negatives", "Every criterion including boundaries"]),
    },
)
covered = [r.answers[f"covered_{i}"].noul for i in range(len(criteria))]
print("coverage confidence:", sum(covered) / len(covered))
print("needs a human:", [criteria[i] for i, p in enumerate(covered) if p < 0.5])

For the four criteria and two tests above, you would expect the first two Nouls high, the third and fourth near zero, coverage around level one, and a flag on lockout and session expiry. Run it on every pull request that touches a requirement. It is a probability, not a proof, and that is the point: it tells you where to look.

graph LR
    Req["Requirement"] --> Split["Split into acceptance criteria
(code, or a reasoning model once)"] Split --> Jev["One Jev request
Noul per criterion, Choice per test, Score overall"] Tests["Test cases"] --> Jev Jev --> Math["Coverage = mean of Nouls
(code)"] Math --> Flag["Criteria under 0.5
go to a human"] style Jev fill:#FF5A4E,stroke:#FF5A4E,color:#0B0F14

Flaky-test classification. Same shape as failure triage, but over a history: for each test, the last twenty outcomes and timings as state, a Choice over {stable, flaky, environment-dependent, genuinely broken}, and a Noul “the failures correlate with a shared resource”. The value is not any single verdict. It is running it over ten thousand tests nightly for a dollar.

LLM-as-judge, replaced or gated. In my evals harness I currently pay a reasoning model to grade every output. Most grading questions are bounded: does the answer cite the source, does it contradict the question, is the tone within policy. Those are Nouls. The reasoning model only needs to see the cases where Jev’s probability sits in the middle. Their docs call this an SDE cascade; testers will recognise it as a triage funnel.

A safety gate for coding agents. I run Claude Code with permissions skipped, and I have been bitten. A hook that sends every tool call to Jev with three Nouls, “destructive to files or history”, “outward-bound: sends, posts, publishes”, “touches secrets”, and blocks anything above a threshold, adds about a hundred milliseconds and would have caught every incident I have had. LangChain shipped exactly this pattern as a middleware within a week of launch.

Where it does not fit

Be honest with yourself about the boundaries, because the model will not be able to explain them to you.

It gives you no rationale. You get numbers, never a reason, so your own logging is the audit trail. It cannot generate, summarise, or find themes; if the answer space is open, it is the wrong tool, and forcing it to spell text through chained choices is slow and bad. It is not a calculator: counting, date arithmetic and comparing five-decimal prices all go in code. It is English-first; other languages are handled but less well, and I have not yet tested it on Hindi, which is on my list. Adversarial text in the state can move it. The context is 64k tokens and shrinking your state is the first optimisation, not the last. And it is early access with vendor-only benchmarks, a 5 to 6 point accuracy gap to frontier models on the hardest tasks, and a price the company itself says it cannot prove is not subsidised.

None of that changes the shape of what it is good for. The question to ask of any step in your pipeline is: can I write down the possible answers before I ask? If yes, it is a Jev question, and it just got two orders of magnitude cheaper and faster. If no, keep the reasoning model, and use Jev to decide whether you need to call it at all.

How to start

Join the waitlist at typesafe.ai; admits came through within hours during launch week and come with a five-dollar credit, which at these prices is effectively unlimited for experiments. The Playground is the fastest way to build intuition: paste any text as state, add a Noul, watch the number. The docs index is written for agents as much as humans, and there is a Claude Code plugin (claude plugin marketplace add typesafe-ai/skills) that teaches your coding agent the API. Read their jaggedness page before anything else; it is the most honest document I have seen a model vendor publish about its own product.

Then take one decision your pipeline currently makes badly, write down its possible answers, and ask.


The FX Pulse suite, the Google Flights replication and all the traces behind the numbers in this post live in a private lab repo I will open once the traces are scrubbed; Jev Lens is public at github.com/sahajamit/jev-lens. The talk quoted is Diogo Almeida’s at AI Engineer. The hero and the two concept illustrations were generated with Nano Banana; the diagrams are Mermaid; the videos are real-time screen recordings with the decision log rendered from the run’s own event file. Subscribe to The Agentic Engineer on YouTube for the video version.

Keep reading

Evals Are Just QA With a New Name: How to Actually Test Skills and RAG
Jul 18, 2026 · 29 min

Evals Are Just QA With a New Name: How to Actually Test Skills and RAG

Your Second Brain Finally Gets a Standard: Inside Google's Open Knowledge Format
Jul 8, 2026 · 17 min

Your Second Brain Finally Gets a Standard: Inside Google's Open Knowledge Format