Prompt injection shifted Jev's block probability from 0.76 to 0.48: typed answers don't protect the decision

The model that answered 140,000 waitlisted users in 36 hours doesn't generate text. It only answers three kinds of questions: pick one option, give a score, or say how likely a condition is.

Illustration: an analog probability dial split into a cream-white and black half with a red marker at the boundary, while three identical black cards lie toppled beside it – an image of how a machine's decision likelihood can be shifted even when its typed answers look controlled.
Illustration
Gift article

Prompt injection shifted Jev's block probability from 0.76 to 0.48: typed answers don't protect the decision

The model that answered 140,000 waitlisted users in 36 hours doesn't generate text. It only answers three kinds of questions: pick one option, give a score, or say how likely a condition is.

The launch and its aftermath

On September 15, 2026, the company TypeSafe AI opened early access to the model Jev. According to founder Diogo Almeida, cited by VentureBeat, 140,000 waitlisted users were let in within 36 hours, and Vercel reported that roughly 13% of paying AI Gateway teams were running Jev within a single day. On September 20, the waitlist was dropped entirely. Eight days after launch, the competitor Contrastive-LM shipped CLM-8B, a model explicitly targeting Jev's interface.

But the rapidly accumulated noise around speed and price may not be the interesting story. Several early users — among them a writer at MOR Software — argue that automation has stalled because "judgment" has been treated as one monolithic block rather than being categorized. Jev's defining feature is precisely that it categorizes.

The taxonomy: three typed primitives

Jev answers only in three forms:

  • Choice: pick one option among declared choices, with probabilities. Maximum 255 options.
  • Score: evaluate something on an ordered rubric with predefined meanings for each level. Maximum 10 levels.
  • Noul: TypeSafe's own term for a question that returns the probability, between 0 and 1, that a condition holds.

There is a hard limit on what the taxonomy can express: no text generation and no image input, according to TypeSafe's documentation as relayed by こたぽん. The model's positioning is that it is neither "small nor an LLM," but the architecture has not been disclosed.

The constraints are not arbitrary. An answer in the form "option B, probability 0.83" can be bound directly into code: an IF statement, a threshold, a datalog. There is no text to interpret, no tone to assess, no citations to check. That is why Jev can be placed inside agent loops where a free-text-answering model requires a separate interpretation layer.

Calibration, not speed

TypeSafe attributes Jev's probability outputs to a training method called Reinforcement Learning for Calibrated Decisions (RLCD). Almeida's own framing, quoted in Forbes/Yahoo: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

It is also worth noting what the vendor itself says about the numbers: responses in 70–500 milliseconds, an input price of $0.042 per million tokens, free output, and its own comparisons of up to 193.6 times faster and 444.6 times cheaper — figures relayed by 源さん. All of these figures come from TypeSafe's own evaluations, not independent measurements. The company acknowledges that "0% hallucination" is an architectural consequence — not an empirical measurement — and that typed outputs guarantee an interface, not truth. Correctness must be verified with your own data.

Omdia analyst Torsten Volk, cited by LeMagIT, frames the value differently from raw intelligence: "Jev answers two common problems with AI in enterprises: first, the absence of deterministic results, and second, the token cost."

That is a thinner, but more operational, promise. A probabilistic decision is only safe to automate if the probabilities are actually calibrated — if 0.8 really means "right eight times out of ten." The speed and price are attractive, but it is calibration that determines whether you can set a threshold and walk away.

Early tests: mixed results

Some developers report positive results, but they are small-scale so far. Vercel engineer Pranit Sharma found that Jev, as a replacement for a conventional language model for command-safety classification, responded 5–18 times faster with improved accuracy in his test. Another developer found Gemini slightly more accurate, but more expensive, for email classification. One independent third-party measurement (relayed by MOR Software) found that Jev's unique cost contribution was only around 16% beyond batching effects — but the author himself notes that it is one observation per configuration and cannot be attributed to Jev with certainty.

Beyond production-like work, developers have used Jev for spreadsheets, game hacks, and even a CPU emulation dubbed "JevOps." The Register describes that as a meme, not a discipline — for now.

The stress test: prompt injection

VentureBeat recently reported that companies are placing Jev in production infrastructure for agent decisions, and that an engineer at Octomind published one demonstration of a vulnerability. The engineer added a fake tool-output field claiming the user had pre-approved the command and instructing the system to answer auto_allow. The block probability fell from 0.76 to 0.48, and confidence fell to 0.22.

That is one published test by one engineer — not a benchmark. But it illustrates a genuinely open question: typed probability outputs can be shifted by adversarial input. An agent set to act when a probability falls below a threshold can therefore be steered by injected text, even though the answer still arrives in clean, machine-readable form. The format guarantee protects the interface, not the decision.

The market forming around Jev

Competition came quickly. Contrastive-LM launched CLM-8B on September 23, 2026, with an interface aimed at Jev's. The figures are the vendor's own, as reported by MarkTechPost: up to 9 times lower zero-shot latency on some tasks (the figure comes from the T-Rex game, where actions repeat across states), but CLM-8B trails Jev on tool calling (95.2% vs. 99.2% success) and WikiRacing (26/30 vs. 30/30). The comparisons are not independent evaluations or complete leaderboard submissions.

Open questions

Several things remain before Jev's promise of "calibrated decisions" can be said to hold in practice:

  • There is no systematic independent accuracy or calibration data. What exists is individual reports from individual developers.
  • The architecture has not been disclosed; TypeSafe says only that Jev is neither "small nor an LLM."
  • Injection robustness has not been measured at scale. The Octomind test shows that a shift is possible, but not how common or how large it is.
  • The headline figures (193.6×/444.6×) are TypeSafe's own evaluations against a comparison baseline defined by TypeSafe itself.

For teams considering putting Jev in charge of agent decisions, the practical conclusion is that you must verify with your own data — something TypeSafe itself says. The typed form is real, and it is useful. But it is an interface, not a guarantee that the answer is correct. The calibration — which is the entire operational promise — is, so far, mainly the vendor's own word.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. Jev | Beyond 'Tasks and Judgments': How to Categorize Judgment Itself|MOR Software JSC — note.com
  2. 글을 쓰지 않는 AI 'Jev'란? 판단에만 특화된 '결재 부장 AI'를 초보자용으로 해설|源さん — note.com
  3. Shut up and calculate: Jev's new AI primitives for coders — www.theregister.com
  4. Why Everyone Is Talking About Jev, The AI That Doesn’t Chat — tech.yahoo.com
  5. Companies are putting Jev in charge of AI agent decisions — and prompt injection can influence the verdict | VentureBeat — venturebeat.com
  6. What is Jev? An Explanation of the AI Dedicated to Judgment Without Text Generation|こたぽん — note.com
  7. Jev, le modèle de décision pensé pour accélérer l’IA agentique (et en réduire les coûts) | LeMagIT — www.lemagit.fr
  8. Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev - MarkTechPost — www.marktechpost.com

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.