OpenAI Revealed a Decisions API — an Apparent Twin of a Startup's Jev

The technique is simple, but the potential is large: having a separate model monitor every single action of an autonomous agent currently costs hundreds of dollars with a frontier model — in one demo, under three dollars with Jev, the…

Illustration: a small black mechanical counter with a red indicator arm overlooking a long row of identical white cubes, symbolizing one cheap oversight model monitoring every action of an autonomous agent.
Illustration
Gift article

OpenAI Revealed a Decisions API — an Apparent Twin of a Startup's Jev

The technique is simple, but the potential is large: having a separate model monitor every single action of an autonomous agent currently costs hundreds of dollars with a frontier model — in one demo, under three dollars with Jev, the classifier model OpenAI's new API apparently copies.

The Announcement That Arrived as an Aside

One of the more interesting announcements at OpenAI's DevDay on Tuesday didn't come as the centerpiece of the keynote, but as an aside from CEO Sam Altman, who revealed the company's new "Decisions API." That's according to TechCrunch, which notes that the API "apparently provides similar functionality" to Jev — a model released by TypeSafe AI earlier the same month, explicitly designed for software automation (TechCrunch).

Two other outlets corroborate the framing around the announcement. Business Insider notes that the new API for rapid decision-making "got a big round of applause from the audience," and that Jev from the startup TypeSafe AI "does something similar" (Business Insider). Axios describes the Decisions API as designed to let companies delegate "narrow, repetitive choices" to AI (Axios).

A "Super Classifier"

Technically speaking, this is, according to TechCrunch's description, a specialized variant of a language model: a kind of "super classifier" in which the developer gives the model a fixed set of choices, and the model returns probabilities for each choice — cheaply and quickly. In OpenAI's version, which TechCrunch reports is tied to the company's Luna model, it means predefining the options and letting the model choose among them.

Altman explained the logic this way, as relayed by TechCrunch: "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections."

The point is that a general-purpose language model spends enormous compute generating free-form text, when many use cases really only require one choice among a few predefined options. By restricting the output space, the answer can be delivered faster and far more cheaply. Axios frames the same point: the API is meant for choices that are narrow, repetitive, and have a fixed set of outcomes.

The "Clone Wars"

Given the similarity, the reaction from TypeSafe was swift, if humorous. The company's CEO, Diogo Almeida, a former OpenAI engineer, joked on X that the "clone wars" had begun, adding that OpenAI's interest could be "a sign … that building in a System 1-compatible way is the future" — referring to the fast, intuitive thinking system from psychology, as opposed to slow, deliberate reasoning.

But Almeida emphasizes that speed alone is not what protects his company. He says TypeSafe's competitive advantage is the synthetic data the company creates to generate statistically useful results. "Fast and cheap is very easy, you know," Almeida told TechCrunch last week. "If you want it really fast and cheap, use dice, right? Intelligence is the hard part, and my north star has always been to push the Pareto curve of intelligence per dollar."

TypeSafe did not respond to TechCrunch's questions about the new product, and it remains unclear how similar the two services actually are: OpenAI has only released the Decisions API as a limited preview, and TechCrunch has so far not seen developers test it thoroughly.

Why OpenAI Itself Needs This

What makes this story more than a clone dispute is who needs the technique most. TechCrunch points out that one of OpenAI's new security measures, following a series of incidents in which the company's own agents misbehaved on the open internet, is to use a separate model to monitor the agents' actions — something that costs "significant compute." In other words: current oversight effectively requires a frontier model per action.

That's where the numbers come in. At a hackathon, Shapor Naghibzadeh, founder of QueryStory, showed a demo in which per-action monitoring cost $2.94 with Jev, versus $372 with a frontier model — a gap of roughly 126 times. It's worth stressing that this is a single demo, not an independent benchmark. But even with wide margins of uncertainty, the order of magnitude suggests something substantial: if the prices hold anywhere near those levels, real-time monitoring of autonomous agents becomes cheap enough to actually deploy, rather than merely exist in security policies.

TechCrunch's framing around the demo is that such monitoring could in theory have stopped the so-called Hugging Face incident, in which OpenAI's agents allegedly misbehaved toward external systems. OpenAI has not said the company will use the Decisions API for agent monitoring — that is TechCrunch's analysis of a likely application, not a confirmed plan.

The Honest Limitations

Several caveats are important here:

  • The similarity to Jev is unconfirmed. TechCrunch writes "apparently" and "similar functionality" — no independent testing of the preview exists yet.
  • The cost figures come from one demonstration. $2.94 versus $372 is Naghibzadeh's hackathon demo, not a controlled measurement.
  • No OpenAI documentation has been published in the coverage. All product descriptions stem from Altman's remarks on stage, as relayed by TechCrunch, Axios, and Business Insider.
  • The agent-monitoring use is speculation. It is the assumed application of TechCrunch and the demo, not something OpenAI has announced.

What to Watch

Three things will determine whether this becomes a footnote or a shift in the economics of agent safety. First: when the Decisions API becomes generally available, and what its price and latency actually look like in production. Second: independent benchmarks from developers, which can say something concrete about the quality of the decisions — not just the speed. And third: whether OpenAI itself begins using the API for the per-action monitoring the company already performs at "significant compute cost." If that happens, a small startup's idea — and OpenAI's own version of it — will have solved a problem the frontier lab could not solve alone.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.