Agentic Engineering
OpenAI Decisions API Costs 2.38x Jev's Token Rate
OpenAI's Decisions API prices input at 2.38x Jev's rate; teams need to compare accepted decisions, image handling, and fallback costs.
Routing teams should test OpenAI’s Decisions API, released in public beta on October 6, against their existing classifier before moving production traffic. Its base input tariff is 2.38x Jev’s, a small absolute premium that needs to buy better accepted decisions, simpler integration, or a capability the cheaper route lacks.
A cheap decision still needs to earn its premium
The arithmetic joins two price sheets retrieved on October 7, 2026. OpenAI’s Decisions guide lists $0.10 per million input tokens; TypeSafe’s current Jev model reference lists $0.042. Dividing 0.10 ÷ 0.042 yields 2.38x after rounding. Subtraction gives $0.058 more per million input tokens, or $58 per billion. Those are base-rate comparisons, not measurements of equivalent tasks: tokenization, prompts, retries, and escalation can change the amount each provider bills.
OpenAI uses GPT-6 Luna at a dedicated endpoint that returns predicates, choices, or rubric scores. The guide specifies input-only billing, with no cache-read, cache-write, or output-token charges, while regional and long-context premiums can apply. Do not import an ordinary chat-model budget just because the model name is familiar. The endpoint is the product being purchased; its published billing terms govern the comparison.
That distinction makes this launch more useful than another generic price cut. A program deciding which queue owns a complaint does not necessarily need a paragraph. It needs an allowed answer and a way to handle uncertainty. If a team currently pays for prose, parses a department name out of it, and repairs formatting failures, a narrower interface gives it a specific component to replace. It does not establish that the underlying judgment improves.
The competitive context already exists. Our September Jev analysis separated its input-price advantage from accuracy, while the Clef comparison found a substantial quality trade-off inside one vendor’s cheaper tier. OpenAI’s entry adds another candidate to that decision layer. It gives buyers more reason to preserve an evaluation set and less reason to standardize on whichever launch has the most striking multiplier.
The absolute premium should discipline the procurement discussion. At equal billed input volume, a billion tokens separates the two base tariffs by $58. That is not a forecast of monthly spending, and it supplies no wage rate or error price. It does show what the buyer must compare with integration effort and downstream mistakes. Negotiating over a dramatic relative difference can miss the larger expense of maintaining a second provider or sending ambiguous cases to a costly fallback.
A useful business case therefore starts with the existing workflow. Identify the decision, the evidence it requires, the consequence of selecting the wrong category, and the current handling of incomplete inputs. Then ask whether a typed endpoint reduces the total work. A classifier that is fast in isolation but forces another model to explain every answer may improve latency without delivering the expected budget reduction.
Buy the endpoint’s contract, not the model’s name
The Decisions API reference accepts a plain text string or user messages containing text and inline images. It does not accept ordinary function-call history, file references, audio, or non-user roles; image inputs use data URLs, with a maximum of 128 images per request. It can return a refusal for an individual question. Those details belong in the migration estimate because an existing conversational request cannot simply be forwarded unchanged.
The engineering implication is an adapter with explicit boundaries. Select the evidence for the decision, construct the allowed answers, and route the result into application code. Preserve the original input and question version so a reviewer can reproduce an error. A failed or refused judgment should enter a defined fallback, not silently become a negative answer. These are deployment recommendations, not capabilities we tested in a paid API trial.
TypeSafe’s current interface accepts text, including structured text state, rather than native images. For a visual workflow, preprocessing is therefore part of the comparison. An image-capable decision endpoint may eliminate a separate extraction step, but that saving must be measured. Comparing raw token prices alone would omit the part of the pipeline that turns the picture into evidence.
The distinction between deciding and generating also survives within OpenAI’s own platform. Its Structured Outputs documentation warns that schema-conforming answers can still contain mistakes. If the application requires extracted fields or an explanation, a generated object remains a different requirement from selecting an option. Forcing a rich task into a narrow classifier can move complexity into application code without removing it.
Billing deserves its own integration check. The general API pricing page separates models and processing tiers, while prompt-caching documentation describes reuse of matching prompt prefixes. Neither is permission to substitute a familiar cached-input rate for the Decisions guide’s endpoint-specific terms. Capture the actual usage returned during a pilot and reconcile it with the invoice before projecting savings. A spreadsheet that assumes a discount the purchased endpoint does not promise is not a cost model.
Data handling can justify choosing a provider even when the tariff is higher. OpenAI’s data-controls documentation distinguishes residency, regional processing, retention, and eligibility. The Decisions guide identifies support for eligible customers and particular processing regions; a buyer still needs the applicable configuration and agreements. Treat those requirements as procurement gates rather than as points in a generic leaderboard. A cheaper route that cannot satisfy the intended deployment terms is not an available substitute.
Today’s Nano Banana analysis follows a cost shift between image output and input. The same discipline applies at API scale: identify what is included, what is metered, and what remains elsewhere in the stack. A well-defined unit makes arithmetic possible. It does not make unlike products interchangeable.
Faster answers can make the wrong thing happen sooner
OpenAI describes Decisions as about 10x faster than the Responses API. That is a vendor claim, not our measured end-to-end speedup. The launch documentation reviewed here does not establish a matched latency-and-accuracy comparison with Jev on a buyer’s workload. A ratio without the relevant input sizes, region, concurrency, and acceptance threshold cannot select the production winner.
There is a plausible reason to investigate. OpenAI’s latency guidance identifies token generation as a major contributor to response time. Removing an unnecessary generated answer can shorten the critical path. But the whole application still includes evidence retrieval, network time, code execution, and any later review. A faster decision is valuable only to the extent that the decision is what users are waiting for.
That measurement boundary echoes Intel’s benchmark configuration in today’s compute brief. A controlled result can be real and still answer a narrower question than the sales headline suggests. For routing software, measure from the arrival of a usable request to an accepted result, and retain the component timings. Otherwise an improvement in one stage can conceal an unchanged overall experience.
Uncertainty is the harder issue. TypeSafe’s confidence documentation derives its confidence field from the distribution over possible answers. That is useful information for application logic, but a concentrated distribution is not independent proof that an answer is correct. Nor should a threshold migrate between providers merely because both expose a number with the same name. The meaning and observed behavior of the score must be established on the buyer’s cases.
TypeSafe also publishes specific Jev limitations involving literal interpretation, arithmetic, dates, and irrelevant context. Those are reasons to keep deterministic comparisons in code and to supply focused evidence. They are not evidence that OpenAI wins those tasks. They offer a concrete test design: include the failure modes your present system encounters, then compare actual errors rather than importing either vendor’s confidence about itself.
The source of the evidence matters as much as the decision model. Today’s EmbeddingGemma brief examines the trade-off between retrieval vectors and task scores. A classifier receiving the wrong passage cannot rescue a retrieval evaluation by returning a perfectly typed answer. Test the decision component in isolation to understand it, then test the full path to discover where the operational mistakes actually arise.
The strongest argument against migration is mundane: the existing classifier may already satisfy the product’s needs. If a rule based on a trusted record chooses the correct queue, adding a probabilistic service creates latency and another failure path. If a current model already meets the acceptance bar, a cheaper or faster alternative must recover its integration and maintenance cost. The launch earns a test; it does not create an obligation to change a working system.
Start with a reversible decision
A sensible pilot begins in observation mode. Keep the current process authoritative and ask each candidate to propose the same bounded decision. Use adjudicated examples with routine cases, uncertain cases, and cases outside the allowed categories. Keep all outcomes, including refusals, timeouts, and escalations. Dropping those rows would make the successful calls look economical while hiding the cost of making the workflow dependable.
OpenAI’s accuracy guidance recommends a baseline evaluation set and inspection of failure patterns before more elaborate optimization. Apply that advice to the new endpoint before redesigning the entire agent. Freeze the questions and labels for the comparison, and record changes when a category definition evolves. A model trial cannot be interpreted if the question changes halfway through it.
Set acceptance rules around consequences. A misrouted internal search result and an incorrectly approved transaction do not have the same cost, even if both fit a choice schema. The application should retain authority over consequential actions and review paths. The model’s contribution is evidence for a bounded decision; the product owner remains responsible for deciding which outcomes may proceed automatically.
The financial comparison can stay simple without becoming simplistic. Add actual decision charges, preprocessing, retries, escalations, and the observed cost of correcting mistakes. Compare that total with the existing route at the same acceptance bar. Keep integration work visible as a separate investment instead of amortizing it over an invented traffic forecast. This edition has calculated tariff differences; it has not measured either service’s production accuracy or operating cost.
Expand only when the evidence changes the verdict. A lower fallback rate, a useful native-image path, or better performance under the required regional configuration could justify OpenAI’s premium. Equal quality with a lower total bill could justify Jev. No meaningful gain could justify keeping the current implementation. The important output of the pilot is a decision the team can explain, including why the attractive alternative lost.
- Routing owners: trial a bounded classification step whose current errors and latency are already observable; preserve the existing route until the new one clears the same acceptance bar.
- Platform engineers: budget for input adaptation, refusal handling, model-version logging, and rollback; measure full-path latency and billed usage instead of relying on the advertised speed multiplier.
- Procurement and product leads: compare the $0.058 per-million-input-token premium with measured preprocessing and fallback costs, then verify deployment terms. Watch for beta contract changes before widening access.
The cheapest classifier is the one whose accepted decisions cost least to operate. A smaller invoice for an isolated call is a useful input to that calculation, not its conclusion.
Sources
- OpenAI — October 6 Decisions API public-beta release
- OpenAI — Decisions billing, typed answers, and availability
- TypeSafe — current Jev tariff, input types, and model identifiers
- OpenAI — Decisions request limits and refusal response
- OpenAI — Structured Outputs and remaining semantic mistakes
- OpenAI — general model and processing-tier prices
- OpenAI — prompt caching behavior
- OpenAI — data-control eligibility and regional configuration
- OpenAI — token generation and latency optimization
- TypeSafe — confidence as a statistic over answer probabilities
- TypeSafe — documented Jev 1.13 limitations
- OpenAI — evaluation baselines and failure analysis