How to Connect JEV to Codex: Typed Decisions Without Replacing GPT

Codex is a general coding agent: it reads and writes code, operates tools, works across files and terminals, and keeps track of a user’s objective. TypeSafe AI’s JEV serves a much narrower purpose. Instead of generating an essay or editing a repository, it examines a supplied state and returns a named choice, an ordered score or a yes probability.
Connecting the two is therefore not a model swap. Codex can continue using GPT—or another supported primary model such as DeepSeek—while calling JEV at one bounded decision point. TypeSafe’s coding-agent guide explicitly says JEV is not a replacement for the underlying LLM in Codex, Claude Code or Cursor. It does not generate code, converse with the user or invoke tools.
The useful architecture is straightforward: Codex gathers evidence and acts; JEV advises on a narrowly specified decision. JEV’s probabilities can inform a policy, but they do not create authority to deploy, delete data or send information outside the workspace.
The three-point summary
- Connecting JEV leaves the primary Codex model, conversation, tools and workspace intact.
- TypeSafe’s official agent skill is a knowledge layer; direct calls from a Codex conversation require an MCP server or application integration.
- Choice, Score and Noul work well for routing and risk signals, but arithmetic, dates, free-form generation and high-impact approvals belong to deterministic code, Codex and people.
What JEV is
TypeSafe’s introduction describes JEV as its first System One model. A general LLM produces open-ended language. JEV evaluates named questions against a shared state and returns typed values that software can consume without trying to parse a paragraph.
JEV supports three decision primitives:
Choiceselects one option from a list supplied by the application.Scoreselects one level from an ordered rubric.Noulreturns the probability, from zero to one, that one statement is true.
One request can evaluate several independent questions against the same state. A release candidate, for example, can be evaluated for its best next step, risk band and probability of being ready to deploy. Each question should still represent one decision. TypeSafe recommends splitting compound judgments into atomic questions and combining the outputs in application code.

JEV does not become a Codex model-picker option
The end-to-end path looks like this:
- The user gives Codex a task.
- Codex reads the repository, tests and operational constraints.
- At one bounded decision point, Codex sends only the required state to an MCP tool.
- The MCP server calls JEV and receives a typed response.
- Codex compares that response with direct evidence.
- Codex and the user decide whether to edit, run a command or deploy.
JEV does not click, run shell commands, edit files or deploy. That separation is a security property. It avoids turning one probabilistic judgment into an immediate external action and keeps the primary agent accountable for inspecting the actual evidence.
Two ways to add JEV to a Codex workflow
1. Install TypeSafe’s official agent skill
TypeSafe publishes an official agent skill that teaches coding agents the API, question patterns and evaluation practices. Its documented general installation command is:
npx skills add typesafe-ai/skills --skill typesafe-ai
The skill helps Codex write an API client or review JEV integration code. It does not by itself expose a live JEV tool inside the conversation. A skill is guidance and reusable instructions; an application or MCP server still has to make the network request.
This path makes sense when a development team is building JEV into its own product and wants Codex to follow TypeSafe’s official patterns while writing and testing that code.
2. Add a private MCP plugin for direct calls
If Codex should ask JEV directly during a conversation, it needs a callable MCP server. The environment tested for this article uses a private JEV Decision Helper plugin with a local stdio MCP process and two deliberately narrow tools:
statuschecks whether a credential is configured without revealing it or making a billable request.decidevalidates a state and one or more Choice, Score or Noul questions before sending them to JEV.
This private helper is not an official TypeSafe Codex MCP plugin. It is an independently maintained bridge to TypeSafe’s documented API. Any organization adopting the same pattern should review and own the wrapper code or choose an implementation it can audit.
OpenAI’s plugin documentation says a Codex plugin can bundle skills, MCP configuration or both. Tool schemas should be explicit, credentials should stay out of plugin files and distributable archives, and a stdio server should receive only the environment variables it actually needs.
What the request contains
TypeSafe’s API documentation specifies POST https://api.typesafe.ai/v1/systemone with Bearer authentication and a JSON body. The important fields are:
model: usejev-latestto follow the current stable alias or pin a version for controlled production behavior.state: a string, object or array containing only facts relevant to the decision.questions: a map of independently named questions.
A narrow release-state request can look like this:
{
"model": "jev-latest",
"state": {
"change": "login session middleware and database migration",
"automated_tests_passed": 42,
"security_integration_tests_failed": 1,
"rollback_plan": true,
"deployment_started": false
},
"questions": {
"allow_deploy_now": {
"type": "noul",
"instructions": "Using only the supplied state, estimate the probability that deployment should begin now."
}
}
}
A production MCP tool should also validate the allowed Choice options, the ordered Score levels and the criteria for every question. If downstream code accepts arbitrary strings as commands or paths, typed output no longer provides a meaningful safety boundary.
When to use Choice, Score and Noul

Choice: select a route from named options
Choice is appropriate when the application already knows the legal options, such as fix_and_retest, request_human_review and deploy_with_monitoring. The response includes the selected choice, the probability assigned to every option and a separate confidence value.
Good uses include:
- Routing a bug to front-end, back-end or infrastructure ownership
- Choosing whether to repair, gather more evidence or request review
- Assigning a customer request to one of several established queues
Score: evaluate an ordered rubric
Score uses levels such as 0=low, 1=moderate, 2=high and 3=critical. It returns the selected level, a distribution across levels and confidence. The number is a label in an ordered policy—not the result of a mathematical calculation.
Risk, urgency and review priority fit this primitive. Invoice totals, calendar intervals and exact counts do not. Those should be calculated by deterministic code before the relevant result is added to the state.
Noul: estimate the probability of one proposition
Noul evaluates one explicit yes-or-no statement, such as this change is ready to deploy now, and returns a value from zero to one. It does not have the separate confidence field returned by Choice and Score.
Do not treat a probability attached to one Choice option as mathematically interchangeable with a Noul value. The question structures and calibration can differ, so each primitive requires its own evaluation.
A live Codex-to-JEV connection test
The connection test used no passwords, API keys or customer records. JEV received only a synthetic, non-identifying release state:
- Login-session middleware and a database migration had changed.
- Forty-two automated tests passed.
- One security integration test failed.
- A rollback plan existed.
- Deployment had not started.
One request carried three independent questions. jev-latest, resolving to JEV 1.13.0, returned:
- Choice for the next step:
fix and retest, 98% option probability, confidence 0.97 - Score for risk:
2=highon a four-level rubric, 100% level probability, confidence 1.0 - Noul for deploying immediately:
0.03

The answer is sensible, but this is not an accuracy benchmark. A single example verifies the connection and response shape. Production use requires a representative labeled set that measures false positives, false negatives, class-specific behavior and threshold stability.
Codex still has work to do after receiving the result: inspect the failing test, read the changed files, apply an authorized fix and rerun validation. A Noul value of 0.03 does not replace a deployment block, change-management rule or human approval.
Turn confidence into a policy, not a feeling
TypeSafe’s confidence guide suggests routing high-confidence cases to automation, medium cases to confirmation or more information, and low-confidence cases to human review. It does not prescribe one universal threshold.
A responsible policy process is:
- Build a labeled evaluation set from historical cases.
- Decide the business cost of each false positive and false negative.
- Set different thresholds for low-risk classification and high-impact actions.
- Route close distributions and threshold-edge cases to people.
- Log the question version, model version, input summary, response and final human decision.
- Re-evaluate whenever the prompt, criteria, model or operating domain changes.
An 0.80 confidence threshold may be acceptable for document tagging while a production deployment remains human-approved even at 0.99. Confidence and authorization answer different questions.
U.S. engineering and governance considerations
U.S. teams should treat JEV as a third-party inference service when data crosses the network boundary. Before submitting proprietary code, customer tickets, employee records or regulated information, review current retention, training-use, subprocessors, deletion, incident response and contractual controls. A public model card and price page do not establish that a particular HIPAA, financial-services or government requirement is satisfied.
Operational controls should include:
- Version pinning for any workflow whose thresholds have been calibrated
- Per-action thresholds instead of one global confidence cutoff
- Audit logs that connect a JEV result to the later human or Codex action
- Redaction before state construction, not after the remote request
- A fail-closed path for timeouts, schema errors and rate limits
- Human review for deployments, access changes, deletion and external communication
English is JEV’s primary training language according to the model documentation. That helps many U.S. workloads, but domain-specific language can still shift performance. Legal support tickets, medical abbreviations, security incidents and internal product terminology each need their own evaluation set.
Credential and plugin security
Never place a JEV API key in source code, .mcp.json, an article, logs or a Git repository. The tested helper stores its credential in a separate Codex secret file and refuses requests if the file is readable beyond its owner. Its status command reports only whether configuration exists.
Additional controls are equally important:
- Pass only required environment variables to the MCP subprocess.
- Exclude unrelated conversation history, passwords, tokens and raw customer records from state.
- Redact authorization headers and request identifiers from logs.
- Never execute a returned string directly as a shell command or file path.
- Start a new Codex session after changing a plugin so the current tool schema is reloaded.
- Test configuration status first, then use synthetic data for the smallest live call.
Good and bad fits
JEV is strongest when the legal outputs and decision criteria can be defined before the request:
- Issue routing and ownership classification
- Security, quality and urgency scoring
- Determining whether a case needs review
- Selecting one recovery path from a closed set
- Providing an advisory signal to a larger guardrail
- Prioritizing cases before a human queue
Use Codex, deterministic code or a qualified reviewer instead for:
- Writing code or open-ended explanations
- Exact arithmetic, counting and date comparison
- Long chains of indirect reasoning
- Reading images, audio or video directly
- Granting authority or making legal and medical final decisions
- Deploying, deleting, charging or sending external messages
TypeSafe’s JEV 1.13 limitations page documents literal interpretation, degraded results with oversized irrelevant state, conflicts between instructions and criteria, and susceptibility to adversarial content or prompt injection. Media must first be converted into verified text or structured fields by another tool.
Context limits, versions and pricing
According to TypeSafe’s model page, the stable model on September 27, 2026 is jev-1.13.0, and jev-latest currently points to it. The total request context limit is 64K. State plus the longest single question is limited to 32K. The input is text-only.
Input costs $42 per billion tokens, equivalent to $0.042 per million tokens, and output tokens are free. The low unit price is not a reason to send entire repositories or conversation histories. Larger state can raise privacy risk, reduce relevance and make failures harder to reproduce.
Published rate limits can change. Direct HTTP clients should implement exponential backoff for 429 and 529 responses and should not convert a retry storm into multiple external actions.
jev-latest is convenient during development. A production policy calibrated to response distributions should pin jev-1.13.0 or another evaluated version because the alias can move when TypeSafe releases an update.
Post-installation checklist
- Create a TypeSafe API credential and store it in an OS secret store or owner-readable file.
- Distinguish the official agent skill from the MCP process that makes live calls.
- Expose only narrowly scoped MCP tools with strict input and output schemas.
- Start a new Codex session and confirm that the expected tools are present.
- Use a non-billable status check without printing the credential.
- Exercise Choice, Score and Noul with synthetic, non-identifying data.
- Verify that timeouts, schema failures, 429s and 529s stop safely.
- Keep questions, criteria and thresholds in one reviewable policy module.
- Measure false positives and false negatives on domain-specific data.
- Require user authorization and direct validation for deployments, deletion and access changes regardless of JEV’s answer.
Bottom line
JEV and Codex are complementary, not competing, models. Codex remains the primary agent that understands the goal, operates tools and is accountable for the completed task. JEV contributes three narrow outputs—Choice, Score and Noul—at points where prose would be harder to validate and automate.
The best first deployment is reversible: routing a case, deciding whether to request review or selecting what evidence to collect next. Keep the questions atomic, minimize state, calibrate distributions on your own data and separate advice from execution authority. Under those constraints, JEV can make a Codex workflow more explicit without pretending that probability is permission.
Trademark and independence notice
JEV and TypeSafe names belong to TypeSafe AI. Codex, GPT and related names belong to OpenAI and their respective owners. This independent editorial article is not sponsored, endorsed or approved by either company. Its images do not reproduce official product interfaces, logos, third-party photography or company charts.
Official sources
- TypeSafe AI: JEV introduction and decision primitives
- TypeSafe AI: JEV in coding-agent workflows
- TypeSafe AI: API endpoint and response types
- TypeSafe AI: Models, context and pricing
- TypeSafe AI: Using confidence
- TypeSafe AI: Known limitations of JEV 1.13
- TypeSafe AI: Official agent skill
- OpenAI Developers: Codex plugin structure
- OpenAI Developers: Docs MCP



