This is the English edition. 한국어판 and 日本語版 are also available.

How to Connect JEV to Codex: Typed Decisions Without Replacing GPT

2026-09-27 · AI · United States · Zoogom Editorial

#JEV#TypeSafe AI#OpenAI Codex#MCP#coding agents#AI decision systems

A coding agent remains in control while a compact decision engine returns three structured judgments

Codex is a general coding agent: it reads and writes code, operates tools, works across files and terminals, and keeps track of a user’s objective. TypeSafe AI’s JEV serves a much narrower purpose. Instead of generating an essay or editing a repository, it examines a supplied state and returns a named choice, an ordered score or a yes probability.

Connecting the two is therefore not a model swap. Codex can continue using GPT—or another supported primary model such as DeepSeek—while calling JEV at one bounded decision point. TypeSafe’s coding-agent guide explicitly says JEV is not a replacement for the underlying LLM in Codex, Claude Code or Cursor. It does not generate code, converse with the user or invoke tools.

The useful architecture is straightforward: Codex gathers evidence and acts; JEV advises on a narrowly specified decision. JEV’s probabilities can inform a policy, but they do not create authority to deploy, delete data or send information outside the workspace.

The three-point summary

  1. Connecting JEV leaves the primary Codex model, conversation, tools and workspace intact.
  2. TypeSafe’s official agent skill is a knowledge layer; direct calls from a Codex conversation require an MCP server or application integration.
  3. Choice, Score and Noul work well for routing and risk signals, but arithmetic, dates, free-form generation and high-impact approvals belong to deterministic code, Codex and people.

What JEV is

TypeSafe’s introduction describes JEV as its first System One model. A general LLM produces open-ended language. JEV evaluates named questions against a shared state and returns typed values that software can consume without trying to parse a paragraph.

JEV supports three decision primitives:

One request can evaluate several independent questions against the same state. A release candidate, for example, can be evaluated for its best next step, risk band and probability of being ready to deploy. Each question should still represent one decision. TypeSafe recommends splitting compound judgments into atomic questions and combining the outputs in application code.

The separation of responsibilities among Codex, an MCP tool, JEV and human review

JEV does not become a Codex model-picker option

The end-to-end path looks like this:

  1. The user gives Codex a task.
  2. Codex reads the repository, tests and operational constraints.
  3. At one bounded decision point, Codex sends only the required state to an MCP tool.
  4. The MCP server calls JEV and receives a typed response.
  5. Codex compares that response with direct evidence.
  6. Codex and the user decide whether to edit, run a command or deploy.

JEV does not click, run shell commands, edit files or deploy. That separation is a security property. It avoids turning one probabilistic judgment into an immediate external action and keeps the primary agent accountable for inspecting the actual evidence.

Two ways to add JEV to a Codex workflow

1. Install TypeSafe’s official agent skill

TypeSafe publishes an official agent skill that teaches coding agents the API, question patterns and evaluation practices. Its documented general installation command is:

npx skills add typesafe-ai/skills --skill typesafe-ai

The skill helps Codex write an API client or review JEV integration code. It does not by itself expose a live JEV tool inside the conversation. A skill is guidance and reusable instructions; an application or MCP server still has to make the network request.

This path makes sense when a development team is building JEV into its own product and wants Codex to follow TypeSafe’s official patterns while writing and testing that code.

2. Add a private MCP plugin for direct calls

If Codex should ask JEV directly during a conversation, it needs a callable MCP server. The environment tested for this article uses a private JEV Decision Helper plugin with a local stdio MCP process and two deliberately narrow tools:

This private helper is not an official TypeSafe Codex MCP plugin. It is an independently maintained bridge to TypeSafe’s documented API. Any organization adopting the same pattern should review and own the wrapper code or choose an implementation it can audit.

OpenAI’s plugin documentation says a Codex plugin can bundle skills, MCP configuration or both. Tool schemas should be explicit, credentials should stay out of plugin files and distributable archives, and a stdio server should receive only the environment variables it actually needs.

What the request contains

TypeSafe’s API documentation specifies POST https://api.typesafe.ai/v1/systemone with Bearer authentication and a JSON body. The important fields are:

A narrow release-state request can look like this:

{
  "model": "jev-latest",
  "state": {
    "change": "login session middleware and database migration",
    "automated_tests_passed": 42,
    "security_integration_tests_failed": 1,
    "rollback_plan": true,
    "deployment_started": false
  },
  "questions": {
    "allow_deploy_now": {
      "type": "noul",
      "instructions": "Using only the supplied state, estimate the probability that deployment should begin now."
    }
  }
}

A production MCP tool should also validate the allowed Choice options, the ordered Score levels and the criteria for every question. If downstream code accepts arbitrary strings as commands or paths, typed output no longer provides a meaningful safety boundary.

When to use Choice, Score and Noul

The appropriate inputs and outputs for Choice, Score and Noul

Choice: select a route from named options

Choice is appropriate when the application already knows the legal options, such as fix_and_retest, request_human_review and deploy_with_monitoring. The response includes the selected choice, the probability assigned to every option and a separate confidence value.

Good uses include:

Score: evaluate an ordered rubric

Score uses levels such as 0=low, 1=moderate, 2=high and 3=critical. It returns the selected level, a distribution across levels and confidence. The number is a label in an ordered policy—not the result of a mathematical calculation.

Risk, urgency and review priority fit this primitive. Invoice totals, calendar intervals and exact counts do not. Those should be calculated by deterministic code before the relevant result is added to the state.

Noul: estimate the probability of one proposition

Noul evaluates one explicit yes-or-no statement, such as this change is ready to deploy now, and returns a value from zero to one. It does not have the separate confidence field returned by Choice and Score.

Do not treat a probability attached to one Choice option as mathematically interchangeable with a Noul value. The question structures and calibration can differ, so each primitive requires its own evaluation.

A live Codex-to-JEV connection test

The connection test used no passwords, API keys or customer records. JEV received only a synthetic, non-identifying release state:

One request carried three independent questions. jev-latest, resolving to JEV 1.13.0, returned:

A live JEV response for a release candidate with a failed security test and the required follow-up controls

The answer is sensible, but this is not an accuracy benchmark. A single example verifies the connection and response shape. Production use requires a representative labeled set that measures false positives, false negatives, class-specific behavior and threshold stability.

Codex still has work to do after receiving the result: inspect the failing test, read the changed files, apply an authorized fix and rerun validation. A Noul value of 0.03 does not replace a deployment block, change-management rule or human approval.

Turn confidence into a policy, not a feeling

TypeSafe’s confidence guide suggests routing high-confidence cases to automation, medium cases to confirmation or more information, and low-confidence cases to human review. It does not prescribe one universal threshold.

A responsible policy process is:

  1. Build a labeled evaluation set from historical cases.
  2. Decide the business cost of each false positive and false negative.
  3. Set different thresholds for low-risk classification and high-impact actions.
  4. Route close distributions and threshold-edge cases to people.
  5. Log the question version, model version, input summary, response and final human decision.
  6. Re-evaluate whenever the prompt, criteria, model or operating domain changes.

An 0.80 confidence threshold may be acceptable for document tagging while a production deployment remains human-approved even at 0.99. Confidence and authorization answer different questions.

U.S. engineering and governance considerations

U.S. teams should treat JEV as a third-party inference service when data crosses the network boundary. Before submitting proprietary code, customer tickets, employee records or regulated information, review current retention, training-use, subprocessors, deletion, incident response and contractual controls. A public model card and price page do not establish that a particular HIPAA, financial-services or government requirement is satisfied.

Operational controls should include:

English is JEV’s primary training language according to the model documentation. That helps many U.S. workloads, but domain-specific language can still shift performance. Legal support tickets, medical abbreviations, security incidents and internal product terminology each need their own evaluation set.

Credential and plugin security

Never place a JEV API key in source code, .mcp.json, an article, logs or a Git repository. The tested helper stores its credential in a separate Codex secret file and refuses requests if the file is readable beyond its owner. Its status command reports only whether configuration exists.

Additional controls are equally important:

Good and bad fits

JEV is strongest when the legal outputs and decision criteria can be defined before the request:

Use Codex, deterministic code or a qualified reviewer instead for:

TypeSafe’s JEV 1.13 limitations page documents literal interpretation, degraded results with oversized irrelevant state, conflicts between instructions and criteria, and susceptibility to adversarial content or prompt injection. Media must first be converted into verified text or structured fields by another tool.

Context limits, versions and pricing

According to TypeSafe’s model page, the stable model on September 27, 2026 is jev-1.13.0, and jev-latest currently points to it. The total request context limit is 64K. State plus the longest single question is limited to 32K. The input is text-only.

Input costs $42 per billion tokens, equivalent to $0.042 per million tokens, and output tokens are free. The low unit price is not a reason to send entire repositories or conversation histories. Larger state can raise privacy risk, reduce relevance and make failures harder to reproduce.

Published rate limits can change. Direct HTTP clients should implement exponential backoff for 429 and 529 responses and should not convert a retry storm into multiple external actions.

jev-latest is convenient during development. A production policy calibrated to response distributions should pin jev-1.13.0 or another evaluated version because the alias can move when TypeSafe releases an update.

Post-installation checklist

  1. Create a TypeSafe API credential and store it in an OS secret store or owner-readable file.
  2. Distinguish the official agent skill from the MCP process that makes live calls.
  3. Expose only narrowly scoped MCP tools with strict input and output schemas.
  4. Start a new Codex session and confirm that the expected tools are present.
  5. Use a non-billable status check without printing the credential.
  6. Exercise Choice, Score and Noul with synthetic, non-identifying data.
  7. Verify that timeouts, schema failures, 429s and 529s stop safely.
  8. Keep questions, criteria and thresholds in one reviewable policy module.
  9. Measure false positives and false negatives on domain-specific data.
  10. Require user authorization and direct validation for deployments, deletion and access changes regardless of JEV’s answer.

Bottom line

JEV and Codex are complementary, not competing, models. Codex remains the primary agent that understands the goal, operates tools and is accountable for the completed task. JEV contributes three narrow outputs—Choice, Score and Noul—at points where prose would be harder to validate and automate.

The best first deployment is reversible: routing a case, deciding whether to request review or selecting what evidence to collect next. Keep the questions atomic, minimize state, calibrate distributions on your own data and separate advice from execution authority. Under those constraints, JEV can make a Codex workflow more explicit without pretending that probability is permission.

Trademark and independence notice

JEV and TypeSafe names belong to TypeSafe AI. Codex, GPT and related names belong to OpenAI and their respective owners. This independent editorial article is not sponsored, endorsed or approved by either company. Its images do not reproduce official product interfaces, logos, third-party photography or company charts.

Official sources

Source: TypeSafe AI · Includes original screenshots or graphics