top of page
検索

Jev Reopens the Boundary Between AI Decisions and Execution

執筆者の写真: kanna qed
kanna qed
9月23日
読了時間: 9分

Probability Is Not Permission — What It Takes to Connect Typed AI Decisions to Real-World Operations

GhostDrift Mathematical Institute | Perspective | September 20, 2026

An AI system says that a shipment may be released. Under what conditions does that answer become an organizational authorization to act?

On September 15, 2026, TypeSafe AI released Jev in early access, a model designed to return typed decisions that software can use directly. What the company calls “System One Models” places less emphasis on generating prose and more on returning choices, scores, and probabilities that can be integrated into programmatic workflows.[1][2]

This article is not about whether Jev will replace existing models. The question is more fundamental: what must connect a model’s decision to an organization’s adoption of that decision, its execution, and the verification of the resulting outcome?



1. Jev Makes the Division of Labor Between Decision and Control Explicit

According to Jev’s official documentation, a system provides state and a question, then receives a value corresponding to predefined choices or rating scales. When several questions are involved, they are decomposed, while the logic for combining their outputs is controlled in code.[2]

This does not mean that AI has suddenly become capable of “making decisions” for the first time. Probabilistic classifiers have existed for years, and TypeSafe itself compares Jev against LLMs wrapped to produce typed decisions. Jev alone therefore does not justify a sharp claim that “generative AI” and “decision AI” have become completely separate categories.[1][3]

What matters is that the division between the model’s decision output and the code that turns that output into action is now made explicit as a product architecture.

TypeSafe’s official materials also describe changing thresholds according to the risk of an action and using those thresholds to distinguish autonomous execution, review, and hold states. The separation between probability and execution permission is therefore not merely an external criticism of Jev. It is a question that follows naturally from the architecture Jev itself presents.[2]


2. Type, Probability, and Permission Answer Different Questions

First, a valid output type is not the same thing as a correct decision.

TypeSafe’s public explanation of “Zero Hallucinations” makes clear that the claim is tied to schema compliance rather than an empirical demonstration that the semantic content of every decision is correct. Remaining within predefined choices is not the same thing as always selecting the right choice about the real world.[1]

There is also an important distinction in Jev’s API between probabilities assigned to choices and confidence. According to the official specification, confidence is a metric derived from the concentration or dispersion of the probability distribution. A value such as confidence = 0.94 should therefore not be read directly as “94% probability of being correct.”[2]

More broadly, probability estimation itself requires validation. Guo et al. showed that classification accuracy and probability calibration — the extent to which predicted probabilities correspond to observed correctness frequencies — are distinct properties. Calibration is a statistical property evaluated over a distribution; it does not prove that any individual decision is correct. This study did not evaluate Jev specifically.[3]

But the argument here does not depend only on imperfect calibration.

Even if probability estimates are well validated for a given use case, the correct action can still differ depending on the cost of error, execution authority, required approvals, applicable business rules, or mandatory evidence. Turning probability into action requires a policy that defines what may proceed and what must not.

Probability may be part of the evidence for permission. It is not itself the rule or authority that grants permission.

This is not an argument against threshold-based automation. In low-risk and well-specified settings, thresholds combined with ordinary business logic may be entirely appropriate. The important question is what those thresholds mean and what other conditions they must be combined with.


3. A Model May Recommend Release Without Being Authorized to Release

Consider a simplified pharmaceutical temperature-controlled logistics example. Assume that an internal procedure requires a shipment to be held and referred for quality review whenever a required temperature record cannot be confirmed.

A model may recommend release, while the operational process still correctly returns:

Model recommendation: RELEASEEstimated probability assigned to that choice: 0.94Required temperature record: partially unavailableBusiness procedure: HOLD / QA REVIEW

The values and workflow above are illustrative only. They are not Jev benchmark results or an actual quality-control procedure.

The shipment is not held because the organization “does not trust AI.” It is held because a condition required for operational authorization has not yet been established. Information that may be inferred probabilistically must be distinguished from conditions that require confirmation through designated evidence.

The problem also does not end at adoption. If the target shipment, destination, or relevant state changes after approval, the original approval may no longer apply. Likewise, a record showing that a command was sent is not the same thing as proof that the intended real-world outcome occurred.

Functionally, the chain can be separated as follows:

Obtain a decision candidate → verify adoption conditions → verify execution-time conditions → certify the outcome

This does not imply that four separate products or services are required. All of these functions may exist within a single application. What matters is that the object of verification and the conditions for moving to the next stage are not conflated.


4. Existing AI Governance Has Never Been Only About Model Outputs

It would be inaccurate to suggest that this governance problem was created for the first time by Jev.

NIST’s AI Risk Management Framework 1.0 states that human judgment should be used when establishing metrics and thresholds for trustworthiness. Its companion Playbook, in MEASURE 2.8, recommends documenting human oversight, downstream actions and overrides following AI outputs, policy exceptions, and responsible parties’ go/no-go decisions. These are voluntary risk-management materials; they do not mandate a particular product architecture or a separate control gate for every AI decision.[4]

The Hiroshima Process International Code of Conduct likewise addresses the AI lifecycle beyond development alone. Where applicable, it covers design, development, deployment, and use, and includes post-deployment monitoring as well as information needed for users to interpret and appropriately use model outputs.[5]

The EU AI Act takes a similar but legally distinct approach in the context of high-risk AI systems. Article 14 addresses human oversight, including the ability to disregard or override outputs and intervene in operation, while Article 15 separately addresses accuracy, robustness, and cybersecurity. Oversight measures are intended to reflect the risk, level of autonomy, and context of use rather than imposing a uniform human-approval requirement for all AI systems.[6]

These frameworks differ in legal status, scope, and intended audience. But none is built on the assumption that evaluating model performance automatically determines how every output should be used.

The four-stage distinction used in this article is not an official taxonomy from NIST, the G7, or the EU. It is a way to organize the practical question of where, how, and by what mechanism already-recognized governance requirements should be verified in an operational system.


5. Existing Research and Tools Already Address Parts of This Problem

The idea of placing conditions between a decision and an action should not be presented as a novel invention belonging to one organization.

Open Policy Agent (OPA), for example, allows authorization and policy conditions to be expressed as code and separates policy evaluation from the mechanism that enforces the result. Structured AI outputs can, in principle, be treated as inputs to such policy decisions. Enforcement still has to be implemented in the execution path.[7]

In formal methods, Bloem et al.’s work on Shield Synthesis studies runtime mechanisms that monitor system inputs and outputs and modify outputs when needed to preserve specified properties. Mitsch and Platzer’s ModelPlex addresses runtime validation of cyber-physical system executions against verified models, helping connect model-level proofs with actual execution. In both cases, the guarantees depend on explicitly specified properties, assumptions, and models.[8][9]

The research question therefore cannot be reduced to “put one more check around the model.”

The deeper issue is:

What must be checked? What must prevent execution when the check fails? How can the system ensure that the enforcement mechanism is not bypassed? And how can the evidence used for a decision remain correctly bound to the specific execution and outcome it is supposed to justify?

If existing approval workflows, policy engines, formal verification, and runtime monitoring can be combined to satisfy those requirements, that may be entirely sufficient. The important issue is not the label attached to the architecture, but whether the necessary properties actually hold for the target operation.


6. Where GhostDrift’s Research Fits

The research conducted by GhostDrift Mathematical Institute under concepts such as Responsibility OS, ADIC, and execution-outcome closure is not intended to replace a model’s decision-making capability. Its focus is the verifiable connection between decision evidence, adoption conditions, execution conditions, and outcome confirmation — and the use of those verification results as conditions for allowing the next operation to proceed.[10]

For example:

  • Was an approval for shipment A reused for shipment B?

  • Was a decision meaning “execution permitted only under condition X” later interpreted as unconditional approval?

  • Was “command sent” silently treated as equivalent to “intended outcome achieved”?

These forms of mismatch are part of the connection layer we study.

Our public Lean 4 formalization, Physical AI Outcome Assurance, addresses a boundary on outcome certification: if, under the stated assumptions and available evidence, both a successful and an unsuccessful execution history remain possible, then certain success cannot be certified from that evidence alone. This does not reject probabilistic inference. It distinguishes inference from what may be treated as a definitively established outcome. The formalization also includes cases in which prior verification can be sufficient without new post-execution evidence, provided the verified conditions fully cover the execution in question.[11]

A second public formalization, Physical AI Verified Composition, studies the fact that guarantees for individual process stages do not automatically compose into an end-to-end guarantee. Explicit bridge conditions are needed between the result of one stage and the assumptions of the next.[12]

These formalizations do not prove that our approach is the only valid architecture, nor do they replace existing work in policy enforcement, runtime assurance, or formal verification. They establish information-theoretic boundaries and sufficient conditions under explicitly stated definitions and assumptions. They do not, by themselves, certify the safety of a specific machine, the authenticity of real-world evidence, legal compliance, or the correctness of an entire implementation. Those require separate validation, including validation of the correspondence between model and reality and of the enforcement path itself.[11][12]

Practical effectiveness must also be evaluated beyond the existence of a formal proof. A useful implementation should be assessed not only on whether it blocks invalid executions, but also on whether it unnecessarily blocks valid ones, whether its latency is acceptable, and whether a third party can reconstruct the basis of the decision.


Conclusion — Not to Stop Automation, but to Define When It Can Be Trusted

A reversible, low-risk digital action should not necessarily be governed with the same controls as an irreversible physical action. Nor is stopping always safe; some systems require safe continuation or fallback behavior. Risk-proportionate control is consistent with the logic found in NIST’s framework and in runtime-verification research.[4][9]

The important question raised by Jev is therefore not whether AI should be trusted less. It is how to make explicit the conditions under which an AI decision may safely become an organizational action.

Probability is not permission. Permission to execute is not proof that the intended outcome occurred.

Keeping those distinctions explicit — and connecting the layers between them in a verifiable way — is an increasingly important systems-design problem, regardless of which model ultimately produces the decision.

Primary Sources and Public Materials

The numbering below corresponds to references in the text. Web materials were accessed on September 20, 2026. Vendor documentation, academic research, public policy documents, and GhostDrift materials are listed separately because they provide different kinds of evidence.

[1] Diogo Almeida / TypeSafe AI (September 15, 2026)“Introducing System One Models & Jev.”Vendor-authored description of Jev’s release, typed outputs, benchmark conditions, and schema-compliance claims. This is not an independent performance evaluation.https://typesafe.ai/blog/introducing-system-one-models-and-jev

[2] TypeSafe AI, Official Documentation“Introduction”; “Confidence”; “Choice.”Official specification covering typed questions, division of responsibilities between model and code, the distinction between `probabilities` and `confidence`, and the use of thresholds according to risk.https://docs.typesafe.ai/

[3] Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger (2017)“On Calibration of Modern Neural Networks.” Proceedings of the 34th International Conference on Machine Learning, PMLR 70:1321–1330.Original research on the distinction between predictive accuracy and probability calibration. It does not evaluate Jev.https://proceedings.mlr.press/v70/guo17a.html

[4] National Institute of Standards and Technology (NIST)*Artificial Intelligence Risk Management Framework (AI RMF 1.0)*, NIST AI 100–1 (2023); AI RMF Playbook, MEASURE 2.8.See AI RMF p.12 on metrics, thresholds, and human judgment, and Playbook guidance on oversight, downstream actions, policy exceptions, and go/no-go decisions. Both are voluntary risk-management resources.https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

[5] G7 (2023)*Hiroshima Process International Code of Conduct for Organizations Developing Advanced AI Systems.*See the preamble and Actions 1–3 regarding lifecycle risk management, deployment, post-deployment monitoring, and information necessary for appropriate use of outputs.https://www.mofa.go.jp/mofaj/files/100573473.pdf

[6] European Union*Regulation (EU) 2024/1689 — Artificial Intelligence Act*, Articles 14–15.Consolidated EUR-Lex version dated July 27, 2026. Article 14 addresses risk-proportionate human oversight, including disregarding, overriding, and intervening in outputs; Article 15 separately addresses accuracy, robustness, and cybersecurity. This article does not attempt to determine applicability to any specific use case or implementation date.https://eur-lex.europa.eu/legal-content/EN/AUTO/?uri=CELEX%3A02024R1689-20260727

[7] Open Policy Agent Project“Open Policy Agent (OPA).” Official Documentation.Official description of policy evaluation, policy-as-code, and separation between policy decision and enforcement.https://www.openpolicyagent.org/docs

[8] Roderick Bloem, Bettina Könighofer, Robert Könighofer, Chao Wang (2015)“Shield Synthesis: Runtime Enforcement for Reactive Systems.” arXiv:1501.02573v2.Public version of the original research on runtime monitoring and output correction to enforce specified properties.https://arxiv.org/abs/1501.02573

[9] Stefan Mitsch, André Platzer (2016)“ModelPlex: Verified Runtime Validation of Verified Cyber-Physical System Models.” Formal Methods in System Design, 49:33–74. DOI: 10.1007/s10703–016–0241-z.Original research on runtime validation of real executions against verified cyber-physical system models.https://link.springer.com/article/10.1007/s10703-016-0241-z

[10] GhostDrift Mathematical Institute (2026)“Taking Manufacturing AI from PoC to Real-World Execution: GhostDrift Files Five International Applications Covering AI Assurance for Physical AI Execution and Outcomes.”Company R&D announcement used here only to identify the scope of GhostDrift’s research and public formalization work. It is not third-party validation or certification.https://prtimes.jp/main/html/rd/p/000000008.000182721.html

[11] GhostDrift Mathematical Institute*Physical AI Outcome Assurance.*Public Lean 4 formalization on outcome certifiability, ambiguity of evidence, and conditions for transferring model-level guarantees to a specific execution. See the repository README for scope and explicit non-claims.https://github.com/GhostDriftTheory/physical-ai-outcome-assurance

[12] GhostDrift Mathematical Institute*Physical AI Verified Composition.*Public Lean 4 formalization concerning information loss and the composition of guarantees across finite process stages under explicit bridge conditions. It is distinct from claims of physical-system safety or performance at operational scale.https://github.com/GhostDriftTheory/physical-ai-verified-composition


 
 
 

コメント


bottom of page