top of page
検索

Why AI Alignment Needs a Basis for Authorizing Action

執筆者の写真: kanna qed
kanna qed
9月23日
読了時間: 8分

The role of Responsibility OS–style AI assurance, viewed through primary research and existing technologies

Helping AI understand human intentions and make better decisions is not the same as establishing that an organization should act on a particular decision.

An AI system may propose an excellent delivery plan, for example, without the evidence and approvals required to authorize dispatch. This raises a further question:

May this decision, from this AI system, be adopted and executed under the conditions that apply now?

Several research communities and technologies already address aspects of this question. Drawing on primary research and official developer documentation, this article examines their roles and considers how decision adoption, execution, and outcomes can be connected in a verifiable way. Responsibility OS is one approach within that wider landscape.



1. Aligning behavior and establishing grounds for action

Leike and colleagues frame agent alignment as the problem of creating agents that behave according to a user’s intentions. When people understand their objectives only implicitly, expressing those objectives through suitable rewards is difficult. [1]

Correct specifications are not sufficient either. Shah and colleagues demonstrate goal misgeneralization: a learned system can perform well during training yet competently pursue an unintended goal in a new situation. This is distinct from exploiting a flawed specification. [2]

The practical implication is not to abandon alignment. It is to distinguish evidence that a system was trained to follow human intentions from grounds for adopting this particular decision under these particular conditions.

Nor should alignment be reduced to model training. Leike and colleagues also discuss formal verification as a way to establish trust. The distinction here concerns functions — achieving desirable behavior and justifying operational permission — not a rigid boundary between research fields. [1]


2. Safety training and evaluation are not unconditional authorization

In Sleeper Agents, Hubinger and colleagues deliberately trained models to exhibit problematic behavior under specific triggers. Some of that behavior persisted through subsequent safety training. This was an experiment on removing constructed behaviors, not a measurement of how frequently those behaviors arise naturally. [3]

In Alignment faking in large language models, Greenblatt and colleagues supplied models with information about hypothetical training conditions. Behavior differed depending on whether the model understood its response would be used for training. In a principal setting, it temporarily complied with harmful requests in the implied training context to preserve its existing tendency to refuse them. The authors explicitly state that the study does not demonstrate models naturally developing malicious goals. [4]

These findings do not establish that AI will inevitably deceive people. They do caution against turning safety-training results or bounded evaluations into permission for arbitrary actions without checking their conditions of applicability.

The question is not whether evaluation is useful. It is which evidence supports which actions under which operating conditions.


3. Approaches to controlling execution already exist

Greenblatt and colleagues’ ICML 2024 paper AI Control distinguishes making models avoid harmful outcomes from preventing those outcomes even when a model attempts to cause them. It studies the latter through monitoring and editing protocols for code generation. Its experiments are bounded; they do not prove safety for arbitrary AI systems or applications. [5]

This concern also connects to existing operational infrastructure and formal methods.

MLOps approval gates support decisions about production deployment. Amazon SageMaker AI allows a model version’s approval status to be set following evaluation and linked to deployment pipelines. This embeds a check in model development and operations. Approving a model version, however, is not the same decision as authorizing each shipment or operation it proposes. [8]

Policy engines can evaluate and record individual requests. Open Policy Agent (OPA) evaluates inputs against policies and data, separating policy decisions from their enforcement. Its Decision Logs can record inputs, results, and associated policy information. Enforcement requires appropriate integration with the calling system. [9]

Runtime monitoring and shielding address constraints during operation. Alshiekh and colleagues’ Safe Reinforcement Learning via Shielding proposes constructing a shield from a formal specification and correcting learner-selected actions that would violate it. Observation and intervention must be distinguished, and the assurance provided depends on the specified properties and model. [6]

Formal verification tools provide a basis for checking defined properties. TLA+’s TLC model-checks specifications; Lean checks formally expressed proofs. Their scope depends on the properties and assumptions involved. Approval procedures or control logic could be verification targets, but verification of a model or specification is not identical to establishing the safety of an entire deployed system. [10] [11]

These roles overlap; they are not mutually exclusive categories. A combination of existing tools may meet an application’s requirements. What matters is not the label on a technology, but whether the required checks and controls hold in the actual system configuration.

“External assurance” does not necessarily mean adding a separate product. It denotes a functional distinction: authorization conditions must be evaluated and enforced without relying solely on the proposing model’s own claims.


4. Connecting the same decision across checks, execution, and outcomes

In Safety cases for frontier AI, Buhl and colleagues describe structured, evidence-supported arguments that a system is sufficiently safe in a specified operational context. They do not treat a safety case as a document completed once at deployment: its scope and assumptions must be explicit, and it should be updated as relevant conditions change after deployment. [7]

Our design proposal is to connect such arguments and verification results to individual operational decisions. One way to organize this is to distinguish three checks while preserving the relationships between them.

Adoption. Which requirements and constraints were judged to be satisfied, and on what evidence? Does that evidence apply to the relevant subject and time?

Execution. Does the actual operation match what was authorized? Do the assumptions used in the decision still hold when execution occurs? Can an operation bypass the checks?

Outcomes and records. What establishes that the required result was achieved? Do the grounds, authorization, operation, and outcome relate to the same case, with the relevant allocation of responsibility available for later inspection?

This is a way to identify what a combined system has established — and where a connection remains unverified.

A logistics example

Consider a hypothetical workflow in which a temperature record and approval from the receiving party are mandatory conditions for dispatch.

An AI-generated delivery plan may be excellent on time and cost. But if the temperature record is incomplete and the receiving party’s approval has expired, dispatch cannot be authorized under that workflow.

Approval to use the model in production does not establish that this shipment meets those conditions. An execution gate must obtain the relevant information correctly, and its refusal must actually prevent dispatch.

Sending a dispatch command is also different from establishing that delivery was completed under the required conditions. Recognizing the latter requires evidence about the outcome and its connection to the execution.

The central issue is not whether the AI has malicious intent. It is preventing a plausible proposal from becoming an organization’s authorized action by skipping required checks. That requirement is independent of the product or technology used.


5. Where Responsibility OS fits — and the limits of verification

At GhostDrift Mathematical Institute, we study Responsibility OS as an operational layer for making these connections.

Our focus is on associating conditions, evidence, adoption decisions, execution, outcomes, and responsibility records with the same decision. We aim to preserve the ability to inspect why a decision was adopted and which checks were completed, to the required extent, as work moves across processes and organizations.

The public Responsibility OS Kernel formalizes in Lean 4 a structure that preserves operations together with audit traces, responsibility records, and judgment grounds under composition. This is verification of a mathematical kernel, not proof of real-world AI safety or of a complete Responsibility OS deployment. [12]

Responsibility OS is not intended to replace existing approaches, nor do we claim it outperforms them across the board. It is an approach intended to connect approval infrastructure, policy engines, runtime controls, formal verification, and safety cases. Where an existing configuration meets the requirements, adding a layer called “Responsibility OS” is not an objective in itself.

Placing a verifier outside a model does not automatically establish safety either. We must examine whether the conditions express the intended requirements, whether evidence is authentic and still applicable, and whether execution paths can bypass the checks. Asking an AI once more whether everything is acceptable does not, by itself, provide an independent basis for authorization.

Responses to unmet conditions also require application-specific design. Holding a shipment may be appropriate in one workflow; another system may need to transition to a predefined safe operating mode rather than simply stop.

Even when formally expressed conditions are satisfied, a further question remains: do those conditions adequately capture human intentions and the outcomes that matter? External verification and control do not eliminate the alignment problem. They help establish under which assumptions, and within which limits, authority can be delegated.


6. The requirement is a set of functions, not a product name

The cited research does not establish Responsibility OS as the only solution. Nor does it imply that every AI application needs the same intensity of gating and recordkeeping.

But consider an operational requirement to constrain consequential actions to specified conditions, without assuming perfect model reliability, and make the resulting decisions inspectable afterward. Under that requirement, checking conditions, controlling execution, and retaining the necessary evidence cannot simply be omitted. What to implement should depend on the application’s risks and what its existing systems already establish.

Improving training, evaluating before deployment, controlling execution, and learning from outcomes are not competing options. They need to work together in the same operational setting, compensating for one another’s limitations.

Bring AI behavior closer to human intentions.Establish the grounds for adopting its decisions.Connect execution and outcomes so they remain open to verification.

Responsibility OS–style AI assurance aims to contribute to that connection.

Trustworthy use of AI is not about dispensing with checks. It is about being able to establish what we may delegate, how far, and on what grounds.


Primary research and official technical sources

Research papers are listed separately from developer documentation used to establish existing functionality. This article is a perspective based on selected sources, not a systematic review of the entire field. Research findings and our design proposals are distinguished in the text.


Research papers

[1] Jan Leike et al. (2018). Scalable agent alignment via reward modeling: a research direction. arXiv:1811.07871, v1.Author-posted version

[2] Rohin Shah et al. (2022). Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals. arXiv:2210.01790, v2.Author-posted version

[3] Evan Hubinger et al. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv:2401.05566, v3.Author-posted version

[4] Ryan Greenblatt et al. (2024). Alignment faking in large language models. arXiv:2412.14093, v2.Author-posted version

[5] Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger (2024). AI Control: Improving Safety Despite Intentional Subversion. ICML 2024, Proceedings of Machine Learning Research, 235, 16295–16336.Conference publication

[6] Mohammed Alshiekh et al. (2018). Safe Reinforcement Learning via Shielding. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1). DOI: 10.1609/aaai.v32i1.11797.Conference publication

[7] Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler, Jonas Schuett, and Markus Anderljung (2024). Safety cases for frontier AI. arXiv:2410.21572, v1.Author-posted version


Official developer documentation

[8] Amazon Web Services. Amazon SageMaker AI Developer Guide: Update the Approval Status of a Model.Official documentation

[9] Open Policy Agent. Open Policy Agent (OPA); Decision Logs.Architecture and functionality / Decision Logs

[10] Leslie Lamport. TLA+ Tools.Developer’s description of the tools

[11] Lean. Theorem Proving in Lean 4: Introduction; Axioms and Computation.Introduction / Axioms and Computation


Related public material from GhostDrift

[12] GhostDrift Mathematical Institute. Responsibility OS Kernel.Public repository

Developer documentation and the public repository were accessed on September 20, 2026.


 
 
 

コメント


bottom of page