The next AI race is government operations, not bigger models
A public agency can buy access to the same frontier model as a private company and still end up with a radically different AI system. The important variable is not only what the model can say. It is what the agent may read, which tools it can invoke, who must approve an action, and whether an investigator can reconstruct the result later.
That distinction grows once an assistant becomes an agent. A chatbot can produce a draft for an employee to copy. An agent can search records, assemble a file, call an internal service and submit a transaction. The model may be identical, but the operational risk is not. Official developments in Korea, Japan and the United States point toward an AI contest measured increasingly by operating discipline rather than a leaderboard score.

This is an original conceptual image for the article. It does not depict a named agency, product, meeting or deployed government system.
Korea put service delivery, data and privacy on one agenda
The Korea Policy Briefing says the fifth council of government chief AI officers met on October 2, 2026. Participation expanded from 28 ministerial-level bodies to 42 central administrative bodies. Four items were discussed: a national AI legislative framework, launch preparation for Everyone’s AI, a government-wide data pipeline, and stronger public-sector privacy management.
The operational detail matters. Ahead of a planned December launch, ministries were asked to coordinate connections to open public APIs, consider additional access to nonpublic APIs and data, develop criteria for an agent’s processing and use of personal information, and conduct security and safety checks. The notice reports preparation and requested cooperation; it does not establish that every capability or safeguard has already passed production testing.
“Connect the API” is therefore an incomplete requirement. Reading public benefit guidance, checking a resident’s eligibility and submitting an application all use interfaces, but the consequences differ. Read-only access should not silently turn into permission to change a record. A transaction with legal or financial effect needs a separate identity, a narrower scope and an explicit approval rule.
Japan is connecting a workforce-scale pilot to agent tooling
Japan’s Digital Agency overview of Government AI Gennai says the fiscal 2026 large-scale demonstration is intended to make generative AI available to about 180,000 employees across ministries. The environment includes general tasks such as drafting and translation, administrative applications, a common government dataset and trials of domestic language models. The page separates the demonstration from full-scale use planned from fiscal 2027 with ministry budgeting.
On September 18, the agency described a plan for Gennai OSS version 2. Proposed materials include an agent chat environment that combines web research, document generation, code execution and reusable skills; a personal coding environment; and components for exporting and analyzing usage data. A cross-organization marketplace for skills was still at the concept stage.
The agency explicitly says the version 2 materials were under development and could change, with publication targeted for around February 2027. It also described the in-government agent environment as planned for around October 2026. As of October 3, a planned introduction should not be reported as confirmed government-wide production use.
The useful lesson is broader than any feature list. Availability for 180,000 people is a distribution metric. Safe, repeatable use by 180,000 people is an operating result. Training, access policy, application ownership, evaluation, usage visibility and support all sit between the two.
The United States is treating agent identity as infrastructure
A May 2026 NIST analysis of responses on AI-agent security reports broad agreement among commenters that agents introduce novel threats and that those concerns create an adoption barrier. Established cybersecurity principles remain relevant, but respondents said they need adjustment for agents. Suggested government roles included implementation guidance, information sharing and standards promotion. This is a summary of an information request, not a final mandatory rule.
A September 24 NIST National Cybersecurity Center of Excellence update makes the identity problem concrete. The center was scoping an implementation in which agents used in a software-development lifecycle can be identified, authenticated and authorized. Its October 28 webinar and November 9 comment deadline were still future dates on October 3.
An agent should not inherit a vague superuser role simply because it acts for a senior employee. It needs a machine-verifiable identity and a permission set tied to a particular task. If it delegates to another agent or tool, the organization must be able to follow that chain without granting the child process more authority than the original task allowed.
Begin with three verbs: read, recommend and act
A practical permission model can start with three levels.
- Read: Search approved sources and return evidence. The agent cannot change an external state.
- Recommend: Produce a decision or response draft. It cannot save, send or submit the result without review.
- Act: Invoke an approved service to update a record, send a notice or complete a transaction. The action is subject to a task-specific approval and confirmation rule.
These levels should be attached to a workflow, not merely to a job title. The same benefits officer may allow an agent to draft a response but not to approve a payment. The same developer may permit an agent to inspect a repository but not to change deployment credentials. Short-lived task tokens are safer than a standing broad credential.
Permission checks also need to cover indirect paths. An agent that lacks database access might call a reporting tool that has it. An agent that cannot send email might write into a queue that triggers a message. The policy engine must evaluate the ultimate effect, not just the name of the first tool.
A complete operating model has seven layers
Public-sector agent design becomes easier to audit when it is separated into seven layers.
- Mission boundary: State the job the agent performs and the jobs it must refuse.
- Data boundary: Classify public data, ordinary internal material, personal information and restricted records; define permitted joins.
- Identity: Distinguish the employee, the agent instance, service accounts and external tools.
- Authorization: Split discovery, drafting, storage, communication and transaction permissions.
- Human control: Name the reviewer for decisions involving money, rights, personal data or an official response.
- Traceability: Record model and skill versions, source versions, tool calls, approvals, edits and outcomes.
- Containment and recovery: Revoke credentials, stop work, notify owners and reverse reversible actions when behavior crosses a boundary.
No single vendor dashboard constitutes this operating model. It has to connect with the agency’s identity provider, records schedule, audit function, incident-response team and program owner. A new model release should not bypass those controls merely because its benchmark is better.
Human approval must show the decision, not just a button
A confirmation dialog that says “Approve?” encourages rubber-stamping. The reviewer needs enough context to recognize the effect of an action. A useful approval packet shows:
- the source records and their revision dates;
- the proposed value beside the current value;
- the people, money, deadline or legal status affected;
- missing evidence and unresolved conflicts;
- whether the action can be reversed and for how long;
- recent errors or interventions in the same workflow.
Approval intensity should follow risk. A reversible calendar hold may use a lightweight confirmation. A benefits denial, public statement, transfer of personal data or expenditure should require a named reviewer and a durable reason. Requiring the same click for every low-risk step creates fatigue; requiring none for high-risk steps creates scale without accountability.
Agencies should also measure review quality. The useful number is not how many approvals were collected. It is how often reviewers changed or rejected a recommendation, how long they had to inspect it, and whether later audits found evidence that the approval screen did not surface.
An audit trail is larger than a chat transcript
Saving the conversation does not reveal everything an agent did. The record should connect a user and agent identity to the permission issued, the model and skill versions, source identifiers, tool calls, returned status, human changes, final output and any recovery action.
At minimum, log a work-item ID, user ID, agent ID, credential issue and expiration times, source version, requested operation, tool response code, approval or rejection, result, exception and rollback status. Sensitive source content does not need to be duplicated in every log. A protected reference can point investigators to the authorized original.
The audit system itself needs access controls and a retention schedule. Otherwise a security record becomes a new collection of personal information. Separate the people who administer agents from those who can erase or alter their evidence, and alert on gaps in the event sequence.
Evaluate a service, not a model in isolation
Model accuracy still matters, but an operational scorecard needs additional measures:
- unsupported answer rate when no approved source was found;
- human edits and rejections per 100 work items;
- attempts to exceed a permission boundary and successful blocks;
- unnecessary exposure of personal or restricted data;
- incorrect external actions and successful reversals;
- time from detection to credential revocation and workflow containment;
- employee supervision time added or removed;
- performance drift after a model, prompt, skill or source update.
An agent can improve average processing time while making rare errors more expensive. Report both the median case and the worst credible failure. For services used by vulnerable residents, accessibility, appeal and alternative human channels belong in the operational evaluation too.
A 90-day rollout can expose control gaps early
During days 1 through 30, choose one read-only task and record the existing turnaround time, rework and error rate. Inventory data sources, identify the owner of each, and test whether the agent refuses unapproved material. Do not start with a workflow that sends money or changes legal status.
During days 31 through 60, allow recommendations while requiring review of every item. Classify failures rather than fixing them one at a time: wrong source, stale source, missing context, excessive access, invalid tool call, poor escalation or reviewer confusion. Run tabletop exercises for a leaked credential and a faulty update.
During days 61 through 90, permit only a small set of reversible actions. Use daily caps, short-lived credentials, explicit approval and automated confirmation. Perform an actual emergency stop drill. Expansion should require that both service results and control metrics pass their thresholds.
If speed improves but unexplained actions, reviewer burden or recovery time worsens, the pilot has found an operational problem—not a reason to hide the control data.
Procurement questions should survive a vendor change
An agency contract should cover more than whether customer data is used to train a model. Ask whether every agent can have a separate service identity, whether permissions can expire, whether all tool calls can be exported, and whether a result can be reproduced after a model or skill update.
Define who may access prompts, retrieved files and logs during vendor support. Require an incident-notification window, investigation cooperation, and deletion of working copies and derived logs at the end of the contract. If a subcontractor or external model provider is involved, document its access separately.
Portability matters. Work definitions, approval policies, evaluations and audit events should be exportable in a documented format. Otherwise the organization may discover that changing the model also means losing the evidence needed to govern the service.
The durable advantage is a smaller permission set
Korea is bringing a public service, government data connections, privacy and security coordination into the same discussion. Japan is combining a workforce-scale pilot with administrative applications, shared datasets and planned agent tooling. NIST work in the United States is framing identity, authentication and authorization as implementation problems that agents cannot bypass.
The three countries are not using one common system, and none of the cited materials proves that every proposed safeguard has been deployed. What they do show is that a stronger model does not eliminate operating responsibility. Public trust grows when an agency can prevent an unauthorized act, explain an approved one and recover from a failed one. The next advantage in government AI is not unlimited agency. It is the smallest useful authority, attached to evidence and a person who remains accountable.



