This is the English edition. 한국어판 and 日本語版 are also available.

Physical AI after the demo: how robots become services

2026-10-03 · AI · United States · Zoogom Editorial

#physical AI#service robots#robot safety#field operations#automation

Robot demonstrations are optimized for the moment when the machine succeeds. A service has to survive all the time around that moment: a blocked hallway, a dirty sensor, a late elevator, a weak wireless signal, a drained battery and a person who moves in an unexpected direction.

That is why the physical-AI contest after the chatbot is not simply a race to publish the most capable video. The useful question is whether a robot can perform a bounded job repeatedly, fail without causing harm, call for the right help and return to work at an affordable cost. Developments in Korea, the United States and Japan illuminate different parts of that transition—from public experience, to measurement, to foundation-model research.

A human operator performing a safety check beside an unbranded transport robot in an original vertical editorial photograph

This is an original conceptual image for the article. It is not a photograph of a named robot, hospital, warehouse or deployment.

Seoul’s show is a starting point, not operating evidence

According to Seoul’s city administration, Smart Life Week is set to run at COEX from October 6 through October 8, 2026. On this article’s October 3 review date, the event had not begun. The stated target is a crowd above 70,000, with over 500 organizations and city delegations taking part. The source was published by Seoul Metropolitan Government. Planned content includes physical AI, humanoid, quadruped and autonomous robots, city services and hands-on programs.

Those figures are targets, not verified attendance. A successful exhibition scene is not a long-term uptime result. The show can still be valuable: it reveals what residents understand, where a device fits into a city service and what makes people uncomfortable. The evidence level simply needs a correct label.

Visitors and buyers should ask what happens outside the scripted route. Where does the robot go when a battery alarm appears? Can it recognize a wheelchair and a child in the same corridor? What does it do when a door integration fails? Does a remote technician see video, and how long is that video retained? These questions turn spectacle into an operations conversation.

NIST is working on the laboratory-to-factory gap

The NIST Physical AI and Data Generation for Robotics project says a large gap exists between embodied AI in academic research and what manufacturers and robot-system integrators can implement in the real world. The project seeks metrics, test methods, software, prototypes and datasets for evaluating AI-enabled robotic systems.

The combination matters. An AI algorithm, a robot body and a task jointly determine cost and performance. A perception model that works with one camera position, gripper, material and lighting condition may fail when any one of those changes. NIST describes work expanding beyond simple pick-and-place to assembly, drilling and dexterous manipulation.

The project page is not a certification of a commercial robot. It does not promise that a hospital delivery machine or store robot is safe. It does show why a buyer needs a test condition, a failure definition and repeatable measurements alongside a headline success rate.

Japan moved robot foundation models into an implementation program

On September 9, Japan’s New Energy and Industrial Technology Development Organization, or NEDO, announced the planned implementing organizations for a robot foundation-model research program. NEDO says it reviewed 29 applications. The program is scheduled for one year from fiscal 2026 in principle; projects meeting specified requirements and passing outside expert evaluation may run for as long as three years.

The number 29 refers to applications reviewed, not the number of finished robot services and not necessarily the number of selected organizations. The detailed list is a separate attachment. The announcement demonstrates an implementation structure for foundation-model research, while commercial reliability still has to be established in the field.

A broader model does not erase the physical limits of a new machine. Sensor placement, joint range, braking distance, payload and emergency-stop behavior change when software moves to another body. Generalization testing and site-specific safety acceptance remain separate obligations.

Four evidence levels keep claims honest

Robot projects become easier to compare when evidence is sorted into four levels.

  1. Capability demonstration: The system completes a task in a controlled setting.
  2. Repeated test: The same task runs tens or hundreds of times and produces a failure distribution.
  3. Limited field operation: The robot works in part of a real site or schedule under close human supervision.
  4. Ongoing service: The system supports shifts, maintenance, seasonal changes, crowds and failures under a service owner.

Success at one level does not prove success at the next. A demo may present a selected take. A pilot may benefit from engineers standing beside the robot. A production service needs a night contact, spare parts, an escalation procedure and a recovery commitment.

Procurement documents should state the level of evidence for every performance claim. “Works in hospitals” is too broad. “Completed 2,000 nonurgent tote deliveries on two mapped floors during staffed hours, with specified intervention and stop rates” is the kind of statement that can be examined.

Choose the task with the clearest boundary

The hardest or least popular job is not automatically the best first deployment. A good opening task has a known start and end, standardized objects, a route or workspace that can be mapped, and a safe failure state.

Candidates may include moving nonurgent supplies along a hospital back corridor, transporting standardized totes inside a warehouse, photographing shelves before a store opens, inspecting a closed facility, or moving known parts between factory stations. Emergency care, lifting a person’s body and unsupervised operation in dense road traffic require much stronger evidence and controls.

Write the task definition in operational terms. Specify origin and destination, hours, payload, priority at intersections, restricted zones, allowable wait time, safe parking location and the person to contact. The robot should not have to invent policy in order to finish a trip.

The site is part of the robot system

Many failures begin with a threshold, reflective floor, narrow turn, elevator interface or wireless dead zone rather than the foundation model. Sometimes changing the site is cheaper and safer than trying to make the robot infer every irregularity.

A readiness survey should cover aisle width and slope, automated doors and elevators, fire egress, charging power, network coverage, cleaning, storage, noise limits and the path used by maintenance staff. Any construction or system-integration cost belongs in the business case, separate from the purchase price.

The survey should map what sensors can observe. If cameras or microphones record employees, patients or visitors, define notice, retention, access, remote-support use and whether the vendor may reuse data for model improvement. Movement data collected for routing can also become sensitive operational or personal information.

Intervention rate is more revealing than average success

A 99 percent success rate can hide a large operational burden. At 1,000 missions per day, a 1 percent failure rate produces 10 failures. If each requires 15 minutes of employee time, the service creates two and a half hours of rescue work before maintenance is counted.

Failures also differ. An automatic retry that clears in five seconds is not equivalent to a robot blocking a doorway until someone walks across the building. Separate at least these measures:

Do not use “no collision occurred” as the sole safety result. A person who jumps away or an employee who presses stop may have prevented the event. Near-miss and human-save records are valuable evidence for redesign.

A 90-day field plan can expose the expensive exceptions

During days 1 through 30, measure the current job without the robot. Record travel, wait time, errors, complaints and safety events. Inspect the site, define prohibited areas and rehearse the physical and software stop procedures.

During days 31 through 60, run in shadow mode with a nearby employee checking every mission. Classify failures into perception, routing, site integration, human behavior, maintenance and operating procedure. The objective is to complete a useful exception catalog, not to maximize throughput for a presentation.

During days 61 through 90, assign a real but restricted service. Limit the hours, zone and cargo. Predefine both success and stop thresholds. If repeated faults, recovery delays, safety behavior or data handling crosses a threshold, reduce the scope. Do not add more robots merely because the last week looked better than the first.

At the end, compare employee workload and total cost with the baseline. A machine that moves more objects but demands more supervision has not necessarily improved the service.

Acceptance tests need adverse conditions

A route demonstration on a quiet morning is not an acceptance test. Include conditions the service is likely to face:

Document the expected safe behavior before the test. Passing should not mean that an engineer improvised a workaround. The system and ordinary site staff should follow a repeatable procedure.

The service agreement matters more than the robot price

A hardware quote may exclude remapping, software subscriptions, fleet management, elevator integration, night support, batteries, wheels and sensor replacement. The contract should define response and restoration times by incident class, replacement equipment, log delivery, retesting after updates and data deletion at termination.

Updates deserve explicit control. If a new model changes navigation or person-detection behavior, the operator should receive a change description and test evidence and should be able to delay deployment. Remote vendor accounts should be temporary, approved for a task and fully logged.

Ask what remains functional if the subscription or network fails. Basic stopping and physical safety should not disappear with a cloud-service outage. Also define ownership and export formats for maps, task definitions, annotations and incident records so a site is not trapped by its own operating history.

Work changes instead of simply disappearing

Robots can reduce repetitive travel while creating dispatch, exception handling, cleaning, charging, inventory and complaint tasks. The number of robots one employee can supervise depends on the mission and site. It should be measured in the local pilot rather than copied from a sales deck.

Frontline workers are not just observers. They know why a cart appears in the wrong place, when a route must yield and which exceptions matter. Give them authority to define stop conditions and categorize failures. Count the time removed from the old job and the time added by the new system in the same ledger.

Training should include more than normal operation. Employees need to know when not to move a stalled robot, how to protect an incident scene, how to report a privacy concern and who authorizes a restart.

Dependable physical AI looks ordinary

Seoul is preparing a public stage for physical AI and future-city services. NIST is building measurement methods around the real-world manufacturing gap. NEDO has established an implementation structure for robot foundation-model research. These are meaningful signals, but none is evidence that every displayed or funded system has become a dependable service.

The winning robot may be the least cinematic one: it performs a narrow job, pauses correctly, asks for help early and returns to service through a documented process. Physical AI becomes infrastructure when bounded work, a prepared site, measurable exceptions, human stop authority and a recoverable service agreement make an ordinary day repeatable.

Source: Seoul Metropolitan Government · Includes original screenshots or graphics