Physical AI Is Entering the Systems Phase

Physical AI systems phase editorial cover showing perception, modeling, planning, control, assurance and learning

By Dr. Barak Or, CEO and Founder, STATE16

Physical AI is moving into its systems-integration phase, and my assessment is that this market is still being framed too narrowly.

Physical AI is not primarily a competition to build the most “human-looking” machine. It is a competition to engineer the highest-performance physical stack: one that can perceive a dynamic environment, maintain state, reason under uncertainty, generate and execute actions, operate inside a verifiable safety envelope, and continuously improve from historical fleet data.

In generative AI, the model can often be presented as the product. In Physical AI, the model is only one component inside a tightly coupled operating system.

The emerging architecture connects multimodal perception, world representation, planning, action policy, real-time control, safety supervision, deployment telemetry, and retraining.



Diagram of the Physical AI operating stack from multimodal perception through safety supervision and retraining

Figure 1. A production Physical AI system is a closed operating stack, not a standalone model.

The strategic question is no longer simply: Which robot can perform the task?

The more relevant question is: Which system can perform the task repeatedly, economically, safely, and with less human intervention after every deployment cycle?

Sovereign Physical AI infrastructure is being built

The first major signal is that governments are no longer treating robotics exclusively as an R&D category. They are treating Physical AI as national infrastructure.

China launched a national initiative targeting more than 100 high-value humanoid application scenarios and the commercial deployment of more than 10,000 humanoid robots by the end of 2026. The targeted environments include manufacturing, logistics, retail, healthcare, and other operational sectors. These figures are deployment objectives, not completed installations, but the direction is strategically significant. [1]

South Korea has positioned Physical AI as one of its three national mega-projects, alongside semiconductors and AI data centers. Its objectives include becoming one of the world’s top three AI-robotics powers by 2030, commercializing industry-specific humanoids across 10 major sectors by 2028, creating Physical AI data factories, developing domestic foundation models, and training 10,000 AI-robotics specialists. [2]

Japan’s revised AI Robotics Strategy maintains a target of introducing approximately 10 million robots across 18 application areas by 2040. The target refers to robotics broadly, not exclusively humanoids, and includes healthcare, food service, manufacturing, and other labor-constrained sectors. [3]



Diagram comparing national Physical AI deployment initiatives in China, South Korea and Japan

Figure 2. National deployment programs can create sovereign data and capability flywheels.

My interpretation is that these countries are not simply funding robot manufacturers. They are attempting to create sovereign deployment flywheels. Every real deployment can generate environment-specific perception data, observation-action trajectories, intervention and failure data, contact and manipulation data, safety-validation evidence, and operational feedback for policy improvement.

In other words, deployment is not only the outcome of Physical AI development. Deployment is also part of the training infrastructure.

Countries that accelerate commercial adoption may create a compounding data advantage that becomes increasingly difficult for slower markets to replicate.

Robot foundation models are becoming complete systems

The second major shift is occurring at the model layer. The industry is moving beyond the assumption that one monolithic vision-language-action model will become a universal robot brain.

Alibaba’s Qwen team introduced the Qwen-Robot suite, separating the stack into three connected foundation-model capabilities:

  • Qwen-RobotNav for embodied navigation

  • Qwen-RobotManip for manipulation across robot embodiments

  • Qwen-RobotWorld for action-conditioned world modeling

This architecture matters. It recognizes that navigation, physical manipulation, and forward prediction are related problems, but not necessarily identical learning problems. [4]

Mistral has also entered embodied AI with Robostral Navigate, an 8-billion-parameter navigation model that processes natural-language instructions and RGB images from a single camera, without requiring LiDAR or a dedicated depth sensor. Mistral reports a 76.6 percent success rate on unseen R2R-CE environments after training entirely in simulation on approximately 2.4 million trajectories across 350,000 scenes. These are developer-reported benchmark results and should not be treated as equivalent to independently validated commercial performance. [5]

Xiaomi Robotics introduced Xiaomi-Robotics-U0, a 38-billion-parameter multimodal world foundation model. The system combines image generation, image editing, multi-view embodied-scene synthesis, embodiment transfer, and action-conditioned video generation. In the authors’ experiments, synthetic data generated by the model increased out-of-distribution manipulation success from 36.9 percent to 63.2 percent. [6]

NVIDIA and Hugging Face are simultaneously working to standardize the development layer. Their LeRobot integration brings NVIDIA Isaac GR00T 1.7, Isaac Teleop, datasets, and robot-development workflows into an open ecosystem, with Cosmos 3 integration planned. The objective is to provide a more consistent pipeline for collecting data, post-training robot models, evaluating policies, and deploying them across different embodiments. [7]

My expectation is that the winning robot architecture will be hierarchical rather than monolithic. It will likely combine:

  • A multimodal semantic model for instruction understanding and high-level reasoning

  • A world model for predicting state transitions, generating synthetic experience, and evaluating possible actions

  • A task policy for producing embodiment-specific action sequences

  • A real-time control layer for low-latency motion, force, balance, and contact control

  • An independent safety supervisor capable of constraining or terminating unsafe behavior



Hierarchical Physical AI system with an intelligence stack, independent safety lane and admissible action gate

Figure 3. A hierarchical architecture separates reasoning, prediction, policy, control, and safety responsibilities.

This is why model benchmarks alone are not sufficient. A model can perform well in an offline evaluation and still fail once it encounters latency, sensor noise, actuator variation, contact uncertainty, moving people, or an environment outside its training distribution. The real product is the integrated control and learning system.

The data engine is becoming the core strategic asset

At STATE16, I see the robot-data layer as one of the most important and least understood parts of the market. Web-scale AI benefited from enormous volumes of relatively accessible text, image, audio, and video data. Physical AI does not have an equivalent native dataset.

Robot learning requires much richer information:

  • What the robot observed

  • What action was selected

  • What command was sent to each actuator

  • What forces and contacts occurred

  • What changed in the environment

  • Whether the task succeeded

  • Whether a human intervened

  • Why the policy failed

  • Whether the trajectory was safe

This data is expensive, embodiment-specific, and difficult to capture at scale.

Apptronik’s expanded Robot Park illustrates how companies are responding. The nearly 90,000-square-foot facility operates fleets of Apollo 2 robots in bipedal and wheeled configurations, collecting data across manufacturing, logistics, retail, and other customer-driven tasks. Apptronik says the resulting data supports its collaboration with Google DeepMind and the development of Gemini Robotics models. [8]

This is more than a robot-testing facility. It is a Physical AI data factory. The strategic loop is: deploy, collect demonstrations and failures, retrain, validate, update the fleet, and collect higher-value data.



Robot data flywheel connecting deployment, observation, failure capture, retraining, validation and fleet updates

Figure 4. Physical AI compounds through a closed data, validation, and fleet-update loop.

The strongest long-term advantage may therefore belong to the company with the fastest learning loop, not necessarily the company with the most advanced robot at one specific moment. A hardware lead can narrow. A proprietary dataset generated across thousands of deployed machines can compound.

Industrial evidence is beginning to separate from demo theater

One of the biggest problems in robotics is the gap between demonstrated capability and sustained operational performance. A carefully selected video can show that a robot completed a task once. It does not tell us:

  • How often the task succeeds

  • How much human supervision is required

  • How long the robot remains operational

  • How frequently components fail

  • How long integration takes

  • Whether the economics are competitive

  • Whether the customer wants to expand the deployment



Industrial evidence ladder from curated robot demo to controlled pilot, sustained operation and customer expansion

Figure 5. Evidence becomes more meaningful as performance moves from a curated demo to repeatable expansion.

This is why BMW’s continued work with Figure is important. BMW reports that Figure 02 supported the production of more than 30,000 BMW X3 vehicles over a ten-month deployment at its Spartanburg plant. Figure 03 is now being introduced for a more complex logistics-sequencing workflow in which it will pick unsorted components, organize them into the required production sequence, and prepare them for transportation to the assembly line. [9]

The critical signal is not that Figure completed a task. The critical signal is that a major industrial customer completed an initial deployment, evaluated the system, returned with the next robot generation, and expanded the scope of work.

AGIBOT reported another type of scale signal when its 15,000th robot rolled off the production line. The company also reported approximately 100 cumulative hours of factory livestreaming involving a G2 robot operating inside a tablet-production quality-inspection process. These are company-reported figures and should be evaluated accordingly, but they indicate progress across supply chains, standardized manufacturing, quality control, and engineering delivery. [10]

The next competitive barrier is the ability to manufacture, calibrate, validate, deploy, maintain, and update thousands of machines with consistent performance. That requires a fundamentally different organization from a robotics research laboratory.

Safety is becoming part of the core architecture

The transition from demonstration to deployment changes the engineering requirements. A robot operating alone in a controlled environment can rely on procedural safeguards. A robot operating beside employees must have safety built into the architecture itself.

NVIDIA introduced Halos for Robotics as a full-stack safety system spanning compute hardware, sensor connectivity, system software, safety applications, inspection, and preparation for third-party certification. Agility Robotics was announced as the first humanoid-robot company incorporating elements of the platform into its own safety system. [11]

This development is strategically important because safety cannot be treated as a final compliance step. A production-grade Physical AI system requires:

  • Redundant sensing where necessary

  • Deterministic fallback behavior

  • Safe operating envelopes

  • Runtime policy monitoring

  • Hardware and software fault detection

  • Traceable decision and event logs

  • Independent emergency-stop pathways

  • Validation under edge cases and distribution shift

  • Evidence suitable for certification and customer review



Physical AI safety architecture with an independent runtime supervisor and admissibility gate

Figure 6. The action policy proposes behavior; an independent assurance layer determines admissibility.

The most capable policy should not also be the sole authority determining whether an action is safe. In mature Physical AI architectures, intelligence and safety will be tightly integrated but operationally separable.

Capital is moving from software-scale funding to infrastructure-scale funding

Physical AI is capital-intensive by design.

The sector requires model development, simulation infrastructure, data-collection facilities, specialized hardware, manufacturing lines, component supply chains, deployment teams, maintenance networks, safety validation, and customer integration.

Physical AI companies increasingly combine the economics of AI infrastructure, semiconductors, advanced manufacturing, and industrial automation.



Physical AI capital stack combining AI infrastructure, semiconductors, advanced manufacturing and industrial automation

Figure 7. Physical AI inherits the capital and operating requirements of four infrastructure-heavy industries.

That requirement is now visible in financing activity.

NEURA Robotics announced a Series C with a potential total size of up to $1.4 billion. Its investors include Amazon, NVIDIA, Qualcomm, Bosch, Schaeffler, Tether, and the European Investment Bank. NEURA says the capital will support its Physical AI platform, serial production, and a network of real-world robot-training environments called NEURA Gyms. [12]

Agility Robotics announced a proposed public-market transaction at a pre-money equity value of approximately $2.5 billion. The transaction is expected to provide more than $620 million in gross proceeds, assuming no redemptions and subject to shareholder, regulatory, and other closing conditions. [13]

Unitree received regulatory approval for a Shanghai STAR Market IPO through which it plans to raise approximately 4.2 billion yuan, or about $619 million. Approval is an important milestone, but it is not the same as a completed offering. [14]

The quality of the technology will remain essential, but capital efficiency, production yield, component costs, service requirements, and customer payback periods will become equally important.

Labor and organizational adoption are part of the deployment equation

Technical capability does not automatically translate into organizational permission to deploy. Hyundai workers in Ulsan, South Korea, entered a partial strike during negotiations that included concerns about AI, automation, employment security, and the possible future introduction of Boston Dynamics’ Atlas humanoid robot. Hyundai has not announced a fixed Atlas deployment date for its South Korean factories, but the union has argued that humanoid systems should not be introduced without worker agreement. [15]

This is an early indication of a constraint that technology companies frequently underestimate. Physical AI deployment will affect:

  • Job design

  • Workforce planning

  • Compensation structures

  • Training and reskilling

  • Productivity-sharing agreements

  • Operational responsibility

  • Workplace safety

  • Labor relations

  • Public legitimacy



Workforce adoption system connecting workflow design, job design, accountability and labor agreement

Figure 8. Workforce adoption is a socio-technical system, not an external communications task.

The most sophisticated deployment strategy will not present the robot as a standalone replacement unit. It will redesign the operating model around humans, machines, workflows, accountability, and measurable productivity. Workforce adoption is not external to the product. It is part of the product’s deployment architecture.

My conclusion

Physical AI is evolving from a robotics category into a systems platform for converting intelligence into physical productivity.

The long-term winners will not simply manufacture robots. They will not simply train foundation models. They will operate closed-loop systems that combine intelligence, embodiment, action data, simulation, control, safety, manufacturing, and distribution.

Every deployment will create data. Every high-quality data point will improve the policy. Every policy improvement will increase autonomy, reliability, and economic output. And every improvement deployed across a fleet will strengthen the underlying platform.



Physical AI systems phase compounding loop from deployments and action data to policy gains, safety evidence, fleet scale and economics

Figure 9. The systems-phase advantage compounds across learning, assurance, manufacturing, and deployment.

That is the compounding mechanism.

My position is clear: the market will not be decided by the robot that produces the best demonstration. It will be decided by the organization that builds the best learning, safety, manufacturing, and deployment loop.

That is the Physical AI transition we are tracking at STATE16.

Sources

  1. China Ministry of Industry and Information Technology: Humanoid robots and embodied intelligence initiative

  2. Republic of Korea policy briefing: Three national mega-projects

  3. Japan Ministry of Economy, Trade and Industry: AI Robotics Strategy

  4. Qwen-Robot Suite: A Foundation Model Suite for Physical World Intelligence

  5. Mistral AI: Robostral Navigate

  6. Xiaomi Robotics: Xiaomi-Robotics-U0

  7. NVIDIA and Hugging Face: GR00T, Isaac Teleop, and LeRobot integration

  8. Apptronik: Robot Park and Apollo 2

  9. BMW Group: Figure 03 deployment at Plant Spartanburg

  10. AGIBOT: 15,000th robot production milestone

  11. NVIDIA: Halos for Robotics

  12. NEURA Robotics: Series C of up to $1.4 billion

  13. Agility Robotics: Proposed public-market transaction

  14. Reuters: Unitree receives approval for planned Shanghai IPO

  15. Semafor: Hyundai workers in South Korea strike over humanoids