Why AI Assistants Are Leaving the Screen and Learning to Act in the Real World

Tuesday, August 18, 2026

SAEDNEWS: AI is entering a new phase in which systems are being designed not just to answer questions, but to perceive environments, plan actions and operate physical machines. The shift toward “embodied” or “physical” AI could change how robots work in factories, homes, laboratories and other human spaces.

Why AI Assistants Are Leaving the Screen and Learning to Act in the Real World

According to SaedNews, the next major chapter of artificial intelligence may be less about what appears on a screen and more about what happens beyond it.

For years, the most visible AI systems have lived inside phones, browsers and computers. They write emails, summarize documents, generate images, answer questions and help users navigate digital information. But increasingly, researchers and technology companies are trying to give AI another capability: the ability to understand and act in the physical world.

This emerging field is often described as embodied AI or physical AI. Its central idea is simple but technically difficult: an intelligent system should be able to perceive its surroundings, reason about what it sees, decide what to do and then physically carry out that decision.

That is a very different challenge from generating a paragraph of text.

 AI systems

A chatbot can describe a cup. A robot has to pick it up.

Imagine asking an AI assistant to make you a cup of coffee.

A conventional chatbot can provide a recipe. A more advanced digital agent might search for the recipe, create a shopping list and schedule a reminder.

A physical AI system would have to do something much harder. It would need to identify the coffee machine, locate the cup, understand where objects are positioned, grasp them without dropping them, operate unfamiliar controls and react if something is moved or spills.

The difficulty is that the physical world does not behave like a webpage.

Objects can be slippery, damaged, heavy, transparent or unexpectedly positioned. People can walk into the robot's path. Lighting can change. A door that was open yesterday may be closed today.

Researchers therefore increasingly view perception, reasoning and physical action as parts of the same intelligence problem. Nature Machine Intelligence noted in 2026 that robustly perceiving, acting and adapting in the real world remains an open challenge, partly because physical interaction generates much more uncertainty than many digital tasks.

Why AI suddenly has a chance to make this leap

One reason is the rise of powerful multimodal foundation models.

Modern AI systems can process combinations of text, images, audio and video. The next step is to connect those abilities to sensors and robot controls.

Google DeepMind's Gemini Robotics illustrates this direction. Its vision-language-action models are designed to turn visual information and instructions into actions for robots, while its embodied-reasoning models focus on understanding physical spaces, planning tasks and making decisions.

In September 2025, Google DeepMind described Gemini Robotics 1.5 as part of a move toward “physical agents”—systems that can perceive, plan, use tools and act on multi-step tasks. One example involved sorting objects into recycling, compost and trash based on local rules, requiring the system to combine online information with visual perception and physical action.

The important change is not simply that robots are becoming better programmed.

It is that researchers are attempting to build general-purpose AI models for robots, much as language models can be adapted to many digital tasks.

The robot needs more than a “brain”

Giving a robot a powerful AI model does not automatically make it useful.

A physical machine also needs cameras, depth sensors, force or tactile sensing, motors, batteries, processors and control systems. The AI has to operate within the limitations of the body carrying it.

NVIDIA's robotics platform, for example, describes humanoid systems as requiring AI, simulation, sensors, embedded computing and mechanical engineering to work together. Its Isaac GR00T platform is designed as a foundation for developing, training, testing and deploying AI-powered humanoid robots.

This helps explain why the current robotics race looks different from the race to build larger language models.

A language model can produce an answer in a fraction of a second. A robot must make decisions while balancing its body, avoiding obstacles and controlling physical forces. A mistake is no longer merely an incorrect sentence—it can mean dropping an object, damaging equipment or colliding with a person.

Training robots is one of the biggest obstacles

There is another problem: data.

Large language models benefited from enormous quantities of digital information. Robots do not have an equally convenient supply of high-quality physical-world training data.

A robot learning to open a drawer, for example, needs information about movement, friction, grip, position, force and the many ways a drawer can behave. Collecting such experiences with real machines is expensive and slow.

A 2026 University of California, Berkeley technical report highlights this problem: robot data is expensive to collect, often tied to particular hardware and environments, and difficult to scale in the same way as the datasets used for modern foundation models.

That is why simulation has become so important.

Instead of making a physical robot perform millions of risky experiments, researchers can train and test systems inside virtual environments. NVIDIA's robotics learning materials describe simulation as a way to generate large numbers of training scenarios, test hazardous situations more safely and reduce some of the cost of real-world data collection.

The remaining challenge is what researchers often call the sim-to-real gap: a robot that succeeds in a virtual environment still has to cope with the messy details of the real one.

Why humanoid robots keep appearing in the AI race

Humanoid robots are receiving particular attention because much of the world has already been designed for human bodies.

Stairs, shelves, doors, tools, kitchen counters and factory workstations are generally built around human reach, height and movement. A machine with a similar physical form could potentially operate in these spaces without requiring everything around it to be redesigned.

NVIDIA says general-purpose humanoids are being developed for human-centered environments, including factories, warehouses, healthcare and retail.

Google DeepMind has similarly demonstrated its robotics models across different robot embodiments, including humanoid systems, emphasizing the goal of transferring intelligence across different physical forms.

But humanoid shape alone is not the breakthrough.

The more important development is the combination of a versatile body with AI that can interpret instructions and adapt to situations it was not explicitly programmed to handle.

The surprising part: intelligence may depend on the body

There is a deeper scientific reason researchers are interested in physical AI.

For decades, it was tempting to think of intelligence primarily as computation happening inside a “brain.” Embodied-intelligence research challenges that separation.

A physical body changes what an intelligent system can sense and how it can learn. A robot's hands, sensors, joints, materials and interaction with its environment all affect what it can accomplish. Nature Machine Intelligence describes this broader convergence around embodiment, sensorimotor interaction, world models and physical AI.

This means the robot's body is not simply a container for software.

It is part of the problem the AI has to solve.

A robot that reaches for a glass must understand not only what the glass looks like, but where its hand is, how far away the object is, how much force is appropriate and what could happen if its fingers slip.

Researchers working on embodied AI are therefore trying to connect perception directly with action. A 2025 study in Nature Machine Intelligence demonstrated an embodied large-language-model framework in which language reasoning was combined with visual and force feedback to perform multi-step tasks in an unpredictable environment.

Does this mean household robot assistants are almost here?

Not necessarily.

The gap between impressive demonstrations and reliable everyday autonomy remains substantial.

Current systems can perform increasingly sophisticated tasks, but the physical world contains an almost endless number of unusual situations. Opening a specific type of door in a controlled laboratory is very different from safely navigating a crowded apartment, recognizing fragile objects and understanding a person's changing intentions.

Even recent research emphasizes that robust real-world behavior remains an unresolved problem.

The direction, however, is becoming clearer.

Google DeepMind's 2026 Gemini Robotics 2 work describes systems aimed at whole-body control, dexterity, multi-step planning and collaboration between robots. NVIDIA is simultaneously developing robot foundation models, simulation systems and onboard computing intended to move physical AI from training environments toward real-world deployment.

The real shift is from answering to acting

The most important change may not be the arrival of a particular humanoid robot.

It is the changing definition of what an AI assistant can be.

A digital assistant can tell you how to assemble a piece of furniture. A physical assistant could eventually assemble it. A chatbot can explain how to organize a warehouse. A physical agent could potentially move objects, inspect shelves and respond to changing conditions.

That transition—from knowing about the world to acting in it—is one of the central challenges of the next generation of AI.

And it explains why the AI race is increasingly moving beyond screens.

The future assistant may not always be an app waiting for your next prompt. In some settings, it could be a machine that sees the same room you do, understands what you ask, and has the physical ability to help.