iantoons-logo

The Robots That Help Us Won’t Look Like Us

We build machines in our own image and measure them against ourselves, which means we often miss the moment they discover a better way.

Testing-Time-iantoons
Testing Times – Human IQ, Robot IQ, and two ways to fail the same door.

I. The Crowd

In 1851, the Great Exhibition put steam engines and power looms behind ropes in Hyde Park, and six million people came to look at them, roughly a third of the population of Britain. Many of those visitors worked in trades that the machines on display were already beginning to destroy, yet they queued to see them anyway.

The same instinct is alive today. Boston Dynamics clips attract tens of millions of views, Chinese humanoid races draw audiences around the world, and when Figure ran a 24-hour livestream of robots sorting warehouse parcels, more than three million people watched 28,000 packages go by. Sorting parcels is hardly compelling entertainment, but put two arms and two legs around the activity and suddenly people want to watch.

The Victorians were fascinated by machines because they could do things people could not. A loom worked faster than human hands and a steam engine could keep going through the night. Nobody cared whether a loom looked like a weaver.

With robots, we often react in the opposite way. A machine walking across a stage attracts attention because walking is something we recognise from ourselves, while a conveyor system moving ten times as many parcels barely registers as technology at all. Resemblance draws us in, and before long it starts shaping what we build and how we judge it.

passing-the-baton-iantoons
Passing the Baton  One astonishing machine, or enough machines to compound.

II. The Yardstick

Many of the benchmarks we now use for machines began as ways of measuring people.

We test machine intelligence using instruments descended from one Alfred Binet developed in 1905 to identify French schoolchildren who needed additional help. We race humanoids over 100 metres, a distance inherited from human athletics, and assess warehouse robots in picks per hour against the productivity of a human worker. The machine keeps finding itself placed on a scale whose reference point is the human mind or body.

IQ illustrates the problem particularly well. On public intelligence tests, frontier models can score around 136, comfortably inside the range we associate with highly intelligent people, while private versions containing questions that have never appeared online produce scores closer to 96 to 116. Whatever proportion of that difference comes from training exposure, it makes a single human benchmark look like a fairly shaky description of machine intelligence.

The cartoon makes the problem clearer. Outside the Mensa Testing Centre are two identical doors marked PULL TO OPEN. On the left, the man pushes his, the familiar cartoon shorthand for somebody unlikely to trouble the upper reaches of an intelligence test. On the right, the robot reads the instruction literally and pulls the sign itself off the wall.

Neither candidate gets through the door, but their mistakes have almost nothing in common. The man knows exactly what a door is and fails to follow the instruction, while the robot follows the instruction with perfect literal accuracy but fails to understand what the sign is referring to.

AI is incredibly powerful, but it still lacks a deep understanding of context that humans possess naturally.

Andrew Ng

An IQ of 136 makes the machine sound like an unusually clever person, even though its strengths and weaknesses are distributed in ways no person shares. Put a human and a machine in front of the same door and they can fail at exactly the same task for completely different reasons.

Once the benchmark is human, the temptation is to build around human assumptions as well, even when those assumptions have little to do with the problem the machine is actually trying to solve.

III. Left to Themselves

At the World Humanoid Robot Games in Beijing last month, Tiangong Ultra ran 100 metres in 8.64 seconds against Usain Bolt’s 9.58, while its 38.15-second 400 metres was faster than Wayde van Niekerk’s human record of 43.03. The starting protocols were different and no record was ratified, so the headline comparison is more theatrical than scientific, but something that happened to another robot in the competition is considerably more interesting.

Engineers initially taught Tiangong Omni to run with a normal human arm swing because human movement was the obvious starting point. As reinforcement learning iterated on the gait in simulation, it found that the familiar arm motion generated too much heat and stress through the robot’s shoulder joints over a sustained effort. The system gradually converged on a different movement in which the arms remained high and close to the head, reducing the demands on joints that occupy roughly the same place as ours but do not work in the same way.

When footage of the gait circulated online, people quickly concluded that the robot had been trained to imitate a little girl running, and the engineers eventually had to explain publicly that the movement was “fully self-developed via reinforcement learning” and had not been designed to imitate anyone.

What is revealing is how quickly observers reached for a human explanation. Faced with a machine moving in a way that did not fit our expectations, many people still assumed there must be another person somewhere in the story for it to be copying.

The human gait had been a perfectly sensible place to begin, but there was no reason it had to be the place where the machine ended up. Once the system was allowed to optimise around its own mechanics, it found a movement that made more sense for the machine than the one inherited from us.

That kind of discovery becomes more common when robots leave demonstrations and start doing real work. Of the 542,000 industrial robots installed worldwide in 2024, China installed 295,000, or 54% of the total, compared with 34,200 in the United States. China’s operational stock has passed two million units, while domestic manufacturers now supply 57% of its home market, up from roughly 28% a decade earlier.

Every one of those machines is another opportunity to find where an assumption breaks, where a component fails, or where a better movement emerges. Large-scale deployment produces more than output. It produces the awkward cases that demonstrations are designed to avoid.

The relay cartoon captures some of that difference. An American robot is out in front performing an impressive leap, while four Chinese robots behind it pass a baton. The advantage may come less from producing one astonishing machine than from having enough machines operating that improvements can compound from one generation to the next.

Robotic-iantoons
Robotic – Do you know where we vote?

IV. Why We Keep Building Ourselves

America’s strength in software has naturally encouraged a mind-first view of robotics in which intelligence is treated as the difficult problem and the body as something that can be added once the intelligence is ready. From that perspective, the humanoid form makes plenty of practical sense because the physical world we inhabit has already been designed around human proportions.

Our stairs assume legs of roughly our length, our handles assume hands, and our work surfaces sit at heights determined by the reach of human arms. A machine shaped like us can operate inside existing infrastructure instead of requiring factories, houses and public spaces to be redesigned around it.

Humanoids therefore make obvious sense in places built for humans. The leap comes when that practical fit starts being treated as evidence that the human body is also the best general form for a machine.

Recent disappointment around enterprise AI has made robotics even more attractive. Models that are wrong 10 or 15% of the time can become liabilities inside important workflows, and Salesforce executives have publicly discussed reducing reliance on large language models in some core processes after encountering problems with reliability. Robotics appears to offer something more tangible because performance can be measured in physical outcomes. Amazon’s warehouse robots, for example, have reduced order cycle times by roughly 20%, while autonomous guided vehicles can remove substantial amounts of repetitive labour from logistics operations.

But putting AI into a physical body does not make unreliability disappear. A wrong answer in a chat window can usually be corrected without much consequence, while the same error translated into movement by a machine with weight, motors and momentum becomes an action in the real world.

The current enthusiasm for humanoids therefore mixes two ideas together. Giving AI a body may make it far more useful, but it does not follow that the most useful body will usually resemble ours.

In the cartoon, everyone is climbing onto the robotics wagon while the horse has stopped to graze, which feels about right for a technology cycle in which conviction is moving slightly faster than certainty.

band-wagon-iantoons
Band Wagon – I’m not sure if they’ve modeled this out properly.

V. A Different Template

Once the human body is no longer treated as the obvious starting point, the range of possible solutions becomes much wider.

Collapsed-building rescue is a useful example because the objective is brutally simple. A 2021 study found that 96% of people trapped under debris for more than 24 hours did not survive, compared with 26% of those reached within the first day, so the problem is getting eyes and sensors into unstable spaces as quickly as possible without putting more people at risk.

Thermal imaging can help locate survivors, but the harder challenge is reaching deep into rubble through gaps that conventional ground robots cannot navigate. Researchers in Japan and Singapore approached the problem by fitting Madagascar hissing cockroaches with lightweight circuit boards carrying sensors and rechargeable batteries, effectively turning the insects into mobile platforms rather than trying to reproduce their capabilities mechanically.

The insects are not controlled step by step. Researchers nudge them loosely in the required direction and let them negotiate the terrain themselves. The technique requires roughly half as many control interventions as earlier approaches, consumes little energy because the insect provides its own locomotion, and avoids having to engineer a machine capable of climbing across fractured surfaces because the animal already knows how to do it.

Unlike robots, insects do not behave as we intend them to. However, instead of forcibly trying to control them precisely, we found that taking a more relaxed and rough approach worked better.

Ryo Wakamiya, Osaka University

The researchers did not make the cockroach behave more like a machine. They changed the system around the way the cockroach already behaved.

Something similar happened with the running robot in Beijing. Tiangong’s engineers began with a human movement and allowed the machine to depart from it, while the cockroach researchers began with an animal they could not precisely control and designed around that limitation. Neither solution came from forcing the system back toward something familiar.

VI. So Much For Two Legs

Nature offers an enormous library of solutions, but evolutionary history also shows why it should be used carefully.

Around 25 million years ago our ancestors lost their tails and left behind the coccyx. Much later, as our lineage adapted to permanent bipedal movement, the arms we already possessed became part of the balancing system used in walking and running. The human body was not designed from the outset for efficient movement on two legs. It emerged through a long sequence of modifications to structures inherited from earlier forms.

Natural selection does not work as an engineer works. It works like a tinkerer.

François Jacob

Evolution works with what is already there. That produces remarkable solutions, but it also leaves plenty of historical baggage behind. A cockroach is extremely good at negotiating difficult terrain, although its anatomy reflects hundreds of millions of years of inherited constraints rather than an attempt to create the ideal rescue platform.

Machines are not bound by the same lineage. An engineer can take one useful principle from an insect, another from a vehicle and something else entirely from software, without needing a plausible sequence of ancestors connecting one form to the next. A machine designed for rubble does not need to look like the animal that inspired its movement any more than an aircraft needs to flap its wings because birds do.

Humanoids will still have an important role because so much of the physical world was built for people, and matching our general proportions gives a machine immediate access to stairs, tools, vehicles and workplaces that already exist. That is a very good reason to build humanoids. It is not a reason to assume that two arms, two legs and a head are where embodied intelligence is ultimately heading.

Some useful machines will look like us because the environment rewards that shape, while others may borrow from insects, vehicles or structures for which there is no biological equivalent at all. What appears awkward or even ridiculous to us may simply reflect the fact that the machine is solving a problem under a different set of physical constraints.

The strange running gait in Beijing looked wrong because we were judging it against our own movement, even though the robot’s shoulder joints were not ours and there was no particular reason the motion that suited a human body should also suit a machine. Once the engineers stopped insisting on that resemblance, a different answer emerged.

We built machines in our own image and then handed them our own tests, but the interesting part starts when they no longer need either.

Evolution – So much for two legs.