4 min read

Brain waves join the race to train better robots

Encord and Zander Labs are testing brain-wave data as a way to improve robot training and address the shortage of physical manipulation data.

Image: TechCrunch

The next bottleneck for humanoid robots may not be model architecture. It may be the lack of enough high-quality data showing machines how to manipulate the physical world.

At Encord’s warehouse in San Leandro, California, pilot Andrew Ceja carefully removes wooden blocks from a Jenga tower. He is wearing a camera-equipped headset, a common setup for collecting robot-training data. This one also measures his brain waves.

Encord builds data tooling for AI-model training, but it is now expanding into the production of physical training data. The company is working with Zander Labs, a German neuroscience startup, to test whether brain activity can reveal mental states such as error, intent and surprise. The trial aims to create an initial brain-wave-tagged data set, run it through customer robotics models and measure whether performance improves before Encord decides whether to scale the effort.

Lucas Gehrke, a Zander neuroscientist supervising the work, says the amount of brain activity during a task could help model builders determine when to use their highest-effort models.

Recommended reading

Kimi reignites the panic over Chinese AI

“The bleeding edge” of the effort to solve the robotics data bottleneck.

Vineeth Velmurugan, Encord’s head of robot learning

Why robotics needs manufactured data

Velmurugan, a veteran of OpenAI’s robot lab and Berkshire Grey, joined Encord to create its internal data-generation team. Encord was founded to help companies building machine-vision applications annotate data and evaluate models. But as its customers began using end-to-end learning for robotic manipulation, the company concluded that it would need to produce training data itself.

“The data simply does not exist.”

Vineeth Velmurugan, Encord’s head of robot learning

The comparison between generative AI for robots and generative AI for chatbots runs into a fundamental difference. Large language models were trained on the text of the entire internet and more. There is no equivalent, readily available corpus for physical manipulation. Self-driving companies collect real-world data themselves, but that approach is difficult to scale. Video can help, but it lacks the fidelity of physical interaction.

Velmurugan estimates that robotics may need a data set roughly five times the size of YouTube’s video corpus to make a breakthrough. That scale is helping turn data generation into a business rather than only a research problem.

Encord’s robot-training data experiments

Encord is pursuing two main sources of data:

  • Egocentric video captured by workers wearing cameras, sometimes combined with additional camera angles and other measurements.
  • Teleoperated robot data, collected when people remotely control robots as they perform tasks.

The company uses both approaches at factories around the globe. Its San Leandro facility is also testing new formats, including brain-wave and forearm-sensor data, and creating task-specific data sets for fine-tuning.

During TechCrunch’s visit, pilots used leader-follower rigs: paired robotic arms in which a human directly controls one arm while the second mimics its movements. The tasks included pouring coffee into mugs, stacking poker chips and plugging and unplugging Ethernet cables from the back of a server.

The facility’s supplies included fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires. These objects are used to train robots for household and industrial manipulation. Encord says every humanoid company has asked for this type of data, although it has not named its customers.

Another experiment uses sensors strapped to the forearm to detect electrical signals in muscles. Human-hand video often fails to capture the entire hand, so Velmurugan hopes the sensors can help create a 3D depiction of the hand’s position over time.

Encord also adds dense physical descriptions to its videos, such as “right hand tightens bolt,” to help language-model-based systems understand each action. Velmurugan estimates that this annotation is worth 100 times more than “junky ego data” for specific training tasks, while costing 20 times more to produce.

That cost remains the central problem. Scraping text from Stack Overflow and the wider internet was nearly free for frontier AI labs. Physical data must be manufactured, making the economics of building robotics models fundamentally different.

Encord currently has roughly a dozen pilots working at its facility. Among them are Ceja and Sofia Infante, both formerly of Scale, another AI data-annotation firm. Ceja previously maintained a robotic trash sorter at a waste-management company. Now he is helping build the data needed to make robots handle tasks that human fingers still perform with far greater dexterity.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via TechCrunch

/ Keep reading