Artificial intelligence July 27, 2026

Encord tests brain-wave headsets for physical AI data collection

A warehouse in San Leandro is quietly turning into a robotics data lab. Encord, the AI data tooling company, is using humans to do awkward robot jobs while wearing sensors, cameras, and now a brain-wave headset from German startup Zander Labs. The po...

Encord tests brain-wave headsets for physical AI data collection

Brain waves may become the most expensive label in robotics

A warehouse in San Leandro is quietly turning into a robotics data lab. Encord, the AI data tooling company, is using humans to do awkward robot jobs while wearing sensors, cameras, and now a brain-wave headset from German startup Zander Labs. The point is straightforward: if humanoids and warehouse robots are going to get useful, someone has to produce training data that looks enough like the real world to matter.

That’s where the industry keeps running into the wall. The model side gets the attention. The data side is the bottleneck.

Encord’s setup shows how far robotics training has moved past generic video collection. Workers, or “pilots” in Encord’s terminology, manipulate objects with camera rigs, leader-follower arms, and extra sensors that capture motion, muscle signals, and now brain activity. The company is testing whether those extra channels produce better labels for robot policies, especially around moments that are hard to read from video alone: intent, surprise, error, hesitation.

Why brain waves are in the mix

The brain-wave experiment sounds odd until you look at what robotics teams are trying to solve.

Language models got the internet. Robots didn’t get anything close to that. Physical manipulation has to be collected in the world, by people moving real objects through real space. That means grasping, slipping, twisting, occlusion, timing, force, and all the small failures video often misses.

Zander Labs thinks EEG-style signals can add another layer of supervision. If a worker’s brain activity changes when a task gets harder, or when they notice an error before the video makes it obvious, that could help builders mark the exact parts of a sequence where a model needs more capacity or a different policy. Lucas Gehrke, the neuroscientist supervising the work, says the amount of brain activity during a task can hint at when systems should switch to higher-effort models.

It’s a decent idea. It’s also messy.

Brain signals are noisy, person-specific, and sensitive to setup quality. EEG data rarely survives contact with the real world without heavy preprocessing, calibration, and caveats. If this turns into anything useful, it won’t be because brain waves are magical. It’ll be because they help label a small slice of high-value robotics data more precisely than video and manual annotation alone.

The product is data, not software

Encord started as a company for annotating machine-vision data and evaluating models. That’s already a crowded business. The robotics angle is different because customers increasingly need data they don’t have, not just better tooling for data they already collected.

Vineeth Velmurugan, Encord’s head of robot learning and a former OpenAI robotics lab and Berkshire Grey engineer, says some customers moved toward end-to-end learning on manipulation tasks and hit the same wall: the data simply doesn’t exist. Not at the scale they need, anyway.

That’s the shift. In robotics, data generation is becoming a production line.

Encord is pulling in two main types of data:

  • Egocentric video, usually from people wearing cameras while doing tasks
  • Robot-generated data, often collected through teleoperation or leader-follower rigs

The company is also experimenting with additional sensors, including forearm electrodes that pick up muscle activity. That makes more sense than it sounds. Video of human hands is often partial, occluded, or too low-resolution to reconstruct fine motor control. Muscle signals can help infer hand pose and motion when the camera can’t see the fingers clearly.

The goal is a denser picture of action, not just a prettier clip archive.

Dense labels matter more than more clips

At the San Leandro site, Encord’s pilots were training robotic arms to pour coffee, stack poker chips, and plug and unplug ethernet cables from server racks. That last task gets waved off a lot because it sounds repetitive and therefore easy. It isn’t. Cable insertion is a precision problem with ugly failure modes. Humans do it casually because our fingers, wrists, and tactile feedback are absurdly good.

Robots are not there yet.

Encord annotates these videos with physical descriptions like “right hand tightens bolt.” That sounds almost quaint until you think about how robot models actually learn. The more precise the description of action and state, the easier it is to train systems that can map observation to control.

Velmurugan says this dense annotation can be worth 100 times as much as “junky ego data” for specific tasks, while costing about 20 times more to produce. That math is the whole argument, and also the risk. If the data is really 100x better, spending 20x more starts to make sense. But the payoff depends on the task, the model, and whether the labels transfer beyond one narrow workflow.

This is where robotics diverges from LLM training. Text scraping was cheap because the web was already there. Physical training data has to be staged, filmed, synchronized, cleaned, and often repeated until the model sees enough variation to generalize. That’s labor. Hardware. Space. Time. Real money.

The scale problem is ugly

Velmurugan says the industry may need a dataset around five times the size of YouTube’s video corpus to break through. Treat that as a directional warning, not a law of nature. Even so, it says enough.

If physical AI really needs that much data, the winners won’t just be model builders. They’ll be the companies that can industrialize data production, label quality, and sensor fusion without blowing up costs.

That also explains why Encord’s pitch is part tooling company, part data factory, part industry wiretap. It sits across a bunch of robotics programs and can see which collection methods are actually getting traction. That matters because robotics teams are still guessing at what signals improve model performance. One company’s clever sensor stack can easily be another company’s expensive dead end.

A lot of robotics progress still feels experimental because that’s what it is. Teams are trying video-only learning, teleoperation, leader-follower rigs, multi-angle capture, muscle sensors, and now EEG. Some of those approaches will stick. Some won’t. The market won’t reward elegance here. It’ll reward whatever produces better policies faster.

The caveat nobody should ignore

There’s a temptation to read this as proof that richer biosignals will solve robotics. That would be premature.

Brain-wave data is attractive because it promises access to hidden state: intent, uncertainty, surprise, fatigue. But hidden state only helps if you can measure it reliably enough, at scale, across different people and tasks. If the signal is too fragile, it becomes a research demo with a headset.

There’s also a privacy issue robotics teams can’t shrug off. Once you start collecting physiological signals from workers, you’re not just logging performance. You’re creating data that could be used to infer cognitive state or stress. That’s a different class of dataset, and companies will need clear consent, access controls, retention policies, and probably more restraint than most AI teams are used to.

Muscle sensors may end up being the more practical path. They’re less exotic, easier to explain, and probably less fraught. But the broader direction is clear: robot training data is getting more multimodal, more instrumented, and more expensive.

Why engineers should care

For teams building robotic systems, this is the part that matters:

  • Video alone is often too lossy for fine manipulation tasks.
  • Teleoperation and leader-follower setups are becoming standard data factories.
  • Dense action labels can be more valuable than raw clip volume.
  • Additional sensors like EMG or EEG may help with edge cases, but they add calibration, cost, and privacy headaches.
  • Data quality now shapes model strategy. If you can detect intent or uncertainty better, you can route tasks to higher-capacity models or different control policies.

That’s where physical AI starts to look more like systems engineering than model wizardry. The hard part isn’t building another policy network. It’s collecting enough good data to make the policy worth deploying.

Ceja, one of Encord’s pilots, used to work in waste management and helped keep a robotic trash sorter running. Now he spends his days toppling Jenga towers and training manipulator systems. That feels oddly fitting. Robotics still depends on people doing tedious, careful work so machines can eventually avoid it.

The machines are still far from that. The data pipeline is where the real action is.

Keep going from here

Useful next reads and implementation paths

If this topic connects to a real workflow, these links give you the service path, a proof point, and related articles worth reading next.

Relevant service
AI agents development

Design controlled AI systems that reason over tools, environments, and operational constraints.

Related proof
Field service mobile platform

How field workflows improved throughput and dispatch coordination.

Related article
CES 2026 puts physical AI, robotics, and edge silicon at the center

CES 2026 made one point very clearly: AI demos have moved past chatbots and image generators. This year, the loudest signal was physical AI. Robots, autonomous machines, sensor-heavy appliances, warehouse systems, and a lot of silicon built to run pe...

Related article
Periodic Labs raises $300M seed to build autonomous scientific labs

Periodic Labs has raised a $300 million seed round to build autonomous labs that can design experiments, run them with robotics, measure results, and use that data to plan the next round. For a seed round, that number is wild. The team helps explain ...

Related article
Nvidia’s Cosmos push is really a robotics and physical AI stack

Nvidia’s latest Cosmos release can look like another model announcement if you skim it. It’s really a stack play. The pieces are familiar on their own: a new 7B-parameter Cosmos Reason world model, a Cosmos Transfer-2 synthetic data system, neural 3D...