Less than three months after emerging from stealth, XDOF, a startup that collects real-world teleoperation data for training general-purpose robots, is in late-stage talks to raise a Series B at a valuation of about $1.2 billion valuation led by 8VC, several people with knowledge of the deal said.
XDOF was co-founded by UC Berkeley researchers Philipp Wu (CEO) and Fred Shentu (CTO) in 2024. TechCrunch reported on the startup’s $70 million Series A in June, with participation from Thrive Capital, Andreessen Horowitz, Lux, and Spark Capital. XDOF wasn’t planning to raise again so soon after that round. But the company’s rapid growth — with annualized revenue approaching $50 million — prompted VCs to approach it about a new round, the people said.
TechCrunch was unable to learn the total capital being raised or whether the valuation includes the new funding. The terms of the deal are not final and could still change.
XDOF and 8VC didn’t respond to our request for comment.
The startup aims to build the data pipelines, collection tools, and annotation systems that frontier AI labs and robotics companies can’t easily build themselves, essentially acting as an outsourced data-supply chain for the robotics industry.
As a PhD student, Wu was studying how robots learn from large datasets. One big impediment to his research was the lack of “large-scale data to work with,” he told TechCrunch in June.
So he teamed up with Shentu on a project called GELLO, a low-cost teleoperation system that allows a human operator to control a robotic arm remotely in order to generate training data. Their work led to an influential paper in robotics.
That research formed the foundation for XDOF, which investors now describe as the Scale AI or Mercor for physical robotics, a reference to the data-labeling giants that helped fuel the AI boom. Unlike LLMs, which initially trained on the entirety of the internet, physical robots don’t have an equivalent real-world dataset to draw from, making data collection a critical bottleneck to building general-purpose machines.
XDOF is partnering with UC Berkeley’s AI Research lab to release what it believes is the largest collection of high-quality robot training data ever assembled, dubbed ABC.
To capture this data, XDOF combines remote robot teleoperation with human collectors who wear sensors to record everyday tasks like folding clothes and flattening boxes.
The startup plans to hire and train teams of data collectors worldwide, including teleoperators who steer robots remotely and egocentric operators who wear body sensors to capture movement data.
XDOF previously told TechCrunch that it is already working with 20 customers, including several frontier AI labs.
Other startups attempting to collect real-world data for robot training include Mecka AI, as well as human-data platforms expanding beyond LLMs, such as Scale AI and Micro1.
Facts Only
* XDOF is in late-stage talks to raise a Series B valuation of about $1.2 billion, led by 8VC.
* The company was co-founded in 2024 by Philipp Wu (CEO) and Fred Shentu (CTO).
* The company raised a $70 million Series A in June.
* Annualized revenue is approaching $50 million.
* XDOF aims to build data pipelines, collection tools, and annotation systems for the robotics industry.
* The startup is partnering with UC Berkeley’s AI Research lab to release a dataset called ABC.
* Data collection involves remote robot teleoperation and human collectors wearing sensors.
* XDOF is working with 20 customers, including several frontier AI labs.
* The work builds upon research by Wu and Shentu on low-cost teleoperation systems (GELLO).
Executive Summary
Full Take
The narrative positions the bottleneck in general-purpose robotics as the lack of real-world data, contrasting it with LLMs which rely on internet text rather than physical experience. This frames XDOF not merely as a data collection service but as an essential infrastructure layer for achieving embodied AI—a bridge between theoretical large models and physical agency. The creation of XDOF stems from an understanding that physical robots require experiential datasets analogous to the massive text corpora available to language models. The ambition to assemble the ABC dataset, leveraging both remote control and egocentric human sensing, suggests a move toward creating foundational physics-aware training data, which has significant implications for the generalization capabilities of future physical AI systems. The investment and partnership structure indicate that establishing this specialized data infrastructure is viewed as a crucial, high-leverage step in the AI ecosystem, potentially positioning XDOF as an essential supply chain entity rather than just a niche data provider.
BRIDGE QUESTIONS: What are the specific limitations of the ABC dataset's quality or scope compared to proprietary or synthetic physical datasets? How will this focus on real-world tasks constrain or enable the eventual development of general-purpose robot intelligence compared to purely language-based AI? What long-term governance and ethical frameworks are being established around the data collection practices involving global teleoperators and human collectors?
Sentinel — Human
The text reads like standard financial or technology journalism, focusing on presenting reported facts about a startup's development and funding strategy.
