
The LG CLOiD Home Robot is demonstrated for a zero labor home during the LG press conference ahead of the annual Consumer Electronics Show (CES) in Las Vegas, Nevada on January 5, 2026.
Patrick T. Fallon/AFP via Getty Images
Before most US newsrooms had opened on August 18, 2026, LG Electronics issued a press release confirming a commitment that has no close parallel in consumer electronics: the company and Nvidia had just agreed, in person on a Seoul construction site, to generate 100,000 hours of robot training data before the end of the year. The figure — equivalent to roughly 12 years of continuous robot operation — surpasses Ant Group Robbyant’s published 60,000-hour pre-training corpus for its LingBot-VLA 2.0 model, and makes LG one of the best-resourced physical AI data programs among major consumer electronics manufacturers.
The meeting that produced the commitment was brief — roughly two hours — and operational in character. Madison Huang, Nvidia’s Senior Director of Product Marketing for Omniverse and Robotics and the eldest daughter of Nvidia CEO Jensen Huang, met LG Electronics CEO Lyu Jae-cheol at LG’s under-construction DataFactory on the company’s Yangjae R&D campus in Seocho district, southern Seoul. The two reviewed the facility’s readiness, discussed data-collection strategy, and closed with a specific year-end target. Huang described the outcome as “Incredible,” according to Seoul Economic Daily reporter Kim Yoon-soo.
Five Days from Handshake to Construction Site
The August 18 meeting followed by just five days the formal memorandum of understanding that LG Group Chairman Koo Kwang-mo and Nvidia CEO Jensen Huang signed at Nvidia’s Santa Clara headquarters on August 13. That MOU converted a high-level cooperation framework — first announced publicly in June when the two companies described joint work on humanoid robots and AI infrastructure — into named programs with delivery dates. The pace of the follow-up underscores how urgently both companies are treating the physical AI data race. LG and Nvidia did not wait to schedule a working session: they were on the construction site inside a week.
The August 13 agreement covers four program tracks:
A next-generation bipedal humanoid robot built on Nvidia’s Isaac GR00T foundation model, with Jetson Thor compute modules and the Halos for Robotics safety system, is planned for a public unveiling in the first quarter of 2027. Hardware will come from across the LG group — LG Innotek contributing sensors, LG Energy Solution supplying batteries, LG Electronics supplying actuators — while Nvidia provides the AI brain.
Before that, LG plans to deploy its wheel-based CLOiD robot on the washing machine manufacturing line at LG Electronics’ Tennessee plant before the end of 2026, using LG CNS’s PhysicalWorks platform to run the data pipeline.
An AI factory reference site powered by Nvidia’s Vera Rubin architecture is targeted for the first half of 2027, followed by an 80-megawatt AI factory in Cheonan, South Korea, planned for completion in the first half of 2028.
A next-generation AI-defined vehicle computing platform on Nvidia DRIVE Hyperion will integrate LG’s infotainment and automotive software capabilities.
What the DataFactory Actually Does
LG’s DataFactory — the company’s preferred branding — is not a data center in the conventional sense. Spanning four floors (one basement level through the third floor above ground) and covering 10,000 square meters (approximately 108,000 square feet) on the Yangjae campus, it is designed as what LG calls a “robot bootcamp”: a controlled environment where robots perform structured tasks, sensor data is collected, and the resulting training corpus is iteratively refined.
The facility contains at least four training environments: a replicated home space where CLOiD robots practice cleaning tasks; a manufacturing simulation modeled on LG’s Tennessee washing machine plant, where robots learn to transport parts, stack them, and perform assembly operations; a logistics space built around LG CNS’s automation solutions; and a robotic-hand training area for LG Innotek. LG CNS leadership and LG Sciencepark leadership attended the August 18 meeting alongside the two CEOs, reflecting the cross-affiliate scope of the program.
By year-end, LG plans to deploy several hundred CLOiD units at the facility.
How the 100,000 Hours Get Made
Why is 100,000 hours achievable with several hundred robots in less than five months? The answer lies in how LG and Nvidia have structured the data pipeline — and it is the architecturally significant detail that the headline figure alone does not convey.
Robot training data is not like text or images. A vision-language-action (VLA) model — the architecture class that now dominates industrial robot AI, including Nvidia’s Isaac GR00T N1 and LG’s own Robot Foundation Model under development — requires time-aligned streams of camera images, joint positions, force-sensor readings, and motor commands captured while a robot actually performs a physical task. This data does not exist on the internet. Every useful hour has to be physically generated or synthetically derived from something that was.
Generating robot demonstration data through teleoperation — the highest-fidelity method, in which a human operator remotely drives a robot while every sensor stream records — now costs approximately $118 per hour in fully loaded costs, according to the Silicon Valley Robotics Center’s State of Robotics 2026 report, down from roughly $340 per hour in early 2024. At that rate, 100,000 hours of pure teleoperation would cost approximately $11.8 million in collection labor alone.
LG and Nvidia’s pipeline does not work that way. According to LG’s official press release, the 100,000-hour target will be met through a combination of “training data collected directly at the facility” and “data synthetically generated and augmented using NVIDIA Cosmos open world models.” Nvidia’s Cosmos is an open world foundation model designed to take real-world sensor recordings and generate physically plausible synthetic variations — different lighting conditions, object placements, material textures, and partial-failure scenarios — multiplying a single real demonstration into many training episodes. Research suggests that roughly eight synthetic samples, properly generated, deliver the training value of one real teleoperation sample for in-domain tasks; the ratio drops for contact-rich manipulation tasks, which is why real-world factory data remains indispensable.
This is LG’s structural advantage over programs that rely primarily on teleoperation: by anchoring on real manufacturing-environment data (contact dynamics, material variation, actual factory lighting and floor conditions that no simulator replicates perfectly) and then amplifying that real corpus through Cosmos augmentation, LG can reach 100,000 hours at a fraction of the pure-teleoperation cost. The real-world factory data LG has accumulated over decades of appliance manufacturing — sensor readings, assembly sequences, logistics operations from its subsidiaries — provides the foundation layer that a greenfield robotics startup cannot buy. As an LG spokesperson put it: the combination of that manufacturing legacy data with Nvidia’s robotics software stack “is expected to be the secret to building a world-class data cycle that sets it apart from other companies.”
The specific technical pipeline runs as follows: real task demonstrations at the DataFactory feed into Nvidia Omniverse libraries and Cosmos for augmentation and synthetic generation; the resulting dataset is processed through the Isaac robotics development platform for policy training and validation; the validated policies feed back into LG’s Robot Foundation Model, which improves over successive collection cycles — a data flywheel that continuously improves as better-trained robots generate cleaner, more informative demonstrations.
This approach directly addresses the governing technical constraint in robot AI: the sim-to-real gap. Policies trained purely in simulation frequently degrade when deployed on real hardware, because simulators cannot yet fully replicate contact dynamics and manipulation physics — friction variation, material deformation, the micro-adjustments a robot’s joints make when gripping an object under load. By grounding its corpus in real factory environments first and then amplifying through Cosmos, LG’s hybrid approach mitigates the gap more robustly than simulation-only pipelines.
How the Target Compares
The 100,000-hour figure is meaningful, but context matters for interpreting it. Not all robot training data is equivalent.
Ant Group’s Robbyant division published a 60,000-hour pre-training corpus for LingBot-VLA 2.0 — approximately 50,000 hours of robot interaction trajectories across 20 hardware configurations, plus 10,000 hours of egocentric human video — in July 2026. LG’s 100,000-hour target is 67% larger, though LG’s is focused on a single embodiment (CLOiD) rather than spread across 20 robot platforms, which makes its per-embodiment data density considerably higher.
Dyna Robotics claimed a 1-million-hour corpus for its DYNA-2 model. But Dyna’s corpus consists of first-person human video rather than robot demonstrations — a structurally different input that requires cross-embodiment transfer algorithms to translate human motion into robot action commands. LG’s 100,000 hours of combined real robot demonstration and Cosmos-augmented data is more directly usable for robot policy training without that translation overhead.
Google’s RT-1 robot model (2022) used 130,000 demonstration trajectories collected by 13 robots over 17 months — the foundational large-scale real-world robot dataset, now surpassed in volume by newer programs.
The competitive infrastructure context is broader than any single corpus. The ISF Voices 2026 analysis put it plainly: “Dedicated robotic data collection facilities will become critical bottlenecks in the physical AI supply chain… whoever establishes the trusted, well-managed pipelines will hold leverage comparable to a semiconductor fab.”
What the Timeline Requires
For both companies, the next five months are the period that determines whether the 100,000-hour commitment holds. The DataFactory is still under construction; full operations are scheduled for year-end 2026. The Tennessee CLOiD deployment on the washing machine line — a working factory validation, not a demo — must also occur before December 31. Both milestones are prerequisites for the humanoid robot program, which is on the clock for a Q1 2027 public unveiling.
“Through the synergy built on ‘One LG’ — bringing together core capabilities across the Group — and strategic collaboration with global partners, we will secure our competitiveness in physical AI and become a comprehensive robotics solutions provider,” said Lyu Jae-cheol, LG Electronics CEO, in the official August 18 announcement.
Whether the 100,000-hour target proves to be a real structural advantage or a first-mover benchmark that better-resourced competitors surpass within months will depend on what LG and Nvidia build with it. The data flywheel only works if the data quality holds — and the DataFactory’s construction timeline means that the bulk of that quality will be determined in the months ahead.
Frequently Asked QuestionsWhy does robot AI need so much training data when other AI learns from the internet?
A language model or image generator can train on text and photos that already exist online — billions of pages and images are freely available. A robot learning to grip objects, open doors, or move parts on an assembly line needs something that does not exist at internet scale: synchronized recordings of physical task performance, capturing camera images, joint positions, force-sensor readings, and motor commands in tight time alignment during real interactions. Simulated versions of these recordings exist but produce robots that frequently fail in real environments because simulators cannot fully replicate contact dynamics — the friction, material variation, and micro-corrections that real physical handling requires. This gap between simulated and real performance is called the sim-to-real gap, and filling it requires real-world data. For a deeper exploration of this constraint, see TechTimes’ embodied AI data shortage explainer.
How does Nvidia’s Cosmos model help LG reach 100,000 hours?
Nvidia Cosmos is an open world foundation model that takes real sensor recordings from robot demonstrations and generates physically plausible synthetic variations: different lighting, different object placements, different material textures, partial-failure scenarios. A single real demonstration — a CLOiD robot stacking washing machine parts under normal factory conditions — can be expanded into many synthetic training episodes with varied conditions, multiplying the data volume without requiring a human operator to drive the robot through each variation. This Cosmos-augmented synthetic data, combined with LG’s real factory recordings, is how the 100,000-hour target becomes achievable with several hundred robots before year-end, rather than requiring thousands of robots or years of collection.
How does LG’s 100,000-hour target compare to other robot AI programs?
The target is larger than Ant Group Robbyant’s published 60,000-hour pre-training corpus for LingBot-VLA 2.0 (July 2026) and meaningfully larger than Google’s RT-1 foundational dataset (approximately 130,000 demonstration trajectories collected over 17 months in 2022). Dyna Robotics claimed a 1-million-hour corpus for DYNA-2, but that corpus consists of first-person human video rather than robot demonstrations, which is a structurally different (and cheaper) input that requires additional cross-embodiment translation before it can train a robot policy. LG’s corpus — combining real factory demonstrations with Cosmos-augmented synthetic data — is more directly usable for training the company’s Robot Foundation Model without that translation step.
When will LG’s humanoid robot be unveiled?
Under the August 13 MOU with Nvidia, LG is planning a public unveiling of its next-generation bipedal humanoid robot in the first quarter of 2027. The robot is designed to run on Nvidia’s Isaac GR00T foundation model, with Jetson Thor compute modules and the Halos for Robotics safety system, while drawing hardware from across the LG group — actuators from LG Electronics, sensors from LG Innotek, and batteries from LG Energy Solution. Before the humanoid reaches a stage, LG is planning to validate its wheel-based CLOiD robot on a working washing machine production line at its Tennessee plant before the end of 2026, generating factory-floor training data that will inform the humanoid program. Full details of the MOU’s named program tracks are available in the Unite.AI coverage of the August 13 announcement.