PHYSICAL AI TRAINING DATA

Japanese worksite data for physical AI development.

We record egocentric video and the tacit knowledge behind skilled work at sites in Japan where people are actually working, with a first focus on industry and infrastructure. Capture follows your model specification, at the volume you need.

Contact us

Director: Toshiyuki Yamamoto, founder of Chatwork

Reference image of egocentric footage of welding work
EGOCENTRIC VIEWTACIT KNOWLEDGEREAL FIELD
Capture reference

01 / PROBLEM

You have the model. The field data is missing.

01

Real-world data sets the ceiling

Simulation and public datasets take a policy only so far on real tasks. Data recorded in real environments is what decides how a model performs in the field.

02

Public datasets do not differentiate

When every team can train on the same corpus, it is hard to build an edge on it. Differentiation is moving toward data that only your team can access.

03

Opening up worksites is not R&D work

Securing consent to film, clearing rights, building trust with the people doing the work. The stage before collection is a different job from model development.

02 / WHY EGOCENTRIC

Why egocentric video, and why tacit knowledge.

01

Manipulation data does not exist on the web

Language models learned from web text, and image models learned from web images. Operational data from real work is not on the web: it has to be recorded one session at a time in the physical world. The quality and diversity of those demonstrations is what ultimately decides how imitation learning and VLA models behave on site.

02

The egocentric view is close to what a robot has to reproduce

A skilled worker's head-mounted view records the relationship between gaze, both hands, tool and workpiece as it happens. Occlusion, gaze shifts and fine hand detail, all lost to a fixed camera, stay in the recording. And because no robot hardware is involved, capture can go where the work actually happens instead of into a lab.

03

Skill does not show up on video alone

Why the worker stopped there. What they looked at before calling it good or bad. Judgment is hard to learn from footage alone, so we interview the worker and attach their intent as timestamped annotation, letting behavior and decision criteria be learned together.

Public datasets such as Open X-Embodiment and DROID are collected largely in research settings. Egocentric data recorded at manufacturing and service sites in Japan is available only to a limited extent in public form.

03 / DATA

What we deliver

EGOCENTRIC VIDEO

Egocentric video

Head-mounted capture of a skilled worker: gaze, hands and tool operation. The world as the operator sees it, which is the view a robot has to reproduce.

TACIT KNOWLEDGE

Verbalized tacit knowledge

Why the worker paused here. What they looked at before deciding. We interview the worker, turn intent and decision criteria into annotation, and deliver it linked to the video.

CUSTOM COLLECTION

Collection to your specification

Camera setup, resolution, frame rate, synchronization, metadata. We design and collect against a written specification aligned with your model and training pipeline.

What the data specification fixes

Before collection starts, we fix the following in a written data specification. We can propose a setup based on how the data will be used: pre-training, fine-tuning or evaluation.

Head-mounted camera placement, count and mounting method; optional hand-level or fixed auxiliary views
Resolution, frame rate, field of view, and how stabilization is handled
Time synchronization across viewpoints; metadata fields such as environment, material, tooling and working conditions
Task segmentation granularity (process, task, motion), format for tacit-knowledge annotation (intent and decision criteria), pass/fail labels
File format, folder structure and delivery method; sample delivery, your review, then full collection
Exclusive or non-exclusive license, scope of use (training, evaluation, redistribution, derivatives), scope of consent from the site and the worker

We set licensing — exclusive or non-exclusive — and the scope of use per contract, to match your intended use.

04 / FIELD ACCESS

We source sites in Japan to match your specification.

We are not limited to a single industry. You tell us the work you want captured and the conditions it has to meet; we look for a matching site in Japan, clear the rights to film and to use the footage, and then record. That sourcing and clearing is the part we take on. We are glad to talk before the target task is settled.

FOCUS

Industry and infrastructure

This is where we are focusing first. Hand skill decides the outcome, and passing that skill on is an open problem across much of this work. We are starting from in-factory manufacturing work and designing capture around it.

Example tasks we expect to capture

Joint fit-upTack and full weldingWeld bead visual checkGougingGrinding and finishingJig setup

Construction and building services

Work where the procedure shifts with each individual object, such as installation and connection.

Equipment installationPipe connectionInterior finishingInspection and adjustment

Retail and hospitality

Work with parallel task flows, such as a restaurant kitchen and floor.

Prep and cookingPlatingServing flowCustomer serviceCleaning and reset

Household and domestic work

Work where conditions differ in every home and the sequence is planned on the spot, such as domestic help.

Tidying and storageCleaningLaundry and ironingFood prep

Agriculture

Work where season and weather change the conditions, and the next action is chosen by reading how the crop is growing.

Seedling raisingRice plantingWater managementCrop monitoringHarvesting

These are areas we expect to cover. They are not sites we own or operate. For each engagement we source a site in Japan that matches the specification at that time.

A site that has agreed to filming

A factory producing steel water pipes and formed fittings has agreed to filming of manual welding and gouging. We are preparing a sample collection there.

Not a mock research setup, but sites where people are actually working.

05 / PROCESS

From first conversation to delivery

  1. 01

    Discovery

    We go through the target tasks, how the data will be used, and the specification you have in mind. It is fine if the specification is not settled yet.

  2. 02

    Data specification

    We fix camera setup, annotation format and granularity, and delivery format in a written document.

  3. 03

    Sample collection

    We collect a small volume so you can check the content against the specification.

  4. 04

    POC collection

    We collect and deliver at volume against the agreed specification.

  5. 05

    Ongoing supply

    We keep supplying data while widening the range of sites and tasks covered.

06 / COMPLIANCE

Rights clearance is part of data quality.

Data with unclear rights cannot be used for training. We settle this in writing with the site before collection begins.

  • Written agreement with the site

    Filming and data-use rights are agreed in writing with the operating company before collection starts.

  • Consent from the worker

    We obtain consent from each worker on the use of their likeness and on the scope of data use.

  • Care for trade secrets

    We manage what appears on camera and can mask areas that need it.

  • NDA

    We can sign an NDA from the specification discussion onward.

07 / COMPANY

Company

Chuo-Sogo Inc.
January 29, 2026
Shion Seki
Toshiyuki Yamamoto (founder of Chatwork)
1-5-18-3 Shinoharakita, Kohoku-ku, Yokohama, Kanagawa 222-0026, Japan
Collection and supply of training data for physical AI / website production

08 / FAQ

Frequently asked questions

Camera position and count, resolution, frame rate, the target work, annotation format and granularity, and delivery format, agreed against a written specification. It is fine to come to us before the specification is settled.

It depends on what is collected, the volume, and the license model (exclusive or non-exclusive), so we quote individually once the data specification is fixed.

We deliver only data for which rights have been cleared in writing with both the operating company and the worker. The scope of use is set out in the contract.

We are not limited to a single industry. Our first focus is manufacturing work in industry and infrastructure, and we also expect to cover construction and building services, retail and hospitality, household work, and agriculture. Our role is to find a site in Japan that matches the specification for each engagement, clear the rights to film and to use the data, and record. If the work you need is not listed, tell us what you want captured and under what conditions.

Public datasets such as Open X-Embodiment and DROID are collected largely in research settings, and the same data is available to every other team. We record at sites in Japan where people are actually working, to your specification, and attach tacit-knowledge annotation drawn from interviews with the worker.

You can start with a small sample collection. The sequence is specification alignment, sample collection, your evaluation, then POC collection, so you can check the content as you go.

Task segmentation (process, task, motion), tacit-knowledge annotation (intent and decision criteria drawn from worker interviews), pass/fail labels, and other formats defined in the specification. We can also deliver in the format your existing pipeline uses.

Camera setup, resolution, frame rate, synchronization method and delivery format can be specified on your side. We can also collect using hardware and formats you provide.

09 / CONTACT

Contact

Get in touch about a data specification, a sample collection, or anything else you want to discuss. We reply within two business days.