Japanese worksite data for physical AI development.
Manipulation data does not exist on the web. We record egocentric video and the tacit knowledge behind skilled work at sites in Japan where people are actually working, with a first focus on manufacturing, to your model specification.
A sample request is one email. It is fine to come to us before the specification is settled.
We are building toward 1,000 recorded hours per week.
Director: Toshiyuki Yamamoto, founder of Chatwork
ILLUSTRATIVERECORDING RIG
How a session is recorded
Capture is designed so the worker does not change how they work. Hardware and recording formats can be specified on your side.
IN 01
Head-mounted camera | egocentric view
Records gaze, both hands, tool and workpiece as they relate in real time.
IN 02
Audio
In-task narration and post-task interviews become the source of tacit-knowledge annotation.
IN 03
Auxiliary views (optional)
Hand-level or fixed cameras added per specification.
SUBJECT
Worker (doing their usual job)
OUT
Delivered data
- Egocentric video
- Tacit-knowledge annotation (timestamped)
- Metadata (environment, material, tooling, conditions)
- Delivery format fixed in the specification
DATA
What we deliver
EGOCENTRIC VIDEO
Egocentric video
Head-mounted capture of a skilled worker: gaze, hands and tool operation. The world as the operator sees it, which is the view a robot has to reproduce.
TACIT KNOWLEDGE
Verbalized tacit knowledge
Why the worker paused here. What they looked at before deciding. We interview the worker, turn intent and decision criteria into annotation, and deliver it linked to the video.
CUSTOM COLLECTION
Collection to your specification
Camera setup, resolution, frame rate, synchronization, metadata. We design and collect against a written specification aligned with your model and training pipeline.
SPEC SHEET
What the data specification fixes
Before collection starts, we fix the following in a written data specification. We can propose a setup based on how the data will be used: pre-training, fine-tuning or evaluation.
- Viewpoint and rig
- Head-mounted camera placement, count and mounting method; optional hand-level or fixed auxiliary views
- Video
- Resolution, frame rate, field of view, and how stabilization is handled
- Sync and metadata
- Time synchronization across viewpoints; metadata fields such as environment, material, tooling and working conditions
- Annotation
- Task segmentation granularity (process, task, motion) and the format for tacit-knowledge annotation (intent and decision criteria)
- Delivery
- File format, folder structure and delivery method; sample delivery, your review, then full collection
- Rights
- Exclusive or non-exclusive license, scope of use (training, evaluation, redistribution, derivatives), scope of consent from the site and the worker
We set licensing — exclusive or non-exclusive — and the scope of use per contract, to match your intended use.
WHY EGOCENTRIC
Why egocentric video, and why tacit knowledge.
Manipulation data does not exist on the web
Language models learned from web text, and image models learned from web images. Operational data from real work is not on the web: it has to be recorded one session at a time in the physical world. The quality and diversity of those demonstrations is what ultimately decides how imitation learning and VLA models behave on site.
The egocentric view is close to what a robot has to reproduce
A skilled worker's head-mounted view records the relationship between gaze, both hands, tool and workpiece as it happens. Occlusion, gaze shifts and fine hand detail, all lost to a fixed camera, stay in the recording. And because no robot hardware is involved, capture can go where the work actually happens instead of into a lab.
Skill does not show up on video alone
Why the worker stopped there. What they looked at before calling it good or bad. Judgment is hard to learn from footage alone, so we interview the worker and attach their intent as timestamped annotation, letting behavior and decision criteria be learned together.
Public datasets such as Open X-Embodiment and DROID are collected largely in research settings. Egocentric data recorded at manufacturing and service sites in Japan is available only to a limited extent in public form.
FIELD ACCESS
We source sites in Japan to match your specification.
The center of our work is manufacturing. You tell us the work you want captured and the conditions it has to meet; we arrange a matching site in Japan, clear the rights to film and to use the footage, and then record. That sourcing and clearing is the part we take on. We are glad to talk before the target task is settled.
FOCUS
Manufacturing
This is where we are focusing first. Hand skill decides the outcome, and passing that skill on is an open problem across much of this work.
Factory processes available for capture
Construction
Work where the procedure shifts with each individual object, such as installation and connection. This is the area we take on next, and we welcome conversations about it.
Construction and other domains not listed are sourced per engagement, with a site in Japan matched to the specification at that time. They are not sites that Chuo-Sogo holds today.
A site that has agreed to filming
A factory producing steel water pipes and formed fittings has agreed to filming of manual welding and gouging. We are preparing a sample collection there.
PROCESS
From first conversation to delivery
- 01
Discovery
We go through the target tasks, how the data will be used, and the specification you have in mind. It is fine if the specification is not settled yet.
- 02
Data specification
We fix camera setup, annotation format and granularity, and delivery format in a written document.
- 03
Sample collection
We collect a small volume so you can check the content against the specification.
- 04
POC collection
We collect and deliver at volume against the agreed specification.
- 05
Ongoing supply
We keep supplying data while widening the range of sites and tasks covered.
Pricing is quoted individually once the data specification is fixed.
COMPLIANCE
Rights clearance is part of data quality.
Data with unclear rights cannot be used for training. We settle this in writing with the site before collection begins.
Written agreement with the site
Filming and data-use rights are agreed in writing with the operating company before collection starts.
Consent from the worker
We obtain consent from each worker on the use of their likeness and on the scope of data use.
Care for trade secrets
We manage what appears on camera and can mask areas that need it.
Storage and scope set by contract
Data storage, access and any third-party provision are set out in the contract.
NDA
We can sign an NDA from the specification discussion onward.
COMPANY
Company
- Company name
- Chuo-Sogo Inc.
- Founded
- January 29, 2026
- CEO
- Shion Seki
- Director
- Toshiyuki Yamamoto (founder of Chatwork)
- Location
- 1-5-18-3 Shinoharakita, Kohoku-ku, Yokohama, Kanagawa 222-0026, Japan
- Business
- Collection and supply of training data for physical AI
FAQ
Frequently asked questions
Yes — start with one email via "Request a sample" above, and we will set up a sample collection conversation. The sequence is specification alignment, sample collection, your evaluation, then POC collection, so you can check the content as you go.
Camera position and count, resolution, frame rate, the target work, annotation format and granularity, and delivery format, agreed against a written specification. It is fine to come to us before the specification is settled.
It depends on what is collected, the volume, and the license model (exclusive or non-exclusive), so we quote individually once the data specification is fixed.
We deliver only data for which rights have been cleared in writing with both the operating company and the worker. The scope of use is set out in the contract.
Manufacturing is our center. Factory processes such as bending, gas cutting, plate fabrication and welding, machining, industrial painting and woodworking are available for capture. Construction comes next. Our role is to arrange a site in Japan that matches the specification for each engagement, clear the rights to film and to use the data, and record. If the work you need is not listed, tell us what you want captured and under what conditions.
Public datasets such as Open X-Embodiment and DROID are collected largely in research settings, and the same data is available to every other team. We record at sites in Japan where people are actually working, to your specification, and attach tacit-knowledge annotation drawn from interviews with the worker.
Standard delivery is egocentric video and audio, tacit-knowledge annotation (timestamped intent and decision criteria drawn from worker interviews), and metadata. Machine annotation such as depth or pose estimation is not part of standard delivery — raise it in the specification discussion if you need it.
Camera setup, resolution, frame rate, synchronization method and delivery format can be specified on your side. We can also collect using hardware and formats you provide.
CONTACT
Contact
Get in touch about a data specification, a sample collection, or anything else you want to discuss. We reply within two business days.