中文 | EN
OctoMind OctoH-Hand OctoSense About Us Join Us
OctoMind
OctoMind

Physical AI Development Platform

One Universal Brain, Powering Intelligence Across the Physical World

SYNWorld | Embodied-Native General-Purpose Foundation Model

One Foundation Model. Across Embodiments. Across Tasks.

SYNWorld is a 10B-parameter-scale embodied world foundation model, trained on hundreds of millions of high-quality samples. With powerful generalization capabilities, it can rapidly adapt across diverse tasks and robotic embodiments with only a few dozen samples.

10B+ Parameters

Embodied-Native World Foundation Model

100M+ Training Samples

High-quality training data spanning internet videos, embodied Ego data, and UMI data

Dozens of Samples

Rapid cross-embodiment and cross-task adaptation with only a few dozen samples

SYNWorld

Cross-Embodiment · Cross-Task · Multimodal · Rapid Adaptation

SYNWorld | Core Foundation Model Technologies

Unified Multimodal Modeling, Integrating Understanding and Generation

At its core, SYNWorld employs a unified multimodal foundation model that integrates vision, touch, force sensing, language, system states, actions, and other modalities. It jointly models environment understanding, future prediction, action generation, value assessment, and task progress, forming a closed-loop embodied intelligence framework of “Understand – Generate – Execute.”

SYNWorld
1

Unified Multimodal Architecture

Multimodal inputs—including multi-view images, tactile sensing, force sensing, and actions—are encoded independently and mapped into a unified representation space. MoT enables cross-modal interaction and collaborative modeling.

2

Multi-Task Joint Training

Scene understanding, world prediction, inverse dynamics, action generation, value assessment, and task progress prediction are unified within a single training framework. Through multi-task collaboration and capability transfer, SYNWorld achieves a more comprehensive understanding of environments, actions, and task objectives.

3

Multi-View Spatial Understanding

Natively supports both pinhole and fisheye cameras. Through 3D geometric alignment, SYNWorld integrates multiple viewpoints from head-mounted, end-effector, and other cameras to build a unified global spatial representation.

4

Integrated Data, Model & Infra Design

Data systems, model architectures, and AI infrastructure are co-designed to continuously scale data volume and model capabilities, enabling:

  • SYNAction: Built on SYNWorld-WAM as the pretrained foundation, SYNAction provides post-training capabilities for embodied AI policy models.
  • SYNEval: Built with SYNWorld-ACWM as one of its core capability foundations, SYNEval provides infrastructure for comprehensive, full-spectrum evaluation.
SYNWorld Architecture

SYNAction | Embodied Policy Post-Training

SYNAction builds on SYNWorld-WAM (World Action Model) as its pretrained foundation model. Through training data generation, post-training and policy optimization, as well as hierarchical evaluation and refinement, it continuously enhances policy model performance across diverse embodied tasks and scenarios.

Data Recipe

Cleaned, segmented, and annotated multimodal and action data are organized into versioned training datasets to ensure full traceability and reproducibility.

Post-Training & Policy Optimization

Supports supervised fine-tuning (SFT), DAgger, offline/online reinforcement learning, as well as full-parameter fine-tuning, LoRA, Adapter, and other parameter-efficient tuning methods.

AI-Powered Automated Tuning

Analyzes training curves, evaluation results, and failure trajectories to identify performance bottlenecks and recommend data recipes and training configurations for the next iteration.

SYNAction Pipeline Steps SYNAction Pipeline 500+ Participants, Including Leading Universities and Technology Companies Worldwide

SYNEval

Evaluation-Driven Training · Full-Spectrum Evaluation Infrastructure

SYNEval integrates offline evaluation, simulation, and closed-loop real-world evaluation, leveraging evaluation results and reward feedback to support model selection, reinforcement learning, and release validation.

VLM Expert Evaluation System

VLM evaluates and scores the physical plausibility of generated outputs and execution policies, assesses task completion, and generates reward signals.

SYNEval Evaluation

Hierarchical Evaluation Capability

1

Offline Open-Loop Evaluation

Regression Testing and Version Comparison on Independent Evaluation Sets

2

Closed-Loop Simulation Evaluation

Built on SYNWorld-ACWM (Action-Conditioned World Model), it uses policy model actions as conditions to generate future observations and predict task progress, reducing real-world testing costs while supporting iterative reinforcement learning.

3

Closed-Loop Real-World Evaluation

For models that pass simulation pre-screening, large-scale real-world testing is conducted, leveraging VLM-based reward feedback to align real-world performance and optimize policies.

One Universal Brain, Powering Intelligence Across the Physical World