Physical AI Development Platform
One Foundation Model. Across Embodiments. Across Tasks.
SYNWorld is a 10B-parameter-scale embodied world foundation model, trained on hundreds of millions of high-quality samples. With powerful generalization capabilities, it can rapidly adapt across diverse tasks and robotic embodiments with only a few dozen samples.
Embodied-Native World Foundation Model
High-quality training data spanning internet videos, embodied Ego data, and UMI data
Rapid cross-embodiment and cross-task adaptation with only a few dozen samples
Cross-Embodiment · Cross-Task · Multimodal · Rapid Adaptation
Unified Multimodal Modeling, Integrating Understanding and Generation
At its core, SYNWorld employs a unified multimodal foundation model that integrates vision, touch, force sensing, language, system states, actions, and other modalities. It jointly models environment understanding, future prediction, action generation, value assessment, and task progress, forming a closed-loop embodied intelligence framework of “Understand – Generate – Execute.”
Multimodal inputs—including multi-view images, tactile sensing, force sensing, and actions—are encoded independently and mapped into a unified representation space. MoT enables cross-modal interaction and collaborative modeling.
Scene understanding, world prediction, inverse dynamics, action generation, value assessment, and task progress prediction are unified within a single training framework. Through multi-task collaboration and capability transfer, SYNWorld achieves a more comprehensive understanding of environments, actions, and task objectives.
Natively supports both pinhole and fisheye cameras. Through 3D geometric alignment, SYNWorld integrates multiple viewpoints from head-mounted, end-effector, and other cameras to build a unified global spatial representation.
Data systems, model architectures, and AI infrastructure are co-designed to continuously scale data volume and model capabilities, enabling:
SYNAction builds on SYNWorld-WAM (World Action Model) as its pretrained foundation model. Through training data generation, post-training and policy optimization, as well as hierarchical evaluation and refinement, it continuously enhances policy model performance across diverse embodied tasks and scenarios.
Cleaned, segmented, and annotated multimodal and action data are organized into versioned training datasets to ensure full traceability and reproducibility.
Supports supervised fine-tuning (SFT), DAgger, offline/online reinforcement learning, as well as full-parameter fine-tuning, LoRA, Adapter, and other parameter-efficient tuning methods.
Analyzes training curves, evaluation results, and failure trajectories to identify performance bottlenecks and recommend data recipes and training configurations for the next iteration.
Evaluation-Driven Training · Full-Spectrum Evaluation Infrastructure
SYNEval integrates offline evaluation, simulation, and closed-loop real-world evaluation, leveraging evaluation results and reward feedback to support model selection, reinforcement learning, and release validation.
Hierarchical Evaluation Capability
Regression Testing and Version Comparison on Independent Evaluation Sets
Built on SYNWorld-ACWM (Action-Conditioned World Model), it uses policy model actions as conditions to generate future observations and predict task progress, reducing real-world testing costs while supporting iterative reinforcement learning.
For models that pass simulation pre-screening, large-scale real-world testing is conducted, leveraging VLM-based reward feedback to align real-world performance and optimize policies.