MEGA Hub

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

Authors

Do you know Haoran Sun?You can claim authorship or link another user.Do you know Wentao Zhang?You can claim authorship or link another user.Do you know Junyang Hua?You can claim authorship or link another user.Do you know Hedan Yang?You can claim authorship or link another user.Do you know Yongjian Guo?You can claim authorship or link another user.Do you know Yifei Zhang?You can claim authorship or link another user.Do you know Xiaolong Xiang?You can claim authorship or link another user.Do you know Mingxi Luo?You can claim authorship or link another user.Do you know Jing Long?You can claim authorship or link another user.Do you know Chen Zhao?You can claim authorship or link another user.Do you know Chen Zhou?You can claim authorship or link another user.Do you know Wanting Xu?You can claim authorship or link another user.Do you know Qiming Yang?You can claim authorship or link another user.Do you know Hui Zhang?You can claim authorship or link another user.Do you know Song Wang?You can claim authorship or link another user.Do you know Xiaodong Bai?You can claim authorship or link another user.Do you know Shuai Di?You can claim authorship or link another user.Do you know Xu Chu?You can claim authorship or link another user.Do you know Xiaotie Deng?You can claim authorship or link another user.Do you know Yicheng Gong?You can claim authorship or link another user.Do you know Junwu Xiong?You can claim authorship or link another user.

Abstract

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.

Community

00