MEGA Hub

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Authors

Do you know Yunlong Lin?You can claim authorship or link another user.Do you know Zixu Lin?You can claim authorship or link another user.Do you know Zhaohu Xing?You can claim authorship or link another user.Do you know Biqiang Li?You can claim authorship or link another user.Do you know Chenxin Li?You can claim authorship or link another user.Do you know Haonan Wang?You can claim authorship or link another user.Do you know Haitao Wu?You can claim authorship or link another user.Do you know Hengyu Liu?You can claim authorship or link another user.Do you know Jianghai Chen?You can claim authorship or link another user.Do you know Kaituo Feng?You can claim authorship or link another user.Do you know Kaixin Li?You can claim authorship or link another user.Do you know Shawn Chen?You can claim authorship or link another user.Do you know Shijue Huang?You can claim authorship or link another user.Do you know Sixiang Chen?You can claim authorship or link another user.Do you know Tsung-Yi Ho?You can claim authorship or link another user.Do you know Wenxuan Huang?You can claim authorship or link another user.Do you know Xiangyan Liu?You can claim authorship or link another user.Do you know Xiaomeng Hu?You can claim authorship or link another user.Do you know Xuanhua He?You can claim authorship or link another user.Do you know Yan Sun?You can claim authorship or link another user.Do you know Yunqing Zhao?You can claim authorship or link another user.Do you know Zhiqin Yang?You can claim authorship or link another user.Do you know Zehan Wang?You can claim authorship or link another user.Do you know Zhengyang Tang?You can claim authorship or link another user.Do you know Tianyu Pang?You can claim authorship or link another user.Do you know Xiangyu Yue?You can claim authorship or link another user.

Abstract

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.

Community

00