#UAE #worldmodels - Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) has unveiled WorldGuide, a video world model that generates procedural tasks such as folding origami, assembling furniture or cooking. The model uses its own generated visuals to choose the next action and decide when to stop, as closed-loop execution. MBZUAI also released WorldGuide Bench to evaluate task execution and video quality of video world models. Using the new benchmark, WorldGuide achieved 33.33 percent task success against 29.90 percent recorded for MiniMax-H3.
SO WHAT? - Most video generators fix their instructions before they start, so a step that includes errors can’t be corrected. The WorldGuide model checks what it has actually produced before choosing what comes next. The margin in performance over MiniMax-H3 is small, and a 33% success rate shows how hard long procedures remain. However, this research proposes a new method: training planning, execution and stopping, then judging progress from the model’s own output. Thus, WorldGuide helps advance intelligent video production.
KEY POINTS:
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) has unveiled WorldGuide, a video world model for goal-directed procedural tasks. Starting from an image and a task goal, the model repeatedly selects an action, generates a short clip and uses the result to guide the next step.
Three components make up the WorldGuide system:
A ContextPlanner that picks the next action or signals completion;
An Executor that renders each action; and
Hierarchical visual memory that maintains context over long sequences.
The ContextPlanner is built on Qwen2.5-VL-7B and the Executor on HunyuanVideo-1.5. The planner was trained first and frozen, and its action embeddings then guided Executor training on the same step-level demonstrations.
Researchers also developed a new benchmark for video world models: WorldGuide Bench. The benchmark contains about 59,000 procedural videos across 245 tasks and 27 categories. The videos are segmented into atomic action clips with explicit completion signals.
WorldGuide scored 33.33% task success on the benchmark, against 29.90% for MiniMax-H3, which was given reference action plans.
Closed-loop execution improved task success by 21.62% over an open-loop version. Visual feedback added 18.61% to task success and cut repeated or skipped steps by 4.46%.
Visual memory improved results by 5.77% on WorldGuide Bench. Older history is stored at progressively coarser resolution, which limits memory costs as sequences grow.
MBZUAI notes that execution and stopping errors remain. Existing approaches such as ORCA, CollabVR and SPIRAL pair planners with frozen executors or separate critics, which is the gap the researchers say WorldGuide addresses.
WorldGuide achieved 47.69% task success on Video-CraftBench compared with 32.73% for MiniMax-H3, while WorldGuide visual memory improved results by 17.91% on Video-CraftBench.
The MBZUAI research team includes: Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, and Salman Khan.
[Written and edited with the assistance of AI]
Source: MBZUAI
LINKS
WorldGuide landing page (GitHub)
WorldGuide research paper (arXiv)
WorldGuide code (GitHub)
WorldGuide checkpoint (Hugging Face)
Data set (coming soon)
Read about more MBZUAI research:
Causal AI for autonomous biomedical discovery (Middle East AI News)
MBZUAI’s IFM releases world’s largest fully open AI model (Middle East AI News)
MBZUAI builds first cultural benchmark for Arabic AI (Middle East AI News)
UAE lab breaks the speed barrier in AI video generation (Middle East AI News)
MBZUAI researchers advance medical AI, cut training costs (Middle East AI News




