Meta Details Muse Spark 1.2 Multimodal Reasoning, Robotics and WildArtifactBench
Meta AI Research published new Muse Spark 1.2 demonstrations and evaluation details spanning visual coding, robotics, audio-video understanding and a preview of WildArtifactBench.
Meta expands the picture of Muse Spark 1.2
Meta AI Research published a new technical overview on August 20, 2026 describing the multimodal capabilities of Muse Spark 1.2, a coding-focused update to Muse Spark 1.1. The post emphasizes that the model's largest multimodal gains appear when it can use tools rather than only inspect visual inputs passively.
Visual coding and physical action
Meta demonstrates Muse Spark 1.2 translating images and video into working digital artifacts, then re-examining what it generated as part of an iterative improvement loop. The research post also describes a robotics setup in which a specialized Muse Spark variant acts as a high-level planner: it interprets a goal, reasons about a scene, breaks the task into subtasks and coordinates lower-level vision-language-action policies. Meta presents these robotics examples as demonstrations of spatial reasoning and orchestration, not as proof of general-purpose physical autonomy.
Audio-video workflows and agent evaluation
The model combines video understanding and dense captioning with tools such as web development, search and spatial grounding. Meta also previews WildArtifactBench, an internal evaluation designed for real-world agentic tasks where deliverables can be difficult to score with a single ground-truth answer. The company says the benchmark can compare generated artifacts using agentic or human judges and is releasing 10 preview tasks.
Availability and how to interpret the claims
Muse Spark 1.2 is available through Meta Model API and Muse Code. The performance and capability statements in Meta's post are first-party research claims, so developers should inspect the linked evaluation report and test the model on their own workflows before drawing broad conclusions about production reliability. The useful signal is the direction of travel: coding models are increasingly being designed as multimodal tool-using agents that move from perception to software creation and, in constrained settings, physical-task planning.
This article is built from the source material below. Open the originals for full context and the latest updates.