Meta’s Muse Spark Is the Best AI at Using Your Computer. Almost Nobody Noticed

3 min read

The most capable AI agent of the month is not from OpenAI or Anthropic, and it arrived with far less noise than it deserved. Meta’s Muse Spark 1.1, out of its Superintelligence Labs, has climbed to the top of the benchmarks that actually measure whether a model can do multi-step work rather than merely talk about it. Mark Zuckerberg came back to X after three years just to post about it. And the headline capability is the one that matters most for real automation: it can use your computer.

What it can actually do

Muse Spark 1.1 takes text, images, video, audio and PDFs in and produces text out, with a 1-million-token context window that it manages itself, compacting long sessions so it keeps the steps that matter instead of forgetting what it was doing. It can operate a desktop application, a browser and a mobile interface, and it can dispatch parallel subagents, a lead model handing work to specialist workers at once. On the benchmarks that measure tool use rather than trivia, it leads the field: 88.1 on MCP Atlas for scaled tool use, and top marks on JobBench and Finance Agent V2 for multi-step professional tasks.

Why “computer use” is the whole game

A model that can drive a desktop app, a browser and a phone can automate the work that previously needed a human clicking through screens, which is the actual bottleneck in most office automation. Answering questions well is table stakes now; finishing a real job across three apps without hand-holding is the frontier. That Meta is topping tool-use benchmarks rather than knowledge quizzes is the tell that it is aiming at work people would genuinely pay to hand off.

The strategy shift underneath

There is a plot twist worth noting. Meta built its whole AI reputation on open weights, on giving models away. Muse Spark arrives instead as a paid Model API, a real turn away from the open-only stance that defined the company. Whether you cheer or mourn that depends on your politics, but it signals Meta has decided the money in agents is in the service, not the giveaway.

The regulatory wrinkle

One more thing, because we covered it this week: Meta is not part of the White House framework that would have the government review frontier models before release. So the lab shipping arguably the strongest computer-use agent going is, for now, the one operating outside the review the other three big labs have accepted. Deliberate positioning or oversight, it is an odd gap, given that “an AI that can operate any computer” is exactly the capability a national-security reviewer would want a look at.

If you are building agents, this is the model to evaluate alongside the ones getting the headlines. It is not the flashiest release of the month. It may well be the most useful.

The AI tool stack actually worth paying for

One email a week. The tools, models and moves that matter, minus the hype and the horseshit filter set to maximum. Free.

Get the free stack →

Did you know: the whole industry has spent 2026 pivoting from “chatbots that answer” to “agents that do”, and the benchmarks moved with it. Older leaderboards rewarded knowing facts; the ones that now matter, like MCP Atlas and JobBench, reward completing a real multi-step task across tools. A model can ace the old tests and flunk the new ones, which is why the leaderboard you cite tells you what someone is actually optimising for.

Sources

Share this: X  ·  LinkedIn  ·  Facebook

Edgar Friendly

Top Tool Stack’s resident cynic, filtering the hype out of AI, tech, quantum and investing. More from Edgar →

Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top