Writer’s Palmyra X6 Bets Your Real Problem Is the Bill, Not the Benchmark

The AI industry spent two years chasing benchmark crowns. Writer’s new flagship is built around a less glamorous number: the invoice.

What Writer released

On Thursday 13 August 2026, Writer, which sells AI tools and agents aimed at marketing and enterprise teams, launched a new flagship model called Palmyra X6. The twist is what it is built from. Palmyra X6 is a post-training variation on Z.ai’s open-source model GLM-5.2, tuned by Writer for deployment rather than leaderboard glory. Alongside it, the company released a serious upgrade to its agentic harness. Both are available to Writer clients from the same day.

Writer’s estimate: the new model plus the harness changes make its agent product run at roughly 52% lower cost, with a 48% speed gain and a 10% quality bump.

The harness is the lever

The word harness is doing a lot of work here, so let us unpack it. The harness is the scaffolding around a model: how it plans, how it calls tools, how many tokens it burns getting to an answer. Writer’s argument, backed by a paper from its own researchers, is that tuning the harness is often a more reliable way to cut costs than swapping models. Their testing found the optimised harness cut the blended cost per task by 41%, from 21 cents to 12 cents, largely by cutting tokens per task 38%, from 14,200 to 8,800. As they put it, the harness is the one component whose efficiency multiplies across every model you run, present and future.

For complex, multi-step jobs that compounding matters. Fewer tokens per step, across dozens of steps, is where the bill actually lives.

Read the asterisk

Now the honest part, because a big round cost-cut percentage is a marketing number wearing a lab coat. Those figures are Writer’s own, and your mileage on gnarly multi-step work will differ. This is also an enterprise product. Palmyra X6 is not an API you will casually poke at over the weekend. It is aimed at companies already deploying Writer, where it sits beside Writer’s other models or outside models imported through Azure or Amazon Bedrock. The experience stays model-agnostic, which is the genuinely useful part. For reference, X6 is priced at $2 per million input tokens and $8 per million output.

The bigger tell

CEO May Habib was blunt about the mood. She told TechCrunch the enterprise is absolutely sick of chasing the next benchmark and wants flattening cost that nobody seems able to deliver. She framed the push to cut token use as a growing distrust of the big labs, who make more money the more tokens you burn. That is a pointed thing for a vendor to say out loud, and it lines up with what plenty of CIOs have been muttering all year.

The takeaway

Whether or not the 52% holds for your workload, the direction is the story. Open-weight base model, tuned for deployment, wrapped in a harness optimised to waste fewer tokens. If 2025 was about who had the smartest model, a good chunk of 2026 is turning into who can run a decent one cheaply. That is a healthier fight for buyers.

Did you know: Palmyra X6 descends from GLM-5.2, an open model from the Chinese lab Z.ai, a neat reminder that open weights are increasingly the raw material behind the enterprise AI you pay for.

Sources

Related on Top Tool Stack: Run Local Models Inside JetBrains: GitHub Copilot Now Speaks Ollama · AI on Reddit This Week: Astra’s Cyber Red Line, Open Video Goes Feral, and the Great Model Deflation · AI on a Budget: cut your AI costs

The free stack. One email a week: the AI tools and moves that actually matter, hype filtered out. Subscribe free →
Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top