Skip to main content
AI model cost reduction: a person at a computer, charts on screen, symbolizing 52% cost savings.

Editorial illustration for Writer says its new AI model cuts costs by 52% amid token spending surge

Writer's AI Model Cuts Enterprise Costs by 52%

4 min read

Writer put a number on the problem everyone in enterprise AI has been dancing around: token spending. The company's new flagship model, Palmyra X6, launched today alongside a rebuilt agent orchestration system and governance tools aimed at IT leaders watching their AI bills climb. Writer says the combination cuts agent costs by 52% on average, speeds things up by 48%, and improves output quality by 10%.

Those numbers matter less than how Writer got them. Palmyra X6 isn't a from-scratch build. It's a post-trained version of GLM-5.2, an open-weight model from Beijing-based Z.ai, formerly Zhipu AI.

Writer discloses this in its own technical report, which puts the San Francisco company squarely inside a fight the industry hasn't resolved: whether American businesses should build critical infrastructure on Chinese open-source foundations. Writer's leadership isn't dodging the question. Matan-Paul Shetrit, the company's director of product management, and Dan Bikel, who runs Writer's AI research, both addressed it directly in interviews with VentureBeat ahead of today's announcement.

The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6.

Why this matters

Writer is selling into a real pain point: enterprises watching token bills climb as agents chain more calls together, and a 52% cost cut paired with a 48% speed gain is the kind of number that gets a CFO's attention fast. But every figure here comes from Writer's own benchmarking, graded by Writer, on Writer's model. Bikel's answer about publishing methodology is a start, treating public benchmarks as "sanity checks rather than targets" is a reasonable framing, but it also means the headline stats aren't the ones anyone should actually trust yet.

For developers and technical buyers, the move to watch is whether Accenture, Uber, and Vanguard, or any third party, publish their own numbers once Palmyra X6 runs in production under real workloads. Governance tools for controlling token spending matter more than the marketing math here; runaway agent costs are a genuine operational problem, and tooling that gives IT leaders visibility and caps is worth testing regardless of whose benchmark you believe. Verify before you budget around it.

Common Questions Answered

What specific performance improvements does Writer's Palmyra X6 model claim to deliver?

Writer claims that Palmyra X6, when paired with their rebuilt agent orchestration system, delivers a 52% reduction in agent costs on average, a 48% improvement in speed, and a 10% improvement in output quality. These metrics represent significant gains for enterprises looking to optimize their AI spending and performance.

How did Writer achieve the cost reduction with Palmyra X6 compared to previous models?

Palmyra X6 was not built from scratch but rather represents an optimization of Writer's existing technology focused on reducing token spending. The model was specifically designed to address the growing problem of token consumption costs that enterprises face as AI agents chain multiple calls together.

What governance tools did Writer introduce alongside Palmyra X6?

Writer launched a rebuilt agent orchestration system and governance tools specifically aimed at IT leaders who are monitoring and managing climbing AI bills. These tools are designed to help enterprises control and optimize their AI spending while maintaining performance standards.

Why is the token spending problem significant for enterprise AI adoption?

Token spending has become a major pain point for enterprises as AI agents chain more function calls together, causing costs to escalate rapidly. A 52% cost reduction paired with a 48% speed improvement represents the kind of financial impact that attracts CFO attention and makes AI solutions more economically viable for organizations.

How does Writer validate the performance claims for Palmyra X6?

Writer's performance metrics come from the company's own benchmarking process, which means all figures are graded and tested by Writer on their own model. The company treats public benchmarks as "sanity checks rather than targets," indicating they use external validation as a reference point rather than the primary measure of success.

LIVE01:25GLM-5.3 Scores 66.9 on DeepSWE v1.1, Trails Behind GPT-5 and Claude