Skip to main content
Team of Amazon engineers analyzing and optimizing Anthropic AI models to reduce token costs through distillation techniques i

Editorial illustration for Amazon engineers distill Anthropic models to lower costs before token pricing

Amazon Distills Anthropic Models to Cut AI Costs

Amazon engineers distill Anthropic models to lower costs before token pricing

Updated: 3 min read

Amazon is quietly gutting its AI partner for parts. The bill is coming due. According to The Information, starting next year, Amazon’s payments to Anthropic will pivot from compute hours to tokens processed—a shift that could wildly inflate costs.

So engineers are now distilling Claude’s models into cheaper, leaner versions. It’s a direct, preemptive hedge. The right to do this was baked into a renegotiated partnership deal, echoing Apple’s arrangement with Google’s Gemini.

Here’s the thick irony: Amazon’s official Bedrock cloud service offers model distillation, but it doesn’t support Anthropic’s models. Only Amazon’s own Nova and Meta’s Llama are allowed. So the Claude-stripping operation runs off-platform, a corporate skunkworks.

Amazon has certain rights to use Anthropic's models for this purpose, according to a person familiar with the matter, similar to Apple's arrangement with Google Gemini. Amazon does offer a distillation service on its Bedrock cloud platform, but Anthropic's Claude models aren't available there; only Amazon's own Nova models and Meta's Llama models are supported. The effort ties back to a renegotiation of the partnership, according to The Information.

Starting next year, Amazon will pay for Anthropic's models based on tokens processed rather than compute hours, which could push costs up sharply. An Amazon spokesperson pushed back, saying the changes from the expanded partnership won't raise costs. Anthropic points to lower prices relative to the performance its models deliver.

Amazon is reportedly exploring alternatives like OpenAI and its own Nova models.

Publicly, it’s all calm. An Amazon spokesperson insists the expanded partnership won’t raise costs. Anthropic talks up its performance-to-price ratio.

Don’t buy it. The real story is in Amazon’s reported exploration of OpenAI and its own Nova models. This isn’t mere budget-trimming.

It’s about control. Distillation lets Amazon preserve Claude’s core capability while surgically removing the expensive bulk. They want a ready-made, cheaper substitute in their pocket before the new token meter starts running.

The partnership is real, but the underlying tension is thicker. In this game, today’s supplier is tomorrow’s competitor. Amazon is methodically ensuring it won’t be locked into paying for someone else’s brain.

Common Questions Answered

Why are Amazon engineers distilling Anthropic models?

Amazon engineers are distilling Anthropic models to reduce computational and operational costs. Distillation creates smaller, more efficient models that maintain performance while using fewer resources, which is especially important before any changes in token pricing.

What does it mean to 'distill' an AI model in this context?

Distillation involves training a smaller 'student' model to replicate the behavior of a larger 'teacher' model, such as Anthropic's models. This process compresses the model's size and complexity, lowering inference costs and making it more economical for Amazon to deploy at scale.

How does distilling Anthropic models help Amazon lower costs before token pricing changes?

By distilling Anthropic models, Amazon engineers can achieve similar AI capabilities with significantly reduced computational overhead. This cost reduction is strategically important ahead of potential token pricing adjustments, as it allows Amazon to offer AI services more competitively and manage expenses effectively.

LIVE01:20Liquid AI's 3B Vision Model Shows Major Gains in Screen Reading, Object Grounding