Skip to main content
Perplexity's portable AI computer running Qwen 3.8 27B locally, showcasing zero token cost.

Editorial illustration for Perplexity's Portable AI Computer Runs Qwen 3.8 27B Locally at Zero Token Cost

Perplexity's AI Computer Runs Qwen Locally Free

4 min read

Perplexity started shipping Portable Computer this week, a local-first version of its agentic Computer platform built to run on NVIDIA's DGX Spark hardware. The pitch is simple: every task starts on the device, local models handle it for free, and only when a job needs live web access or heavier cloud reasoning does the system pause and ask permission before routing that single step to one of more than 15 cloud models.

That's a real shift from how most AI agent platforms bill and behave. Instead of metering every token through a cloud API, Perplexity packages the model, inference engine, agent harness, tool sandbox and app connectors into one system that runs on hardware the customer already owns or buys once. The catch is the entry price: a GB10-class machine or an RTX GPU with 24GB of VRAM, which puts this out of reach for casual users and squarely in the lap of enterprises, mid-market teams, and startups with money for workstations.

The industries most likely to care are the ones where sending data to the cloud is the actual problem, not the convenience. Before getting into what Perplexity built and why it drew the line where it did, here's how the company frames the tradeoff.

Perplexity has released Portable Computer, a local-first build of its agentic Computer platform that runs the agent harness, orchestrator, planner, tool router and post-trained models directly on NVIDIA DGX Spark. The local model, inference engine, tool sandbox and app connectors ship as one packaged system, every task begins on the device, and work handled by local models carries no per-token charge.

Why this matters Perplexity's move is less about a flashy on-device AI box and more about admitting where local models actually break. Advertising a 260K-token window while quietly capping working context at 100K, then building the entire harness around that ceiling, tells us more than any benchmark chart would. For developers and founders building agents, the real lesson is architectural: compact system prompts, on-demand skill loading, and CLI-style connectors instead of full MCP definitions are becoming necessary workarounds, not nice-to-haves, once you try to run serious agent work on local hardware.

The zero-token-cost pitch is real, but it's tied to NVIDIA's DGX Spark specifically, so anyone evaluating this needs to weigh hardware lock-in against savings on inference bills. The "orchestrator asks before escalating to the web or frontier models" pattern is worth watching closely, since it's the actual cost and privacy boundary in practice, not the marketing headline. We'd want to see the full 53-point benchmark set before trusting Qwen 3.8 27B's local performance claims at face value.

Common Questions Answered

What is Perplexity's Portable Computer and how does it differ from traditional AI agent platforms?

Perplexity's Portable Computer is a local-first version of its agentic Computer platform that runs on NVIDIA's DGX Spark hardware, with tasks beginning on the device using local models at zero token cost. Unlike most AI agent platforms that route requests to cloud services, Portable Computer only asks for permission to access cloud models when a task requires live web access or heavier cloud reasoning for a specific step.

What model does Perplexity's Portable Computer run locally and what components are included?

Perplexity's Portable Computer runs the Qwen 3.8 27B model locally on NVIDIA DGX Spark hardware. The packaged system includes the agent harness, orchestrator, planner, tool router, post-trained models, local inference engine, tool sandbox, and app connectors all shipped as one integrated system.

How does Perplexity's billing model work with Portable Computer compared to cloud-based AI agents?

Work handled by Portable Computer's local models carries no per-token charge, making it cost-effective for routine tasks. The system only incurs costs when tasks are routed to one of more than 15 cloud models for operations requiring live web access or advanced cloud reasoning capabilities.

What architectural lessons does Perplexity's Portable Computer reveal about building AI agents?

The platform demonstrates that successful agent architecture requires compact system prompts, on-demand skill loading, and CLI-style connectors instead of full MCP definitions. This design approach acknowledges the practical limitations of local models, such as working context windows being significantly smaller than advertised specifications, and builds the entire system around realistic operational constraints.

LIVE22:36Sotheby's Data Scientist Builds Algorithms to Predict Art Prices