Skip to main content
NVIDIA Jetson module with Nemotron 3.5 Lightning & Qwen3.8-27B logos, signifying edge AI advancements.

Editorial illustration for NVIDIA Jetson Gains Nemotron 3.5 Lightning, Qwen3.8-27B for Edge AI

NVIDIA Jetson Runs Reasoning Models at the Edge

NVIDIA Jetson Gains Nemotron 3.5 Lightning, Qwen3.8-27B for Edge AI

4 min read

Six months ago, a robot working in a mine shaft or a truck cab out of cell range couldn't run a model smart enough to reason through a multi-step problem on its own. The hardware existed. The models didn't fit. Anything with real reasoning power needed a data center behind it, which meant a network connection, a round-trip delay, and data leaving the device whether you wanted it to or not.

That's changed fast. A run of compact open model releases this summer has pushed reasoning and agentic capabilities that used to require server racks down onto edge hardware, and NVIDIA Jetson can run them now. Nemotron 3.5 Lightning and Qwen3.8-27B are two of the models making that possible, small enough to deploy locally but built to handle the kind of multi-step decision-making that in-cab assistants, anomaly detection systems, and field robots actually need.

Getting there on Jetson isn't just a matter of downloading a checkpoint. Picking the right architecture, applying quantization and decoding tricks to squeeze out performance, and serving the model correctly all matter. Here's what to know before deploying.

Several model families released throughout the summer have collectively marked a turning point for edge AI. This new generation of compact open models now delivers reasoning and agentic capabilities that required large data center systems only a few months ago, and NVIDIA Jetson can run them today.

Why this matters

For developers who've been forced to choose between capable models and local deployment, this closes that gap on Jetson AGX Orin and Thor. That's a real shift, not a marketing footnote: multi-step reasoning running on-device means agentic AI applications in robotics, industrial inspection, or offline field tools no longer need a round trip to a data center for every decision. Cutting that network dependency matters for cost, latency, and for teams handling data that legally or contractually can't leave the device.

We'd temper the enthusiasm on one point. Quantized checkpoints and optimized runtimes get you deployable models, not free performance. Anyone shipping Nemotron 3.5 Lightning or Qwen3.8-27B on Jetson still has to validate accuracy after quantization and benchmark against their actual workload, not NVIDIA's demo numbers. The "turning point" framing is fair given how fast these model families have moved this summer, but the real test is whether founders building on this stack see reasoning quality hold up once it's running on a box in a warehouse instead of a GPU cluster.

Common Questions Answered

What reasoning capabilities can NVIDIA Jetson now run with Nemotron 3.5 Lightning and Qwen3.8-27B?

NVIDIA Jetson can now run compact open models that deliver multi-step reasoning and agentic capabilities directly on edge devices, eliminating the need for data center connectivity. These models enable on-device AI applications to perform complex reasoning tasks that previously required cloud-based systems, representing a significant shift in edge AI capabilities.

How does deploying reasoning models on Jetson AGX Orin and Thor improve latency compared to cloud-based approaches?

By running models locally on Jetson devices, developers eliminate round-trip delays to data centers for every decision, dramatically reducing latency for time-sensitive applications. This on-device processing is particularly beneficial for robotics, industrial inspection, and offline field tools that cannot rely on consistent network connectivity.

What was the primary limitation for edge AI applications six months before this release?

Six months ago, hardware capable of running edge AI existed, but compact models with real reasoning power were unavailable for on-device deployment. Any model with advanced reasoning capabilities required a data center backend, forcing developers to choose between capable models and local deployment on edge devices.

Why does running agentic AI applications on Jetson matter for teams handling sensitive data?

Deploying models locally on Jetson devices eliminates the need to send data to remote servers for processing, addressing legal and privacy concerns for teams handling sensitive information. This on-device reasoning capability allows organizations to maintain data sovereignty while still leveraging advanced AI capabilities for autonomous decision-making.

LIVE14:21Researchers Propose "AI Psychosis" Diagnosis to Speed Treatment