Blog
Engineering

Nemotron on GB10: FP8 experts that would not load, now 64% faster than BF16

TensorRT-LLM refused FP8 block-scale checkpoints with non-gated experts on GB10, the chip inside DGX Spark. Our pull request makes them run: 64.3% more throughput and 44.9% less weight memory than BF16, with accuracy unchanged.

Nemotron on GB10: FP8 experts that would not load, now 64% faster than BF16
8 Oct 2026 · 7 min read

TensorRT-LLM PR #19806 lets FP8 block-scale mixture-of-experts checkpoints with non-gated experts run on GB10. On Nemotron 3.5 Lightning it delivers 64.3% more output throughput and 44.9% smaller model weights than the BF16 model, with MMLU unchanged. A TensorRT-LLM reviewer approved it on 8 October 2026, and it is waiting to merge. It comes from Dorian Magasic on our team.

Why GB10 matters

GB10 is NVIDIA's Grace Blackwell superchip for the desk. It puts a 20-core Arm CPU and a Blackwell GPU in one package, sharing 128 GB of LPDDR5x memory, with up to a petaflop of FP4 compute. NVIDIA ships it as DGX Spark, and the same silicon is in the ASUS Ascent GX10, the Dell Pro Max with GB10, the HP ZGX Nano and the Lenovo ThinkStation PGX. NVIDIA positions these machines for running and fine-tuning models of up to around 200 billion parameters locally.

On a machine like this, memory decides what fits and how fast it runs. Every gigabyte of weights saved is room for a larger model, a longer context or more concurrent users, and less data to move for each token. Lower-precision formats such as FP8 are how you get there, but only if the inference stack supports them on this chip, for this kind of model.

The model

Nemotron 3.5 Lightning 30B-A3B is one of NVIDIA's own open models: a hybrid of Mamba-2, mixture-of-experts and attention layers, with 30 billion parameters of which about 3 billion are active per token. It is built for agentic work, where many small calls run back to back.

Its experts are non-gated. Most mixture-of-experts models use a gated activation, two projections multiplied together. The Nemotron-H family uses a single projection followed by a squared ReLU. That difference sounds small, but it changes the shapes the inference engine has to load and compute.

The problem

On TensorRT-LLM release 1.3.0rc28, an FP8 block-scale checkpoint with non-gated experts did not load on GB10 (compute capability SM121). The server stopped with CutlassFusedMoE: sm_unsupported. Three things stood in the way:

  • SM121 was missing from the FP8 block-scale support entry.
  • The FP8 block-scale weight loaders assumed gated experts, with two projections per expert.
  • Nemotron's expert intermediate size, 1,856, is not a multiple of the 128-row blocks that FP8 block scaling uses.

So the cheaper format existed, the model existed and the hardware existed, but they did not meet.

The fix

The pull request changes the MoE path in three places:

  • Dispatch. SM121 joins the FP8 block-scale support entry, and SM120 and SM121 route to the Triton FP8 block-scale path.
  • Loading. In the FP8 block-scale MoE method only, non-gated experts load one projection and one block scale. With MoE tensor parallelism of one, an unaligned intermediate size is padded up to a multiple of 128, and the padded tail is zeroed. A strict shape check guards the block-scale grids.
  • Compute. The Triton kernel detects non-gated experts from the weight shapes and applies the squared ReLU. Backends that only handle gated experts now reject non-gated ones with a clear error, instead of failing somewhere deeper.

The shared loaders used by every other quantisation path are untouched, so aligned and gated models load exactly as before.

The receipts

Measured on GB10 with TensorRT-LLM 1.3.0rc28, trtllm-serve at tensor parallelism 1, using NVIDIA's AIPerf 0.13.0 on 512 input and 512 output tokens with two repeats per concurrency. Accuracy is MMLU five-shot over all 14,042 questions.

BF16, stockFP8, this PRChange
Output throughput, concurrency 32175.4 tok/s288.2 tok/s+64.3%
Time per output token, concurrency 457.7 ms39.0 ms−32.3%
Model weights58.82 GB32.41 GB−44.9%
MMLU five-shot0.77970.7787McNemar p = 0.42

The unit tests show no regressions over 2,528 test IDs, and the pull request adds tests for the padded and non-gated cases.

Two honest notes on reading the table. The throughput and time-per-token rows come from different concurrency levels, so they describe two operating points, not one. And the comparison is FP8 against BF16: the gain comes from making the FP8 path available on this chip, not from a faster kernel at the same precision.

What it does not cover

Padding is limited to MoE tensor parallelism of one; unaligned sizes across several GPUs still raise an error. The SM90 and SM100 paths were not tested, because we did not have that hardware.

GB10 so far

This is our second change for GB10 in TensorRT-LLM. PR #18313, still open, adds INT4 weight-only support to the fused MoE path: on Nemotron-3-Nano-30B-A3B it more than doubles output throughput (199.8 to 414.7 tokens per second) and cuts weight memory by 69.7%. Before that, PR #15550, merged in August, opened the INT8 weight-only path to non-gated models and gave Nemotron 40% more throughput on an A100.

The pattern is the same each time. The fast path already exists in the engine, and something narrow keeps a whole class of models off it. Finding those places, and proving the change keeps the model's answers the same, is what Artemis is built for.

More blogs

Discover the ROI hiding in your stack.

Point Artemis at a system you already run, and see the improvement it finds, validated, before you change a thing.