We went looking for a faster model. The speedup was already in the checkpoint.

For three consecutive research cycles we scanned the open-weight frontier looking for a better model to serve on our GB10 nodes, and each time concluded the same thing: nothing new in the 20 to 120B class, stay where we are. That conclusion was correct and completely beside the point. The speedup we wanted was already sitting on disk, in the checkpoint we had been serving for months, behind a flag we never set. ...

September 2, 2026 · 7 min · Conselara Labs

Tokens per second is the wrong benchmark for picking a local model

Every local-inference writeup leads with tokens per second. It is the number on the vendor slide, the number in the forum post, and the number people compare when choosing what to run on their own hardware. It is also close to useless on its own, and we have the measurements to show why. We evaluated two models on one GB10 node (128 GB unified memory, SM121), same stock NVIDIA NGC vLLM container, same serve flags, same quantization class. One was Qwen3.6-35B-A3B, our current daily driver. The other was Ornith-1.5-35B-A3B, a continued-pretrain of that exact base, which we confirmed by diffing the configs: identical 40 layers, 256 experts with 8 active, 2048 hidden, same 248320 vocabulary, same GatedDeltaNet split of 30 linear and 10 full attention layers. ...

August 28, 2026 · 5 min · Conselara Labs

DGX Spark Benchmark Results: vLLM on SM121

Measured throughput and latency on DGX Spark GB10 (SM121) hardware. All results use vLLM 0.19.0 (NGC container nvcr.io/nvidia/vllm:26.04-py3) unless noted. Qwen3-235B-A22B-GPTQ-Int4: Two-node cluster Date: 2026-05-03 Config: TP=2, EP=2, Ray cluster over QSFP-DD RoCE direct interconnect, --attention-backend=TRITON_ATTN, --quantization=gptq_marlin, --kv-cache-dtype=fp8, --gpu-memory-utilization=0.87 Batch Avg completion tokens tok/s per request Aggregate tok/s 1 (serial) 256 17.0 17.0 2 (concurrent) 256 12.1 24.1 4 (concurrent) 256 9.1 36.4 Prefix cache: 97% delta hit rate on repeated system prompt. Startup to first inference: ~15 minutes (Ray init + weight load across two nodes + compile). Weight resident per node: 57.64 GiB. ...

May 9, 2026 · 2 min · Conselara Labs

DGX Spark Model Comparison: What Fits and What Runs (SM121, 128 GB)

Quick-reference comparison of open-weight models for a single DGX Spark GB10 (SM121, 128 GB unified LPDDR5X memory). Based on tested configurations and community results as of May 2026. Model Architecture Quantization Memory Expected tok/s SM121 notes Qwen3.6-35B-A3B GDN-hybrid MoE (qwen3_5_moe, 3B active) FP8 (~35 GB) ✅ easily 100+ GDN-hybrid MoE, low active params → fully supported on stock NGC Qwen3.6-27B Dense hybrid (GDN) FP8 (~28 GB) ✅ easily 14–21 (stock) / 136–200 (fork) GDN kernel gap; experimental fork needed for full speed Qwen3-30B-A3B Pure MoE (3.3B active) NVFP4 / FP8 / BF16 (~16–60 GB) ✅ easily 32–50 Solid single-node option; no GDN gpt-oss-120b Sparse MoE (5.1B active) mxfp4 (~61 GB) ✅ 32–60 128K context; proprietary quant format Qwen3.5-122B-A10B Pure MoE (10B active) NVFP4 only (~75 GB) ✅ up to 51 BF16 is 234 GB and does not fit; NVFP4 is the only path Qwen3-235B-A22B Pure MoE (22B active) GPTQ-Int4 (~60 GB/node) ✅ (two nodes) 17–36 agg Requires two DGX Sparks; best quality available Qwen3.5-397B-A17B Pure MoE (17B active) NVFP4 (TP=2) ✅ (two nodes) Unknown SM121 MoE kernel not yet optimized; not recommended Key observations Throughput vs quality tradeoff at single-node: Qwen3.6-35B-A3B gives the highest throughput (100+ tok/s) with pure MoE architecture. Qwen3.5-122B-A10B gives the most capable model (10B active parameters) that fits on one node, at 51 tok/s. For most agentic workloads the bottleneck is tool latency, not token generation, so 51 tok/s is more than sufficient. ...

May 9, 2026 · 2 min · Conselara Labs

vLLM Model Selection for DGX Spark (SM121)

The DGX Spark GB10 SoC (SM121) has specific constraints that determine which models run well and which don’t. This is a practical guide based on what we’ve tested in production. The key constraint: SM121 kernel compatibility Not all model architectures run well on SM121 with the NGC vLLM container. The main constraint is the MoE kernel: Marlin kernel: stable, fast, supports GPTQ-Int4 and mxfp4 CUTLASS FP4: broken on SM121, produces garbage outputs silently; never use GDN (GatedDeltaNet): kernel gap on SM121, 14–21 tok/s with stock NGC; requires experimental fork for full speed Prefer low-active-param MoE models when using the NGC container. A GDN-hybrid MoE like Qwen3.6-35B-A3B (qwen3_5_moe, ~3B active) runs fully through Marlin and is well-tested on SM121. The architecture to avoid is not GDN itself but dense GDN and specific unsupported model_types (see below). ...

May 9, 2026 · 4 min · Conselara Labs

Running Qwen3.5-122B on a Single DGX Spark

The NVIDIA DGX Spark (GB10 SoC, 128GB unified LPDDR5X memory) can run Qwen3.5-122B-A10B, a 122B parameter MoE model, at usable throughput for production workloads. Here’s what it actually takes. The key constraint: NVFP4 only Qwen3.5-122B-A10B at full precision is ~250GB. In NVFP4 quantization it’s ~75GB, which fits comfortably in 128GB unified memory. There is no other quantization path that both fits and runs correctly on the GB10. The only verified checkpoint we’ve found: bjk110/SPARK_Qwen3.5-122B-A10B-NVFP4 on HuggingFace, which includes 15 patches for the SM121 architecture. Use this; don’t try to quantize the base model yourself unless you’re prepared to debug SM121-specific kernel failures. ...

May 5, 2026 · 3 min · Conselara Labs