
A local assistant should be able to write at a useful speed, work through a long document and serve more than one person. Our latest Paiton release brings Qwen3.8 Flash Next to two Radeon AI PRO R9700 cards, with 216.3 tokens per second in the reported weighted decode score, 565.1 tokens per second across eight concurrent requests, and a separate 200,000-token context mode.1
The hardware is two cards with 32 GB of VRAM each. The large routed experts use 3-bit weights, while more sensitive parts retain higher precision. The model runs through regular vLLM 0.29, the Paiton plugin and native AMD kernels, not a separate serving engine.23
There are important qualifications. The 200K mode generates at about 100 tokens/s rather than the default mode's roughly 200, because speculative decoding is off. A large n-gram embedding table remains in system RAM. Quality checks cover agreement with the BF16 model and bounded long-context retrieval, not broad math, coding, knowledge, multilingual or tool-calling benchmarks. Here is what we measured, what the numbers mean and how to try it.
Fast generation on two workstation cards
Decode measures generation after the first token. In the release-image validation run on 10 October, BetterBench 0.6.0's standard profile reported a weighted score of 216.3 tokens/s. Individual workloads ranged from 183.3 for prose to 259.2 for JSON.1
Release-image validation, two R9700 cards, 98,304-token configured context and speculative decoding enabled. Generated tokens/s; higher is better. The weighted score is not an equal-weight average of the eight rows; the workload weights are explained below. These are measurements of this model and configuration, not a comparison with our earlier 27B releases.
| Workload | Generated tokens/s |
|---|---|
| Chat | 202.1 |
| Code | 214.3 |
| File edit | 236.0 |
| JSON | 259.2 |
| Math | 255.5 |
| Prose | 183.3 |
| Reasoning | 190.1 |
| Summarization | 239.8 |
| Reported weighted score | 216.3 |
The short-prompt median time to first token was 117 ms. The stream-update p99 was 13.9 ms, but that is the gap between updates, not a per-token latency: speculation can deliver several accepted tokens in one update.1
More output when requests overlap
Aggregate generation throughput across simultaneous requests; higher is better. BetterBench ran 48 requests at each concurrency level. These are combined server rates, not the speed each user receives. The single-request concurrency result is a different workload from the weighted decode score above.
| Concurrent requests | Aggregate tokens/s |
|---|---|
| 1 | 204.6 |
| 2 | 315.1 |
| 4 | 449.2 |
| 8 | 565.1 |
All 48 of 48 requests completed at the eight-request level. That is useful for a shared local assistant, but eight overlapping benchmark requests do not mean eight full-length 98K conversations fit at once. Prompts, generated output and per-request state all use the available memory.1
Long prompts without a steep throughput drop
Prefill is processing the input before generation begins. The separate long-context mode, with speculation disabled, processed the tested prompt depths at roughly 7,900 to 8,300 tokens/s.1
Input tokens processed per second; higher is better. BetterBench standard profile with an extended 128K sweep, two R9700 cards, 200,000-token configured context, speculation off and 2,048-token prefill chunks. Depth labels are nominal settings, not exact prompt-token counts. These results come from the long mode, not the speculative decode configuration.
| BetterBench prompt-depth setting | Input tokens/s |
|---|---|
| 2K | 7,925 |
| 8K | 8,321 |
| 16K | 8,311 |
| 32K | 8,232 |
| 64K | 8,161 |
| 128K | 7,928 |
The 128K result remains close to the 8K result. A separate cold-cache probe with exactly 64,000 prompt tokens measured 8,212 tokens/s across the whole prompt, or 8,204 in steady state after the first chunk. Prompt throughput is not the same as the complete wait for an answer: a fresh 190K prompt took 24.8 seconds to its first token.1
Choose the mode that matches the job
| Mode | Configured context window | Speculative decoding |
|---|---|---|
decode, default | 98,304 tokens | On, depth 3 |
prefill-long | 200,000 tokens | Off |
These are total context windows: the prompt, chat formatting and generated answer must fit together. The checkpoint's native window is 262,144 tokens, but this release does not establish that full window on these two cards.32
The 200K mode needs the memory otherwise used by the speculative drafter's per-request state. Without speculation, single-stream generation after 32K, 100K and 190K prompts measured 105.3, 100.6 and 99.9 tokens/s, respectively, using 256-token continuations. Long context therefore remains usable, with a different speed tradeoff from the default mode.1
Reuse a long prompt with optional prefix caching
If you repeatedly ask questions about the same document, processing the unchanged prefix again can dominate the wait. The opt-in prefix cache stores attention-cache and recurrent-state checkpoints at aligned 2,048-token boundaries.2
Separate repeated-prompt checks in long-context mode. Time to first token in seconds; lower is better. A cache hit reuses an unchanged prefix still present in the cache. These are not the cold-cache BetterBench measurements above.
| Repeated prompt | Cold time to first token | Prefix-cache hit |
|---|---|---|
| 64K | 9.6 s | 0.35 s |
| 128K | 20 s | 0.41 s |
In these checks, the hit produced byte-identical output to the uncached run. That is a result for the tested cached and uncached requests, not a promise of identical output across all sampling settings. Prefix caching is optional, and the headline throughput benchmarks were measured with it off.1
What 3-bit means here
This is a mixed-precision model, not a claim that every tensor uses three bits.
The routed experts, which contain 120.8 billion weights, use 3.125 bits per weight including group scales. Those experts occupy about 47.2 GB in total. A rotated input basis helps the low-precision representation, and the native kernels read the packed weights directly rather than first expanding the entire model to a larger format. The W3A8 name refers to the 3-bit expert weights and 8-bit expert activations.3
The main non-expert projections use 8-bit weights. The speculative MTP layer uses 4-bit expert weights and a 2-bit draft head, with other tensors retaining higher precision. Routers, normalization and the retained vision tower also have their own precision choices. The public model card records the tensor formats; the validated serving modes are text-only, even though the repository contains the vision tower.32
The weights come from our own sequential GPTQ calibration on two million tokens of prose, code and assistant conversations. We tested two calibration replicates and separated selection from confirmation data. Their confirmation KL results, 0.0630 and 0.0686, differed enough that small gaps between candidate formats should not be overinterpreted.3
Where the memory goes
The large decoder and expert weights remain on the GPUs, split across two tensor-parallel ranks. Each card holds 24.9 GiB of text weights plus 0.7 GiB for the MTP layer. Both benchmark configurations peaked at 30.5 to 31.5 GiB per card, including cache, recurrent state and execution graphs.3
There is no expert offload, but this is not a GPU-only setup. The 48.9 GiB n-gram table stays pinned in system RAM, split across the ranks, and the runtime reads 16 of its rows per token. The test host had 251 GiB of system RAM. The checkpoint repository occupies 108 GiB on disk, before allowing room for the container and runtime cache.32
Quality: close distributions are not task benchmark scores
We evaluated how much the quantized model's full-vocabulary next-token distribution differs from BF16. KL divergence is lower-is-better; zero would mean identical distributions. Top-1 agreement is how often both versions prefer the same next token. It is not the percentage of math problems or coding tasks solved.3
Text and assistant positions on the same evaluation corpus. Free routing lets each model choose experts; locked routing forces BF16's expert choices to separate arithmetic error from routing changes. KL and top-1 agreement measure different things and have different units. The BF16 control shows variation between two correct implementations.
| Evaluation | KL to BF16, free routing | KL, BF16 routing forced | Top-1 agreement |
|---|---|---|---|
| BF16 against BF16, implementation floor | 0.0077 | Not applicable | 96.77% |
| This 3-bit model | 0.0604 | 0.0413 | 90.74% |
The corpus contains 294,912 positions, including 170,884 text and assistant positions used for the headline metrics. Separately, the release container passed its runtime served-KL gate at 0.0616, against a stated budget of 0.0600 plus 0.002. That gate uses top-64 bucket KL on the confirmation half, so it is not the same full-vocabulary metric as the table above. It validates the running server rather than only an emulation.34
Long documents and bounded retrieval
A held-out set of 44 documents from 8K to 32K tokens, totaling 524,288 positions, provides a separate view of error with document depth. The document-level text-and-assistant result is 0.0561 KL with free routing, 0.0408 with locked routing and 90.39% top-1 agreement.3
| Position in document | KL, free routing | KL, routing locked | Text/assistant top-1 agreement |
|---|---|---|---|
| Below 2K | 0.1554 | 0.1043 | 89.46% |
| 2K–4K | 0.6300 | 0.2931 | 89.66% |
| 4K–8K | 0.6408 | 0.2984 | 90.06% |
| 8K and beyond | 0.4333 | 0.2150 | 91.65% |
The bucket KL values above cover all positions, while their top-1 values cover text and assistant positions. They should not be compared directly with the text-and-assistant headline KL. The final bucket contains 12 documents; the other buckets contain 44.
On text and assistant positions, document-level negative log-likelihood was 1.7574 nats/token for BF16 and 1.7671 for this model, a paired difference of +0.0097. This is a small measured loss on that held-out set, not proof of identical task quality.3
| Served-stack retrieval check | This model | BF16 control |
|---|---|---|
| Needle retrieval at 131,072 tokens | 40 / 40 | 40 / 40 |
| Needle retrieval at 200,000 tokens | 40 / 40 | 40 / 40 |
Both versions returned the same answers in these tests. The drafter also remained close in a teacher-forced greedy check at depth 3: 1.768 expected accepted drafts per step, against 1.789 for the BF16 MTP layer on the BF16 model. Neither the retrieval checks nor draft acceptance establish general reasoning quality.3
Math, code, general knowledge, multilingual quality and tool calling were not task-benchmarked for this release. Image input is not validated. Test the workload you intend to deploy rather than treating distribution agreement as a substitute for that evaluation.
The runtime, without a custom engine
vLLM supplies the serving framework and OpenAI-compatible API. Paiton's native kernels handle the packed experts, tensor-parallel execution and model-specific operations. Speculative decoding drafts three tokens and verifies them together; the runtime also handles the recurrent layers' state when a draft is rejected. The release run accepted a median of 2.85 tokens per stream update.23
A faster prefill path, or reproducible arithmetic
The default fast prefill path uses 8-bit weights and 8-bit activations in the prompt-processing trunk. That changes the arithmetic precision on prompt rows. It passed the served-KL budget and the retrieval checks above, but the long-document control found it not bit-reproducible between runs.21
The publicly labelled exact Gated DeltaNet prefill path remains available; it retains the other quantization choices rather than turning all prompt arithmetic into BF16. A development-stack measurement returned 7,607 tokens/s at 64K, versus 8,161 for the release image's fast path. The exact path was bit-reproducible in its checks; it was not rerun on the release image, so this is not a same-image benchmark comparison. Decode arithmetic is unchanged by that choice.1
| Prompt-depth setting | Exact GDN prefill, development stack |
|---|---|
| 2K | 7,575 tokens/s |
| 8K | 7,785 tokens/s |
| 16K | 7,762 tokens/s |
| 32K | 7,711 tokens/s |
| 64K | 7,607 tokens/s |
| 128K | 7,422 tokens/s |
How we measured
The main tables use BetterBench 0.6.0, standard profile, against the release image through its own launcher and OpenAI-compatible endpoint. Sampling used temperature 0.7, top-p 0.95 and top-k 20. Each decode category had 20 timed passes after three warmups; concurrency used 48 requests at each of 1, 2, 4 and 8 streams; prefill used eight runs per depth and 2,048-token chunks.1
The weighted decode score assigns 30% to code, 20% to reasoning, 15% each to prose and JSON, and 10% each to file editing and summarization. Chat and math are measured separately and carry no weight in that score.5
Every benchmark request carried a unique nonce so its prefix was cold. Decode, short-prompt latency and concurrency use the 98,304-token speculative mode. Prefill uses the 200,000-token non-speculative mode. The two GPUs exchanged uncompressed BF16 activations. The host was quiet, with no builds or uploads running, and startup took roughly three to five minutes, mainly for weight loading, with compile caches included in the image.12
These measurements describe our host, prompts and release configuration. They do not measure whole-system energy, cost per token or a speedup over a different model. The public benchmark report preserves the configuration and detailed results.
Try it on two R9700 cards
Use a Linux host with two Radeon AI PRO R9700 cards, a ROCm 10 host driver, Python 3, the Hugging Face CLI and Docker with access to /dev/kfd and /dev/dri. Allow substantial system RAM beyond the 48.9 GiB n-gram table and disk space beyond the 108 GiB checkpoint. Our test host's 251 GiB is a measured configuration, not a stated minimum requirement.2
The commands below pin the public launcher and explicitly download the complete model revision, including the runtime drafter. This matters: the pinned launcher's automatic download still points at an older model revision. Download the revision below first and pass its directory with --weights, rather than relying on that automatic download. If you already have the plugin repository, use a separate checkout instead of replacing local work.23
git clone https://github.com/Eliovp-BV/paiton-vllm-plugin.git
cd paiton-vllm-plugin
git checkout be0f1a53bd60ab1bf77131121f9cacdcf46d2863
cd models/Qwen3.8-Flash-Next
export PAITON_FLASHNEXT_DIR="$PWD/model-cache/qwen38-flash-next-w3a8"
hf download EliovpAI/Qwen3.8-Flash-Next-W3A8-Paiton-RDNA4 \
--revision 829b089bf6636af9ffed1f333b103e4e383f7b48 \
--local-dir "$PAITON_FLASHNEXT_DIR"
(cd "$PAITON_FLASHNEXT_DIR" && sha256sum -c SHA256SUMS)
python3 launch-flashnext.py --weights "$PAITON_FLASHNEXT_DIR" --mode decode
The launcher selects ghcr.io/eliovp/paiton-vllm-plugin:qwen38-flashnext-rocm10-vllm029-20261010-r1, pinned by digest in the runtime lock. The model-specific setup guide documents the Python launcher modes. Stop the running server before selecting another mode:
# 200K context, speculation off
python3 launch-flashnext.py --weights "$PAITON_FLASHNEXT_DIR" --mode prefill-long
# Exact Gated DeltaNet prefill, default context
python3 launch-flashnext.py --weights "$PAITON_FLASHNEXT_DIR" --mode decode-nopf
# 200K mode with optional prefix caching
python3 launch-flashnext.py --weights "$PAITON_FLASHNEXT_DIR" \
--mode prefill-long --prefix-caching
Run one mode at a time. prefill-long-nopf selects the exact Gated DeltaNet prefill path at 200K; --dry-run prints the Docker command without starting the server. Both released base modes use a BF16 attention cache. Once ready, the endpoint is http://127.0.0.1:18982/v1, serving the model name Qwen3.8-Flash-Next.2
In another terminal, send a streaming request:
curl --fail http://127.0.0.1:18982/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen3.8-Flash-Next","messages":[{"role":"user","content":"Write a short Python function that removes duplicates while preserving order."}],"temperature":0.7,"max_tokens":256,"stream":true}'
To reproduce the performance workload, use BetterBench 0.6.0 and its standard profile against that endpoint, with the release report's extended prefill sweep. Keep the decode and prefill mode results separate, as in the tables above.
What comes next
The public roadmap includes RAM and SSD cache tiers and a higher-precision 4-bit build. Image input also needs its own validation. These are next steps, not features promised by the two validated text modes.2
For now, the result is a larger local model with useful generation speed, a measured 200K option and explicit quality limits, on two workstation GPUs. Regular vLLM, our plugin, our kernels.
Credits and licensing
Qwen3.8 Flash Next is by the Qwen team. vLLM provides the serving framework, BetterBench the performance workload suite, and Paiton the quantized weights and native execution in this release.
The upstream checkpoint uses Qwen Community License 1.0, not Apache-2.0. It includes commercial-use conditions; review the upstream license at the checkpoint revision before deployment. The public plugin adapter, model weights, packaged runtime and proprietary compiler are distinct artifacts with distinct terms. This post does not imply unrestricted commercial reuse.
Explore Paiton, read the model card and evaluation details, or contact us to discuss a local AMD AI workload.
Sources
- 10 October release-image benchmarks, pinned checkout. 2 3 4 5 6 7 8 9 10 11 12
- Model-specific release and setup guide, pinned checkout. 2 3 4 5 6 7 8 9 10 11 12
- Qwen3.8 Flash Next W3A8 model card, pinned revision. 2 3 4 5 6 7 8 9 10 11 12 13 14
- Public quality report, pinned model revision.
- BetterBench 0.6.0 workload weights, pinned defaults.
