[{"data":1,"prerenderedAt":2254},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":1756},{"id":4,"title":5,"body":6,"categories":1737,"date":1742,"description":1743,"extension":1744,"heading":1745,"image":1746,"meta":1747,"navigation":1150,"originalUrl":1748,"path":1749,"seo":1750,"slug":1751,"socialImage":1752,"stem":1753,"updated":1754,"__hash__":1755},"blog\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700.md","Qwen3.8 27B on 1 × Radeon AI PRO R9700: 3-bit weights, 20% faster decode",{"type":7,"value":8,"toc":1716},"minimark",[9,18,45,60,65,68,83,95,99,103,115,306,309,314,317,322,395,398,402,405,410,515,533,617,621,624,629,677,690,697,700,705,766,779,783,787,799,802,806,809,828,832,839,848,852,859,874,878,881,888,892,895,900,984,997,1004,1033,1037,1047,1070,1079,1082,1086,1101,1401,1408,1416,1423,1441,1452,1488,1505,1509,1516,1520,1543,1546,1563,1712],[10,11,12,13,17],"p",{},"A local assistant is more useful when it writes faster, starts processing a long document sooner and has room for more people to work at once. Our new Paiton release brings those improvements to ",[14,15,16],"strong",{},"Qwen3.8 27B on 1 × Radeon AI PRO R9700 (32 GB)",", without adding another GPU.",[10,19,20,21,24,25,28,29,32,33],{},"Against our previous 24 September MXFP4 release, weighted generation speed rises from ",[14,22,23],{},"153.8 to 184.4 tokens per second: 19.9% faster decode",". Eight concurrent requests produce ",[14,26,27],{},"492.1 tokens per second in aggregate",", up from 425.3. Smaller model weights also leave room for ",[14,30,31],{},"43.5% more reported cache-token capacity"," in the 65K serving profile.",[34,35,36],"sup",{},[37,38,44],"a",{"href":39,"ariaDescribedBy":40,"dataFootnoteRef":42,"id":43},"#user-content-fn-model",[41],"footnote-label","","user-content-fnref-model","1",[10,46,47,48,51,52,55,56,59],{},"The change is our own rotated ",[14,49,50],{},"3-bit weights for the large decoder projections",", with ",[14,53,54],{},"4-bit activation math for selected layers during prompt processing",". Other tensors retain their existing formats. The tradeoff is real: the MMLU-Pro general-knowledge subset drops by about three percentage points. Math and code differences are not distinguishable from noise in these tests, and all 80 bounded long-context retrieval checks pass. If knowledge accuracy matters most, the ",[14,57,58],{},"MXFP4 option remains one switch away",".",[61,62,64],"h2",{"id":63},"what-changed-for-everyday-use","What changed for everyday use",[10,66,67],{},"The previous release already combined MXFP4 weights, an FP8 cache, DFlash2 speculative decoding and Paiton's native AMD execution inside vLLM. This update keeps that serving foundation and reduces the amount of model data the GPU needs to read. The result is faster output, faster prompt processing and additional cache space on the same card.",[10,69,70,71,74,75,78,79,82],{},"Those benefits are different measurements. ",[14,72,73],{},"Decode"," is generation after the first token. ",[14,76,77],{},"Aggregate throughput"," is output shared across simultaneous requests. ",[14,80,81],{},"Prefill"," is processing the prompt before generation begins. A higher token rate in one phase should not be presented as an identical reduction in total request time.",[10,84,85,86,89,90,94],{},"The headline 19.9% compares this release with the ",[14,87,88],{},"24 September MXFP4 image",", not the older Radiance comparison in our ",[37,91,93],{"href":92},"\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","earlier Qwen3.8 article",". That article's original 57% result remains a separate benchmark with its own configuration.",[61,96,98],{"id":97},"faster-generation-across-different-workloads","Faster generation, across different workloads",[100,101],"w3a4-figure",{"kind":102},"decode",[10,104,105],{},[106,107,108,109,59],"em",{},"BetterBench 0.6.0, quick preset. Generation after the first token, in tokens\u002Fs; higher is better. Values are means of two complete runs. The weighted result rises 19.9% against the previous release. Workload gains vary, so this is not a promise that every prompt generates 20% faster. ",[37,110,114],{"href":111,"rel":112},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fa44044105287da3d040652ea8bcd3927a94a1bf2\u002Fmodels\u002FQwen3.8-MXFP4-DFlash2\u002Fbenchmarks\u002F2026-09-26-w3a4\u002FREADME.md",[113],"nofollow","Published benchmark report",[116,117,118,141],"table",{},[119,120,121],"thead",{},[122,123,124,128,132,135,138],"tr",{},[125,126,127],"th",{},"Workload",[125,129,131],{"align":130},"right","Previous MXFP4 release",[125,133,134],{"align":130},"This round, MXFP4",[125,136,137],{"align":130},"W3A4 release",[125,139,140],{"align":130},"Change vs previous",[142,143,144,162,179,194,211,228,245,262,279],"tbody",{},[122,145,146,150,153,156,159],{},[147,148,149],"td",{},"Chat",[147,151,152],{"align":130},"121.2 tok\u002Fs",[147,154,155],{"align":130},"122.9 tok\u002Fs",[147,157,158],{"align":130},"135.8 tok\u002Fs",[147,160,161],{"align":130},"+12.0%",[122,163,164,167,170,173,176],{},[147,165,166],{},"Code",[147,168,169],{"align":130},"179.9 tok\u002Fs",[147,171,172],{"align":130},"182.7 tok\u002Fs",[147,174,175],{"align":130},"226.0 tok\u002Fs",[147,177,178],{"align":130},"+25.6%",[122,180,181,184,186,188,191],{},[147,182,183],{},"File edit",[147,185,169],{"align":130},[147,187,172],{"align":130},[147,189,190],{"align":130},"195.0 tok\u002Fs",[147,192,193],{"align":130},"+8.5%",[122,195,196,199,202,205,208],{},[147,197,198],{},"JSON",[147,200,201],{"align":130},"217.5 tok\u002Fs",[147,203,204],{"align":130},"220.8 tok\u002Fs",[147,206,207],{"align":130},"269.0 tok\u002Fs",[147,209,210],{"align":130},"+23.7%",[122,212,213,216,219,222,225],{},[147,214,215],{},"Math",[147,217,218],{"align":130},"183.8 tok\u002Fs",[147,220,221],{"align":130},"186.5 tok\u002Fs",[147,223,224],{"align":130},"228.9 tok\u002Fs",[147,226,227],{"align":130},"+24.5%",[122,229,230,233,236,239,242],{},[147,231,232],{},"Prose",[147,234,235],{"align":130},"78.5 tok\u002Fs",[147,237,238],{"align":130},"79.6 tok\u002Fs",[147,240,241],{"align":130},"94.7 tok\u002Fs",[147,243,244],{"align":130},"+20.7%",[122,246,247,250,253,256,259],{},[147,248,249],{},"Reasoning",[147,251,252],{"align":130},"117.8 tok\u002Fs",[147,254,255],{"align":130},"119.4 tok\u002Fs",[147,257,258],{"align":130},"133.5 tok\u002Fs",[147,260,261],{"align":130},"+13.3%",[122,263,264,267,270,273,276],{},[147,265,266],{},"Summarization",[147,268,269],{"align":130},"138.4 tok\u002Fs",[147,271,272],{"align":130},"140.5 tok\u002Fs",[147,274,275],{"align":130},"158.4 tok\u002Fs",[147,277,278],{"align":130},"+14.4%",[122,280,281,286,291,296,301],{},[147,282,283],{},[14,284,285],{},"Weighted",[147,287,288],{"align":130},[14,289,290],{},"153.8 tok\u002Fs",[147,292,293],{"align":130},[14,294,295],{},"156.1 tok\u002Fs",[147,297,298],{"align":130},[14,299,300],{},"184.4 tok\u002Fs",[147,302,303],{"align":130},[14,304,305],{},"+19.9%",[10,307,308],{},"The middle column separates the smaller runtime improvements available with MXFP4 from the larger gain of the new weight format. Weighted decode on this round's MXFP4 path is about 1.5% higher than the previous release.",[310,311,313],"h3",{"id":312},"more-output-when-requests-overlap","More output when requests overlap",[100,315],{"kind":316},"concurrency",[10,318,319],{},[106,320,321],{},"Aggregate generated tokens\u002Fs, with 48 requests per concurrency level; higher is better. These are combined server rates, not the speed each user receives. One R9700, 65,536-token configured context, up to eight scheduled sequences.",[116,323,324,338],{},[119,325,326],{},[122,327,328,331,333,335],{},[125,329,330],{},"Concurrent requests",[125,332,131],{"align":130},[125,334,137],{"align":130},[125,336,337],{"align":130},"Change",[142,339,340,353,367,381],{},[122,341,342,344,347,350],{},[147,343,44],{},[147,345,346],{"align":130},"122.0 tok\u002Fs",[147,348,349],{"align":130},"148.8 tok\u002Fs",[147,351,352],{"align":130},"+22.0%",[122,354,355,358,361,364],{},[147,356,357],{},"2",[147,359,360],{"align":130},"204.2 tok\u002Fs",[147,362,363],{"align":130},"249.3 tok\u002Fs",[147,365,366],{"align":130},"+22.1%",[122,368,369,372,375,378],{},[147,370,371],{},"4",[147,373,374],{"align":130},"308.2 tok\u002Fs",[147,376,377],{"align":130},"368.3 tok\u002Fs",[147,379,380],{"align":130},"+19.5%",[122,382,383,386,389,392],{},[147,384,385],{},"8",[147,387,388],{"align":130},"425.3 tok\u002Fs",[147,390,391],{"align":130},"492.1 tok\u002Fs",[147,393,394],{"align":130},"+15.7%",[10,396,397],{},"For a shared local assistant, this means the GPU can deliver more output while requests overlap. The exact improvement depends on the prompt lengths, output lengths and available cache.",[310,399,401],{"id":400},"prompt-processing-improves-too","Prompt processing improves too",[100,403],{"kind":404},"prefill",[10,406,407],{},[106,408,409],{},"Input tokens processed per second; higher is better. The depth labels are nominal BetterBench settings, not exact prompt lengths. The largest nominal 64K workload has a median actual prompt length of 47,016.5 tokens. Measured time to first token is shown separately below.",[116,411,412,428],{},[119,413,414],{},[122,415,416,419,422,424,426],{},[125,417,418],{},"Nominal depth",[125,420,421],{"align":130},"Median actual prompt tokens",[125,423,131],{"align":130},[125,425,137],{"align":130},[125,427,337],{"align":130},[142,429,430,447,464,481,498],{},[122,431,432,435,438,441,444],{},[147,433,434],{},"2K",[147,436,437],{"align":130},"1,516.5",[147,439,440],{"align":130},"3,689 tok\u002Fs",[147,442,443],{"align":130},"4,156 tok\u002Fs",[147,445,446],{"align":130},"+12.7%",[122,448,449,452,455,458,461],{},[147,450,451],{},"8K",[147,453,454],{"align":130},"5,894.5",[147,456,457],{"align":130},"3,834 tok\u002Fs",[147,459,460],{"align":130},"4,165 tok\u002Fs",[147,462,463],{"align":130},"+8.6%",[122,465,466,469,472,475,478],{},[147,467,468],{},"16K",[147,470,471],{"align":130},"11,802",[147,473,474],{"align":130},"3,871 tok\u002Fs",[147,476,477],{"align":130},"4,103 tok\u002Fs",[147,479,480],{"align":130},"+6.0%",[122,482,483,486,489,492,495],{},[147,484,485],{},"32K",[147,487,488],{"align":130},"23,549.5",[147,490,491],{"align":130},"3,751 tok\u002Fs",[147,493,494],{"align":130},"3,958 tok\u002Fs",[147,496,497],{"align":130},"+5.5%",[122,499,500,503,506,509,512],{},[147,501,502],{},"64K",[147,504,505],{"align":130},"47,016.5",[147,507,508],{"align":130},"3,455 tok\u002Fs",[147,510,511],{"align":130},"3,629 tok\u002Fs",[147,513,514],{"align":130},"+5.0%",[10,516,517,518,521,522,525,526],{},"The tested prefill rates improve by ",[14,519,520],{},"5.0–12.7%",". Client time to first token falls by ",[14,523,524],{},"4.8–10.8%",", comparing the mean of the two runs' median TTFTs. A throughput increase and a latency reduction are different percentages; scheduling and other overheads can also affect the wait.",[34,527,528],{},[37,529,357],{"href":530,"ariaDescribedBy":531,"dataFootnoteRef":42,"id":532},"#user-content-fn-bench",[41],"user-content-fnref-bench",[116,534,535,550],{},[119,536,537],{},[122,538,539,541,544,547],{},[125,540,418],{},[125,542,543],{"align":130},"Previous TTFT, mean of run medians",[125,545,546],{"align":130},"W3A4 TTFT, mean of run medians",[125,548,549],{"align":130},"Less waiting",[142,551,552,565,578,591,604],{},[122,553,554,556,559,562],{},[147,555,434],{},[147,557,558],{"align":130},"407.453 ms",[147,560,561],{"align":130},"363.395 ms",[147,563,564],{"align":130},"10.8%",[122,566,567,569,572,575],{},[147,568,451],{},[147,570,571],{"align":130},"1,548.047 ms",[147,573,574],{"align":130},"1,416.748 ms",[147,576,577],{"align":130},"8.5%",[122,579,580,582,585,588],{},[147,581,468],{},[147,583,584],{"align":130},"3,054.554 ms",[147,586,587],{"align":130},"2,877.005 ms",[147,589,590],{"align":130},"5.8%",[122,592,593,595,598,601],{},[147,594,485],{},[147,596,597],{"align":130},"6,270.297 ms",[147,599,600],{"align":130},"5,943.109 ms",[147,602,603],{"align":130},"5.2%",[122,605,606,608,611,614],{},[147,607,502],{},[147,609,610],{"align":130},"13,617.798 ms",[147,612,613],{"align":130},"12,962.238 ms",[147,615,616],{"align":130},"4.8%",[61,618,620],{"id":619},"more-room-for-long-conversations","More room for long conversations",[100,622],{"kind":623},"memory",[10,625,626],{},[106,627,628],{},"Model memory falls from 19.18 to 15.89 GiB. The launcher reallocates the freed memory to the conversation cache, increasing reported cache-token slots from 174,634 to 250,578. Cache capacity is an estimate from the serving configuration, not a count of fully measured conversations.",[116,630,631,642],{},[119,632,633],{},[122,634,635,638,640],{},[125,636,637],{},"Resource",[125,639,131],{"align":130},[125,641,137],{"align":130},[142,643,644,655,666],{},[122,645,646,649,652],{},[147,647,648],{},"Model memory",[147,650,651],{"align":130},"19.18 GiB",[147,653,654],{"align":130},"15.89 GiB",[122,656,657,660,663],{},[147,658,659],{},"Reported cache-token capacity",[147,661,662],{"align":130},"174,634",[147,664,665],{"align":130},"250,578",[122,667,668,671,674],{},[147,669,670],{},"Equivalent capacity at 65,536 tokens per sequence",[147,672,673],{"align":130},"About 2.7",[147,675,676],{"align":130},"About 3.8",[10,678,679,680,683,684],{},"The last row is a ",[14,681,682],{},"fractional capacity estimate",", not a claim that 3.8 requests can run. Prompt text, chat\u002Ftool formatting and generated output all consume context. Eight scheduled sequences also does not mean eight full-length 65K conversations fit simultaneously.",[34,685,686],{},[37,687,44],{"href":39,"ariaDescribedBy":688,"dataFootnoteRef":42,"id":689},[41],"user-content-fnref-model-2",[10,691,692,693,696],{},"We tested the practical effect separately with ",[14,694,695],{},"four prompts of about 61,400 tokens",", each requesting 512 output tokens. Actual prompt lengths range from 61,390 to 61,426 tokens. The original cache budget fits only two of these requests at a time. Giving W3A4 the freed memory as cache lets all four decode together:",[100,698],{"kind":699},"long-context",[10,701,702],{},[106,703,704],{},"Four long requests on one R9700. Complete batch wall time is lower-is-better. This control uses the 26 September MXFP4 build, not the 24 September image. Expanded-cache W3A4 changes the resource budget and uses the shipped calibration; the same-cache W3A4 check used an earlier calibration.",[116,706,707,723],{},[119,708,709],{},[122,710,711,714,717,720],{},[125,712,713],{},"Long-context measurement",[125,715,716],{"align":130},"MXFP4",[125,718,719],{"align":130},"W3A4, same cache budget",[125,721,722],{"align":130},"W3A4, expanded cache",[142,724,725,739,752],{},[122,726,727,730,733,736],{},[147,728,729],{},"Decode while all active requests run",[147,731,732],{"align":130},"69 tok\u002Fs, 2 active",[147,734,735],{"align":130},"117 tok\u002Fs, 2 active",[147,737,738],{"align":130},"199 tok\u002Fs, 4 active",[122,740,741,744,747,749],{},[147,742,743],{},"All four decoding together",[147,745,746],{"align":130},"No, two at a time",[147,748,746],{"align":130},[147,750,751],{"align":130},"Yes",[122,753,754,757,760,763],{},[147,755,756],{},"Wall time, 4 × 61K prompts + 512 tokens each",[147,758,759],{"align":130},"105.6 s",[147,761,762],{"align":130},"92.9 s",[147,764,765],{"align":130},"86.2 s",[10,767,768,769,772,773],{},"The new weights make the two-request decode faster even without extra cache. Reallocating memory then reduces the wait for the whole four-request batch. The 199 tok\u002Fs figure is the aggregate decode rate ",[14,770,771],{},"while all four are active",", not throughput across the complete batch. Whole-device peak usage remains nearly the same: 31.39 GiB with the default W3A4 cache versus 31.37 GiB for the MXFP4 control, leaving little headroom on this card. This is one long-context workload, not a universal concurrency guarantee.",[34,774,775],{},[37,776,357],{"href":530,"ariaDescribedBy":777,"dataFootnoteRef":42,"id":778},[41],"user-content-fnref-bench-2",[61,780,782],{"id":781},"what-we-changed-and-why","What we changed, and why",[310,784,786],{"id":785},"first-make-the-existing-runtime-predictable","First, make the existing runtime predictable",[10,788,789,790,794,795,798],{},"Before changing weights, we investigated server starts that alternated between roughly 28 and 36 ms per speculative step despite an unchanged configuration. Pinning the runtime to one hardware compute queue with ",[791,792,793],"code",{},"GPU_MAX_HW_QUEUES=1"," kept those starts in the faster mode. The image includes that setting. ",[14,796,797],{},"All comparison arms already used it",", so its gain is not included in the headline improvement.",[10,800,801],{},"We also combined the GatedDeltaNet speculative verification operation into one native kernel. That operation measured 23–28% faster, while the combined runtime fixes improve end-to-end weighted MXFP4 decode by about 1.5%. An individual kernel gain and the overall generation gain are not the same thing.",[310,803,805],{"id":804},"read-fewer-bytes-when-generating","Read fewer bytes when generating",[10,807,808],{},"Decode on this card is largely limited by memory bandwidth. The previous kernels were already close to the card's available bandwidth, so reducing weight traffic offered more room than scheduling changes alone.",[10,810,811,812,815,816,822],{},"The large linear projections move from approximately ",[14,813,814],{},"4.25 bits per weight with MXFP4 to 3.125 bits with grouped INT3",". This affects 24.3 billion decoder-projection weights, not every tensor in the 27B model, and is roughly a quarter less data for those weights. Native Paiton kernels use the packed values directly. Decode keeps 8-bit FP8 activations; in the serving comparison, median single-stream forward time falls from about 28.3–28.4 ms to 22.4–22.5 ms.",[34,817,818],{},[37,819,44],{"href":39,"ariaDescribedBy":820,"dataFootnoteRef":42,"id":821},[41],"user-content-fnref-model-3",[34,823,824],{},[37,825,357],{"href":530,"ariaDescribedBy":826,"dataFootnoteRef":42,"id":827},[41],"user-content-fnref-bench-3",[310,829,831],{"id":830},"use-different-precision-for-prompt-processing","Use different precision for prompt processing",[10,833,834,835,838],{},"Prompt processing has a different bottleneck. Simply using 3-bit weights with the old 8-bit activation path made it slower in our checks. The release therefore uses ",[14,836,837],{},"4-bit activation math only in selected prefill projections",", while generation keeps 8-bit activations.",[10,840,841,842],{},"A block-wise Hadamard rotation spreads large activation outliers across small groups, making lower-precision prompt processing more useful. The weights are calibrated in that rotated basis, and the runtime applies the corresponding activation transform. Grouped scaling retains more local precision than one scale for an entire token. This is the public method behind the W3A4 name, not a new model architecture.",[34,843,844],{},[37,845,44],{"href":39,"ariaDescribedBy":846,"dataFootnoteRef":42,"id":847},[41],"user-content-fnref-model-4",[310,849,851],{"id":850},"calibrate-our-own-weights-and-check-the-data-licenses","Calibrate our own weights and check the data licenses",[10,853,854,855,858],{},"These are our own GPTQ-calibrated weights, rather than a required third-party 3-bit checkpoint. Calibration uses ",[14,856,857],{},"292,864 tokens"," across math, code, science, web text, multilingual material, agentic traces and long documents. The layer-by-layer process takes about 13 minutes on a single AMD Instinct MI355X.",[10,860,861,862,867,868,873],{},"Before release, we replaced calibration sources carrying noncommercial or share-alike terms with permissively licensed sources. The shipped weights use that final calibration. The ",[37,863,866],{"href":864,"rel":865},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-27B-W3Rot-INT3-Paiton-RDNA4\u002Fblob\u002F002684a90ea9a879c2e071217f460f24b3ecf6ec\u002FCALIBRATION.md",[113],"calibration report"," lists dataset revisions and licenses; the ",[37,869,872],{"href":870,"rel":871},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-27B-W3Rot-INT3-Paiton-RDNA4\u002Fblob\u002F002684a90ea9a879c2e071217f460f24b3ecf6ec\u002FTHIRD_PARTY_NOTICES.md",[113],"third-party notices"," preserve required attributions.",[310,875,877],{"id":876},"keep-startup-and-integrity-checks-explicit","Keep startup and integrity checks explicit",[10,879,880],{},"The runtime loads the rotated replacements alongside the pinned base checkpoint and checks their files by SHA-256. No separate third-party 3-bit model is needed. We also fixed a warm-up issue where uninitialized dummy inputs could trigger a non-finite-value flag before the first real request: the flag is cleared after warm-up, while strict checks remain enabled for served requests.",[10,882,883,884,887],{},"Startup is about ",[14,885,886],{},"230 seconds",", compared with about 210 seconds previously. The warm serving gains do not remove this initial loading time.",[61,889,891],{"id":890},"quality-the-benefit-comes-with-a-choice","Quality: the benefit comes with a choice",[100,893],{"kind":894},"accuracy",[10,896,897],{},[106,898,899],{},"Served-model accuracy with greedy decoding, thinking disabled and identical questions. The paired 95% confidence intervals include zero for GSM8K and HumanEval, but not for the MMLU-Pro subset. The 80 needle tests are a bounded retrieval check, not broad long-context validation.",[116,901,902,917],{},[119,903,904],{},[122,905,906,909,911,914],{},[125,907,908],{},"Benchmark",[125,910,716],{"align":130},[125,912,913],{"align":130},"W3A4",[125,915,916],{},"Difference and paired 95% confidence interval",[142,918,919,937,954,971],{},[122,920,921,924,927,930],{},[147,922,923],{},"GSM8K, 5-shot, 1,319 questions",[147,925,926],{"align":130},"95.68%",[147,928,929],{"align":130},"95.30%",[147,931,932,933],{},"−0.38 pts ",[934,935,936],"span",{},"−1.44, +0.68",[122,938,939,942,945,948],{},[147,940,941],{},"HumanEval pass@1, 164 tasks",[147,943,944],{"align":130},"95.12%",[147,946,947],{"align":130},"93.90%",[147,949,950,951],{},"−1.22 pts ",[934,952,953],{},"−4.99, +2.56",[122,955,956,959,962,965],{},[147,957,958],{},"MMLU-Pro subset, 0-shot, 14 × 100 questions",[147,960,961],{"align":130},"62.57%",[147,963,964],{"align":130},"59.71%",[147,966,967,968],{},"−2.86 pts ",[934,969,970],{},"−4.81, −0.90",[122,972,973,976,979,981],{},[147,974,975],{},"Needle retrieval at 61,440 tokens, 80 checks",[147,977,978],{"align":130},"100%",[147,980,978],{"align":130},[147,982,983],{},"No observed difference",[10,985,986,987,990,991],{},"Math and code stay close in the evaluated sets; the intervals do ",[14,988,989],{},"not"," prove identical quality. Knowledge-heavy multiple choice shows a measurable reduction of about three points. DFlash2 acceptance also remains close, with task-level changes from −0.4% to +2.8%.",[34,992,993],{},[37,994,44],{"href":39,"ariaDescribedBy":995,"dataFootnoteRef":42,"id":996},[41],"user-content-fnref-model-5",[10,998,999,1000,1003],{},"For knowledge-heavy work, choose ",[791,1001,1002],{},"--weights mxfp4",". It uses the higher-precision weight path and still receives this round's MXFP4 runtime improvements. W3A4 is useful when faster generation and more conversation cache are worth that measured accuracy tradeoff. Outputs can differ; we do not promise identical answers.",[10,1005,1006,1007,1010,1011,1014,1015,1018,1019,1025],{},"Multilingual, tool-calling and longer-context quality were ",[14,1008,1009],{},"not measured"," in this evaluation. The released W3A4 package is ",[14,1012,1013],{},"text-only and R9700-specific",", and these performance and quality results cover the ",[14,1016,1017],{},"65K profile",". Other launcher profile combinations have not been measured; the separate 200K image continues to use MXFP4. These replacement weights require the Paiton runtime; stock vLLM, Transformers and llama.cpp cannot load them on their own.",[34,1020,1021],{},[37,1022,44],{"href":39,"ariaDescribedBy":1023,"dataFootnoteRef":42,"id":1024},[41],"user-content-fnref-model-6",[34,1026,1027],{},[37,1028,1032],{"href":1029,"ariaDescribedBy":1030,"dataFootnoteRef":42,"id":1031},"#user-content-fn-release",[41],"user-content-fnref-release","3",[61,1034,1036],{"id":1035},"how-to-read-these-results","How to read these results",[10,1038,1039,1040,1043,1044,59],{},"The performance comparison uses ",[14,1041,1042],{},"one R9700 at its 300 W cap",", vLLM 0.29, ROCm 10, a 65,536-token context, up to eight sequences, DFlash2 and thinking disabled. Automatic prefix caching and n-gram co-drafting are off. Each image ran in two fresh server processes, interleaved, and the tables show the mean of the two runs. BetterBench used version 0.6.0 and its quick preset. The reference image is ",[791,1045,1046],{},"qwen38-rocm10-vllm029-65k-20260924-r3",[10,1048,1049,1050,1053,1054,1057,1058,1064],{},"The public performance runs used the ",[14,1051,1052],{},"earlier v3 calibration",". The accuracy figures and expanded-cache four-request run use the ",[14,1055,1056],{},"final, permissively calibrated weights that are shipped",". Their tensor format and runtime are identical, but these are not all tests of one identical set of weight bytes. Sampled text and speculative acceptance can differ: W3A4 changed seven of twelve greedy control outputs relative to MXFP4. These are serving-throughput measurements, not identical-output timing. Percentage changes come from unrounded measurements, so recalculating from the rounded tables can differ slightly.",[34,1059,1060],{},[37,1061,44],{"href":39,"ariaDescribedBy":1062,"dataFootnoteRef":42,"id":1063},[41],"user-content-fnref-model-7",[34,1065,1066],{},[37,1067,357],{"href":530,"ariaDescribedBy":1068,"dataFootnoteRef":42,"id":1069},[41],"user-content-fnref-bench-4",[10,1071,1072,1073,1078],{},"BetterBench's weighted decode emphasizes code (30%), reasoning (20%), prose and JSON (15% each), then file editing and summarization (10% each). Chat and math are shown separately and have no weight in that headline. Each run scores five requests per decode category, 48 per concurrency level and eight per nominal prefill depth, after fixed warmups. File-edit results vary most between the two W3A4 runs: 179.0 and 211.1 tok\u002Fs. The ",[37,1074,1077],{"href":1075,"rel":1076},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Ftree\u002Fa44044105287da3d040652ea8bcd3927a94a1bf2\u002Fmodels\u002FQwen3.8-MXFP4-DFlash2\u002Fbenchmarks\u002F2026-09-26-w3a4",[113],"full report and source data"," retain those measurement details.",[10,1080,1081],{},"The 300 W figure is the configured GPU power cap. We did not measure whole-system energy or cost per token, and faster generation alone is not an electricity-savings measurement.",[61,1083,1085],{"id":1084},"try-it-on-your-r9700","Try it on your R9700",[10,1087,1088,1089,1094,1095],{},"Use Linux x86-64, Python 3, Docker, the Hugging Face CLI and AMD GPU device access. The commands below pin the public release checkout and the target, drafter and rotated-weight revisions together. If you already have the plugin repository, use a separate checkout rather than replacing local work. The ",[37,1090,1093],{"href":1091,"rel":1092},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fa44044105287da3d040652ea8bcd3927a94a1bf2\u002Fmodels\u002FQwen3.8-MXFP4-DFlash2\u002FREADME.md",[113],"release setup guide"," documents existing downloads and other profiles.",[34,1096,1097],{},[37,1098,1032],{"href":1029,"ariaDescribedBy":1099,"dataFootnoteRef":42,"id":1100},[41],"user-content-fnref-release-2",[1102,1103,1107],"pre",{"className":1104,"code":1105,"language":1106,"meta":42,"style":42},"language-bash shiki shiki-themes github-light github-dark","git clone https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\ncd paiton-vllm-plugin\ngit checkout a44044105287da3d040652ea8bcd3927a94a1bf2\n\nexport PAITON_TARGET_DIR=\"$PWD\u002Fmodel-cache\u002Fqwen38-nvfp4\"\nexport PAITON_DRAFT_DIR=\"$PWD\u002Fmodel-cache\u002Fqwen38-dflash2\"\nexport PAITON_W3ROT_DIR=\"$PWD\u002Fmodel-cache\u002Fqwen38-w3rot-int3\"\nexport PAITON_CACHE_DIR=\"$PWD\u002Fruntime-cache\u002Fqwen38-rocm10-65k-w3a4\"\nmkdir -p \"$PAITON_TARGET_DIR\" \"$PAITON_DRAFT_DIR\" \"$PAITON_W3ROT_DIR\" \"$PAITON_CACHE_DIR\"\n\nhf download unsloth\u002FQwen3.8-27B-NVFP4 \\\n  --revision f0b7c9e722f5565102fff8481c99e4d86ae099c7 --local-dir \"$PAITON_TARGET_DIR\"\nhf download tcclaviger\u002FQwen3.8-27B-DFlash2-FP8 \\\n  --revision ee0cb26a8279b7910cc28d82a8a3e15e4728d56f --local-dir \"$PAITON_DRAFT_DIR\"\nhf download EliovpAI\u002FQwen3.8-27B-W3Rot-INT3-Paiton-RDNA4 \\\n  --revision 278486debe64e21e5e9d45ac8d02798d72fbdf83 --local-dir \"$PAITON_W3ROT_DIR\"\n(cd \"$PAITON_W3ROT_DIR\" && sha256sum -c SHA256SUMS)\n\nbash models\u002FQwen3.8-MXFP4-DFlash2\u002Frun-rocm10-65k.sh\n","bash",[791,1108,1109,1124,1134,1145,1152,1175,1192,1209,1226,1265,1270,1285,1303,1315,1331,1343,1359,1388,1393],{"__ignoreMap":42},[934,1110,1113,1117,1121],{"class":1111,"line":1112},"line",1,[934,1114,1116],{"class":1115},"sScJk","git",[934,1118,1120],{"class":1119},"sZZnC"," clone",[934,1122,1123],{"class":1119}," https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\n",[934,1125,1127,1131],{"class":1111,"line":1126},2,[934,1128,1130],{"class":1129},"sj4cs","cd",[934,1132,1133],{"class":1119}," paiton-vllm-plugin\n",[934,1135,1137,1139,1142],{"class":1111,"line":1136},3,[934,1138,1116],{"class":1115},[934,1140,1141],{"class":1119}," checkout",[934,1143,1144],{"class":1119}," a44044105287da3d040652ea8bcd3927a94a1bf2\n",[934,1146,1148],{"class":1111,"line":1147},4,[934,1149,1151],{"emptyLinePlaceholder":1150},true,"\n",[934,1153,1155,1159,1163,1166,1169,1172],{"class":1111,"line":1154},5,[934,1156,1158],{"class":1157},"szBVR","export",[934,1160,1162],{"class":1161},"sVt8B"," PAITON_TARGET_DIR",[934,1164,1165],{"class":1157},"=",[934,1167,1168],{"class":1119},"\"",[934,1170,1171],{"class":1161},"$PWD",[934,1173,1174],{"class":1119},"\u002Fmodel-cache\u002Fqwen38-nvfp4\"\n",[934,1176,1178,1180,1183,1185,1187,1189],{"class":1111,"line":1177},6,[934,1179,1158],{"class":1157},[934,1181,1182],{"class":1161}," PAITON_DRAFT_DIR",[934,1184,1165],{"class":1157},[934,1186,1168],{"class":1119},[934,1188,1171],{"class":1161},[934,1190,1191],{"class":1119},"\u002Fmodel-cache\u002Fqwen38-dflash2\"\n",[934,1193,1195,1197,1200,1202,1204,1206],{"class":1111,"line":1194},7,[934,1196,1158],{"class":1157},[934,1198,1199],{"class":1161}," PAITON_W3ROT_DIR",[934,1201,1165],{"class":1157},[934,1203,1168],{"class":1119},[934,1205,1171],{"class":1161},[934,1207,1208],{"class":1119},"\u002Fmodel-cache\u002Fqwen38-w3rot-int3\"\n",[934,1210,1212,1214,1217,1219,1221,1223],{"class":1111,"line":1211},8,[934,1213,1158],{"class":1157},[934,1215,1216],{"class":1161}," PAITON_CACHE_DIR",[934,1218,1165],{"class":1157},[934,1220,1168],{"class":1119},[934,1222,1171],{"class":1161},[934,1224,1225],{"class":1119},"\u002Fruntime-cache\u002Fqwen38-rocm10-65k-w3a4\"\n",[934,1227,1229,1232,1235,1238,1241,1243,1245,1248,1250,1252,1255,1257,1259,1262],{"class":1111,"line":1228},9,[934,1230,1231],{"class":1115},"mkdir",[934,1233,1234],{"class":1129}," -p",[934,1236,1237],{"class":1119}," \"",[934,1239,1240],{"class":1161},"$PAITON_TARGET_DIR",[934,1242,1168],{"class":1119},[934,1244,1237],{"class":1119},[934,1246,1247],{"class":1161},"$PAITON_DRAFT_DIR",[934,1249,1168],{"class":1119},[934,1251,1237],{"class":1119},[934,1253,1254],{"class":1161},"$PAITON_W3ROT_DIR",[934,1256,1168],{"class":1119},[934,1258,1237],{"class":1119},[934,1260,1261],{"class":1161},"$PAITON_CACHE_DIR",[934,1263,1264],{"class":1119},"\"\n",[934,1266,1268],{"class":1111,"line":1267},10,[934,1269,1151],{"emptyLinePlaceholder":1150},[934,1271,1273,1276,1279,1282],{"class":1111,"line":1272},11,[934,1274,1275],{"class":1115},"hf",[934,1277,1278],{"class":1119}," download",[934,1280,1281],{"class":1119}," unsloth\u002FQwen3.8-27B-NVFP4",[934,1283,1284],{"class":1129}," \\\n",[934,1286,1288,1291,1294,1297,1299,1301],{"class":1111,"line":1287},12,[934,1289,1290],{"class":1129},"  --revision",[934,1292,1293],{"class":1119}," f0b7c9e722f5565102fff8481c99e4d86ae099c7",[934,1295,1296],{"class":1129}," --local-dir",[934,1298,1237],{"class":1119},[934,1300,1240],{"class":1161},[934,1302,1264],{"class":1119},[934,1304,1306,1308,1310,1313],{"class":1111,"line":1305},13,[934,1307,1275],{"class":1115},[934,1309,1278],{"class":1119},[934,1311,1312],{"class":1119}," tcclaviger\u002FQwen3.8-27B-DFlash2-FP8",[934,1314,1284],{"class":1129},[934,1316,1318,1320,1323,1325,1327,1329],{"class":1111,"line":1317},14,[934,1319,1290],{"class":1129},[934,1321,1322],{"class":1119}," ee0cb26a8279b7910cc28d82a8a3e15e4728d56f",[934,1324,1296],{"class":1129},[934,1326,1237],{"class":1119},[934,1328,1247],{"class":1161},[934,1330,1264],{"class":1119},[934,1332,1334,1336,1338,1341],{"class":1111,"line":1333},15,[934,1335,1275],{"class":1115},[934,1337,1278],{"class":1119},[934,1339,1340],{"class":1119}," EliovpAI\u002FQwen3.8-27B-W3Rot-INT3-Paiton-RDNA4",[934,1342,1284],{"class":1129},[934,1344,1346,1348,1351,1353,1355,1357],{"class":1111,"line":1345},16,[934,1347,1290],{"class":1129},[934,1349,1350],{"class":1119}," 278486debe64e21e5e9d45ac8d02798d72fbdf83",[934,1352,1296],{"class":1129},[934,1354,1237],{"class":1119},[934,1356,1254],{"class":1161},[934,1358,1264],{"class":1119},[934,1360,1362,1365,1367,1369,1371,1373,1376,1379,1382,1385],{"class":1111,"line":1361},17,[934,1363,1364],{"class":1161},"(",[934,1366,1130],{"class":1129},[934,1368,1237],{"class":1119},[934,1370,1254],{"class":1161},[934,1372,1168],{"class":1119},[934,1374,1375],{"class":1161}," && ",[934,1377,1378],{"class":1115},"sha256sum",[934,1380,1381],{"class":1129}," -c",[934,1383,1384],{"class":1119}," SHA256SUMS",[934,1386,1387],{"class":1161},")\n",[934,1389,1391],{"class":1111,"line":1390},18,[934,1392,1151],{"emptyLinePlaceholder":1150},[934,1394,1396,1398],{"class":1111,"line":1395},19,[934,1397,1106],{"class":1115},[934,1399,1400],{"class":1119}," models\u002FQwen3.8-MXFP4-DFlash2\u002Frun-rocm10-65k.sh\n",[10,1402,1403,1404,1407],{},"The released W3A4 image is ",[791,1405,1406],{},"ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin:qwen38-rocm10-vllm029-65k-20260926-w3a4-r1",". Its immutable digest is:",[1102,1409,1414],{"className":1410,"code":1412,"language":1413,"meta":42},[1411],"language-text","sha256:c4134aba665f6dd3b89354a43be2b5b814f7078db456351647a3f1b106a0da49\n","text",[791,1415,1412],{"__ignoreMap":42},[10,1417,1418,1419,1422],{},"When ",[791,1420,1421],{},"PAITON_W3ROT_DIR"," is set, the model-specific launcher selects the rotated 3-bit weights. After stopping that server, select MXFP4 with:",[1102,1424,1426],{"className":1104,"code":1425,"language":1106,"meta":42,"style":42},"bash models\u002FQwen3.8-MXFP4-DFlash2\u002Frun-rocm10-65k.sh --weights mxfp4\n",[791,1427,1428],{"__ignoreMap":42},[934,1429,1430,1432,1435,1438],{"class":1111,"line":1112},[934,1431,1106],{"class":1115},[934,1433,1434],{"class":1119}," models\u002FQwen3.8-MXFP4-DFlash2\u002Frun-rocm10-65k.sh",[934,1436,1437],{"class":1129}," --weights",[934,1439,1440],{"class":1119}," mxfp4\n",[10,1442,1443,1444,1447,1448,1451],{},"Once ready, the server exposes the OpenAI-compatible API at ",[791,1445,1446],{},"http:\u002F\u002F127.0.0.1:18982\u002Fv1"," using the model name ",[791,1449,1450],{},"Qwen3.8",". In another terminal, send a streaming request:",[1102,1453,1455],{"className":1104,"code":1454,"language":1106,"meta":42,"style":42},"curl --fail http:\u002F\u002F127.0.0.1:18982\u002Fv1\u002Fchat\u002Fcompletions \\\n  -H 'Content-Type: application\u002Fjson' \\\n  -d '{\"model\":\"Qwen3.8\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a short Python function that removes duplicates while keeping their original order.\"}],\"temperature\":0,\"max_tokens\":256,\"stream\":true,\"chat_template_kwargs\":{\"enable_thinking\":false}}'\n",[791,1456,1457,1470,1480],{"__ignoreMap":42},[934,1458,1459,1462,1465,1468],{"class":1111,"line":1112},[934,1460,1461],{"class":1115},"curl",[934,1463,1464],{"class":1129}," --fail",[934,1466,1467],{"class":1119}," http:\u002F\u002F127.0.0.1:18982\u002Fv1\u002Fchat\u002Fcompletions",[934,1469,1284],{"class":1129},[934,1471,1472,1475,1478],{"class":1111,"line":1126},[934,1473,1474],{"class":1129},"  -H",[934,1476,1477],{"class":1119}," 'Content-Type: application\u002Fjson'",[934,1479,1284],{"class":1129},[934,1481,1482,1485],{"class":1111,"line":1136},[934,1483,1484],{"class":1129},"  -d",[934,1486,1487],{"class":1119}," '{\"model\":\"Qwen3.8\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a short Python function that removes duplicates while keeping their original order.\"}],\"temperature\":0,\"max_tokens\":256,\"stream\":true,\"chat_template_kwargs\":{\"enable_thinking\":false}}'\n",[10,1489,1490,1491,1494,1495,1498,1499],{},"These commands run the ",[14,1492,1493],{},"model-specific W3A4 Docker profile with DFlash2",". The shorter native ",[791,1496,1497],{},"paiton serve qwen38-nvfp4"," preset is a different, non-speculative configuration and does not load these 3-bit weights.",[34,1500,1501],{},[37,1502,1032],{"href":1029,"ariaDescribedBy":1503,"dataFootnoteRef":42,"id":1504},[41],"user-content-fnref-release-3",[61,1506,1508],{"id":1507},"what-is-next-and-what-is-not-shipped","What is next, and what is not shipped",[10,1510,1511,1512,1515],{},"A 4-bit conversation cache passed our bounded accuracy gate, but currently reduces attention traffic without adding usable cache capacity. Its engineering gain was about 4–5% with two 61K requests. Doubling capacity in the same memory is still work in progress, ",[14,1513,1514],{},"not a benefit of this release",". A longer speculative block is another investigation, not a published setting here.",[61,1517,1519],{"id":1518},"credits-and-licensing","Credits and licensing",[10,1521,1522,1523,1526,1527,1530,1531,1537],{},"Qwen3.8 27B is by the Qwen team. The pinned target checkpoint is provided by Unsloth, and the DFlash2 drafter by tcclaviger, building on z-lab's DFlash2 work. The released rotated weights are an Apache-2.0 quantized derivative; the model card records Apache-2.0 terms for those upstream checkpoints. The public method builds on ",[14,1524,1525],{},"GPTQ and GSQ",". vLLM provides the serving foundation, and BetterBench the generation workload suite. ",[14,1528,1529],{},"Radiance and StillDeadcode\u002Flibr4d are credited for adapted kernel techniques"," in the public runtime guide.",[34,1532,1533],{},[37,1534,44],{"href":39,"ariaDescribedBy":1535,"dataFootnoteRef":42,"id":1536},[41],"user-content-fnref-model-8",[34,1538,1539],{},[37,1540,1032],{"href":1029,"ariaDescribedBy":1541,"dataFootnoteRef":42,"id":1542},[41],"user-content-fnref-release-4",[10,1544,1545],{},"The public adapter, model weights and packaged native runtime have distinct terms; the weights' Apache-2.0 license is not a blanket license for proprietary compiler or implementation code. Calibration datasets also retain their own attribution requirements and, where applicable, Common Crawl terms. Review the published notices before deployment.",[10,1547,1548,1549,1553,1554,1557,1558,1562],{},"Explore ",[37,1550,1552],{"href":1551},"\u002Fproducts\u002Fpaiton","Paiton",", read the ",[37,1555,1556],{"href":92},"previous Qwen3.8 release story",", or ",[37,1559,1561],{"href":1560},"\u002Fcontact","contact us"," to discuss a local AMD AI workload.",[1564,1565,1568,1573],"section",{"className":1566,"dataFootnotes":42},[1567],"footnotes",[61,1569,1572],{"className":1570,"id":41},[1571],"sr-only","Footnotes",[1574,1575,1576,1646,1679],"ol",{},[1577,1578,1580,1585,1586,1593,1594,1593,1601,1593,1608,1593,1615,1593,1623,1593,1631,1593,1639],"li",{"id":1579},"user-content-fn-model",[37,1581,1584],{"href":1582,"rel":1583},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-27B-W3Rot-INT3-Paiton-RDNA4\u002Fblob\u002F002684a90ea9a879c2e071217f460f24b3ecf6ec\u002FREADME.md",[113],"Public W3Rot INT3 model card, reviewed revision",". ",[37,1587,1592],{"href":1588,"ariaLabel":1589,"className":1590,"dataFootnoteBackref":42},"#user-content-fnref-model","Back to reference 1",[1591],"data-footnote-backref","↩"," ",[37,1595,1592,1599],{"href":1596,"ariaLabel":1597,"className":1598,"dataFootnoteBackref":42},"#user-content-fnref-model-2","Back to reference 1-2",[1591],[34,1600,357],{},[37,1602,1592,1606],{"href":1603,"ariaLabel":1604,"className":1605,"dataFootnoteBackref":42},"#user-content-fnref-model-3","Back to reference 1-3",[1591],[34,1607,1032],{},[37,1609,1592,1613],{"href":1610,"ariaLabel":1611,"className":1612,"dataFootnoteBackref":42},"#user-content-fnref-model-4","Back to reference 1-4",[1591],[34,1614,371],{},[37,1616,1592,1620],{"href":1617,"ariaLabel":1618,"className":1619,"dataFootnoteBackref":42},"#user-content-fnref-model-5","Back to reference 1-5",[1591],[34,1621,1622],{},"5",[37,1624,1592,1628],{"href":1625,"ariaLabel":1626,"className":1627,"dataFootnoteBackref":42},"#user-content-fnref-model-6","Back to reference 1-6",[1591],[34,1629,1630],{},"6",[37,1632,1592,1636],{"href":1633,"ariaLabel":1634,"className":1635,"dataFootnoteBackref":42},"#user-content-fnref-model-7","Back to reference 1-7",[1591],[34,1637,1638],{},"7",[37,1640,1592,1644],{"href":1641,"ariaLabel":1642,"className":1643,"dataFootnoteBackref":42},"#user-content-fnref-model-8","Back to reference 1-8",[1591],[34,1645,385],{},[1577,1647,1649,1585,1653,1593,1658,1593,1665,1593,1672],{"id":1648},"user-content-fn-bench",[37,1650,1652],{"href":1075,"rel":1651},[113],"26 September W3A4 benchmark report and machine-readable results",[37,1654,1592],{"href":1655,"ariaLabel":1656,"className":1657,"dataFootnoteBackref":42},"#user-content-fnref-bench","Back to reference 2",[1591],[37,1659,1592,1663],{"href":1660,"ariaLabel":1661,"className":1662,"dataFootnoteBackref":42},"#user-content-fnref-bench-2","Back to reference 2-2",[1591],[34,1664,357],{},[37,1666,1592,1670],{"href":1667,"ariaLabel":1668,"className":1669,"dataFootnoteBackref":42},"#user-content-fnref-bench-3","Back to reference 2-3",[1591],[34,1671,1032],{},[37,1673,1592,1677],{"href":1674,"ariaLabel":1675,"className":1676,"dataFootnoteBackref":42},"#user-content-fnref-bench-4","Back to reference 2-4",[1591],[34,1678,371],{},[1577,1680,1682,1585,1686,1593,1691,1593,1698,1593,1705],{"id":1681},"user-content-fn-release",[37,1683,1685],{"href":1091,"rel":1684},[113],"Public model-specific release and setup guide, pinned checkout",[37,1687,1592],{"href":1688,"ariaLabel":1689,"className":1690,"dataFootnoteBackref":42},"#user-content-fnref-release","Back to reference 3",[1591],[37,1692,1592,1696],{"href":1693,"ariaLabel":1694,"className":1695,"dataFootnoteBackref":42},"#user-content-fnref-release-2","Back to reference 3-2",[1591],[34,1697,357],{},[37,1699,1592,1703],{"href":1700,"ariaLabel":1701,"className":1702,"dataFootnoteBackref":42},"#user-content-fnref-release-3","Back to reference 3-3",[1591],[34,1704,1032],{},[37,1706,1592,1710],{"href":1707,"ariaLabel":1708,"className":1709,"dataFootnoteBackref":42},"#user-content-fnref-release-4","Back to reference 3-4",[1591],[34,1711,371],{},[1713,1714,1715],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html pre.shiki code .szBVR, html code.shiki .szBVR{--shiki-default:#D73A49;--shiki-dark:#F97583}html pre.shiki code .sVt8B, html code.shiki .sVt8B{--shiki-default:#24292E;--shiki-dark:#E1E4E8}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":42,"searchDepth":1126,"depth":1126,"links":1717},[1718,1719,1723,1724,1731,1732,1733,1734,1735,1736],{"id":63,"depth":1126,"text":64},{"id":97,"depth":1126,"text":98,"children":1720},[1721,1722],{"id":312,"depth":1136,"text":313},{"id":400,"depth":1136,"text":401},{"id":619,"depth":1126,"text":620},{"id":781,"depth":1126,"text":782,"children":1725},[1726,1727,1728,1729,1730],{"id":785,"depth":1136,"text":786},{"id":804,"depth":1136,"text":805},{"id":830,"depth":1136,"text":831},{"id":850,"depth":1136,"text":851},{"id":876,"depth":1136,"text":877},{"id":890,"depth":1126,"text":891},{"id":1035,"depth":1126,"text":1036},{"id":1084,"depth":1126,"text":1085},{"id":1507,"depth":1126,"text":1508},{"id":1518,"depth":1126,"text":1519},{"id":41,"depth":1126,"text":1572},[1552,1738,1739,1740,1741],"AMD Radeon","Local AI","Qwen","Quantization","2026-09-27T09:00:00Z","Qwen3.8 27B on one R9700: 19.9% faster weighted decode, more conversation cache and a measured accuracy tradeoff. Benchmarks and Paiton setup.","md","3-bit weights. 20% faster decode. 1 × Radeon AI PRO R9700.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-w3a4\u002Fhero.webp",{},"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700","\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700",{"title":5,"description":1743},"paiton-qwen38-w3a4-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-w3a4\u002Fsocial.webp","blog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700",null,"CD2h9HpgoAZG8HzOVA3Z5fTvfKIe3pOl6d2nm1rGmHw",[1757,1759,1769,1780,1790,1802,1811,1826,1835,1849,1881,1893,1915,1933,1952,1970,1988,2005,2017,2033,2048,2060,2069,2077,2092,2104,2115,2126,2136,2149,2159,2172,2183,2193,2204,2213,2225,2236,2245],{"path":1749,"title":5,"description":1743,"date":1742,"slug":1751,"image":1746,"originalUrl":1748,"categories":1758},[1552,1738,1739,1740,1741],{"path":1760,"title":1761,"description":1762,"date":1763,"slug":1764,"image":1765,"originalUrl":1766,"categories":1767},"\u002Fblog\u002Fpaiton-qwen-image-21-radeon-ai-pro-r9700","Qwen-Image 2.1 on 1 × Radeon AI PRO R9700: 2048×2048 images in 103 seconds","Generate 2048×2048 Qwen-Image 2.1 images locally with 1 × Radeon AI PRO R9700. The released Paiton v1.0.2 container measured 103.29 seconds per warm request, through PNG delivery.","2026-09-23T09:00:00Z","paiton-qwen-image-21-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen-image-21\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen-image-21-radeon-ai-pro-r9700",[1552,1738,1739,1768,1740],"Image Generation",{"path":92,"title":1770,"description":1771,"date":1772,"slug":1773,"image":1774,"originalUrl":1775,"categories":1776},"Qwen3.8: 400.7 tok\u002Fs on R9700 | Paiton","Qwen3.8 on one R9700: 400.7 aggregate tok\u002Fs with ROCm 10 and vLLM 0.29, plus public 200K\u002F220K chat profiles. Benchmarks, limits and launch commands.","2026-09-16T07:30:00Z","paiton-qwen38-mxfp4-dflash2-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fupdate-2026-09-19\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",[1552,1738,1777,1450,1778,1779],"vLLM","Inference Optimization","DFlash2",{"path":1781,"title":1782,"description":1783,"date":1784,"slug":1785,"image":1786,"originalUrl":1787,"categories":1788},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[1552,1738,1739,1789,1777],"GGUF",{"path":1791,"title":1792,"description":1793,"date":1794,"slug":1795,"image":1796,"originalUrl":1797,"categories":1798},"\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","MiniMax H3 on Radeon: 15-Second Video With Native Sound","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","2026-09-09T07:30:00Z","paiton-minimax-h3-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",[1552,1738,1739,1799,1800,1801],"Video Generation","MiniMax H3","ComfyUI",{"path":1803,"title":1804,"description":1805,"date":1806,"slug":1807,"image":1808,"originalUrl":1803,"categories":1809},"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","2026-09-07T09:00:00","paiton-flux2-klein-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",[1552,1738,1739,1768,1810,1801],"FLUX",{"path":1812,"title":1813,"description":1814,"date":1815,"slug":1816,"image":1817,"originalUrl":1818,"categories":1819},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[1552,1820,1738,1821,1822,1823,1778,1824,1777,1825],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Large Language Models","Cost Efficiency",{"path":1827,"title":1828,"description":1829,"date":1830,"slug":1831,"image":1832,"originalUrl":1833,"categories":1834},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[1552,1820,1738,1821,1822,1823,1778,1824,1777,1825],{"path":1836,"title":1837,"description":1838,"date":1839,"slug":1840,"image":1841,"originalUrl":1754,"categories":1842},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",[1843,1844,1845,1846,1847,1848],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1850,"title":1851,"description":1852,"date":1853,"slug":1854,"image":1855,"originalUrl":1856,"categories":1857},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1843,1820,1552,1858,1859,1860,1861,1862,1863,1864,1865,1866,1867,1868,1869,1870,1871,1872,1873,1874,1552,1875,1876,1877,1878,1879,1880],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","GPU","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1882,"title":1883,"description":1884,"date":1885,"slug":1886,"image":1887,"originalUrl":1888,"categories":1889},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1843,1820,1890,1891,1859,1892,1890,1872],"Modular DC","Uncategorized","De Tijd",{"path":1894,"title":1895,"description":1896,"date":1897,"slug":1898,"image":1899,"originalUrl":1900,"categories":1901},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1843,1820,1902,1891,1903,1904,1905,1906,1907,1908,1909,1910,1865,1911,1866,1912,1913,1914],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":1916,"title":1917,"description":1918,"date":1919,"slug":1920,"image":1921,"originalUrl":1922,"categories":1923},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1843,1924,1925,1926,1908,1927,1928,1929,1911,1930,1931,1932,1913],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1934,"title":1935,"description":1936,"date":1937,"slug":1938,"image":1939,"originalUrl":1940,"categories":1941},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1843,1820,1942,1902,1943,1944,1945,1946,1947,1948,1949,1950,1875,1951],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1953,"title":1954,"description":1955,"date":1956,"slug":1957,"image":1958,"originalUrl":1959,"categories":1960},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1843,1820,1902,1844,1961,1962,1963,1964,1965,1966,1967,1968,1969],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1971,"title":1972,"description":1973,"date":1974,"slug":1975,"image":1976,"originalUrl":1977,"categories":1978},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1843,1890,1891,1979,1844,1980,1981,1982,1983,1984,1985,1986,1987],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1989,"title":1990,"description":1991,"date":1992,"slug":1993,"image":1994,"originalUrl":1995,"categories":1996},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1843,1820,1552,1891,1997,1820,1998,1999,2000,2001,2002,2003,1552,2004],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning","ROCm",{"path":2006,"title":2007,"description":2008,"date":2009,"slug":2010,"image":2011,"originalUrl":2012,"categories":2013},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1843,1820,1552,1821,2014,1997,1825,2015,1778,2003,1552,2016,1777],"AMD Instinct","High Throughput","SGLang",{"path":2018,"title":2019,"description":2020,"date":2021,"slug":2022,"image":2023,"originalUrl":2024,"categories":2025},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1843,1820,1552,2026,1997,2027,1778,2028,2029,2030,2031,1552,2032],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":2034,"title":2035,"description":2036,"date":2037,"slug":2038,"image":2039,"originalUrl":2040,"categories":2041},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1843,1820,1942,1891,1943,2042,2043,2044,2045,1948,1950,1875,2046,2047],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":2049,"title":2050,"description":2051,"date":2052,"slug":2053,"image":2054,"originalUrl":2055,"categories":2056},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[1843,1820,1552,1891,2057,1859,1860,2058,2059,1871,1872,1552,1777],"AI","H200","MI300X",{"path":2061,"title":2062,"description":2063,"date":2064,"slug":2065,"image":2066,"originalUrl":2067,"categories":2068},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1843,1820,1942,1943,2042,2043,2044,2045,1948,1950,1875,2046,2047],{"path":2070,"title":2071,"description":2072,"date":2073,"slug":2074,"image":42,"originalUrl":2075,"categories":2076},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1843,1820,1552],{"path":2078,"title":2079,"description":2080,"date":2081,"slug":2082,"image":2083,"originalUrl":2084,"categories":2085},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1843,1820,1552,1891,1821,1997,2086,1999,2087,2088,2089,1552,2090,2091],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":2093,"title":2094,"description":2095,"date":2096,"slug":2097,"image":2098,"originalUrl":2099,"categories":2100},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[1843,1820,1552,1891,1997,2101,1947,1866,1822,1823,1824,2088,2102,2103],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":2105,"title":2106,"description":2107,"date":2108,"slug":2109,"image":2110,"originalUrl":2111,"categories":2112},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1843,1820,1552,1942,1891,1859,2059,1872,2113,2114],"RX7900XTX","tenstorrent",{"path":2116,"title":2117,"description":2118,"date":2119,"slug":2120,"image":2121,"originalUrl":2122,"categories":2123},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1843,1820,1890,1942,1860,2058,2124,1872,2125],"MI325x","pnl calculator",{"path":2127,"title":2128,"description":2129,"date":2130,"slug":2131,"image":2132,"originalUrl":2133,"categories":2134},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1843,1820,1552,2057,1859,2058,2135,1872,1552,1777],"MI300",{"path":2137,"title":2138,"description":2139,"date":2140,"slug":2141,"image":2142,"originalUrl":2143,"categories":2144},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1843,1890,2145,1844,1981,1847,1982,1983,2146,2147,1986,2148],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":2150,"title":2151,"description":2152,"date":2153,"slug":2154,"image":2155,"originalUrl":2156,"categories":2157},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1843,1820,1942,1891,2057,1859,2158],"Healthcare",{"path":2160,"title":2161,"description":2162,"date":2163,"slug":2164,"image":2165,"originalUrl":2166,"categories":2167},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1843,1902,2057,1859,2168,2169,2170,2171],"import","Taiwan","Tariffs","Trump",{"path":2173,"title":2174,"description":2175,"date":2176,"slug":2177,"image":2178,"originalUrl":2179,"categories":2180},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1843,1820,1942,2057,2181,2182],"AI Agents","ERP",{"path":2184,"title":2185,"description":2186,"date":2187,"slug":2188,"image":2189,"originalUrl":2190,"categories":2191},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1843,1820,1902,2192,1859,1867,1872],"AI news",{"path":2194,"title":2195,"description":2196,"date":2197,"slug":2198,"image":2199,"originalUrl":2200,"categories":2201},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1843,1820,1552,2057,1859,2202,2203,2059,1552],"benchmark","LLM",{"path":2205,"title":2206,"description":2207,"date":2208,"slug":2209,"image":2210,"originalUrl":2211,"categories":2212},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1843,1820,1552],{"path":2214,"title":2215,"description":2216,"date":2217,"slug":2218,"image":2219,"originalUrl":2220,"categories":2221},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1843,1859,2222,2223,2224],"Jim Greene","Podcast","Tech Talk",{"path":2226,"title":2227,"description":2228,"date":2229,"slug":2230,"image":2231,"originalUrl":2232,"categories":2233},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1843,1820,1552,1859,2234,2235,2058,2059,2124,1552,1777],"Deepseek","H100",{"path":2237,"title":2238,"description":2239,"date":2240,"slug":2241,"image":2242,"originalUrl":2243,"categories":2244},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1843,1820,1552,1859,2234,2235,2058,2059,2124,1552,1777],{"path":2246,"title":2247,"description":2248,"date":2249,"slug":2250,"image":2251,"originalUrl":2252,"categories":2253},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1843,1820,1552,1859,2235,2058,2059,2124,1552,1777],1790515511541]