[{"data":1,"prerenderedAt":1717},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700":3,"blog-posts-sidebar-en":1238},{"id":4,"title":5,"body":6,"categories":1221,"date":1226,"description":1227,"extension":1228,"heading":1229,"image":1230,"meta":1231,"navigation":75,"originalUrl":1232,"path":1233,"seo":1234,"slug":1235,"stem":1236,"__hash__":1237},"blog\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700.md","57% More Qwen3.8 Throughput on R9700 | Paiton",{"type":7,"value":8,"toc":1201},"minimark",[9,16,22,25,55,62,67,70,84,89,95,99,121,126,132,145,148,152,173,180,185,281,293,299,303,306,323,329,334,337,341,361,367,372,385,388,393,397,410,416,421,441,447,452,455,459,472,549,560,576,587,593,598,604,608,624,644,662,665,668,672,691,705,711,716,721,725,840,849,853,882,901,904,913,916,920,926,929,937,942,950,971,975],[10,11,12],"p",{},[13,14,15],"strong",{},"One Radeon AI PRO R9700. The same Qwen3.8-27B MXFP4 checkpoint. The same 5 GiB cache allocation. Paiton delivers 57% more aggregate throughput than our matched Radiance + DFlash2 baseline at eight concurrent requests and reduces median time to first token from 6.59 seconds to 195 milliseconds.",[10,17,18],{},[19,20,21],"em",{},"Throughput and latency headlines use the full 188-request comparison at eight concurrent requests. Both accelerated engines use the same target and DFlash2 snapshots. GPU artwork is illustrative.",[10,23,24],{},"This is a substantial step forward for serving a 27B model on one workstation GPU. Not just faster generation when one person is waiting. More requests progressing together, almost three times the estimated cache-token capacity, and dramatically less waiting under load.",[10,26,27,28,31,32,35,36,39,40,43],{},"Paiton reaches ",[13,29,30],{},"314.5 aggregate output tokens per second",", against ",[13,33,34],{},"200.3 tok\u002Fs"," for the matched accelerated baseline. Weighted serial decode improves by ",[13,37,38],{},"22%",", and prefill is faster at every tested prompt depth. The execution runs through a plugin on the official vLLM 0.28 ROCm runtime. ",[13,41,42],{},"The installed vLLM library stays unchanged.",[44,45,46],"sup",{},[47,48,54],"a",{"href":49,"ariaDescribedBy":50,"dataFootnoteRef":52,"id":53},"#user-content-fn-1",[51],"footnote-label","","user-content-fnref-1","1",[10,56,57,58,61],{},"The important combination is performance ",[13,59,60],{},"and"," deployment: a much stronger local serving result without moving to a separate inference-engine distribution.",[63,64,66],"h2",{"id":65},"see-it-in-action","See it in action",[10,68,69],{},"Watch Qwen3.8 generate responses on one Radeon AI PRO R9700, with Paiton + DFlash2 running through regular vLLM. The recording shows streamed output alongside live generation figures and GPU activity.",[10,71,72],{},[73,74],"video",{"controls":75,"playsInline":75,"preload":76,"width":77,"height":78,"ariaLabel":79,"ariaDescribedBy":80,"poster":82,"src":83},true,"none",1354,1080,"Screen recording of Qwen3.8 generating responses with Paiton on one Radeon AI PRO R9700",[81],"paiton-live-demo-caption","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Flive-demo-poster.webp","\u002Fasset\u002Fvideos\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fpaiton-qwen38-r9700-demo.mp4",[10,85,86],{"id":81},[19,87,88],{},"70-second screen recording with one active request at a time. The live figures describe the requests shown, not the eight-request aggregate benchmark below.",[10,90,91,94],{},[47,92,93],{"href":83},"Open the recording (MP4, 10.2 MB)",". Use the player's fullscreen control for a closer look.",[63,96,98],{"id":97},"a-strong-baseline-a-bigger-result","A strong baseline. A bigger result.",[10,100,101,102,108,109,112,113],{},"The investigation began with the impressive work around ",[47,103,107],{"href":104,"rel":105},"https:\u002F\u002Fgithub.com\u002Fmagiccodingman\u002Fvllm-radiance",[106],"nofollow","vLLM-Radiance",". Seeing what the team achieved on AMD hardware gave us ideas and encouraged us to push our own R9700 implementation further. Its public documentation describes support for the same AMD Quark MXFP4 model and DFlash2 acceleration; its published performance results use ",[13,110,111],{},"two R9700s",".",[44,114,115],{},[47,116,120],{"href":117,"ariaDescribedBy":118,"dataFootnoteRef":52,"id":119},"#user-content-fn-2",[51],"user-content-fnref-2","2",[10,122,123],{},[13,124,125],{},"Radiance inspired the investigation; Paiton's native implementation is our own.",[10,127,128,129],{},"We wanted to answer a different question: ",[13,130,131],{},"how much could we deliver from this model on one card, while keeping the regular vLLM deployment path?",[10,133,134,135,138,139],{},"We ran Radiance + DFlash2 and Paiton on vLLM + DFlash2 ourselves on ",[13,136,137],{},"one R9700",", using the same checkpoint, target and draft snapshots, FP8 KV cache, 5 GiB cache allocation and 8,192-token context ceiling. Both accelerated profiles used seven speculative tokens, greedy sampling and disabled prefix caching.",[44,140,141],{},[47,142,54],{"href":49,"ariaDescribedBy":143,"dataFootnoteRef":52,"id":144},[51],"user-content-fnref-1-2",[10,146,147],{},"Our percentage gains compare those matched single-GPU runs, not a one-card result against somebody else's two-card screenshot.",[63,149,151],{"id":150},"_57-more-throughput-with-a-growing-lead-under-load","57% more throughput, with a growing lead under load",[10,153,154,155,158,159,165],{},"The headline comes from the ",[13,156,157],{},"full 188-request BetterBench preset",", not the shorter, 128-output-token-cap comparison. We retained the preset's original longer output budgets, ten measured passes per task category, 24 requests at each tested concurrency level and repeated prefill measurements. Both engines completed the full request set.",[44,160,161],{},[47,162,54],{"href":49,"ariaDescribedBy":163,"dataFootnoteRef":52,"id":164},[51],"user-content-fnref-1-3",[44,166,167],{},[47,168,172],{"href":169,"ariaDescribedBy":170,"dataFootnoteRef":52,"id":171},"#user-content-fn-4",[51],"user-content-fnref-4","3",[10,174,175],{},[176,177],"img",{"alt":178,"src":179},"Paiton reaches 314.5 aggregate output tokens per second versus 200.3 for Radiance at eight concurrent requests, leading at every tested concurrency.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F01-full-throughput.webp",[10,181,182],{},[19,183,184],{},"Full 188-request workload, runs 495 \u002F 494. Aggregate throughput includes prefill and queueing. Higher is better.",[186,187,188,208],"table",{},[189,190,191],"thead",{},[192,193,194,198,202,205],"tr",{},[195,196,197],"th",{},"Concurrent requests",[195,199,201],{"align":200},"right","Radiance + DFlash2",[195,203,204],{"align":200},"Paiton on vLLM + DFlash2",[195,206,207],{"align":200},"Paiton uplift",[209,210,211,229,246,264],"tbody",{},[192,212,213,216,219,224],{},[214,215,54],"td",{},[214,217,218],{"align":200},"76.9 tok\u002Fs",[214,220,221],{"align":200},[13,222,223],{},"89.6 tok\u002Fs",[214,225,226],{"align":200},[13,227,228],{},"16.5%",[192,230,231,233,236,241],{},[214,232,120],{},[214,234,235],{"align":200},"142.7 tok\u002Fs",[214,237,238],{"align":200},[13,239,240],{},"165.7 tok\u002Fs",[214,242,243],{"align":200},[13,244,245],{},"16.1%",[192,247,248,251,254,259],{},[214,249,250],{},"4",[214,252,253],{"align":200},"187.1 tok\u002Fs",[214,255,256],{"align":200},[13,257,258],{},"254.8 tok\u002Fs",[214,260,261],{"align":200},[13,262,263],{},"36.2%",[192,265,266,269,271,276],{},[214,267,268],{},"8",[214,270,34],{"align":200},[214,272,273],{"align":200},[13,274,275],{},"314.5 tok\u002Fs",[214,277,278],{"align":200},[13,279,280],{},"57.0%",[10,282,283,286,287],{},[13,284,285],{},"Paiton leads at every tested concurrency level."," The full task-category results also show a lead across code, reasoning, prose, JSON, file editing, summarization, math and chat. The improvement survives the longer workload instead of depending on a single favorable prompt.",[44,288,289],{},[47,290,54],{"href":49,"ariaDescribedBy":291,"dataFootnoteRef":52,"id":292},[51],"user-content-fnref-1-4",[10,294,295,296],{},"These are aggregate serving rates across the active workload. ",[13,297,298],{},"314.5 tok\u002Fs does not mean each of eight users receives 314.5 tok\u002Fs.",[63,300,302],{"id":301},"from-a-659-second-wait-to-a-195-millisecond-first-token","From a 6.59-second wait to a 195-millisecond first token",[10,304,305],{},"The throughput gain is large. The change in responsiveness under load is larger.",[10,307,308,309,312,313,316,317],{},"At four concurrent requests, median time to first token falls from ",[13,310,311],{},"1,627 ms to 180 ms",". At eight, it falls from ",[13,314,315],{},"6,586 ms to 195 ms: 97% lower",". These timings include queueing, so they describe the wait a client experiences before a response begins.",[44,318,319],{},[47,320,54],{"href":49,"ariaDescribedBy":321,"dataFootnoteRef":52,"id":322},[51],"user-content-fnref-1-5",[10,324,325],{},[176,326],{"alt":327,"src":328},"At eight concurrent requests, median time to first token falls from 6,586 milliseconds with Radiance to 195 milliseconds with Paiton.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F02-time-to-first-token.webp",[10,330,331],{},[19,332,333],{},"Full 188-request workload. Lower is better. Radiance retains a 10–11 ms TTFT advantage at one and two concurrent requests; Paiton still leads throughput at both levels.",[10,335,336],{},"For a shared coding assistant or local agent endpoint, throughput alone is not enough. An endpoint that produces more tokens but leaves requests waiting to start can still feel slow. Here, the higher concurrency result comes with a substantially shorter wait for that first token.",[63,338,340],{"id":339},"the-same-5-gib-holds-almost-three-times-the-token-capacity","The same 5 GiB holds almost three times the token capacity",[10,342,343,344,347,348,351,352,112,355],{},"With ",[13,345,346],{},"exactly 5 GiB reserved on each engine",", the runtimes report ",[13,349,350],{},"25,746 cache-token slots for Radiance and 74,430 for Paiton",". That is ",[13,353,354],{},"2.89× estimated token capacity inside the same allocation",[44,356,357],{},[47,358,54],{"href":49,"ariaDescribedBy":359,"dataFootnoteRef":52,"id":360},[51],"user-content-fnref-1-6",[10,362,363],{},[176,364],{"alt":365,"src":366},"A 5 GiB cache allocation reports 25,746 token slots for Radiance and 74,430 for Paiton: 2.89 times the estimated capacity.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F03-cache-capacity.webp",[10,368,369],{},[19,370,371],{},"Runtime-reported shared-cache capacity estimates, not physical VRAM capacity. The per-request context ceiling remains 8,192 tokens.",[10,373,374,375,378,379],{},"The run logs capture up to ",[13,376,377],{},"three active requests for Radiance versus eight for Paiton"," in this matched configuration. Those are observations from these runs, not universal concurrency limits for either engine.",[44,380,381],{},[47,382,54],{"href":49,"ariaDescribedBy":383,"dataFootnoteRef":52,"id":384},[51],"user-content-fnref-1-7",[10,386,387],{},"The extra capacity helps explain the stronger concurrent experience: more requests can make progress inside the same budget. The capacity figures and active-request observations align with the throughput and latency result, although they do not isolate cache handling as the sole cause of the improvement.",[10,389,390],{},[13,391,392],{},"More usable serving capacity, not more VRAM or a larger advertised context window.",[63,394,396],{"id":395},"faster-decode-faster-prefill-across-the-workload","Faster decode. Faster prefill. Across the workload.",[10,398,399,400,403,404],{},"The gains are not confined to admitting more requests. In the full workload, weighted serial decode rises from ",[13,401,402],{},"86.0 to 104.9 output tok\u002Fs: 22% higher",". This measures generation after the first token and is separate from the aggregate complete-workload throughput above.",[44,405,406],{},[47,407,54],{"href":49,"ariaDescribedBy":408,"dataFootnoteRef":52,"id":409},[51],"user-content-fnref-1-8",[10,411,412],{},[176,413],{"alt":414,"src":415},"Full-workload weighted serial decode rises from 86.0 to 104.9 output tokens per second, a 22 percent increase.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F04-serial-decode.webp",[10,417,418],{},[19,419,420],{},"Weighted serial decode, full workload. Same target and draft snapshots on the same R9700.",[10,422,423,424,427,428,431,432,112,435],{},"Prefill improves at every tested prompt depth as well. At approximately ",[13,425,426],{},"1,556 \u002F 3,024 \u002F 5,226 input tokens",", Radiance measures ",[13,429,430],{},"2,828 \u002F 3,003 \u002F 2,926 input tok\u002Fs",". Paiton reaches ",[13,433,434],{},"3,182 \u002F 3,521 \u002F 3,367 input tok\u002Fs",[44,436,437],{},[47,438,54],{"href":49,"ariaDescribedBy":439,"dataFootnoteRef":52,"id":440},[51],"user-content-fnref-1-9",[10,442,443],{},[176,444],{"alt":445,"src":446},"Paiton delivers 3,182, 3,521 and 3,367 input tokens per second at three prompt depths, ahead of Radiance at each.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F05-full-prefill.webp",[10,448,449],{},[19,450,451],{},"Full-workload prefill results: 12.5%, 17.2% and 15.1% higher, calculated from the displayed summary values. Actual prompt tokens divided by HTTP time to first token, not isolated kernel throughput.",[10,453,454],{},"Together, these results show improvement at several points a user notices: getting the prompt processed, starting the answer under load and generating the rest of it.",[63,456,458],{"id":457},"against-stock-vllm-the-whole-engine-gap-is-striking","Against stock vLLM, the whole-engine gap is striking",[10,460,461,462,465,466],{},"We also ran a separate ",[13,463,464],{},"54-request matrix with generation capped at 128 tokens",", comparing stock vLLM O2, Radiance + DFlash2 and Paiton on vLLM + DFlash2. The stock profile used the best O2 settings we tested.",[44,467,468],{},[47,469,54],{"href":49,"ariaDescribedBy":470,"dataFootnoteRef":52,"id":471},[51],"user-content-fnref-1-10",[186,473,474,487],{},[189,475,476],{},[192,477,478,480,483,485],{},[195,479,197],{},[195,481,482],{"align":200},"Stock vLLM O2",[195,484,201],{"align":200},[195,486,204],{"align":200},[209,488,489,504,519,534],{},[192,490,491,493,496,499],{},[214,492,54],{},[214,494,495],{"align":200},"4.5 tok\u002Fs",[214,497,498],{"align":200},"78.0 tok\u002Fs",[214,500,501],{"align":200},[13,502,503],{},"90.0 tok\u002Fs",[192,505,506,508,511,514],{},[214,507,120],{},[214,509,510],{"align":200},"8.8 tok\u002Fs",[214,512,513],{"align":200},"151.2 tok\u002Fs",[214,515,516],{"align":200},[13,517,518],{},"160.5 tok\u002Fs",[192,520,521,523,526,529],{},[214,522,250],{},[214,524,525],{"align":200},"17.4 tok\u002Fs",[214,527,528],{"align":200},"189.2 tok\u002Fs",[214,530,531],{"align":200},[13,532,533],{},"231.5 tok\u002Fs",[192,535,536,538,541,544],{},[214,537,268],{},[214,539,540],{"align":200},"33.7 tok\u002Fs",[214,542,543],{"align":200},"175.6 tok\u002Fs",[214,545,546],{"align":200},[13,547,548],{},"328.5 tok\u002Fs",[10,550,551,554],{},[19,552,553],{},"Separate 54-request matrix, runs 403 \u002F 493 \u002F 492. Values from the published benchmark summary. This table is not the source of the 57% headline.",[44,555,556],{},[47,557,54],{"href":49,"ariaDescribedBy":558,"dataFootnoteRef":52,"id":559},[51],"user-content-fnref-1-11",[10,561,562,563,566,567,112,570],{},"At eight concurrent requests, aggregate throughput rises from ",[13,564,565],{},"33.7 tok\u002Fs on stock to 328.5 tok\u002Fs with Paiton",", nearly tenfold. Weighted serial decode in this shorter comparison rises from ",[13,568,569],{},"4.5 to 113.5 tok\u002Fs",[44,571,572],{},[47,573,54],{"href":49,"ariaDescribedBy":574,"dataFootnoteRef":52,"id":575},[51],"user-content-fnref-1-12",[10,577,578,579,582,583,586],{},"The scope matters: stock uses checkpoint-native ",[13,580,581],{},"W4A4 emulation",", while the accelerated configurations use ",[13,584,585],{},"W4A8 execution plus DFlash2",". These are complete engine configurations on the same weights, not identical activation arithmetic or a claim that replacing one kernel delivers the entire gain. Nor do these figures describe every stock-vLLM model or quantization.",[10,588,589],{},[176,590],{"alt":591,"src":592},"The separate 54-request prefill comparison shows Paiton ahead of stock vLLM O2 and Radiance at all three tested prompt depths.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F07-three-engine-prefill.webp",[10,594,595],{},[19,596,597],{},"Prefill in the separate capped matrix, with actual tokenized prompt lengths. Keep these values separate from the full-workload prefill series above.",[10,599,600,601],{},"That is why the main headline uses the harder comparison: ",[13,602,603],{},"Paiton against an already accelerated Radiance + DFlash2 baseline, confirmed on the longer workload.",[63,605,607],{"id":606},"regular-vllm-the-installed-library-stays-unchanged","Regular vLLM. The installed library stays unchanged.",[10,609,610,611,617],{},"The deployment result is deliberate. Paiton integrates through vLLM's extension mechanisms and supplies native HIP runtime artifacts for the optimized execution. vLLM's plugin system is designed to support extensions without modifying its codebase.",[44,612,613],{},[47,614,54],{"href":49,"ariaDescribedBy":615,"dataFootnoteRef":52,"id":616},[51],"user-content-fnref-1-13",[44,618,619],{},[47,620,250],{"href":621,"ariaDescribedBy":622,"dataFootnoteRef":52,"id":623},"#user-content-fn-3",[51],"user-content-fnref-3",[10,625,626,627,630,636],{},"Paiton's native HIP kernels and integrated DFlash2 support provide the optimized execution on the official vLLM ROCm runtime. The release supplies the runtime artifacts needed for the tested deployment, including the DFlash2 integration, so users can deploy the supported profile without a separate DFlash package. ",[13,628,629],{},"Our Paiton compiler remains proprietary.",[44,631,632],{},[47,633,54],{"href":49,"ariaDescribedBy":634,"dataFootnoteRef":52,"id":635},[51],"user-content-fnref-1-14",[44,637,638],{},[47,639,643],{"href":640,"ariaDescribedBy":641,"dataFootnoteRef":52,"id":642},"#user-content-fn-6",[51],"user-content-fnref-6","5",[10,645,646,647,651,652,655,656],{},"We checked the ordinary ",[648,649,650],"code",{},"vllm.entrypoints.openai.api_server"," entry point independently of the benchmark harness. It passed streaming chat, eight concurrent requests, a request at the configured context boundary and a fresh request afterward. ",[13,653,654],{},"All 2,893 installed vLLM files matched the official base image."," Validation applies to the pinned tested runtime and supported profile, not every vLLM feature or future version.",[44,657,658],{},[47,659,54],{"href":49,"ariaDescribedBy":660,"dataFootnoteRef":52,"id":661},[51],"user-content-fnref-1-15",[10,663,664],{},"Standard vLLM still carries its normal framework dependencies; Paiton's native libraries load independently of those frameworks. This is not a claim that the entire serving stack has become framework-free.",[10,666,667],{},"Completion and API checks are useful reliability checks. They are not an independent model-accuracy evaluation or a guarantee of identical generated text across execution profiles.",[63,669,671],{"id":670},"more-output-from-the-same-active-hour","More output from the same active hour",[10,673,674,675,677,678,681,682,677,684,687,688,112],{},"The full-workload concurrency-eight rates translate into a useful capacity illustration. Sustaining ",[13,676,34],{}," would produce about ",[13,679,680],{},"721,000 output tokens per active hour",". Sustaining ",[13,683,275],{},[13,685,686],{},"1.13 million",", roughly ",[13,689,690],{},"411,000 additional output tokens from the same hour",[10,692,693,694,697,698,701,702,112],{},"Equivalently, one million output tokens would take ",[13,695,696],{},"1.387 active hours"," at the baseline rate or ",[13,699,700],{},"0.883 hours"," at Paiton's rate: ",[13,703,704],{},"36.3% less active time",[10,706,707],{},[176,708],{"alt":709,"src":710},"At sustained measured rates, modeled active time per million output tokens falls from 1.387 hours to 0.883 hours: 36.3 percent less.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F08-active-time-per-million.webp",[10,712,713],{},[19,714,715],{},"Arithmetic illustration using the full-workload concurrency-eight rates. Assumes those rates are sustained; this is not an hour-long measurement, energy measurement or financial-cost claim.",[10,717,718],{},[13,719,720],{},"The hardware does not change. How much useful work it can deliver does.",[63,722,724],{"id":723},"tested-configuration","Tested configuration",[186,726,727,737],{},[189,728,729],{},[192,730,731,734],{},[195,732,733],{},"Component",[195,735,736],{},"Matched profile",[209,738,739,750,760,768,776,784,792,800,808,816,824,832],{},[192,740,741,744],{},[214,742,743],{},"GPU",[214,745,746,747],{},"One Radeon AI PRO R9700, ",[648,748,749],{},"gfx1201",[192,751,752,755],{},[214,753,754],{},"Target checkpoint",[214,756,757],{},[648,758,759],{},"amd\u002FQwen3.8-27B-Quark-AWQ-MXFP4",[192,761,762,765],{},[214,763,764],{},"Accelerated profiles",[214,766,767],{},"Radiance + DFlash2; Paiton on official vLLM + DFlash2",[192,769,770,773],{},[214,771,772],{},"Paiton serving base",[214,774,775],{},"Official vLLM 0.28 ROCm runtime",[192,777,778,781],{},[214,779,780],{},"Cache",[214,782,783],{},"FP8 KV; exactly 5 GiB reserved per engine",[192,785,786,789],{},[214,787,788],{},"Context ceiling",[214,790,791],{},"8,192 tokens per request",[192,793,794,797],{},[214,795,796],{},"Concurrent requests tested",[214,798,799],{},"1, 2, 4 and 8",[192,801,802,805],{},[214,803,804],{},"Speculation",[214,806,807],{},"Same target and draft snapshots; seven speculative tokens",[192,809,810,813],{},[214,811,812],{},"Sampling",[214,814,815],{},"Greedy",[192,817,818,821],{},[214,819,820],{},"Prefix caching",[214,822,823],{},"Disabled",[192,825,826,829],{},[214,827,828],{},"Headline workload",[214,830,831],{},"Full 188-request preset with original output budgets",[192,833,834,837],{},[214,835,836],{},"Supporting stock comparison",[214,838,839],{},"Separate 54-request matrix, output cap 128",[10,841,842,843],{},"These are results for the tested model, runtime and workload. They do not establish a guarantee for other GPUs, larger contexts, different drafters or untested concurrency levels.",[44,844,845],{},[47,846,54],{"href":49,"ariaDescribedBy":847,"dataFootnoteRef":52,"id":848},[51],"user-content-fnref-1-16",[63,850,852],{"id":851},"run-it-on-your-r9700","Run it on your R9700",[10,854,855,856,861,862,867,868,874],{},"The ",[47,857,860],{"href":858,"rel":859},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Ftree\u002Fmain\u002Fmodels\u002FQwen3.8-MXFP4-DFlash2",[106],"Paiton model guide"," is the starting point for the deployment instructions and benchmark evidence. The ",[47,863,866],{"href":864,"rel":865},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-27B-Quark-AWQ-MXFP4-DFlash2-Paiton-RDNA4",[106],"Hugging Face companion repository"," identifies the model release.",[44,869,870],{},[47,871,54],{"href":49,"ariaDescribedBy":872,"dataFootnoteRef":52,"id":873},[51],"user-content-fnref-1-17",[44,875,876],{},[47,877,881],{"href":878,"ariaDescribedBy":879,"dataFootnoteRef":52,"id":880},"#user-content-fn-5",[51],"user-content-fnref-5","6",[10,883,884,887,888,112,893],{},[13,885,886],{},"Runtime package:"," ",[47,889,892],{"href":890,"rel":891},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Freleases\u002Ftag\u002Fqwen38-mxfp4-dflash2-rdna4-v1.0.0",[106],"Get the runtime package and release notes",[44,894,895],{},[47,896,900],{"href":897,"ariaDescribedBy":898,"dataFootnoteRef":52,"id":899},"#user-content-fn-7",[51],"user-content-fnref-7","7",[10,902,903],{},"Container image for this release:",[905,906,911],"pre",{"className":907,"code":909,"language":910,"meta":52},[908],"language-text","ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin:qwen38-mxfp4-dflash2-rdna4-v1.0.0\n","text",[648,912,909],{"__ignoreMap":52},[10,914,915],{},"Use the model guide's pinned configuration to reproduce the result. The 5 GiB cache allocation, target and draft snapshots, speculative settings and context limit are part of the comparison, not incidental defaults.",[63,917,919],{"id":918},"same-hardware-a-much-stronger-serving-result","Same hardware. A much stronger serving result.",[10,921,922,925],{},[13,923,924],{},"57% more aggregate throughput. 97% lower median time to first token at eight concurrent requests. 2.89× estimated cache-token capacity."," Plus faster serial decode and prefill, all on one Radeon AI PRO R9700 through regular vLLM.",[10,927,928],{},"For local developers, that means a stronger shared endpoint from a single workstation GPU. For teams running AMD inference at scale, it is another demonstration of why execution efficiency matters alongside hardware capacity. The R9700 numbers are not a prediction of gains on other AMD platforms.",[10,930,931,932,936],{},"Running an AMD inference workload that should be delivering more? Talk to us about ",[47,933,935],{"href":934},"\u002Fproducts\u002Fpaiton","Paiton",". Bring the model, workload and current baseline.",[938,939,941],"h3",{"id":940},"credit-where-it-belongs","Credit where it belongs",[10,943,944,945,949],{},"Thank you to the ",[47,946,948],{"href":104,"rel":947},[106],"Radiance team"," for their excellent work on AMD inference and for providing the inspiration and a strong comparison point for this investigation.",[10,951,952,953,958,959,964,965,970],{},"We also thank ",[47,954,957],{"href":955,"rel":956},"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm",[106],"vLLM",", ",[47,960,963],{"href":961,"rel":962},"https:\u002F\u002Fcodeberg.org\u002FStillDeadcode\u002Flibr4d",[106],"StillDeadcode\u002Flibr4d",", the Qwen and DFlash2 teams, AMD's Quark checkpoint team and ",[47,966,969],{"href":967,"rel":968},"https:\u002F\u002Fgithub.com\u002FGGZ14\u002FBetterBench",[106],"BetterBench"," for the foundations, models and measurement tools that support this work.",[938,972,974],{"id":973},"sources-and-benchmark-references","Sources and benchmark references",[976,977,980,985],"section",{"className":978,"dataFootnotes":52},[979],"footnotes",[63,981,984],{"className":982,"id":51},[983],"sr-only","Footnotes",[986,987,988,1127,1140,1152,1166,1176,1188],"ol",{},[989,990,992,993,998,999,887,1006,887,1013,887,1020,887,1027,887,1034,887,1041,887,1048,887,1055,887,1063,887,1071,887,1079,887,1087,887,1095,887,1103,887,1111,887,1119],"li",{"id":991},"user-content-fn-1","ElioVP's supplied benchmark summary and charts; accompanying ",[47,994,997],{"href":995,"rel":996},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fmain\u002Fmodels\u002FQwen3.8-MXFP4-DFlash2\u002FBENCHMARKS.md",[106],"Paiton benchmark evidence",". Full-workload runs 495 \u002F 494; capped comparison runs 403 \u002F 493 \u002F 492. The figures in this article come from the supplied summary and charts rather than a raw request log. ",[47,1000,1005],{"href":1001,"ariaLabel":1002,"className":1003,"dataFootnoteBackref":52},"#user-content-fnref-1","Back to reference 1",[1004],"data-footnote-backref","↩",[47,1007,1005,1011],{"href":1008,"ariaLabel":1009,"className":1010,"dataFootnoteBackref":52},"#user-content-fnref-1-2","Back to reference 1-2",[1004],[44,1012,120],{},[47,1014,1005,1018],{"href":1015,"ariaLabel":1016,"className":1017,"dataFootnoteBackref":52},"#user-content-fnref-1-3","Back to reference 1-3",[1004],[44,1019,172],{},[47,1021,1005,1025],{"href":1022,"ariaLabel":1023,"className":1024,"dataFootnoteBackref":52},"#user-content-fnref-1-4","Back to reference 1-4",[1004],[44,1026,250],{},[47,1028,1005,1032],{"href":1029,"ariaLabel":1030,"className":1031,"dataFootnoteBackref":52},"#user-content-fnref-1-5","Back to reference 1-5",[1004],[44,1033,643],{},[47,1035,1005,1039],{"href":1036,"ariaLabel":1037,"className":1038,"dataFootnoteBackref":52},"#user-content-fnref-1-6","Back to reference 1-6",[1004],[44,1040,881],{},[47,1042,1005,1046],{"href":1043,"ariaLabel":1044,"className":1045,"dataFootnoteBackref":52},"#user-content-fnref-1-7","Back to reference 1-7",[1004],[44,1047,900],{},[47,1049,1005,1053],{"href":1050,"ariaLabel":1051,"className":1052,"dataFootnoteBackref":52},"#user-content-fnref-1-8","Back to reference 1-8",[1004],[44,1054,268],{},[47,1056,1005,1060],{"href":1057,"ariaLabel":1058,"className":1059,"dataFootnoteBackref":52},"#user-content-fnref-1-9","Back to reference 1-9",[1004],[44,1061,1062],{},"9",[47,1064,1005,1068],{"href":1065,"ariaLabel":1066,"className":1067,"dataFootnoteBackref":52},"#user-content-fnref-1-10","Back to reference 1-10",[1004],[44,1069,1070],{},"10",[47,1072,1005,1076],{"href":1073,"ariaLabel":1074,"className":1075,"dataFootnoteBackref":52},"#user-content-fnref-1-11","Back to reference 1-11",[1004],[44,1077,1078],{},"11",[47,1080,1005,1084],{"href":1081,"ariaLabel":1082,"className":1083,"dataFootnoteBackref":52},"#user-content-fnref-1-12","Back to reference 1-12",[1004],[44,1085,1086],{},"12",[47,1088,1005,1092],{"href":1089,"ariaLabel":1090,"className":1091,"dataFootnoteBackref":52},"#user-content-fnref-1-13","Back to reference 1-13",[1004],[44,1093,1094],{},"13",[47,1096,1005,1100],{"href":1097,"ariaLabel":1098,"className":1099,"dataFootnoteBackref":52},"#user-content-fnref-1-14","Back to reference 1-14",[1004],[44,1101,1102],{},"14",[47,1104,1005,1108],{"href":1105,"ariaLabel":1106,"className":1107,"dataFootnoteBackref":52},"#user-content-fnref-1-15","Back to reference 1-15",[1004],[44,1109,1110],{},"15",[47,1112,1005,1116],{"href":1113,"ariaLabel":1114,"className":1115,"dataFootnoteBackref":52},"#user-content-fnref-1-16","Back to reference 1-16",[1004],[44,1117,1118],{},"16",[47,1120,1005,1124],{"href":1121,"ariaLabel":1122,"className":1123,"dataFootnoteBackref":52},"#user-content-fnref-1-17","Back to reference 1-17",[1004],[44,1125,1126],{},"17",[989,1128,1130,1134,1135],{"id":1129},"user-content-fn-2",[47,1131,1133],{"href":104,"rel":1132},[106],"vLLM-Radiance project documentation",", including the same AMD Quark MXFP4 target, DFlash2 profile and explicitly dual-R9700 published measurements. Those external measurements are context, not the denominator of our headline. ",[47,1136,1005],{"href":1137,"ariaLabel":1138,"className":1139,"dataFootnoteBackref":52},"#user-content-fnref-2","Back to reference 2",[1004],[989,1141,1143,1146,1147],{"id":1142},"user-content-fn-4",[47,1144,969],{"href":967,"rel":1145},[106],". The request counts, selected workload settings and results above come from our supplied run summary. ",[47,1148,1005],{"href":1149,"ariaLabel":1150,"className":1151,"dataFootnoteBackref":52},"#user-content-fnref-4","Back to reference 3",[1004],[989,1153,1155,1160,1161],{"id":1154},"user-content-fn-3",[47,1156,1159],{"href":1157,"rel":1158},"https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Flatest\u002Fdesign\u002Fplugin_system\u002F",[106],"Official vLLM plugin-system documentation",". ",[47,1162,1005],{"href":1163,"ariaLabel":1164,"className":1165,"dataFootnoteBackref":52},"#user-content-fnref-3","Back to reference 4",[1004],[989,1167,1169,1160,1171],{"id":1168},"user-content-fn-6",[47,1170,935],{"href":934},[47,1172,1005],{"href":1173,"ariaLabel":1174,"className":1175,"dataFootnoteBackref":52},"#user-content-fnref-6","Back to reference 5",[1004],[989,1177,1179,1160,1183],{"id":1178},"user-content-fn-5",[47,1180,1182],{"href":864,"rel":1181},[106],"Paiton release companion on Hugging Face",[47,1184,1005],{"href":1185,"ariaLabel":1186,"className":1187,"dataFootnoteBackref":52},"#user-content-fnref-5","Back to reference 6",[1004],[989,1189,1191,1195,1196],{"id":1190},"user-content-fn-7",[47,1192,1194],{"href":890,"rel":1193},[106],"Paiton Qwen3.8 MXFP4 + DFlash2 runtime release",", including the runtime package and release notes. ",[47,1197,1005],{"href":1198,"ariaLabel":1199,"className":1200,"dataFootnoteBackref":52},"#user-content-fnref-7","Back to reference 7",[1004],{"title":52,"searchDepth":1202,"depth":1202,"links":1203},2,[1204,1205,1206,1207,1208,1209,1210,1211,1212,1213,1214,1215,1220],{"id":65,"depth":1202,"text":66},{"id":97,"depth":1202,"text":98},{"id":150,"depth":1202,"text":151},{"id":301,"depth":1202,"text":302},{"id":339,"depth":1202,"text":340},{"id":395,"depth":1202,"text":396},{"id":457,"depth":1202,"text":458},{"id":606,"depth":1202,"text":607},{"id":670,"depth":1202,"text":671},{"id":723,"depth":1202,"text":724},{"id":851,"depth":1202,"text":852},{"id":918,"depth":1202,"text":919,"children":1216},[1217,1219],{"id":940,"depth":1218,"text":941},3,{"id":973,"depth":1218,"text":974},{"id":51,"depth":1202,"text":984},[935,1222,957,1223,1224,1225],"AMD Radeon","Qwen3.8","Inference Optimization","DFlash2","2026-09-16T07:30:00Z","Paiton delivers 57% more Qwen3.8 throughput, 97% lower median TTFT at eight concurrent requests and 2.89× estimated cache-token capacity on one R9700.","md","57% more Qwen3.8 throughput. Same Radeon. Regular vLLM.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002F00-hero.webp",{},"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",{"title":5,"description":1227},"paiton-qwen38-mxfp4-dflash2-r9700","blog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","D7vdcpM101QAnNF2mDxlp1a443xduvciOPgQSWs1Tmo",[1239,1241,1252,1264,1274,1289,1298,1313,1344,1356,1378,1396,1415,1433,1451,1468,1480,1496,1511,1523,1532,1540,1555,1567,1578,1589,1599,1612,1622,1635,1646,1656,1667,1676,1688,1699,1708],{"path":1233,"title":5,"description":1227,"date":1226,"slug":1235,"image":1230,"originalUrl":1232,"categories":1240},[935,1222,957,1223,1224,1225],{"path":1242,"title":1243,"description":1244,"date":1245,"slug":1246,"image":1247,"originalUrl":1248,"categories":1249},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[935,1222,1250,1251,957],"Local AI","GGUF",{"path":1253,"title":1254,"description":1255,"date":1256,"slug":1257,"image":1258,"originalUrl":1259,"categories":1260},"\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","MiniMax H3 on Radeon: 15-Second Video With Native Sound","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","2026-09-09T07:30:00Z","paiton-minimax-h3-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",[935,1222,1250,1261,1262,1263],"Video Generation","MiniMax H3","ComfyUI",{"path":1265,"title":1266,"description":1267,"date":1268,"slug":1269,"image":1270,"originalUrl":1265,"categories":1271},"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","2026-09-07T09:00:00","paiton-flux2-klein-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",[935,1222,1250,1272,1273,1263],"Image Generation","FLUX",{"path":1275,"title":1276,"description":1277,"date":1278,"slug":1279,"image":1280,"originalUrl":1281,"categories":1282},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[935,1283,1222,1284,1285,1286,1224,1287,957,1288],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Large Language Models","Cost Efficiency",{"path":1290,"title":1291,"description":1292,"date":1293,"slug":1294,"image":1295,"originalUrl":1296,"categories":1297},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[935,1283,1222,1284,1285,1286,1224,1287,957,1288],{"path":1299,"title":1300,"description":1301,"date":1302,"slug":1303,"image":1304,"originalUrl":1305,"categories":1306},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",null,[1307,1308,1309,1310,1311,1312],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1314,"title":1315,"description":1316,"date":1317,"slug":1318,"image":1319,"originalUrl":1320,"categories":1321},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1307,1283,935,1322,1323,1324,1325,1326,1327,1328,1329,1330,743,1331,1332,1333,1334,1335,1336,1337,935,1338,1339,1340,1341,1342,1343],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1345,"title":1346,"description":1347,"date":1348,"slug":1349,"image":1350,"originalUrl":1351,"categories":1352},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1307,1283,1353,1354,1323,1355,1353,1335],"Modular DC","Uncategorized","De Tijd",{"path":1357,"title":1358,"description":1359,"date":1360,"slug":1361,"image":1362,"originalUrl":1363,"categories":1364},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1307,1283,1365,1354,1366,1367,1368,1369,1370,1371,1372,1373,1329,1374,1330,1375,1376,1377],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":1379,"title":1380,"description":1381,"date":1382,"slug":1383,"image":1384,"originalUrl":1385,"categories":1386},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1307,1387,1388,1389,1371,1390,1391,1392,1374,1393,1394,1395,1376],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1397,"title":1398,"description":1399,"date":1400,"slug":1401,"image":1402,"originalUrl":1403,"categories":1404},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1307,1283,1405,1365,1406,1407,1408,1409,1410,1411,1412,1413,1338,1414],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1416,"title":1417,"description":1418,"date":1419,"slug":1420,"image":1421,"originalUrl":1422,"categories":1423},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1307,1283,1365,1308,1424,1425,1426,1427,1428,1429,1430,1431,1432],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1434,"title":1435,"description":1436,"date":1437,"slug":1438,"image":1439,"originalUrl":1440,"categories":1441},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1307,1353,1354,1442,1308,1443,1444,1445,1446,1447,1448,1449,1450],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1452,"title":1453,"description":1454,"date":1455,"slug":1456,"image":1457,"originalUrl":1458,"categories":1459},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1307,1283,935,1354,1460,1283,1461,1462,1463,1464,1465,1466,935,1467],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning","ROCm",{"path":1469,"title":1470,"description":1471,"date":1472,"slug":1473,"image":1474,"originalUrl":1475,"categories":1476},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1307,1283,935,1284,1477,1460,1288,1478,1224,1466,935,1479,957],"AMD Instinct","High Throughput","SGLang",{"path":1481,"title":1482,"description":1483,"date":1484,"slug":1485,"image":1486,"originalUrl":1487,"categories":1488},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1307,1283,935,1489,1460,1490,1224,1491,1492,1493,1494,935,1495],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1497,"title":1498,"description":1499,"date":1500,"slug":1501,"image":1502,"originalUrl":1503,"categories":1504},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1307,1283,1405,1354,1406,1505,1506,1507,1508,1411,1413,1338,1509,1510],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1512,"title":1513,"description":1514,"date":1515,"slug":1516,"image":1517,"originalUrl":1518,"categories":1519},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[1307,1283,935,1354,1520,1323,1324,1521,1522,1334,1335,935,957],"AI","H200","MI300X",{"path":1524,"title":1525,"description":1526,"date":1527,"slug":1528,"image":1529,"originalUrl":1530,"categories":1531},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1307,1283,1405,1406,1505,1506,1507,1508,1411,1413,1338,1509,1510],{"path":1533,"title":1534,"description":1535,"date":1536,"slug":1537,"image":52,"originalUrl":1538,"categories":1539},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1307,1283,935],{"path":1541,"title":1542,"description":1543,"date":1544,"slug":1545,"image":1546,"originalUrl":1547,"categories":1548},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1307,1283,935,1354,1284,1460,1549,1462,1550,1551,1552,935,1553,1554],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1556,"title":1557,"description":1558,"date":1559,"slug":1560,"image":1561,"originalUrl":1562,"categories":1563},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[1307,1283,935,1354,1460,1564,1410,1330,1285,1286,1287,1551,1565,1566],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1568,"title":1569,"description":1570,"date":1571,"slug":1572,"image":1573,"originalUrl":1574,"categories":1575},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1307,1283,935,1405,1354,1323,1522,1335,1576,1577],"RX7900XTX","tenstorrent",{"path":1579,"title":1580,"description":1581,"date":1582,"slug":1583,"image":1584,"originalUrl":1585,"categories":1586},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1307,1283,1353,1405,1324,1521,1587,1335,1588],"MI325x","pnl calculator",{"path":1590,"title":1591,"description":1592,"date":1593,"slug":1594,"image":1595,"originalUrl":1596,"categories":1597},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1307,1283,935,1520,1323,1521,1598,1335,935,957],"MI300",{"path":1600,"title":1601,"description":1602,"date":1603,"slug":1604,"image":1605,"originalUrl":1606,"categories":1607},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1307,1353,1608,1308,1444,1311,1445,1446,1609,1610,1449,1611],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1613,"title":1614,"description":1615,"date":1616,"slug":1617,"image":1618,"originalUrl":1619,"categories":1620},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1307,1283,1405,1354,1520,1323,1621],"Healthcare",{"path":1623,"title":1624,"description":1625,"date":1626,"slug":1627,"image":1628,"originalUrl":1629,"categories":1630},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1307,1365,1520,1323,1631,1632,1633,1634],"import","Taiwan","Tariffs","Trump",{"path":1636,"title":1637,"description":1638,"date":1639,"slug":1640,"image":1641,"originalUrl":1642,"categories":1643},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1307,1283,1405,1520,1644,1645],"AI Agents","ERP",{"path":1647,"title":1648,"description":1649,"date":1650,"slug":1651,"image":1652,"originalUrl":1653,"categories":1654},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1307,1283,1365,1655,1323,743,1335],"AI news",{"path":1657,"title":1658,"description":1659,"date":1660,"slug":1661,"image":1662,"originalUrl":1663,"categories":1664},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1307,1283,935,1520,1323,1665,1666,1522,935],"benchmark","LLM",{"path":1668,"title":1669,"description":1670,"date":1671,"slug":1672,"image":1673,"originalUrl":1674,"categories":1675},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1307,1283,935],{"path":1677,"title":1678,"description":1679,"date":1680,"slug":1681,"image":1682,"originalUrl":1683,"categories":1684},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1307,1323,1685,1686,1687],"Jim Greene","Podcast","Tech Talk",{"path":1689,"title":1690,"description":1691,"date":1692,"slug":1693,"image":1694,"originalUrl":1695,"categories":1696},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1307,1283,935,1323,1697,1698,1521,1522,1587,935,957],"Deepseek","H100",{"path":1700,"title":1701,"description":1702,"date":1703,"slug":1704,"image":1705,"originalUrl":1706,"categories":1707},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1307,1283,935,1323,1697,1698,1521,1522,1587,935,957],{"path":1709,"title":1710,"description":1711,"date":1712,"slug":1713,"image":1714,"originalUrl":1715,"categories":1716},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1307,1283,935,1323,1698,1521,1522,1587,935,957],1789564927915]