[{"data":1,"prerenderedAt":2534},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-qwen38-flash-next-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":2025},{"id":4,"title":5,"body":6,"categories":2006,"date":2011,"description":2012,"extension":2013,"heading":2014,"image":2015,"meta":2016,"navigation":1273,"originalUrl":2017,"path":2018,"seo":2019,"slug":2020,"socialImage":2021,"stem":2022,"updated":2023,"__hash__":2024},"blog\u002Fblog\u002Fpaiton-qwen38-flash-next-radeon-ai-pro-r9700.md","Qwen3.8 Flash Next on 2 × Radeon AI PRO R9700: 216 tokens\u002Fs, 200K context",{"type":7,"value":8,"toc":1983},"minimark",[9,42,69,76,81,97,101,107,204,221,226,229,234,278,291,295,310,313,318,380,400,404,448,466,479,483,495,498,503,540,556,560,563,584,603,620,624,641,672,676,693,696,701,748,780,784,804,880,885,902,938,951,957,961,979,983,1001,1026,1082,1086,1111,1126,1153,1163,1167,1192,1215,1397,1416,1510,1538,1541,1577,1590,1594,1606,1609,1613,1616,1629,1648,1979],[10,11,12,13,17,18,21,22,25,26,29,30],"p",{},"A local assistant should be able to write at a useful speed, work through a long document and serve more than one person. Our latest Paiton release brings ",[14,15,16],"strong",{},"Qwen3.8 Flash Next to two Radeon AI PRO R9700 cards",", with ",[14,19,20],{},"216.3 tokens per second in the reported weighted decode score",", ",[14,23,24],{},"565.1 tokens per second across eight concurrent requests",", and a separate ",[14,27,28],{},"200,000-token context mode",".",[31,32,33],"sup",{},[34,35,41],"a",{"href":36,"ariaDescribedBy":37,"dataFootnoteRef":39,"id":40},"#user-content-fn-bench",[38],"footnote-label","","user-content-fnref-bench","1",[10,43,44,45,48,49,52,53,61],{},"The hardware is two cards with ",[14,46,47],{},"32 GB of VRAM each",". The large routed experts use 3-bit weights, while more sensitive parts retain higher precision. The model runs through ",[14,50,51],{},"regular vLLM 0.29, the Paiton plugin and native AMD kernels",", not a separate serving engine.",[31,54,55],{},[34,56,60],{"href":57,"ariaDescribedBy":58,"dataFootnoteRef":39,"id":59},"#user-content-fn-release",[38],"user-content-fnref-release","2",[31,62,63],{},[34,64,68],{"href":65,"ariaDescribedBy":66,"dataFootnoteRef":39,"id":67},"#user-content-fn-model",[38],"user-content-fnref-model","3",[10,70,71,72,75],{},"There are important qualifications. The 200K mode generates at about 100 tokens\u002Fs rather than the default mode's roughly 200, because speculative decoding is off. A large n-gram embedding table remains in system RAM. Quality checks cover agreement with the BF16 model and bounded long-context retrieval, ",[14,73,74],{},"not broad math, coding, knowledge, multilingual or tool-calling benchmarks",". Here is what we measured, what the numbers mean and how to try it.",[77,78,80],"h2",{"id":79},"fast-generation-on-two-workstation-cards","Fast generation on two workstation cards",[10,82,83,86,87,90,91],{},[14,84,85],{},"Decode"," measures generation after the first token. In the release-image validation run on 10 October, BetterBench 0.6.0's standard profile reported a weighted score of ",[14,88,89],{},"216.3 tokens\u002Fs",". Individual workloads ranged from 183.3 for prose to 259.2 for JSON.",[31,92,93],{},[34,94,41],{"href":36,"ariaDescribedBy":95,"dataFootnoteRef":39,"id":96},[38],"user-content-fnref-bench-2",[98,99],"flash-next-figure",{"kind":100},"decode",[10,102,103],{},[104,105,106],"em",{},"Release-image validation, two R9700 cards, 98,304-token configured context and speculative decoding enabled. Generated tokens\u002Fs; higher is better. The weighted score is not an equal-weight average of the eight rows; the workload weights are explained below. These are measurements of this model and configuration, not a comparison with our earlier 27B releases.",[108,109,110,124],"table",{},[111,112,113],"thead",{},[114,115,116,120],"tr",{},[117,118,119],"th",{},"Workload",[117,121,123],{"align":122},"right","Generated tokens\u002Fs",[125,126,127,136,144,152,160,168,176,184,192],"tbody",{},[114,128,129,133],{},[130,131,132],"td",{},"Chat",[130,134,135],{"align":122},"202.1",[114,137,138,141],{},[130,139,140],{},"Code",[130,142,143],{"align":122},"214.3",[114,145,146,149],{},[130,147,148],{},"File edit",[130,150,151],{"align":122},"236.0",[114,153,154,157],{},[130,155,156],{},"JSON",[130,158,159],{"align":122},"259.2",[114,161,162,165],{},[130,163,164],{},"Math",[130,166,167],{"align":122},"255.5",[114,169,170,173],{},[130,171,172],{},"Prose",[130,174,175],{"align":122},"183.3",[114,177,178,181],{},[130,179,180],{},"Reasoning",[130,182,183],{"align":122},"190.1",[114,185,186,189],{},[130,187,188],{},"Summarization",[130,190,191],{"align":122},"239.8",[114,193,194,199],{},[130,195,196],{},[14,197,198],{},"Reported weighted score",[130,200,201],{"align":122},[14,202,203],{},"216.3",[10,205,206,207,210,211,214,215],{},"The short-prompt median ",[14,208,209],{},"time to first token was 117 ms",". The stream-update p99 was ",[14,212,213],{},"13.9 ms",", but that is the gap between updates, not a per-token latency: speculation can deliver several accepted tokens in one update.",[31,216,217],{},[34,218,41],{"href":36,"ariaDescribedBy":219,"dataFootnoteRef":39,"id":220},[38],"user-content-fnref-bench-3",[222,223,225],"h3",{"id":224},"more-output-when-requests-overlap","More output when requests overlap",[98,227],{"kind":228},"concurrency",[10,230,231],{},[104,232,233],{},"Aggregate generation throughput across simultaneous requests; higher is better. BetterBench ran 48 requests at each concurrency level. These are combined server rates, not the speed each user receives. The single-request concurrency result is a different workload from the weighted decode score above.",[108,235,236,246],{},[111,237,238],{},[114,239,240,243],{},[117,241,242],{},"Concurrent requests",[117,244,245],{"align":122},"Aggregate tokens\u002Fs",[125,247,248,255,262,270],{},[114,249,250,252],{},[130,251,41],{},[130,253,254],{"align":122},"204.6",[114,256,257,259],{},[130,258,60],{},[130,260,261],{"align":122},"315.1",[114,263,264,267],{},[130,265,266],{},"4",[130,268,269],{"align":122},"449.2",[114,271,272,275],{},[130,273,274],{},"8",[130,276,277],{"align":122},"565.1",[10,279,280,281,284,285],{},"All ",[14,282,283],{},"48 of 48 requests"," completed at the eight-request level. That is useful for a shared local assistant, but eight overlapping benchmark requests do not mean eight full-length 98K conversations fit at once. Prompts, generated output and per-request state all use the available memory.",[31,286,287],{},[34,288,41],{"href":36,"ariaDescribedBy":289,"dataFootnoteRef":39,"id":290},[38],"user-content-fnref-bench-4",[77,292,294],{"id":293},"long-prompts-without-a-steep-throughput-drop","Long prompts without a steep throughput drop",[10,296,297,300,301,29,304],{},[14,298,299],{},"Prefill"," is processing the input before generation begins. The separate long-context mode, with speculation disabled, processed the tested prompt depths at roughly ",[14,302,303],{},"7,900 to 8,300 tokens\u002Fs",[31,305,306],{},[34,307,41],{"href":36,"ariaDescribedBy":308,"dataFootnoteRef":39,"id":309},[38],"user-content-fnref-bench-5",[98,311],{"kind":312},"prefill",[10,314,315],{},[104,316,317],{},"Input tokens processed per second; higher is better. BetterBench standard profile with an extended 128K sweep, two R9700 cards, 200,000-token configured context, speculation off and 2,048-token prefill chunks. Depth labels are nominal settings, not exact prompt-token counts. These results come from the long mode, not the speculative decode configuration.",[108,319,320,330],{},[111,321,322],{},[114,323,324,327],{},[117,325,326],{},"BetterBench prompt-depth setting",[117,328,329],{"align":122},"Input tokens\u002Fs",[125,331,332,340,348,356,364,372],{},[114,333,334,337],{},[130,335,336],{},"2K",[130,338,339],{"align":122},"7,925",[114,341,342,345],{},[130,343,344],{},"8K",[130,346,347],{"align":122},"8,321",[114,349,350,353],{},[130,351,352],{},"16K",[130,354,355],{"align":122},"8,311",[114,357,358,361],{},[130,359,360],{},"32K",[130,362,363],{"align":122},"8,232",[114,365,366,369],{},[130,367,368],{},"64K",[130,370,371],{"align":122},"8,161",[114,373,374,377],{},[130,375,376],{},"128K",[130,378,379],{"align":122},"7,928",[10,381,382,383,386,387,390,391,29,394],{},"The 128K result remains close to the 8K result. A separate cold-cache probe with ",[14,384,385],{},"exactly 64,000 prompt tokens"," measured ",[14,388,389],{},"8,212 tokens\u002Fs"," across the whole prompt, or 8,204 in steady state after the first chunk. Prompt throughput is not the same as the complete wait for an answer: a fresh ",[14,392,393],{},"190K prompt took 24.8 seconds to its first token",[31,395,396],{},[34,397,41],{"href":36,"ariaDescribedBy":398,"dataFootnoteRef":39,"id":399},[38],"user-content-fnref-bench-6",[222,401,403],{"id":402},"choose-the-mode-that-matches-the-job","Choose the mode that matches the job",[108,405,406,419],{},[111,407,408],{},[114,409,410,413,416],{},[117,411,412],{},"Mode",[117,414,415],{"align":122},"Configured context window",[117,417,418],{},"Speculative decoding",[125,420,421,435],{},[114,422,423,429,432],{},[130,424,425,428],{},[426,427,100],"code",{},", default",[130,430,431],{"align":122},"98,304 tokens",[130,433,434],{},"On, depth 3",[114,436,437,442,445],{},[130,438,439],{},[426,440,441],{},"prefill-long",[130,443,444],{"align":122},"200,000 tokens",[130,446,447],{},"Off",[10,449,450,451,29,454,460],{},"These are total context windows: the prompt, chat formatting and generated answer must fit together. The checkpoint's native window is 262,144 tokens, but ",[14,452,453],{},"this release does not establish that full window on these two cards",[31,455,456],{},[34,457,68],{"href":65,"ariaDescribedBy":458,"dataFootnoteRef":39,"id":459},[38],"user-content-fnref-model-2",[31,461,462],{},[34,463,60],{"href":57,"ariaDescribedBy":464,"dataFootnoteRef":39,"id":465},[38],"user-content-fnref-release-2",[10,467,468,469,472,473],{},"The 200K mode needs the memory otherwise used by the speculative drafter's per-request state. Without speculation, single-stream generation after 32K, 100K and 190K prompts measured ",[14,470,471],{},"105.3, 100.6 and 99.9 tokens\u002Fs",", respectively, using 256-token continuations. Long context therefore remains usable, with a different speed tradeoff from the default mode.",[31,474,475],{},[34,476,41],{"href":36,"ariaDescribedBy":477,"dataFootnoteRef":39,"id":478},[38],"user-content-fnref-bench-7",[77,480,482],{"id":481},"reuse-a-long-prompt-with-optional-prefix-caching","Reuse a long prompt with optional prefix caching",[10,484,485,486,29,489],{},"If you repeatedly ask questions about the same document, processing the unchanged prefix again can dominate the wait. The opt-in prefix cache stores attention-cache and recurrent-state checkpoints at aligned ",[14,487,488],{},"2,048-token boundaries",[31,490,491],{},[34,492,60],{"href":57,"ariaDescribedBy":493,"dataFootnoteRef":39,"id":494},[38],"user-content-fnref-release-3",[98,496],{"kind":497},"prefix-cache",[10,499,500],{},[104,501,502],{},"Separate repeated-prompt checks in long-context mode. Time to first token in seconds; lower is better. A cache hit reuses an unchanged prefix still present in the cache. These are not the cold-cache BetterBench measurements above.",[108,504,505,518],{},[111,506,507],{},[114,508,509,512,515],{},[117,510,511],{},"Repeated prompt",[117,513,514],{"align":122},"Cold time to first token",[117,516,517],{"align":122},"Prefix-cache hit",[125,519,520,530],{},[114,521,522,524,527],{},[130,523,368],{},[130,525,526],{"align":122},"9.6 s",[130,528,529],{"align":122},"0.35 s",[114,531,532,534,537],{},[130,533,376],{},[130,535,536],{"align":122},"20 s",[130,538,539],{"align":122},"0.41 s",[10,541,542,543,546,547,29,550],{},"In these checks, the hit produced ",[14,544,545],{},"byte-identical output to the uncached run",". That is a result for the tested cached and uncached requests, not a promise of identical output across all sampling settings. Prefix caching is optional, and ",[14,548,549],{},"the headline throughput benchmarks were measured with it off",[31,551,552],{},[34,553,41],{"href":36,"ariaDescribedBy":554,"dataFootnoteRef":39,"id":555},[38],"user-content-fnref-bench-8",[77,557,559],{"id":558},"what-3-bit-means-here","What 3-bit means here",[10,561,562],{},"This is a mixed-precision model, not a claim that every tensor uses three bits.",[10,564,565,566,569,570,573,574,577,578],{},"The ",[14,567,568],{},"routed experts",", which contain 120.8 billion weights, use ",[14,571,572],{},"3.125 bits per weight including group scales",". Those experts occupy about ",[14,575,576],{},"47.2 GB"," in total. A rotated input basis helps the low-precision representation, and the native kernels read the packed weights directly rather than first expanding the entire model to a larger format. The W3A8 name refers to the 3-bit expert weights and 8-bit expert activations.",[31,579,580],{},[34,581,68],{"href":65,"ariaDescribedBy":582,"dataFootnoteRef":39,"id":583},[38],"user-content-fnref-model-3",[10,585,586,587,590,591,597],{},"The main non-expert projections use 8-bit weights. The speculative MTP layer uses 4-bit expert weights and a 2-bit draft head, with other tensors retaining higher precision. Routers, normalization and the retained vision tower also have their own precision choices. The public model card records the tensor formats; ",[14,588,589],{},"the validated serving modes are text-only",", even though the repository contains the vision tower.",[31,592,593],{},[34,594,68],{"href":65,"ariaDescribedBy":595,"dataFootnoteRef":39,"id":596},[38],"user-content-fnref-model-4",[31,598,599],{},[34,600,60],{"href":57,"ariaDescribedBy":601,"dataFootnoteRef":39,"id":602},[38],"user-content-fnref-release-4",[10,604,605,606,609,610,613,614],{},"The weights come from our own sequential GPTQ calibration on ",[14,607,608],{},"two million tokens"," of prose, code and assistant conversations. We tested two calibration replicates and separated selection from confirmation data. Their confirmation KL results, ",[14,611,612],{},"0.0630 and 0.0686",", differed enough that small gaps between candidate formats should not be overinterpreted.",[31,615,616],{},[34,617,68],{"href":65,"ariaDescribedBy":618,"dataFootnoteRef":39,"id":619},[38],"user-content-fnref-model-5",[222,621,623],{"id":622},"where-the-memory-goes","Where the memory goes",[10,625,626,627,630,631,634,635],{},"The large decoder and expert weights remain on the GPUs, split across two tensor-parallel ranks. Each card holds ",[14,628,629],{},"24.9 GiB of text weights plus 0.7 GiB for the MTP layer",". Both benchmark configurations peaked at ",[14,632,633],{},"30.5 to 31.5 GiB per card",", including cache, recurrent state and execution graphs.",[31,636,637],{},[34,638,68],{"href":65,"ariaDescribedBy":639,"dataFootnoteRef":39,"id":640},[38],"user-content-fnref-model-6",[10,642,643,644,647,648,651,652,655,656,659,660,666],{},"There is ",[14,645,646],{},"no expert offload",", but this is not a GPU-only setup. The ",[14,649,650],{},"48.9 GiB n-gram table stays pinned in system RAM",", split across the ranks, and the runtime reads 16 of its rows per token. The test host had ",[14,653,654],{},"251 GiB of system RAM",". The checkpoint repository occupies ",[14,657,658],{},"108 GiB on disk",", before allowing room for the container and runtime cache.",[31,661,662],{},[34,663,68],{"href":65,"ariaDescribedBy":664,"dataFootnoteRef":39,"id":665},[38],"user-content-fnref-model-7",[31,667,668],{},[34,669,60],{"href":57,"ariaDescribedBy":670,"dataFootnoteRef":39,"id":671},[38],"user-content-fnref-release-5",[77,673,675],{"id":674},"quality-close-distributions-are-not-task-benchmark-scores","Quality: close distributions are not task benchmark scores",[10,677,678,679,682,683,686,687],{},"We evaluated how much the quantized model's full-vocabulary next-token distribution differs from BF16. ",[14,680,681],{},"KL divergence is lower-is-better","; zero would mean identical distributions. ",[14,684,685],{},"Top-1 agreement"," is how often both versions prefer the same next token. It is not the percentage of math problems or coding tasks solved.",[31,688,689],{},[34,690,68],{"href":65,"ariaDescribedBy":691,"dataFootnoteRef":39,"id":692},[38],"user-content-fnref-model-8",[98,694],{"kind":695},"quality",[10,697,698],{},[104,699,700],{},"Text and assistant positions on the same evaluation corpus. Free routing lets each model choose experts; locked routing forces BF16's expert choices to separate arithmetic error from routing changes. KL and top-1 agreement measure different things and have different units. The BF16 control shows variation between two correct implementations.",[108,702,703,718],{},[111,704,705],{},[114,706,707,710,713,716],{},[117,708,709],{},"Evaluation",[117,711,712],{"align":122},"KL to BF16, free routing",[117,714,715],{"align":122},"KL, BF16 routing forced",[117,717,685],{"align":122},[125,719,720,734],{},[114,721,722,725,728,731],{},[130,723,724],{},"BF16 against BF16, implementation floor",[130,726,727],{"align":122},"0.0077",[130,729,730],{"align":122},"Not applicable",[130,732,733],{"align":122},"96.77%",[114,735,736,739,742,745],{},[130,737,738],{},"This 3-bit model",[130,740,741],{"align":122},"0.0604",[130,743,744],{"align":122},"0.0413",[130,746,747],{"align":122},"90.74%",[10,749,750,751,754,755,758,759,762,763,766,767,773],{},"The corpus contains ",[14,752,753],{},"294,912 positions",", including ",[14,756,757],{},"170,884 text and assistant positions"," used for the headline metrics. Separately, the release container passed its runtime served-KL gate at ",[14,760,761],{},"0.0616",", against a stated budget of 0.0600 plus 0.002. That gate uses top-64 bucket KL on the confirmation half, so it is ",[14,764,765],{},"not the same full-vocabulary metric as the table above",". It validates the running server rather than only an emulation.",[31,768,769],{},[34,770,68],{"href":65,"ariaDescribedBy":771,"dataFootnoteRef":39,"id":772},[38],"user-content-fnref-model-9",[31,774,775],{},[34,776,266],{"href":777,"ariaDescribedBy":778,"dataFootnoteRef":39,"id":779},"#user-content-fn-quality",[38],"user-content-fnref-quality",[222,781,783],{"id":782},"long-documents-and-bounded-retrieval","Long documents and bounded retrieval",[10,785,786,787,790,791,794,795,29,798],{},"A held-out set of ",[14,788,789],{},"44 documents from 8K to 32K tokens",", totaling 524,288 positions, provides a separate view of error with document depth. The document-level text-and-assistant result is ",[14,792,793],{},"0.0561 KL with free routing",", 0.0408 with locked routing and ",[14,796,797],{},"90.39% top-1 agreement",[31,799,800],{},[34,801,68],{"href":65,"ariaDescribedBy":802,"dataFootnoteRef":39,"id":803},[38],"user-content-fnref-model-10",[108,805,806,822],{},[111,807,808],{},[114,809,810,813,816,819],{},[117,811,812],{},"Position in document",[117,814,815],{"align":122},"KL, free routing",[117,817,818],{"align":122},"KL, routing locked",[117,820,821],{"align":122},"Text\u002Fassistant top-1 agreement",[125,823,824,838,852,866],{},[114,825,826,829,832,835],{},[130,827,828],{},"Below 2K",[130,830,831],{"align":122},"0.1554",[130,833,834],{"align":122},"0.1043",[130,836,837],{"align":122},"89.46%",[114,839,840,843,846,849],{},[130,841,842],{},"2K–4K",[130,844,845],{"align":122},"0.6300",[130,847,848],{"align":122},"0.2931",[130,850,851],{"align":122},"89.66%",[114,853,854,857,860,863],{},[130,855,856],{},"4K–8K",[130,858,859],{"align":122},"0.6408",[130,861,862],{"align":122},"0.2984",[130,864,865],{"align":122},"90.06%",[114,867,868,871,874,877],{},[130,869,870],{},"8K and beyond",[130,872,873],{"align":122},"0.4333",[130,875,876],{"align":122},"0.2150",[130,878,879],{"align":122},"91.65%",[10,881,882],{},[104,883,884],{},"The bucket KL values above cover all positions, while their top-1 values cover text and assistant positions. They should not be compared directly with the text-and-assistant headline KL. The final bucket contains 12 documents; the other buckets contain 44.",[10,886,887,888,891,892,895,896],{},"On text and assistant positions, document-level negative log-likelihood was ",[14,889,890],{},"1.7574 nats\u002Ftoken for BF16 and 1.7671 for this model",", a paired difference of ",[14,893,894],{},"+0.0097",". This is a small measured loss on that held-out set, not proof of identical task quality.",[31,897,898],{},[34,899,68],{"href":65,"ariaDescribedBy":900,"dataFootnoteRef":39,"id":901},[38],"user-content-fnref-model-11",[108,903,904,917],{},[111,905,906],{},[114,907,908,911,914],{},[117,909,910],{},"Served-stack retrieval check",[117,912,913],{"align":122},"This model",[117,915,916],{"align":122},"BF16 control",[125,918,919,929],{},[114,920,921,924,927],{},[130,922,923],{},"Needle retrieval at 131,072 tokens",[130,925,926],{"align":122},"40 \u002F 40",[130,928,926],{"align":122},[114,930,931,934,936],{},[130,932,933],{},"Needle retrieval at 200,000 tokens",[130,935,926],{"align":122},[130,937,926],{"align":122},[10,939,940,941,944,945],{},"Both versions returned the same answers in these tests. The drafter also remained close in a teacher-forced greedy check at depth 3: ",[14,942,943],{},"1.768 expected accepted drafts per step",", against 1.789 for the BF16 MTP layer on the BF16 model. Neither the retrieval checks nor draft acceptance establish general reasoning quality.",[31,946,947],{},[34,948,68],{"href":65,"ariaDescribedBy":949,"dataFootnoteRef":39,"id":950},[38],"user-content-fnref-model-12",[10,952,953,956],{},[14,954,955],{},"Math, code, general knowledge, multilingual quality and tool calling were not task-benchmarked for this release. Image input is not validated."," Test the workload you intend to deploy rather than treating distribution agreement as a substitute for that evaluation.",[77,958,960],{"id":959},"the-runtime-without-a-custom-engine","The runtime, without a custom engine",[10,962,963,964,29,967,973],{},"vLLM supplies the serving framework and OpenAI-compatible API. Paiton's native kernels handle the packed experts, tensor-parallel execution and model-specific operations. Speculative decoding drafts three tokens and verifies them together; the runtime also handles the recurrent layers' state when a draft is rejected. The release run accepted a median of ",[14,965,966],{},"2.85 tokens per stream update",[31,968,969],{},[34,970,60],{"href":57,"ariaDescribedBy":971,"dataFootnoteRef":39,"id":972},[38],"user-content-fnref-release-6",[31,974,975],{},[34,976,68],{"href":65,"ariaDescribedBy":977,"dataFootnoteRef":39,"id":978},[38],"user-content-fnref-model-13",[222,980,982],{"id":981},"a-faster-prefill-path-or-reproducible-arithmetic","A faster prefill path, or reproducible arithmetic",[10,984,985,986,29,989,995],{},"The default fast prefill path uses 8-bit weights and 8-bit activations in the prompt-processing trunk. That changes the arithmetic precision on prompt rows. It passed the served-KL budget and the retrieval checks above, but the long-document control found it ",[14,987,988],{},"not bit-reproducible between runs",[31,990,991],{},[34,992,60],{"href":57,"ariaDescribedBy":993,"dataFootnoteRef":39,"id":994},[38],"user-content-fnref-release-7",[31,996,997],{},[34,998,41],{"href":36,"ariaDescribedBy":999,"dataFootnoteRef":39,"id":1000},[38],"user-content-fnref-bench-9",[10,1002,1003,1004,1007,1008,1011,1012,1015,1016,1019,1020],{},"The publicly labelled ",[14,1005,1006],{},"exact Gated DeltaNet prefill path"," remains available; it retains the other quantization choices rather than turning all prompt arithmetic into BF16. A ",[14,1009,1010],{},"development-stack"," measurement returned ",[14,1013,1014],{},"7,607 tokens\u002Fs at 64K",", versus 8,161 for the release image's fast path. The exact path was bit-reproducible in its checks; ",[14,1017,1018],{},"it was not rerun on the release image",", so this is not a same-image benchmark comparison. Decode arithmetic is unchanged by that choice.",[31,1021,1022],{},[34,1023,41],{"href":36,"ariaDescribedBy":1024,"dataFootnoteRef":39,"id":1025},[38],"user-content-fnref-bench-10",[108,1027,1028,1038],{},[111,1029,1030],{},[114,1031,1032,1035],{},[117,1033,1034],{},"Prompt-depth setting",[117,1036,1037],{"align":122},"Exact GDN prefill, development stack",[125,1039,1040,1047,1054,1061,1068,1075],{},[114,1041,1042,1044],{},[130,1043,336],{},[130,1045,1046],{"align":122},"7,575 tokens\u002Fs",[114,1048,1049,1051],{},[130,1050,344],{},[130,1052,1053],{"align":122},"7,785 tokens\u002Fs",[114,1055,1056,1058],{},[130,1057,352],{},[130,1059,1060],{"align":122},"7,762 tokens\u002Fs",[114,1062,1063,1065],{},[130,1064,360],{},[130,1066,1067],{"align":122},"7,711 tokens\u002Fs",[114,1069,1070,1072],{},[130,1071,368],{},[130,1073,1074],{"align":122},"7,607 tokens\u002Fs",[114,1076,1077,1079],{},[130,1078,376],{},[130,1080,1081],{"align":122},"7,422 tokens\u002Fs",[77,1083,1085],{"id":1084},"how-we-measured","How we measured",[10,1087,1088,1089,1092,1093,1096,1097,1100,1101,1104,1105],{},"The main tables use ",[14,1090,1091],{},"BetterBench 0.6.0, standard profile",", against the release image through its own launcher and OpenAI-compatible endpoint. Sampling used temperature 0.7, top-p 0.95 and top-k 20. Each decode category had ",[14,1094,1095],{},"20 timed passes after three warmups","; concurrency used ",[14,1098,1099],{},"48 requests at each of 1, 2, 4 and 8 streams","; prefill used ",[14,1102,1103],{},"eight runs per depth"," and 2,048-token chunks.",[31,1106,1107],{},[34,1108,41],{"href":36,"ariaDescribedBy":1109,"dataFootnoteRef":39,"id":1110},[38],"user-content-fnref-bench-11",[10,1112,1113,1114,1117,1118],{},"The weighted decode score assigns ",[14,1115,1116],{},"30% to code, 20% to reasoning, 15% each to prose and JSON, and 10% each to file editing and summarization",". Chat and math are measured separately and carry no weight in that score.",[31,1119,1120],{},[34,1121,1125],{"href":1122,"ariaDescribedBy":1123,"dataFootnoteRef":39,"id":1124},"#user-content-fn-betterbench",[38],"user-content-fnref-betterbench","5",[10,1127,1128,1129,1132,1133,1136,1137,1140,1141,1147],{},"Every benchmark request carried a unique nonce so its prefix was cold. Decode, short-prompt latency and concurrency use the ",[14,1130,1131],{},"98,304-token speculative mode",". Prefill uses the ",[14,1134,1135],{},"200,000-token non-speculative mode",". The two GPUs exchanged uncompressed BF16 activations. The host was quiet, with no builds or uploads running, and startup took roughly ",[14,1138,1139],{},"three to five minutes",", mainly for weight loading, with compile caches included in the image.",[31,1142,1143],{},[34,1144,41],{"href":36,"ariaDescribedBy":1145,"dataFootnoteRef":39,"id":1146},[38],"user-content-fnref-bench-12",[31,1148,1149],{},[34,1150,60],{"href":57,"ariaDescribedBy":1151,"dataFootnoteRef":39,"id":1152},[38],"user-content-fnref-release-8",[10,1154,1155,1156,1162],{},"These measurements describe our host, prompts and release configuration. They do not measure whole-system energy, cost per token or a speedup over a different model. The ",[34,1157,1161],{"href":1158,"rel":1159},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fbe0f1a53bd60ab1bf77131121f9cacdcf46d2863\u002Fmodels\u002FQwen3.8-Flash-Next\u002FBENCHMARKS.md",[1160],"nofollow","public benchmark report"," preserves the configuration and detailed results.",[77,1164,1166],{"id":1165},"try-it-on-two-r9700-cards","Try it on two R9700 cards",[10,1168,1169,1170,1173,1174,1177,1178,1181,1182,1185,1186],{},"Use a Linux host with ",[14,1171,1172],{},"two Radeon AI PRO R9700 cards",", a ",[14,1175,1176],{},"ROCm 10 host driver",", Python 3, the Hugging Face CLI and Docker with access to ",[426,1179,1180],{},"\u002Fdev\u002Fkfd"," and ",[426,1183,1184],{},"\u002Fdev\u002Fdri",". Allow substantial system RAM beyond the 48.9 GiB n-gram table and disk space beyond the 108 GiB checkpoint. Our test host's 251 GiB is a measured configuration, not a stated minimum requirement.",[31,1187,1188],{},[34,1189,60],{"href":57,"ariaDescribedBy":1190,"dataFootnoteRef":39,"id":1191},[38],"user-content-fnref-release-9",[10,1193,1194,1195,1198,1199,1202,1203,1209],{},"The commands below pin the public launcher and ",[14,1196,1197],{},"explicitly download the complete model revision, including the runtime drafter",". This matters: the pinned launcher's automatic download still points at an older model revision. Download the revision below first and pass its directory with ",[426,1200,1201],{},"--weights",", rather than relying on that automatic download. If you already have the plugin repository, use a separate checkout instead of replacing local work.",[31,1204,1205],{},[34,1206,60],{"href":57,"ariaDescribedBy":1207,"dataFootnoteRef":39,"id":1208},[38],"user-content-fnref-release-10",[31,1210,1211],{},[34,1212,68],{"href":65,"ariaDescribedBy":1213,"dataFootnoteRef":39,"id":1214},[38],"user-content-fnref-model-14",[1216,1217,1221],"pre",{"className":1218,"code":1219,"language":1220,"meta":39,"style":39},"language-bash shiki shiki-themes github-light github-dark","git clone https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\ncd paiton-vllm-plugin\ngit checkout be0f1a53bd60ab1bf77131121f9cacdcf46d2863\ncd models\u002FQwen3.8-Flash-Next\n\nexport PAITON_FLASHNEXT_DIR=\"$PWD\u002Fmodel-cache\u002Fqwen38-flash-next-w3a8\"\nhf download EliovpAI\u002FQwen3.8-Flash-Next-W3A8-Paiton-RDNA4 \\\n  --revision 829b089bf6636af9ffed1f333b103e4e383f7b48 \\\n  --local-dir \"$PAITON_FLASHNEXT_DIR\"\n(cd \"$PAITON_FLASHNEXT_DIR\" && sha256sum -c SHA256SUMS)\n\npython3 launch-flashnext.py --weights \"$PAITON_FLASHNEXT_DIR\" --mode decode\n","bash",[426,1222,1223,1239,1249,1260,1268,1275,1298,1313,1324,1339,1368,1373],{"__ignoreMap":39},[1224,1225,1228,1232,1236],"span",{"class":1226,"line":1227},"line",1,[1224,1229,1231],{"class":1230},"sScJk","git",[1224,1233,1235],{"class":1234},"sZZnC"," clone",[1224,1237,1238],{"class":1234}," https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\n",[1224,1240,1242,1246],{"class":1226,"line":1241},2,[1224,1243,1245],{"class":1244},"sj4cs","cd",[1224,1247,1248],{"class":1234}," paiton-vllm-plugin\n",[1224,1250,1252,1254,1257],{"class":1226,"line":1251},3,[1224,1253,1231],{"class":1230},[1224,1255,1256],{"class":1234}," checkout",[1224,1258,1259],{"class":1234}," be0f1a53bd60ab1bf77131121f9cacdcf46d2863\n",[1224,1261,1263,1265],{"class":1226,"line":1262},4,[1224,1264,1245],{"class":1244},[1224,1266,1267],{"class":1234}," models\u002FQwen3.8-Flash-Next\n",[1224,1269,1271],{"class":1226,"line":1270},5,[1224,1272,1274],{"emptyLinePlaceholder":1273},true,"\n",[1224,1276,1278,1282,1286,1289,1292,1295],{"class":1226,"line":1277},6,[1224,1279,1281],{"class":1280},"szBVR","export",[1224,1283,1285],{"class":1284},"sVt8B"," PAITON_FLASHNEXT_DIR",[1224,1287,1288],{"class":1280},"=",[1224,1290,1291],{"class":1234},"\"",[1224,1293,1294],{"class":1284},"$PWD",[1224,1296,1297],{"class":1234},"\u002Fmodel-cache\u002Fqwen38-flash-next-w3a8\"\n",[1224,1299,1301,1304,1307,1310],{"class":1226,"line":1300},7,[1224,1302,1303],{"class":1230},"hf",[1224,1305,1306],{"class":1234}," download",[1224,1308,1309],{"class":1234}," EliovpAI\u002FQwen3.8-Flash-Next-W3A8-Paiton-RDNA4",[1224,1311,1312],{"class":1244}," \\\n",[1224,1314,1316,1319,1322],{"class":1226,"line":1315},8,[1224,1317,1318],{"class":1244},"  --revision",[1224,1320,1321],{"class":1234}," 829b089bf6636af9ffed1f333b103e4e383f7b48",[1224,1323,1312],{"class":1244},[1224,1325,1327,1330,1333,1336],{"class":1226,"line":1326},9,[1224,1328,1329],{"class":1244},"  --local-dir",[1224,1331,1332],{"class":1234}," \"",[1224,1334,1335],{"class":1284},"$PAITON_FLASHNEXT_DIR",[1224,1337,1338],{"class":1234},"\"\n",[1224,1340,1342,1345,1347,1349,1351,1353,1356,1359,1362,1365],{"class":1226,"line":1341},10,[1224,1343,1344],{"class":1284},"(",[1224,1346,1245],{"class":1244},[1224,1348,1332],{"class":1234},[1224,1350,1335],{"class":1284},[1224,1352,1291],{"class":1234},[1224,1354,1355],{"class":1284}," && ",[1224,1357,1358],{"class":1230},"sha256sum",[1224,1360,1361],{"class":1244}," -c",[1224,1363,1364],{"class":1234}," SHA256SUMS",[1224,1366,1367],{"class":1284},")\n",[1224,1369,1371],{"class":1226,"line":1370},11,[1224,1372,1274],{"emptyLinePlaceholder":1273},[1224,1374,1376,1379,1382,1385,1387,1389,1391,1394],{"class":1226,"line":1375},12,[1224,1377,1378],{"class":1230},"python3",[1224,1380,1381],{"class":1234}," launch-flashnext.py",[1224,1383,1384],{"class":1244}," --weights",[1224,1386,1332],{"class":1234},[1224,1388,1335],{"class":1284},[1224,1390,1291],{"class":1234},[1224,1392,1393],{"class":1244}," --mode",[1224,1395,1396],{"class":1234}," decode\n",[10,1398,1399,1400,1403,1404,1409,1410,1415],{},"The launcher selects ",[426,1401,1402],{},"ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin:qwen38-flashnext-rocm10-vllm029-20261010-r1",", pinned by digest in the ",[34,1405,1408],{"href":1406,"rel":1407},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fbe0f1a53bd60ab1bf77131121f9cacdcf46d2863\u002Fmodels\u002FQwen3.8-Flash-Next\u002Fruntime.lock.json",[1160],"runtime lock",". The ",[34,1411,1414],{"href":1412,"rel":1413},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Fblob\u002Fbe0f1a53bd60ab1bf77131121f9cacdcf46d2863\u002Fmodels\u002FQwen3.8-Flash-Next\u002FREADME.md",[1160],"model-specific setup guide"," documents the Python launcher modes. Stop the running server before selecting another mode:",[1216,1417,1419],{"className":1218,"code":1418,"language":1220,"meta":39,"style":39},"# 200K context, speculation off\npython3 launch-flashnext.py --weights \"$PAITON_FLASHNEXT_DIR\" --mode prefill-long\n\n# Exact Gated DeltaNet prefill, default context\npython3 launch-flashnext.py --weights \"$PAITON_FLASHNEXT_DIR\" --mode decode-nopf\n\n# 200K mode with optional prefix caching\npython3 launch-flashnext.py --weights \"$PAITON_FLASHNEXT_DIR\" \\\n  --mode prefill-long --prefix-caching\n",[426,1420,1421,1427,1446,1450,1455,1474,1478,1483,1499],{"__ignoreMap":39},[1224,1422,1423],{"class":1226,"line":1227},[1224,1424,1426],{"class":1425},"sJ8bj","# 200K context, speculation off\n",[1224,1428,1429,1431,1433,1435,1437,1439,1441,1443],{"class":1226,"line":1241},[1224,1430,1378],{"class":1230},[1224,1432,1381],{"class":1234},[1224,1434,1384],{"class":1244},[1224,1436,1332],{"class":1234},[1224,1438,1335],{"class":1284},[1224,1440,1291],{"class":1234},[1224,1442,1393],{"class":1244},[1224,1444,1445],{"class":1234}," prefill-long\n",[1224,1447,1448],{"class":1226,"line":1251},[1224,1449,1274],{"emptyLinePlaceholder":1273},[1224,1451,1452],{"class":1226,"line":1262},[1224,1453,1454],{"class":1425},"# Exact Gated DeltaNet prefill, default context\n",[1224,1456,1457,1459,1461,1463,1465,1467,1469,1471],{"class":1226,"line":1270},[1224,1458,1378],{"class":1230},[1224,1460,1381],{"class":1234},[1224,1462,1384],{"class":1244},[1224,1464,1332],{"class":1234},[1224,1466,1335],{"class":1284},[1224,1468,1291],{"class":1234},[1224,1470,1393],{"class":1244},[1224,1472,1473],{"class":1234}," decode-nopf\n",[1224,1475,1476],{"class":1226,"line":1277},[1224,1477,1274],{"emptyLinePlaceholder":1273},[1224,1479,1480],{"class":1226,"line":1300},[1224,1481,1482],{"class":1425},"# 200K mode with optional prefix caching\n",[1224,1484,1485,1487,1489,1491,1493,1495,1497],{"class":1226,"line":1315},[1224,1486,1378],{"class":1230},[1224,1488,1381],{"class":1234},[1224,1490,1384],{"class":1244},[1224,1492,1332],{"class":1234},[1224,1494,1335],{"class":1284},[1224,1496,1291],{"class":1234},[1224,1498,1312],{"class":1244},[1224,1500,1501,1504,1507],{"class":1226,"line":1326},[1224,1502,1503],{"class":1244},"  --mode",[1224,1505,1506],{"class":1234}," prefill-long",[1224,1508,1509],{"class":1244}," --prefix-caching\n",[10,1511,1512,1513,1516,1517,1520,1521,1524,1525,1528,1529,29,1532],{},"Run ",[14,1514,1515],{},"one mode at a time",". ",[426,1518,1519],{},"prefill-long-nopf"," selects the exact Gated DeltaNet prefill path at 200K; ",[426,1522,1523],{},"--dry-run"," prints the Docker command without starting the server. Both released base modes use a BF16 attention cache. Once ready, the endpoint is ",[426,1526,1527],{},"http:\u002F\u002F127.0.0.1:18982\u002Fv1",", serving the model name ",[426,1530,1531],{},"Qwen3.8-Flash-Next",[31,1533,1534],{},[34,1535,60],{"href":57,"ariaDescribedBy":1536,"dataFootnoteRef":39,"id":1537},[38],"user-content-fnref-release-11",[10,1539,1540],{},"In another terminal, send a streaming request:",[1216,1542,1544],{"className":1218,"code":1543,"language":1220,"meta":39,"style":39},"curl --fail http:\u002F\u002F127.0.0.1:18982\u002Fv1\u002Fchat\u002Fcompletions \\\n  -H 'Content-Type: application\u002Fjson' \\\n  -d '{\"model\":\"Qwen3.8-Flash-Next\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a short Python function that removes duplicates while preserving order.\"}],\"temperature\":0.7,\"max_tokens\":256,\"stream\":true}'\n",[426,1545,1546,1559,1569],{"__ignoreMap":39},[1224,1547,1548,1551,1554,1557],{"class":1226,"line":1227},[1224,1549,1550],{"class":1230},"curl",[1224,1552,1553],{"class":1244}," --fail",[1224,1555,1556],{"class":1234}," http:\u002F\u002F127.0.0.1:18982\u002Fv1\u002Fchat\u002Fcompletions",[1224,1558,1312],{"class":1244},[1224,1560,1561,1564,1567],{"class":1226,"line":1241},[1224,1562,1563],{"class":1244},"  -H",[1224,1565,1566],{"class":1234}," 'Content-Type: application\u002Fjson'",[1224,1568,1312],{"class":1244},[1224,1570,1571,1574],{"class":1226,"line":1251},[1224,1572,1573],{"class":1244},"  -d",[1224,1575,1576],{"class":1234}," '{\"model\":\"Qwen3.8-Flash-Next\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a short Python function that removes duplicates while preserving order.\"}],\"temperature\":0.7,\"max_tokens\":256,\"stream\":true}'\n",[10,1578,1579,1580,1585,1586,1589],{},"To reproduce the performance workload, use ",[34,1581,1584],{"href":1582,"rel":1583},"https:\u002F\u002Fgithub.com\u002FGGZ14\u002FBetterBench",[1160],"BetterBench"," ",[14,1587,1588],{},"0.6.0 and its standard profile"," against that endpoint, with the release report's extended prefill sweep. Keep the decode and prefill mode results separate, as in the tables above.",[77,1591,1593],{"id":1592},"what-comes-next","What comes next",[10,1595,1596,1597,29,1600],{},"The public roadmap includes RAM and SSD cache tiers and a higher-precision 4-bit build. Image input also needs its own validation. These are ",[14,1598,1599],{},"next steps, not features promised by the two validated text modes",[31,1601,1602],{},[34,1603,60],{"href":57,"ariaDescribedBy":1604,"dataFootnoteRef":39,"id":1605},[38],"user-content-fnref-release-12",[10,1607,1608],{},"For now, the result is a larger local model with useful generation speed, a measured 200K option and explicit quality limits, on two workstation GPUs. Regular vLLM, our plugin, our kernels.",[77,1610,1612],{"id":1611},"credits-and-licensing","Credits and licensing",[10,1614,1615],{},"Qwen3.8 Flash Next is by the Qwen team. vLLM provides the serving framework, BetterBench the performance workload suite, and Paiton the quantized weights and native execution in this release.",[10,1617,1618,1619,1622,1623,1628],{},"The upstream checkpoint uses ",[14,1620,1621],{},"Qwen Community License 1.0, not Apache-2.0",". It includes commercial-use conditions; review the ",[34,1624,1627],{"href":1625,"rel":1626},"https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.8-Flash-Next\u002Fblob\u002Fde4b8e4d43b917e7706784d8bb445c9af86a3540\u002FLICENSE",[1160],"upstream license at the checkpoint revision"," before deployment. The public plugin adapter, model weights, packaged runtime and proprietary compiler are distinct artifacts with distinct terms. This post does not imply unrestricted commercial reuse.",[10,1630,1631,1632,1636,1637,1642,1643,1647],{},"Explore ",[34,1633,1635],{"href":1634},"\u002Fproducts\u002Fpaiton","Paiton",", read the ",[34,1638,1641],{"href":1639,"rel":1640},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-Flash-Next-W3A8-Paiton-RDNA4\u002Fblob\u002F829b089bf6636af9ffed1f333b103e4e383f7b48\u002FREADME.md",[1160],"model card and evaluation details",", or ",[34,1644,1646],{"href":1645},"\u002Fcontact","contact us"," to discuss a local AMD AI workload.",[1649,1650,1653,1658],"section",{"className":1651,"dataFootnotes":39},[1652],"footnotes",[77,1654,1657],{"className":1655,"id":38},[1656],"sr-only","Footnotes",[1659,1660,1661,1759,1848,1953,1966],"ol",{},[1662,1663,1665,1516,1669,1585,1676,1585,1683,1585,1690,1585,1697,1585,1704,1585,1712,1585,1720,1585,1727,1585,1735,1585,1743,1585,1751],"li",{"id":1664},"user-content-fn-bench",[34,1666,1668],{"href":1158,"rel":1667},[1160],"10 October release-image benchmarks, pinned checkout",[34,1670,1675],{"href":1671,"ariaLabel":1672,"className":1673,"dataFootnoteBackref":39},"#user-content-fnref-bench","Back to reference 1",[1674],"data-footnote-backref","↩",[34,1677,1675,1681],{"href":1678,"ariaLabel":1679,"className":1680,"dataFootnoteBackref":39},"#user-content-fnref-bench-2","Back to reference 1-2",[1674],[31,1682,60],{},[34,1684,1675,1688],{"href":1685,"ariaLabel":1686,"className":1687,"dataFootnoteBackref":39},"#user-content-fnref-bench-3","Back to reference 1-3",[1674],[31,1689,68],{},[34,1691,1675,1695],{"href":1692,"ariaLabel":1693,"className":1694,"dataFootnoteBackref":39},"#user-content-fnref-bench-4","Back to reference 1-4",[1674],[31,1696,266],{},[34,1698,1675,1702],{"href":1699,"ariaLabel":1700,"className":1701,"dataFootnoteBackref":39},"#user-content-fnref-bench-5","Back to reference 1-5",[1674],[31,1703,1125],{},[34,1705,1675,1709],{"href":1706,"ariaLabel":1707,"className":1708,"dataFootnoteBackref":39},"#user-content-fnref-bench-6","Back to reference 1-6",[1674],[31,1710,1711],{},"6",[34,1713,1675,1717],{"href":1714,"ariaLabel":1715,"className":1716,"dataFootnoteBackref":39},"#user-content-fnref-bench-7","Back to reference 1-7",[1674],[31,1718,1719],{},"7",[34,1721,1675,1725],{"href":1722,"ariaLabel":1723,"className":1724,"dataFootnoteBackref":39},"#user-content-fnref-bench-8","Back to reference 1-8",[1674],[31,1726,274],{},[34,1728,1675,1732],{"href":1729,"ariaLabel":1730,"className":1731,"dataFootnoteBackref":39},"#user-content-fnref-bench-9","Back to reference 1-9",[1674],[31,1733,1734],{},"9",[34,1736,1675,1740],{"href":1737,"ariaLabel":1738,"className":1739,"dataFootnoteBackref":39},"#user-content-fnref-bench-10","Back to reference 1-10",[1674],[31,1741,1742],{},"10",[34,1744,1675,1748],{"href":1745,"ariaLabel":1746,"className":1747,"dataFootnoteBackref":39},"#user-content-fnref-bench-11","Back to reference 1-11",[1674],[31,1749,1750],{},"11",[34,1752,1675,1756],{"href":1753,"ariaLabel":1754,"className":1755,"dataFootnoteBackref":39},"#user-content-fnref-bench-12","Back to reference 1-12",[1674],[31,1757,1758],{},"12",[1662,1760,1762,1516,1766,1585,1771,1585,1778,1585,1785,1585,1792,1585,1799,1585,1806,1585,1813,1585,1820,1585,1827,1585,1834,1585,1841],{"id":1761},"user-content-fn-release",[34,1763,1765],{"href":1412,"rel":1764},[1160],"Model-specific release and setup guide, pinned checkout",[34,1767,1675],{"href":1768,"ariaLabel":1769,"className":1770,"dataFootnoteBackref":39},"#user-content-fnref-release","Back to reference 2",[1674],[34,1772,1675,1776],{"href":1773,"ariaLabel":1774,"className":1775,"dataFootnoteBackref":39},"#user-content-fnref-release-2","Back to reference 2-2",[1674],[31,1777,60],{},[34,1779,1675,1783],{"href":1780,"ariaLabel":1781,"className":1782,"dataFootnoteBackref":39},"#user-content-fnref-release-3","Back to reference 2-3",[1674],[31,1784,68],{},[34,1786,1675,1790],{"href":1787,"ariaLabel":1788,"className":1789,"dataFootnoteBackref":39},"#user-content-fnref-release-4","Back to reference 2-4",[1674],[31,1791,266],{},[34,1793,1675,1797],{"href":1794,"ariaLabel":1795,"className":1796,"dataFootnoteBackref":39},"#user-content-fnref-release-5","Back to reference 2-5",[1674],[31,1798,1125],{},[34,1800,1675,1804],{"href":1801,"ariaLabel":1802,"className":1803,"dataFootnoteBackref":39},"#user-content-fnref-release-6","Back to reference 2-6",[1674],[31,1805,1711],{},[34,1807,1675,1811],{"href":1808,"ariaLabel":1809,"className":1810,"dataFootnoteBackref":39},"#user-content-fnref-release-7","Back to reference 2-7",[1674],[31,1812,1719],{},[34,1814,1675,1818],{"href":1815,"ariaLabel":1816,"className":1817,"dataFootnoteBackref":39},"#user-content-fnref-release-8","Back to reference 2-8",[1674],[31,1819,274],{},[34,1821,1675,1825],{"href":1822,"ariaLabel":1823,"className":1824,"dataFootnoteBackref":39},"#user-content-fnref-release-9","Back to reference 2-9",[1674],[31,1826,1734],{},[34,1828,1675,1832],{"href":1829,"ariaLabel":1830,"className":1831,"dataFootnoteBackref":39},"#user-content-fnref-release-10","Back to reference 2-10",[1674],[31,1833,1742],{},[34,1835,1675,1839],{"href":1836,"ariaLabel":1837,"className":1838,"dataFootnoteBackref":39},"#user-content-fnref-release-11","Back to reference 2-11",[1674],[31,1840,1750],{},[34,1842,1675,1846],{"href":1843,"ariaLabel":1844,"className":1845,"dataFootnoteBackref":39},"#user-content-fnref-release-12","Back to reference 2-12",[1674],[31,1847,1758],{},[1662,1849,1851,1516,1855,1585,1860,1585,1867,1585,1874,1585,1881,1585,1888,1585,1895,1585,1902,1585,1909,1585,1916,1585,1923,1585,1930,1585,1937,1585,1945],{"id":1850},"user-content-fn-model",[34,1852,1854],{"href":1639,"rel":1853},[1160],"Qwen3.8 Flash Next W3A8 model card, pinned revision",[34,1856,1675],{"href":1857,"ariaLabel":1858,"className":1859,"dataFootnoteBackref":39},"#user-content-fnref-model","Back to reference 3",[1674],[34,1861,1675,1865],{"href":1862,"ariaLabel":1863,"className":1864,"dataFootnoteBackref":39},"#user-content-fnref-model-2","Back to reference 3-2",[1674],[31,1866,60],{},[34,1868,1675,1872],{"href":1869,"ariaLabel":1870,"className":1871,"dataFootnoteBackref":39},"#user-content-fnref-model-3","Back to reference 3-3",[1674],[31,1873,68],{},[34,1875,1675,1879],{"href":1876,"ariaLabel":1877,"className":1878,"dataFootnoteBackref":39},"#user-content-fnref-model-4","Back to reference 3-4",[1674],[31,1880,266],{},[34,1882,1675,1886],{"href":1883,"ariaLabel":1884,"className":1885,"dataFootnoteBackref":39},"#user-content-fnref-model-5","Back to reference 3-5",[1674],[31,1887,1125],{},[34,1889,1675,1893],{"href":1890,"ariaLabel":1891,"className":1892,"dataFootnoteBackref":39},"#user-content-fnref-model-6","Back to reference 3-6",[1674],[31,1894,1711],{},[34,1896,1675,1900],{"href":1897,"ariaLabel":1898,"className":1899,"dataFootnoteBackref":39},"#user-content-fnref-model-7","Back to reference 3-7",[1674],[31,1901,1719],{},[34,1903,1675,1907],{"href":1904,"ariaLabel":1905,"className":1906,"dataFootnoteBackref":39},"#user-content-fnref-model-8","Back to reference 3-8",[1674],[31,1908,274],{},[34,1910,1675,1914],{"href":1911,"ariaLabel":1912,"className":1913,"dataFootnoteBackref":39},"#user-content-fnref-model-9","Back to reference 3-9",[1674],[31,1915,1734],{},[34,1917,1675,1921],{"href":1918,"ariaLabel":1919,"className":1920,"dataFootnoteBackref":39},"#user-content-fnref-model-10","Back to reference 3-10",[1674],[31,1922,1742],{},[34,1924,1675,1928],{"href":1925,"ariaLabel":1926,"className":1927,"dataFootnoteBackref":39},"#user-content-fnref-model-11","Back to reference 3-11",[1674],[31,1929,1750],{},[34,1931,1675,1935],{"href":1932,"ariaLabel":1933,"className":1934,"dataFootnoteBackref":39},"#user-content-fnref-model-12","Back to reference 3-12",[1674],[31,1936,1758],{},[34,1938,1675,1942],{"href":1939,"ariaLabel":1940,"className":1941,"dataFootnoteBackref":39},"#user-content-fnref-model-13","Back to reference 3-13",[1674],[31,1943,1944],{},"13",[34,1946,1675,1950],{"href":1947,"ariaLabel":1948,"className":1949,"dataFootnoteBackref":39},"#user-content-fnref-model-14","Back to reference 3-14",[1674],[31,1951,1952],{},"14",[1662,1954,1956,1516,1961],{"id":1955},"user-content-fn-quality",[34,1957,1960],{"href":1958,"rel":1959},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FQwen3.8-Flash-Next-W3A8-Paiton-RDNA4\u002Fblob\u002F829b089bf6636af9ffed1f333b103e4e383f7b48\u002FQUALITY.md",[1160],"Public quality report, pinned model revision",[34,1962,1675],{"href":1963,"ariaLabel":1964,"className":1965,"dataFootnoteBackref":39},"#user-content-fnref-quality","Back to reference 4",[1674],[1662,1967,1969,1516,1974],{"id":1968},"user-content-fn-betterbench",[34,1970,1973],{"href":1971,"rel":1972},"https:\u002F\u002Fgithub.com\u002FGGZ14\u002FBetterBench\u002Fblob\u002Fd00ad5ec8098c06584a88ec3468bacd37d5ed098\u002Fconfig\u002Fdefault.json",[1160],"BetterBench 0.6.0 workload weights, pinned defaults",[34,1975,1675],{"href":1976,"ariaLabel":1977,"className":1978,"dataFootnoteBackref":39},"#user-content-fnref-betterbench","Back to reference 5",[1674],[1980,1981,1982],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html pre.shiki code .szBVR, html code.shiki .szBVR{--shiki-default:#D73A49;--shiki-dark:#F97583}html pre.shiki code .sVt8B, html code.shiki .sVt8B{--shiki-default:#24292E;--shiki-dark:#E1E4E8}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .sJ8bj, html code.shiki .sJ8bj{--shiki-default:#6A737D;--shiki-dark:#6A737D}",{"title":39,"searchDepth":1241,"depth":1241,"links":1984},[1985,1988,1991,1992,1995,1998,2001,2002,2003,2004,2005],{"id":79,"depth":1241,"text":80,"children":1986},[1987],{"id":224,"depth":1251,"text":225},{"id":293,"depth":1241,"text":294,"children":1989},[1990],{"id":402,"depth":1251,"text":403},{"id":481,"depth":1241,"text":482},{"id":558,"depth":1241,"text":559,"children":1993},[1994],{"id":622,"depth":1251,"text":623},{"id":674,"depth":1241,"text":675,"children":1996},[1997],{"id":782,"depth":1251,"text":783},{"id":959,"depth":1241,"text":960,"children":1999},[2000],{"id":981,"depth":1251,"text":982},{"id":1084,"depth":1241,"text":1085},{"id":1165,"depth":1241,"text":1166},{"id":1592,"depth":1241,"text":1593},{"id":1611,"depth":1241,"text":1612},{"id":38,"depth":1241,"text":1657},[1635,2007,2008,2009,2010],"AMD Radeon","Local AI","Qwen","Quantization","2026-10-10T10:30:00Z","Paiton runs Qwen3.8 Flash Next with 3-bit experts on two R9700 cards: 216.3 tokens\u002Fs, 565.1 aggregate tokens\u002Fs and a separate 200K mode. Benchmarks, quality and setup.","md","A larger model. Two cards. 216 tokens per second.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-flash-next\u002Fhero.webp",{},"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-flash-next-radeon-ai-pro-r9700","\u002Fblog\u002Fpaiton-qwen38-flash-next-radeon-ai-pro-r9700",{"title":5,"description":2012},"paiton-qwen38-flash-next-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-flash-next\u002Fsocial.webp","blog\u002Fpaiton-qwen38-flash-next-radeon-ai-pro-r9700",null,"46nyEPZJBwQ2NWcp3MgR2U3Fhafxr8qu4tSXG-3dU3k",[2026,2028,2037,2047,2060,2070,2082,2091,2106,2115,2129,2161,2173,2195,2213,2232,2250,2268,2285,2297,2313,2328,2340,2349,2357,2372,2384,2395,2406,2416,2429,2439,2452,2463,2473,2484,2493,2505,2516,2525],{"path":2018,"title":5,"description":2012,"date":2011,"slug":2020,"image":2015,"originalUrl":2017,"categories":2027},[1635,2007,2008,2009,2010],{"path":2029,"title":2030,"description":2031,"date":2032,"slug":2033,"image":2034,"originalUrl":2035,"categories":2036},"\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700","Qwen3.8 27B on 1 × Radeon AI PRO R9700: 3-bit weights, 20% faster decode","Qwen3.8 27B on one R9700: 19.9% faster weighted decode, a new 4-bit cache, 200K context and optional vision. Benchmarks and current Paiton setup.","2026-09-27T09:00:00Z","paiton-qwen38-w3a4-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-w3a4\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-w3a4-radeon-ai-pro-r9700",[1635,2007,2008,2009,2010],{"path":2038,"title":2039,"description":2040,"date":2041,"slug":2042,"image":2043,"originalUrl":2044,"categories":2045},"\u002Fblog\u002Fpaiton-qwen-image-21-radeon-ai-pro-r9700","Qwen-Image 2.1 on 1 × Radeon AI PRO R9700: 2048×2048 images in 103 seconds","Generate 2048×2048 Qwen-Image 2.1 images locally with 1 × Radeon AI PRO R9700. The released Paiton v1.0.2 container measured 103.29 seconds per warm request, through PNG delivery.","2026-09-23T09:00:00Z","paiton-qwen-image-21-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen-image-21\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen-image-21-radeon-ai-pro-r9700",[1635,2007,2008,2046,2009],"Image Generation",{"path":2048,"title":2049,"description":2050,"date":2051,"slug":2052,"image":2053,"originalUrl":2054,"categories":2055},"\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","Qwen3.8: 400.7 tok\u002Fs on R9700 | Paiton","Qwen3.8 on one R9700: 400.7 aggregate tok\u002Fs with ROCm 10 and vLLM 0.29, plus public 200K\u002F220K chat profiles. Benchmarks, limits and launch commands.","2026-09-16T07:30:00Z","paiton-qwen38-mxfp4-dflash2-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fupdate-2026-09-19\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",[1635,2007,2056,2057,2058,2059],"vLLM","Qwen3.8","Inference Optimization","DFlash2",{"path":2061,"title":2062,"description":2063,"date":2064,"slug":2065,"image":2066,"originalUrl":2067,"categories":2068},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[1635,2007,2008,2069,2056],"GGUF",{"path":2071,"title":2072,"description":2073,"date":2074,"slug":2075,"image":2076,"originalUrl":2077,"categories":2078},"\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","MiniMax H3 on Radeon: 15-Second Video With Native Sound","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","2026-09-09T07:30:00Z","paiton-minimax-h3-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",[1635,2007,2008,2079,2080,2081],"Video Generation","MiniMax H3","ComfyUI",{"path":2083,"title":2084,"description":2085,"date":2086,"slug":2087,"image":2088,"originalUrl":2083,"categories":2089},"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","2026-09-07T09:00:00","paiton-flux2-klein-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",[1635,2007,2008,2046,2090,2081],"FLUX",{"path":2092,"title":2093,"description":2094,"date":2095,"slug":2096,"image":2097,"originalUrl":2098,"categories":2099},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[1635,2100,2007,2101,2102,2103,2058,2104,2056,2105],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Large Language Models","Cost Efficiency",{"path":2107,"title":2108,"description":2109,"date":2110,"slug":2111,"image":2112,"originalUrl":2113,"categories":2114},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[1635,2100,2007,2101,2102,2103,2058,2104,2056,2105],{"path":2116,"title":2117,"description":2118,"date":2119,"slug":2120,"image":2121,"originalUrl":2023,"categories":2122},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",[2123,2124,2125,2126,2127,2128],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":2130,"title":2131,"description":2132,"date":2133,"slug":2134,"image":2135,"originalUrl":2136,"categories":2137},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[2123,2100,1635,2138,2139,2140,2141,2142,2143,2144,2145,2146,2147,2148,2149,2150,2151,2152,2153,2154,1635,2155,2156,2157,2158,2159,2160],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","GPU","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":2162,"title":2163,"description":2164,"date":2165,"slug":2166,"image":2167,"originalUrl":2168,"categories":2169},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[2123,2100,2170,2171,2139,2172,2170,2152],"Modular DC","Uncategorized","De Tijd",{"path":2174,"title":2175,"description":2176,"date":2177,"slug":2178,"image":2179,"originalUrl":2180,"categories":2181},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[2123,2100,2182,2171,2183,2184,2185,2186,2187,2188,2189,2190,2145,2191,2146,2192,2193,2194],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":2196,"title":2197,"description":2198,"date":2199,"slug":2200,"image":2201,"originalUrl":2202,"categories":2203},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[2123,2204,2205,2206,2188,2207,2208,2209,2191,2210,2211,2212,2193],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":2214,"title":2215,"description":2216,"date":2217,"slug":2218,"image":2219,"originalUrl":2220,"categories":2221},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[2123,2100,2222,2182,2223,2224,2225,2226,2227,2228,2229,2230,2155,2231],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":2233,"title":2234,"description":2235,"date":2236,"slug":2237,"image":2238,"originalUrl":2239,"categories":2240},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[2123,2100,2182,2124,2241,2242,2243,2244,2245,2246,2247,2248,2249],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":2251,"title":2252,"description":2253,"date":2254,"slug":2255,"image":2256,"originalUrl":2257,"categories":2258},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[2123,2170,2171,2259,2124,2260,2261,2262,2263,2264,2265,2266,2267],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":2269,"title":2270,"description":2271,"date":2272,"slug":2273,"image":2274,"originalUrl":2275,"categories":2276},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[2123,2100,1635,2171,2277,2100,2278,2279,2280,2281,2282,2283,1635,2284],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning","ROCm",{"path":2286,"title":2287,"description":2288,"date":2289,"slug":2290,"image":2291,"originalUrl":2292,"categories":2293},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[2123,2100,1635,2101,2294,2277,2105,2295,2058,2283,1635,2296,2056],"AMD Instinct","High Throughput","SGLang",{"path":2298,"title":2299,"description":2300,"date":2301,"slug":2302,"image":2303,"originalUrl":2304,"categories":2305},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[2123,2100,1635,2306,2277,2307,2058,2308,2309,2310,2311,1635,2312],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":2314,"title":2315,"description":2316,"date":2317,"slug":2318,"image":2319,"originalUrl":2320,"categories":2321},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[2123,2100,2222,2171,2223,2322,2323,2324,2325,2228,2230,2155,2326,2327],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":2329,"title":2330,"description":2331,"date":2332,"slug":2333,"image":2334,"originalUrl":2335,"categories":2336},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[2123,2100,1635,2171,2337,2139,2140,2338,2339,2151,2152,1635,2056],"AI","H200","MI300X",{"path":2341,"title":2342,"description":2343,"date":2344,"slug":2345,"image":2346,"originalUrl":2347,"categories":2348},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[2123,2100,2222,2223,2322,2323,2324,2325,2228,2230,2155,2326,2327],{"path":2350,"title":2351,"description":2352,"date":2353,"slug":2354,"image":39,"originalUrl":2355,"categories":2356},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[2123,2100,1635],{"path":2358,"title":2359,"description":2360,"date":2361,"slug":2362,"image":2363,"originalUrl":2364,"categories":2365},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[2123,2100,1635,2171,2101,2277,2366,2279,2367,2368,2369,1635,2370,2371],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":2373,"title":2374,"description":2375,"date":2376,"slug":2377,"image":2378,"originalUrl":2379,"categories":2380},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[2123,2100,1635,2171,2277,2381,2227,2146,2102,2103,2104,2368,2382,2383],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":2385,"title":2386,"description":2387,"date":2388,"slug":2389,"image":2390,"originalUrl":2391,"categories":2392},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[2123,2100,1635,2222,2171,2139,2339,2152,2393,2394],"RX7900XTX","tenstorrent",{"path":2396,"title":2397,"description":2398,"date":2399,"slug":2400,"image":2401,"originalUrl":2402,"categories":2403},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[2123,2100,2170,2222,2140,2338,2404,2152,2405],"MI325x","pnl calculator",{"path":2407,"title":2408,"description":2409,"date":2410,"slug":2411,"image":2412,"originalUrl":2413,"categories":2414},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[2123,2100,1635,2337,2139,2338,2415,2152,1635,2056],"MI300",{"path":2417,"title":2418,"description":2419,"date":2420,"slug":2421,"image":2422,"originalUrl":2423,"categories":2424},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[2123,2170,2425,2124,2261,2127,2262,2263,2426,2427,2266,2428],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":2430,"title":2431,"description":2432,"date":2433,"slug":2434,"image":2435,"originalUrl":2436,"categories":2437},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[2123,2100,2222,2171,2337,2139,2438],"Healthcare",{"path":2440,"title":2441,"description":2442,"date":2443,"slug":2444,"image":2445,"originalUrl":2446,"categories":2447},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[2123,2182,2337,2139,2448,2449,2450,2451],"import","Taiwan","Tariffs","Trump",{"path":2453,"title":2454,"description":2455,"date":2456,"slug":2457,"image":2458,"originalUrl":2459,"categories":2460},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[2123,2100,2222,2337,2461,2462],"AI Agents","ERP",{"path":2464,"title":2465,"description":2466,"date":2467,"slug":2468,"image":2469,"originalUrl":2470,"categories":2471},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[2123,2100,2182,2472,2139,2147,2152],"AI news",{"path":2474,"title":2475,"description":2476,"date":2477,"slug":2478,"image":2479,"originalUrl":2480,"categories":2481},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[2123,2100,1635,2337,2139,2482,2483,2339,1635],"benchmark","LLM",{"path":2485,"title":2486,"description":2487,"date":2488,"slug":2489,"image":2490,"originalUrl":2491,"categories":2492},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[2123,2100,1635],{"path":2494,"title":2495,"description":2496,"date":2497,"slug":2498,"image":2499,"originalUrl":2500,"categories":2501},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[2123,2139,2502,2503,2504],"Jim Greene","Podcast","Tech Talk",{"path":2506,"title":2507,"description":2508,"date":2509,"slug":2510,"image":2511,"originalUrl":2512,"categories":2513},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[2123,2100,1635,2139,2514,2515,2338,2339,2404,1635,2056],"Deepseek","H100",{"path":2517,"title":2518,"description":2519,"date":2520,"slug":2521,"image":2522,"originalUrl":2523,"categories":2524},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[2123,2100,1635,2139,2514,2515,2338,2339,2404,1635,2056],{"path":2526,"title":2527,"description":2528,"date":2529,"slug":2530,"image":2531,"originalUrl":2532,"categories":2533},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[2123,2100,1635,2139,2515,2338,2339,2404,1635,2056],1791631590342]