[{"data":1,"prerenderedAt":1390},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":968},{"id":4,"title":5,"body":6,"categories":946,"date":956,"description":957,"extension":958,"image":959,"meta":960,"navigation":961,"originalUrl":962,"path":963,"seo":964,"slug":965,"stem":966,"__hash__":967},"blog\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700.md","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700",{"type":7,"value":8,"toc":933},"minimark",[9,16,31,38,49,56,61,64,157,160,163,167,170,173,195,198,201,207,211,214,221,224,273,340,347,350,356,359,363,366,369,372,375,379,382,385,388,392,395,398,401,405,408,517,520,538,541,561,576,586,590,593,768,771,775,885,888,891,895,910,913,917,920,923,929],[10,11,12],"p",{},[13,14,15],"strong",{},"One GPU. The same public source checkpoint. The same benchmark requests. More output.",[10,17,18,19,26,27,30],{},"Paiton served AMD’s public\n",[20,21,25],"a",{"href":22,"rel":23},"https:\u002F\u002Fhuggingface.co\u002Famd\u002FQwen3.8-27B-Quark-Qronos-INT4-W4A16",[24],"nofollow","Qwen3.8-27B-Quark-Qronos-INT4-W4A16","\nat ",[13,28,29],{},"39.77 output tokens per second"," in our batch-one interactive benchmark on a\nsingle 32 GB AMD Radeon AI PRO R9700.",[10,32,33,34,37],{},"Our fastest qualified stock vLLM configuration on the same system reached\n32.86 output tokens per second. That gives Paiton ",[13,35,36],{},"21.0% more output per active\ninference hour"," from the same GPU.",[10,39,40,41,44,45,48],{},"The advantage widened to ",[13,42,43],{},"54.3%"," in the coding workload and reached ",[13,46,47],{},"26.8%","\nwith a 4,096-token input.",[10,50,51],{},[52,53],"img",{"alt":54,"src":55},"Paiton delivers 39.77 output tokens per second on the interactive workload, 37.93 on coding and 25.45 on long context, outperforming the qualified stock vLLM baseline in all three tests.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F01-throughput-by-workload-eliovp.webp",[57,58,60],"h2",{"id":59},"throughput-across-three-workloads","Throughput across three workloads",[10,62,63],{},"These are single-user serving tests with one request at a time, one active\nmodel sequence, temperature zero, thinking disabled and fixed random inputs.",[65,66,67,90],"table",{},[68,69,70],"thead",{},[71,72,73,77,81,84,87],"tr",{},[74,75,76],"th",{},"Workload",[74,78,80],{"align":79},"right","Requested input \u002F output",[74,82,83],{"align":79},"Qualified stock vLLM",[74,85,86],{"align":79},"Paiton",[74,88,89],{"align":79},"Paiton advantage",[91,92,93,115,136],"tbody",{},[71,94,95,99,102,105,110],{},[96,97,98],"td",{},"Interactive",[96,100,101],{"align":79},"256 \u002F 256",[96,103,104],{"align":79},"32.86 tok\u002Fs",[96,106,107],{"align":79},[13,108,109],{},"39.77 tok\u002Fs",[96,111,112],{"align":79},[13,113,114],{},"+21.0%",[71,116,117,120,123,126,131],{},[96,118,119],{},"Coding",[96,121,122],{"align":79},"1,024 \u002F 512",[96,124,125],{"align":79},"24.58 tok\u002Fs",[96,127,128],{"align":79},[13,129,130],{},"37.93 tok\u002Fs",[96,132,133],{"align":79},[13,134,135],{},"+54.3%",[71,137,138,141,144,147,152],{},[96,139,140],{},"Long context",[96,142,143],{"align":79},"4,096 \u002F 256",[96,145,146],{"align":79},"20.07 tok\u002Fs",[96,148,149],{"align":79},[13,150,151],{},"25.45 tok\u002Fs",[96,153,154],{"align":79},[13,155,156],{},"+26.8%",[10,158,159],{},"Each workload was run twice from a fresh server. Every run used three warmups\nfollowed by 12 measured requests, and the table reports the mean of both runs.\nAll six Paiton runs completed 12 of 12 requests with zero failures and returned\nthe full requested output length.",[10,161,162],{},"Chat templating increased the actual prompt lengths to 268–270, 1,036–1,038\nand 4,108–4,110 tokens respectively.",[57,164,166],{"id":165},"latency-faster-streaming-mixed-first-token-results","Latency: faster streaming, mixed first-token results",[10,168,169],{},"Output throughput is only part of the user experience, so we are publishing\nfirst-token and streaming latency as well.",[10,171,172],{},"Paiton reduced median time per output token in every workload:",[174,175,176,183,189],"ul",{},[177,178,179,180],"li",{},"Interactive: ",[13,181,182],{},"30 ms to 24 ms",[177,184,185,186],{},"Coding: ",[13,187,188],{},"39 ms to 25 ms",[177,190,191,192],{},"Long context: ",[13,193,194],{},"37 ms to 28 ms",[10,196,197],{},"Time to first token was workload-dependent. Median interactive TTFT increased\nfrom 258 ms to 290 ms. Coding improved slightly from 757 ms to 737 ms, while\nlong-context TTFT fell from 3,337 ms to 2,984 ms.",[10,199,200],{},"The result is a clear streaming-throughput gain, not a claim that every latency\nmetric improves in every scenario.",[10,202,203],{},[52,204],{"alt":205,"src":206},"Paiton lowers time per output token across all three workloads. Time to first token is higher for the interactive test, slightly lower for coding and lower for the long-context test.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F02-latency-snapshot-eliovp.webp",[57,208,210],{"id":209},"what-the-gain-means-for-cost","What the gain means for cost",[10,212,213],{},"For owned inference hardware, throughput determines how much output a fixed\nhour of GPU time can produce. At equal hourly system cost, a 21.0% throughput\ngain gives 21.0% more tokens for the same active runtime budget.",[10,215,216,217,220],{},"Because cost per token is the inverse of throughput, the corresponding modeled\nreduction in time-based cost per million output tokens is ",[13,218,219],{},"17.4%",".",[10,222,223],{},"The examples below amortize the GPU over 5,000 productive inference hours and\nassume an equal 450 W whole-system draw for both runtimes.",[65,225,226,239],{},[68,227,228],{},[71,229,230,233,236],{},[74,231,232],{},"Assumption",[74,234,235],{"align":79},"US example",[74,237,238],{"align":79},"European example",[91,240,241,252,263],{},[71,242,243,246,249],{},[96,244,245],{},"GPU purchase price",[96,247,248],{"align":79},"$1,299",[96,250,251],{"align":79},"€1,749",[71,253,254,257,260],{},[96,255,256],{},"Electricity",[96,258,259],{"align":79},"$0.17\u002FkWh",[96,261,262],{"align":79},"€0.2558\u002FkWh",[71,264,265,268,271],{},[96,266,267],{},"Productive inference lifetime",[96,269,270],{"align":79},"5,000 hours",[96,272,270],{"align":79},[65,274,275,286],{},[68,276,277],{},[71,278,279,282,284],{},[74,280,281],{},"Interactive token economics",[74,283,83],{"align":79},[74,285,86],{"align":79},[91,287,288,301,314,327],{},[71,289,290,293,296],{},[96,291,292],{},"Hours per million output tokens",[96,294,295],{"align":79},"8.45",[96,297,298],{"align":79},[13,299,300],{},"6.98",[71,302,303,306,309],{},[96,304,305],{},"Modeled cost per million, US example",[96,307,308],{"align":79},"$2.84",[96,310,311],{"align":79},[13,312,313],{},"$2.35",[71,315,316,319,322],{},[96,317,318],{},"Modeled cost per million, European example",[96,320,321],{"align":79},"€3.93",[96,323,324],{"align":79},[13,325,326],{},"€3.25",[71,328,329,332,335],{},[96,330,331],{},"Output over 5,000 active hours",[96,333,334],{"align":79},"591.5 million",[96,336,337],{"align":79},[13,338,339],{},"715.8 million",[10,341,342,343,346],{},"At the measured interactive rates, the same card produces roughly ",[13,344,345],{},"124 million\nadditional output tokens"," over 5,000 productive hours.",[10,348,349],{},"For the same modeled $100 ownership-and-electricity budget, Paiton produces\nabout 42.6 million tokens instead of 35.2 million. In the European example,\n€100 produces about 30.8 million tokens instead of 25.4 million.",[10,351,352],{},[52,353],{"alt":354,"src":355},"At 5,000 productive inference hours, the modeled European cost is €3.93 per million output tokens for stock vLLM and €3.25 for Paiton.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F03-economics-sensitivity-eliovp.webp",[10,357,358],{},"These figures are a transparent time-based ownership model, not a measured\nenergy-efficiency claim. They exclude the host computer, idle time, cooling,\nmaintenance, financing, tax and residual value. Substitute your own purchase\nprice, electricity rate and productive lifetime as needed. The 17.4% relative\nreduction remains the same when both runtimes carry the same hourly cost.",[57,360,362],{"id":361},"a-stock-baseline-worth-comparing-against","A stock baseline worth comparing against",[10,364,365],{},"We did not compare Paiton with an eager or default installation and call the\nresult an optimization win. We spent significant time qualifying the fastest\nstable stock configuration we could obtain on this system.",[10,367,368],{},"The stock baseline used compiled vLLM execution at optimization level 2,\nfull-and-piecewise HIP graph capture, stock RDNA hybrid W4A16 kernels, vLLM’s\nTriton GDN implementation, the ROCm attention backend and the R9700’s supported\nhigh clock policy.",[10,370,371],{},"Both sides used the same GPU, ROCm host, model revision, tokenizer, pinned vLLM\nrevision, request data, random seed, output lengths, concurrency and supported\nGPU clock policy. We retained the complete result files and averaged two\nfresh-server runs per workload rather than selecting one favorable terminal\nresult.",[10,373,374],{},"The stock logs confirm compiled execution and graph capture. They load no\nPaiton model artifact or Paiton compute kernel.",[57,376,378],{"id":377},"what-paiton-does","What Paiton does",[10,380,381],{},"At a high level, Paiton builds a qualified, model- and hardware-specific\nexecution path and integrates it with the serving runtime. The original AMD\ncheckpoint remains unchanged on disk.",[10,383,384],{},"Paiton assembles target-specific runtime artifacts during model loading. The\noptimized path should therefore not be assumed to produce output that is\nbit-for-bit identical to stock W4 execution.",[10,386,387],{},"That is where the implementation detail stops. The model-specific kernel,\nfusion, scheduling and runtime design are closed-source Paiton IP. What we do\npublish is the part customers can validate: the supported target, benchmark\nmethod, throughput, latency, quality gates, package identity and operating\nscope.",[57,389,391],{"id":390},"quality-checks","Quality checks",[10,393,394],{},"We ran a deterministic 12-case suite covering instruction following,\narithmetic, algebra, logic, factual recall, translation, structured JSON,\nPython, SQL, summarization, format constraints and modular reasoning.",[10,396,397],{},"All 12 cases passed. Repeated temperature-zero responses were byte-identical,\nand all recorded log probabilities were finite.",[10,399,400],{},"This is a practical quality gate for the stated local-chat scope. It is not a\nreplacement for a full academic accuracy evaluation, and it should not be read\nas a claim of quality parity across every task or prompt distribution.",[57,402,404],{"id":403},"run-the-local-chatbot","Run the local chatbot",[10,406,407],{},"The public image contains the Paiton runtime plugin and compiled runtime\nartifacts. It downloads the original 19.9 GB AMD checkpoint directly from\nHugging Face into a persistent Docker volume; Paiton does not redistribute the\nmodel weights.",[409,410,415],"pre",{"className":411,"code":412,"language":413,"meta":414,"style":414},"language-bash shiki shiki-themes github-light github-dark","docker run -d \\\n  --name paiton-qwen38 \\\n  --device \u002Fdev\u002Fkfd \\\n  --device \u002Fdev\u002Fdri \\\n  --group-add video \\\n  --ipc=host \\\n  --network host \\\n  --mount type=volume,src=paiton-qwen38-cache,dst=\u002Fmodels\u002Fcache \\\n  ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin@sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9\n","bash","",[416,417,418,438,449,460,470,481,489,500,511],"code",{"__ignoreMap":414},[419,420,423,427,431,435],"span",{"class":421,"line":422},"line",1,[419,424,426],{"class":425},"sScJk","docker",[419,428,430],{"class":429},"sZZnC"," run",[419,432,434],{"class":433},"sj4cs"," -d",[419,436,437],{"class":433}," \\\n",[419,439,441,444,447],{"class":421,"line":440},2,[419,442,443],{"class":433},"  --name",[419,445,446],{"class":429}," paiton-qwen38",[419,448,437],{"class":433},[419,450,452,455,458],{"class":421,"line":451},3,[419,453,454],{"class":433},"  --device",[419,456,457],{"class":429}," \u002Fdev\u002Fkfd",[419,459,437],{"class":433},[419,461,463,465,468],{"class":421,"line":462},4,[419,464,454],{"class":433},[419,466,467],{"class":429}," \u002Fdev\u002Fdri",[419,469,437],{"class":433},[419,471,473,476,479],{"class":421,"line":472},5,[419,474,475],{"class":433},"  --group-add",[419,477,478],{"class":429}," video",[419,480,437],{"class":433},[419,482,484,487],{"class":421,"line":483},6,[419,485,486],{"class":433},"  --ipc=host",[419,488,437],{"class":433},[419,490,492,495,498],{"class":421,"line":491},7,[419,493,494],{"class":433},"  --network",[419,496,497],{"class":429}," host",[419,499,437],{"class":433},[419,501,503,506,509],{"class":421,"line":502},8,[419,504,505],{"class":433},"  --mount",[419,507,508],{"class":429}," type=volume,src=paiton-qwen38-cache,dst=\u002Fmodels\u002Fcache",[419,510,437],{"class":433},[419,512,514],{"class":421,"line":513},9,[419,515,516],{"class":429},"  ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin@sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9\n",[10,518,519],{},"The first start downloads the checkpoint. With the weights already cached,\nmodel assembly takes roughly 10 to 12 minutes on our R9700. Follow startup with:",[409,521,523],{"className":411,"code":522,"language":413,"meta":414,"style":414},"docker logs -f paiton-qwen38\n",[416,524,525],{"__ignoreMap":414},[419,526,527,529,532,535],{"class":421,"line":422},[419,528,426],{"class":425},[419,530,531],{"class":429}," logs",[419,533,534],{"class":433}," -f",[419,536,537],{"class":429}," paiton-qwen38\n",[10,539,540],{},"Once the log reports that the application is ready, open the included streaming\nchat:",[409,542,544],{"className":411,"code":543,"language":413,"meta":414,"style":414},"docker exec -it paiton-qwen38 paiton-chat\n",[416,545,546],{"__ignoreMap":414},[419,547,548,550,553,556,558],{"class":421,"line":422},[419,549,426],{"class":425},[419,551,552],{"class":429}," exec",[419,554,555],{"class":433}," -it",[419,557,446],{"class":429},[419,559,560],{"class":429}," paiton-chat\n",[10,562,563,564,567,568,571,572,575],{},"Use ",[416,565,566],{},"\u002Freset"," to clear the conversation and ",[416,569,570],{},"\u002Fquit"," to exit. Thinking is off by\ndefault for responsive local conversation; pass ",[416,573,574],{},"--thinking"," when required.",[10,577,578,579,582,583,220],{},"The same server exposes an OpenAI-compatible endpoint at\n",[416,580,581],{},"http:\u002F\u002F127.0.0.1:8000\u002Fv1\u002Fchat\u002Fcompletions"," using model name ",[416,584,585],{},"qwen38",[57,587,589],{"id":588},"reproduce-the-interactive-benchmark","Reproduce the interactive benchmark",[10,591,592],{},"After the server is ready:",[409,594,596],{"className":411,"code":595,"language":413,"meta":414,"style":414},"docker exec paiton-qwen38 sh -lc '\nMODEL_DIR=\"$(find \u002Ftmp -maxdepth 2 -type d \\\n  -path \"\u002Ftmp\u002Fpaiton-qwen38-reused-*\u002Fmodel\" -print -quit)\"\nexec vllm bench serve \\\n  --backend openai-chat \\\n  --base-url http:\u002F\u002F127.0.0.1:8000 \\\n  --endpoint \u002Fv1\u002Fchat\u002Fcompletions \\\n  --model qwen38 \\\n  --tokenizer \"$MODEL_DIR\" \\\n  --dataset-name random \\\n  --seed 42 \\\n  --num-warmups 3 \\\n  --num-prompts 12 \\\n  --random-input-len 256 \\\n  --random-output-len 256 \\\n  --random-range-ratio 0 \\\n  --random-prefix-len 0 \\\n  --request-rate inf \\\n  --max-concurrency 1 \\\n  --temperature 0 \\\n  --ignore-eos \\\n  --extra-body '\\''{\"chat_template_kwargs\":{\"enable_thinking\":false}}'\\'' \\\n  --percentile-metrics ttft,tpot,itl,e2el \\\n  --metric-percentiles 50,90,95,99 \\\n  --disable-tqdm\n'\n",[416,597,598,615,620,625,630,635,640,645,650,655,661,667,673,679,685,691,697,703,709,715,721,727,744,750,756,762],{"__ignoreMap":414},[419,599,600,602,604,606,609,612],{"class":421,"line":422},[419,601,426],{"class":425},[419,603,552],{"class":429},[419,605,446],{"class":429},[419,607,608],{"class":429}," sh",[419,610,611],{"class":433}," -lc",[419,613,614],{"class":429}," '\n",[419,616,617],{"class":421,"line":440},[419,618,619],{"class":429},"MODEL_DIR=\"$(find \u002Ftmp -maxdepth 2 -type d \\\n",[419,621,622],{"class":421,"line":451},[419,623,624],{"class":429},"  -path \"\u002Ftmp\u002Fpaiton-qwen38-reused-*\u002Fmodel\" -print -quit)\"\n",[419,626,627],{"class":421,"line":462},[419,628,629],{"class":429},"exec vllm bench serve \\\n",[419,631,632],{"class":421,"line":472},[419,633,634],{"class":429},"  --backend openai-chat \\\n",[419,636,637],{"class":421,"line":483},[419,638,639],{"class":429},"  --base-url http:\u002F\u002F127.0.0.1:8000 \\\n",[419,641,642],{"class":421,"line":491},[419,643,644],{"class":429},"  --endpoint \u002Fv1\u002Fchat\u002Fcompletions \\\n",[419,646,647],{"class":421,"line":502},[419,648,649],{"class":429},"  --model qwen38 \\\n",[419,651,652],{"class":421,"line":513},[419,653,654],{"class":429},"  --tokenizer \"$MODEL_DIR\" \\\n",[419,656,658],{"class":421,"line":657},10,[419,659,660],{"class":429},"  --dataset-name random \\\n",[419,662,664],{"class":421,"line":663},11,[419,665,666],{"class":429},"  --seed 42 \\\n",[419,668,670],{"class":421,"line":669},12,[419,671,672],{"class":429},"  --num-warmups 3 \\\n",[419,674,676],{"class":421,"line":675},13,[419,677,678],{"class":429},"  --num-prompts 12 \\\n",[419,680,682],{"class":421,"line":681},14,[419,683,684],{"class":429},"  --random-input-len 256 \\\n",[419,686,688],{"class":421,"line":687},15,[419,689,690],{"class":429},"  --random-output-len 256 \\\n",[419,692,694],{"class":421,"line":693},16,[419,695,696],{"class":429},"  --random-range-ratio 0 \\\n",[419,698,700],{"class":421,"line":699},17,[419,701,702],{"class":429},"  --random-prefix-len 0 \\\n",[419,704,706],{"class":421,"line":705},18,[419,707,708],{"class":429},"  --request-rate inf \\\n",[419,710,712],{"class":421,"line":711},19,[419,713,714],{"class":429},"  --max-concurrency 1 \\\n",[419,716,718],{"class":421,"line":717},20,[419,719,720],{"class":429},"  --temperature 0 \\\n",[419,722,724],{"class":421,"line":723},21,[419,725,726],{"class":429},"  --ignore-eos \\\n",[419,728,730,733,736,739,741],{"class":421,"line":729},22,[419,731,732],{"class":429},"  --extra-body '",[419,734,735],{"class":433},"\\'",[419,737,738],{"class":429},"'{\"chat_template_kwargs\":{\"enable_thinking\":false}}'",[419,740,735],{"class":433},[419,742,743],{"class":429},"' \\\n",[419,745,747],{"class":421,"line":746},23,[419,748,749],{"class":429},"  --percentile-metrics ttft,tpot,itl,e2el \\\n",[419,751,753],{"class":421,"line":752},24,[419,754,755],{"class":429},"  --metric-percentiles 50,90,95,99 \\\n",[419,757,759],{"class":421,"line":758},25,[419,760,761],{"class":429},"  --disable-tqdm\n",[419,763,765],{"class":421,"line":764},26,[419,766,767],{"class":429},"'\n",[10,769,770],{},"One run is useful as a local check, but consumer GPU power management can make a\ncold run slower. Run a short warmup first and confirm that the memory clock has\nreached its normal loaded state before recording a comparison.",[57,772,774],{"id":773},"tested-configuration","Tested configuration",[65,776,777,787],{},[68,778,779],{},[71,780,781,784],{},[74,782,783],{},"Component",[74,785,786],{},"Tested value",[91,788,789,800,808,818,826,836,844,854,862,869,877],{},[71,790,791,794],{},[96,792,793],{},"GPU",[96,795,796,797],{},"AMD Radeon AI PRO R9700, 32 GB, ",[416,798,799],{},"gfx1201",[71,801,802,805],{},[96,803,804],{},"Model",[96,806,807],{},"AMD Qwen3.8 27B Qronos W4A16 INT4",[71,809,810,813],{},[96,811,812],{},"Model revision",[96,814,815],{},[416,816,817],{},"649ca9d47a7de5364c6fcccc0c1b4f6e542e15e2",[71,819,820,823],{},[96,821,822],{},"ROCm",[96,824,825],{},"7.14",[71,827,828,831],{},[96,829,830],{},"vLLM revision",[96,832,833],{},[416,834,835],{},"39bd959b582c85e78e7e0326d49042ce7c3c07ed",[71,837,838,841],{},[96,839,840],{},"Paiton image",[96,842,843],{},"Qwen3.8 Qronos for Radeon AI PRO R9700",[71,845,846,849],{},[96,847,848],{},"Image digest",[96,850,851],{},[416,852,853],{},"sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9",[71,855,856,859],{},[96,857,858],{},"Tensor parallelism",[96,860,861],{},"1",[71,863,864,867],{},[96,865,866],{},"Maximum active sequences",[96,868,861],{},[71,870,871,874],{},[96,872,873],{},"Maximum context",[96,875,876],{},"8,192 tokens",[71,878,879,882],{},[96,880,881],{},"Qualified scope",[96,883,884],{},"Text-only, single-user inference",[10,886,887],{},"The package fails closed on an unsupported GPU identity, architecture, runtime\ncontract, artifact checksum, model contract or serving configuration. Multiple\nsimultaneous requests queue.",[10,889,890],{},"This article does not claim support for other GPUs, ROCm versions,\ntensor-parallel configurations, multimodal input or production continuous\nbatching.",[57,892,894],{"id":893},"the-practical-result","The practical result",[10,896,897,898,901,902,905,906,909],{},"On one Radeon AI PRO R9700, Paiton increased interactive Qwen3.8 output\nthroughput from 32.86 to 39.77 tokens per second. That translates into ",[13,899,900],{},"21.0%\nmore output per active hour",", a ",[13,903,904],{},"17.4% reduction in modeled time-based cost\nper million output tokens",", and roughly ",[13,907,908],{},"124 million additional tokens over\n5,000 productive hours"," at the measured interactive rate.",[10,911,912],{},"The coding and long-context results show that the gain is not limited to one\nprompt shape. The latency data also makes the boundary clear: streaming became\nfaster across all three workloads, while first-token latency remained\nworkload-dependent.",[57,914,916],{"id":915},"optimize-your-inference-workload-with-paiton","Optimize your inference workload with Paiton",[10,918,919],{},"This is a qualified Radeon AI PRO R9700 result for one model and one serving\nscope. Paiton’s broader commercial work also targets AMD Instinct CDNA\naccelerators for larger inference deployments, where throughput, fleet\nutilization and cost per generated token compound across the infrastructure.",[10,921,922],{},"We benchmark the real workload, identify the runtime bottleneck, build the\nqualified AMD path and measure the delivered result against an agreed baseline.",[10,924,925,926,220],{},"Learn more about ",[20,927,86],{"href":928},"\u002Fproducts\u002Fpaiton",[930,931,932],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":414,"searchDepth":440,"depth":440,"links":934},[935,936,937,938,939,940,941,942,943,944,945],{"id":59,"depth":440,"text":60},{"id":165,"depth":440,"text":166},{"id":209,"depth":440,"text":210},{"id":361,"depth":440,"text":362},{"id":377,"depth":440,"text":378},{"id":390,"depth":440,"text":391},{"id":403,"depth":440,"text":404},{"id":588,"depth":440,"text":589},{"id":773,"depth":440,"text":774},{"id":893,"depth":440,"text":894},{"id":915,"depth":440,"text":916},[86,947,948,949,950,951,952,953,954,955],"Artificial Intelligence","AMD Radeon","AI Inference","GPU Performance","Inference Latency","Inference Optimization","Large Language Models","vLLM","Cost Efficiency","2026-09-04T09:00:00","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","md","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp",{},true,"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",{"title":5,"description":957},"paiton-qwen38-radeon-ai-pro-r9700","blog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","sMjv9vdYcOWRJniI9yX_U_i1eSjuspgxcBHT_Gz88rM",[969,971,986,1017,1029,1052,1070,1089,1107,1125,1141,1153,1169,1184,1196,1205,1213,1228,1240,1251,1262,1272,1285,1295,1308,1319,1329,1340,1349,1361,1372,1381],{"path":963,"title":5,"description":957,"date":956,"slug":965,"image":959,"originalUrl":962,"categories":970},[86,947,948,949,950,951,952,953,954,955],{"path":972,"title":973,"description":974,"date":975,"slug":976,"image":977,"originalUrl":978,"categories":979},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",null,[980,981,982,983,984,985],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":987,"title":988,"description":989,"date":990,"slug":991,"image":992,"originalUrl":993,"categories":994},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Paiton Returns to Its Diffusion Roots: Optimizing Wan2.2-T2V-A14B on AMD MI355X","When we first started building Paiton, one of our earliest focus areas was optimizing diffusion models. Stable Diffusion XL was one of the first large models where we showed that fused operators, efficient execution, and hardware-aware kernels could make a real difference. Now we are returning to those origins.With the growing interest in text-to-video generation, ...","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[980,947,86,995,996,997,998,999,1000,1001,1002,1003,793,1004,1005,1006,1007,1008,1009,1010,86,1011,1012,1013,1014,1015,1016],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1018,"title":1019,"description":1020,"date":1021,"slug":1022,"image":1023,"originalUrl":1024,"categories":1025},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","From the Attic to the Front Page: ElioVP Recognized as a Pioneer in Chip Optimization & Data Center Infrastructure","It has been some incredible weeks for the team here at Eliovp. We are extremely proud to share that our company was recently featured on the front page of De Tijd, Belgium’s leading business newspaper. Seeing our story, from our founder’s early days tinkering with wires in an attic to generating €215 million in revenue, ...","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[980,947,1026,1027,996,1028,1026,1008],"Modular DC","Uncategorized","De Tijd",{"path":1030,"title":1031,"description":1032,"date":1033,"slug":1034,"image":1035,"originalUrl":1036,"categories":1037},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","Privacy is geen IT-probleem meer, het is een strategische prioriteit: ","Privacyrisico’s, Psychologische Valkuilen en de Operationele Realiteit van Generatieve AI in de Benelux De recente verschijning van een frontpage-artikel over ons bedrijf in het gerespecteerde dagblad “De Tijd” heeft onze zichtbaarheid aanzienlijk vergroot, wat de aanleiding is voor dit artikel. Deze mediabelangstelling, gecombineerd met de talrijke uitnodigingen voor spreekbeurten die we hebben ontvangen, fungeert als ...","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[980,947,1038,1027,1039,1040,1041,1042,1043,1044,1045,1046,1002,1047,1048,1049,1050,1051],"Trending","AI Act","Antropomorfisme","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Generatieve AI","Microsoft Copilot","Privacy","Shadow AI",{"path":1053,"title":1054,"description":1055,"date":1056,"slug":1057,"image":1058,"originalUrl":1059,"categories":1060},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Itsme? Bij ons is het “it’s not me”, en dit is waarom.","1. Inleiding: De Strategische Noodzaak van Weigering In het hedendaagse digitale landschap wordt de keuze voor een identiteit leverancier (IdP) vaak gereduceerd tot een discussie over User Experience (UX) en conversieratio’s. Deze reductionistische benadering verhult echter de diepgaande strategische, juridische en operationele risico’s die gepaard gaan met het uitbesteden van de “Sleutels tot het Koninkrijk”, ...","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[980,1061,1062,1063,1044,1064,1065,1066,1047,1067,1068,1069,1050],"AWS","Belgian Mobile ID","Cloud Act","Data Soevereiniteit","Digitale Identiteit","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1071,"title":1072,"description":1073,"date":1074,"slug":1075,"image":1076,"originalUrl":1077,"categories":1078},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","From Hype to Sovereign Infrastructure Nederlandse Versie Summary The narrative surrounding “Agentic AI” in 2025 is defined by a sharp contrast between market expectations and engineering reality. While the general public, conditioned by the ease of ChatGPT, expects “miracles” and instant integration, the reality of building autonomous agents for enterprise workflows is a discipline of ...","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[980,947,1079,1038,1080,1081,1082,1083,1084,1085,1086,1087,1011,1088],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1090,"title":1091,"description":1092,"date":1093,"slug":1094,"image":1095,"originalUrl":1096,"categories":1097},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","How AI Neocloud Infrastructure Turns Circular Cash Flows Into Fake Growth Executive Summary The interval between 2023 and 2025 has birthed a capital allocation phenomenon arguably without precedent: the “Synthetic Bubble.” Driven by the scramble for AI dominance, the venture capital apparatus has directed billions into the “Neocloud” ecosystem. However, a rigorous analysis suggests that ...","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[980,947,1038,981,1098,1099,1100,1101,1102,1103,1104,1105,1106],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1108,"title":1109,"description":1110,"date":1111,"slug":1112,"image":1113,"originalUrl":1114,"categories":1115},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","Building the Engine for the AI Race: The 4-Month Path to NVIDIA GB300 NVL72 Power","In artificial intelligence infrastructure, speed is the foundation of competitive differentiation. From model training velocity to inference latency, every millisecond matters. But before any workload executes, there is a critical prerequisite that often determines success or failure: time to market. Traditional builds typically require 18–24 months. In the current AI cycle, that is simply too ...","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[980,1026,1027,1116,981,1117,1118,1119,1120,1121,1122,1123,1124],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1126,"title":1127,"description":1128,"date":1129,"slug":1130,"image":1131,"originalUrl":1132,"categories":1133},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Every few years, a new solution pops up promising the same dream: On paper, that sounds perfect. Take your existing CUDA applications, swap out the toolchain, and suddenly you’re “portable.” And to be fair: if you’re running research code or trying to get an internal tool to compile on a non-NVIDIA box, that can absolutely ...","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[980,947,86,1027,1134,947,1135,1136,1137,1138,1139,1140,86,822],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning",{"path":1142,"title":1143,"description":1144,"date":1145,"slug":1146,"image":1147,"originalUrl":1148,"categories":1149},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Let’s be honest, we’re not the marketing type.We’ve never taken a cent of outside investment, never burned cash on ad campaigns, and never hired a sales army.We just build things that work. In today’s world, it seems the companies shouting the loudest often get the spotlight, while the ones doing the actual engineering quietly build ...","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[980,947,86,949,1150,1134,955,1151,952,1140,86,1152,954],"AMD Instinct","High Throughput","SGLang",{"path":1154,"title":1155,"description":1156,"date":1157,"slug":1158,"image":1159,"originalUrl":1160,"categories":1161},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Stop Overpaying: Paiton MI300X MoE Beats H200\u002FB200 on $\u002F1M Tokens","Short summary: We benchmarked Paiton with our new MoE support on Qwen\u002FQwen3-30B-A3B-Instruct-2507 to compare inference performance across several setups. Each configuration was run five times per batch size and we report the mean across runs. Why this benchmark Most published numbers use synthetic prompts or toy datasets. We focused on realistic conversational workloads (we always ...","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[980,947,86,1162,1134,1163,952,1164,1165,1166,1167,86,1168],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1170,"title":1171,"description":1172,"date":1173,"slug":1174,"image":1175,"originalUrl":1176,"categories":1177},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Agentic AI, But Make It Local: From Inbox to Insight to Action","(Nederlandse versie) We’ve built production-ready, local-first agentic AI that plugs into your existing email stack, auto-creates tickets, classifies messages, extracts multi-question threads, reads PDFs, spots invoices\u002Fquotes, analyzes images (yes, damage detection), and pushes structured reports into your systems, no dependency on OpenAI, Google, or Microsoft unless you want it. Tailor-made models trained on your data, ...","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[980,947,1079,1027,1080,1178,1179,1180,1181,1085,1087,1011,1182,1183],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1185,"title":1186,"description":1187,"date":1188,"slug":1189,"image":1190,"originalUrl":1191,"categories":1192},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Data‑Parallel Benchmarks (8–64 GPUs): H200 Left Behind, B200 Within Reach","At ElioVP, we’re all about pushing AI inference past the limits, and packaging every squeeze of performance into a plug‑and‑play runtime.  Remember our last blog, where Paiton’s FP8 pipeline on AMD’s MI300X completely outclassed NVIDIA’s H200? Well, buckle up, because we’ve gone back to the drawing board. This time, we’re loading Llama-3.1-8B-Instruct-FP8-KV, the leaner, meaner ...","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[980,947,86,1027,1193,996,997,1194,1195,1007,1008,86,954],"AI","H200","MI300X",{"path":1197,"title":1198,"description":1199,"date":1200,"slug":1201,"image":1202,"originalUrl":1203,"categories":1204},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Here at Eliovp, we continuously innovate when it comes to building practical solutions. If there’s one core strength, it’s our team’s ability to think outside the box. One key area of focus for us is developing applicable AI solutions, everyday usable AI implementations tailored specifically for our clients’ needs. In this blog, we’ll explore our ...","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[980,947,1079,1080,1178,1179,1180,1181,1085,1087,1011,1182,1183],{"path":1206,"title":1207,"description":1208,"date":1209,"slug":1210,"image":414,"originalUrl":1211,"categories":1212},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Introduction AI is rapidly transforming every industry, but running large models efficiently remains a major technical and financial challenge. At ElioVP, we specialize in optimizing for AMD accelerators, helping organizations unlock the full potential of their hardware. Today, we’re excited to announce a new offering: free evaluation models that let you test our cutting-edge optimizations ...","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[980,947,86],{"path":1214,"title":1215,"description":1216,"date":1217,"slug":1218,"image":1219,"originalUrl":1220,"categories":1221},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Paiton: Dramatically Faster Startup and Performance for Llama-3.1-405B","With Paiton, we’re not merely pursuing peak inference speeds, we’re fundamentally reshaping the entire lifecycle of large language model (LLM) deployment. Our latest endeavor pairs AMD’s cutting-edge MI300X GPUs with the colossal Llama-3.1-405B-Instruct-FP8-KV model, achieving groundbreaking milestones: Visual Demonstration: Startup Speed Showcase We’re excited to share a visual demonstration of Paiton’s revolutionary startup performance. Watch ...","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[980,947,86,1027,949,1134,1222,1136,1223,1224,1225,86,1226,1227],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1229,"title":1230,"description":1231,"date":1232,"slug":1233,"image":1234,"originalUrl":1235,"categories":1236},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","The world of AI is moving at an unprecedented pace, and efficient inference is key to deploying powerful models in real-world applications. At Eliovp, we’ve consistently pushed the boundaries of AI performance, as highlighted in our previous blogs showcasing significant inference speedups when benchmarking with fp16\u002Fbf16. Now, we’re thrilled to announce a further significant leap ...","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[980,947,86,1027,1134,1237,1084,1003,950,951,953,1224,1238,1239],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1241,"title":1242,"description":1243,"date":1244,"slug":1245,"image":1246,"originalUrl":1247,"categories":1248},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","As large language models (LLMs) become a foundational part of modern applications, picking the right server for deployment is more important than ever. Whether you’re an enterprise scaling up inference, a startup optimizing for cost, or a researcher pushing throughput boundaries. This blog compares two high-profile server setups and two not so high-profile setups which ...","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[980,947,86,1079,1027,996,1195,1008,1249,1250],"RX7900XTX","tenstorrent",{"path":1252,"title":1253,"description":1254,"date":1255,"slug":1256,"image":1257,"originalUrl":1258,"categories":1259},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Empowering GPU Cluster Investors with Real-World Financial Insights","At Eliovp BV, we’ve spent years on the cutting edge of GPU cluster deployment and optimization across Europe. Our team supports leading organizations in AI, finance, and research, architecting, building, and scaling high-performance infrastructure. Over time, our customers, both newcomers and seasoned adopters, repeatedly asked the same question: “Can you help us build a P&L ...","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[980,947,1026,1079,997,1194,1260,1008,1261],"MI325x","pnl calculator",{"path":1263,"title":1264,"description":1265,"date":1266,"slug":1267,"image":1268,"originalUrl":1269,"categories":1270},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","Cranking Out Faster Tokens for Fewer Dollars: AMD MI300X vs. NVIDIA H200","Qwen3-32B on Paiton + AMD MI300x vs.NVIDIA H200 1. Introduction “While we’re actively training models for local customers, automating and streamlining critical business processes, we still found time to push our Paiton framework to the limit on Qwen3-32B.” In the competitive realm of LLMs, next-gen hardware like the NVIDIA H200 often steals the headlines. But ...","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[980,947,86,1193,996,1194,1271,1008,86,954],"MI300",{"path":1273,"title":1274,"description":1275,"date":1276,"slug":1277,"image":1278,"originalUrl":1279,"categories":1280},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Power Meets Precision: High-Density Modular Data Center for NVIDIA NVL Deployments (1–2 MW)","Purpose-Built High-Density Infrastructure for Blackwell-Class AI Workloads At Eliovp, we’re engineering a new class of AI infrastructure. Our advanced modular platform is built to also support NVIDIA’s cutting-edge NVL architecture, from the efficient NVL4 to the ultra-scale NVL72, enabling deployments that range from distributed edge inference to full-stack model training at hyperscale. Designed to meet ...","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[980,1026,1281,981,1118,984,1119,1120,1282,1283,1123,1284],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1286,"title":1287,"description":1288,"date":1289,"slug":1290,"image":1291,"originalUrl":1292,"categories":1293},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","At Eliovp, we’re constantly keeping up with the newest AI trends. Consequently, we have been looking into AI agents and have created a medical agent designed to seamlessly interact with DICOM servers inside hospitals. This isn’t just another chatbot or AI tool. This is an intelligent assistant that understands the language of radiology and is ...","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[980,947,1079,1027,1193,996,1294],"Healthcare",{"path":1296,"title":1297,"description":1298,"date":1299,"slug":1300,"image":1301,"originalUrl":1302,"categories":1303},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","Eliovp BV: Your Trusted Partner for Supply Chain Resilience Amidst New U.S. Tariffs","In today’s rapidly evolving global trade landscape, businesses face unprecedented challenges in maintaining efficient and cost-effective IT infrastructure. The recent U.S. tariff adjustments have created waves of uncertainty across international markets, particularly for companies relying on high-performance computing and AI solutions. At Eliovp BV, we want to assure our valued clients that our comprehensive end-to-end ...","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[980,1038,1193,996,1304,1305,1306,1307],"import","Taiwan","Tariffs","Trump",{"path":1309,"title":1310,"description":1311,"date":1312,"slug":1313,"image":1314,"originalUrl":1315,"categories":1316},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","1. Versatile IntegrationAI Agents are designed to integrate seamlessly with your existing software stack. This includes ERP, CRM, and marketing automation platforms. Instead of disrupting current systems, they complement and enhance them, all while learning from and adapting to your specific operational needs. 2. Intelligent Decision-MakingConventional automation scripts handle if-then scenarios, but they fall short ...","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[980,947,1079,1193,1317,1318],"AI Agents","ERP",{"path":1320,"title":1321,"description":1322,"date":1323,"slug":1324,"image":1325,"originalUrl":1326,"categories":1327},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","In the rapidly evolving landscape of artificial intelligence (AI), open-source solutions are emerging as pivotal drivers of innovation and performance enhancement. These community-driven platforms democratize access to cutting-edge technologies, fostering collaboration and accelerating advancements in AI model optimization.​ The Open-Source Revolution in AI Open-source AI models have transformed the development and deployment of machine learning ...","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[980,947,1038,1328,996,793,1008],"AI news",{"path":1330,"title":1331,"description":1332,"date":1333,"slug":1334,"image":1335,"originalUrl":1336,"categories":1337},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","1. Introduction Benchmarking is an essential part of optimizing AI models and software applications. Whether you’re testing AI model inference speeds, profiling different hardware configurations, or ensuring system performance over time, having a reliable benchmarking tool is crucial. However, many existing tools suffer from issues like inconsistent environments, difficult configuration setups, and lack of automation. ...","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[980,947,86,1193,996,1338,1339,1195,86],"benchmark","LLM",{"path":1341,"title":1342,"description":1343,"date":1344,"slug":1345,"image":1346,"originalUrl":1347,"categories":1348},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","1. Introduction In the world of large language models (LLMs), most benchmarks center on Llama or DeepSeek derivatives. We decided to diversify by adding the Qwen2 architecture, using our Paiton framework. This 32-billion-parameter model pushes GPU resources to the limit, perfect for comparing NVIDIA’s new H200 to our AMD MI300X, which leverages Paiton for advanced ...","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[980,947,86],{"path":1350,"title":1351,"description":1352,"date":1353,"slug":1354,"image":1355,"originalUrl":1356,"categories":1357},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","We’re excited to share that Eliovp was recently featured on AMD’s “Tech Talk” podcast! In this episode, our CEO, Elio Van Puyvelde sits down with Jim greene to talk about the origins of Eliovp, the passion and expertise that brought the company to life, and the innovative full end-to-end solutions we offer today. From our ...","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[980,996,1358,1359,1360],"Jim Greene","Podcast","Tech Talk",{"path":1362,"title":1363,"description":1364,"date":1365,"slug":1366,"image":1367,"originalUrl":1368,"categories":1369},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Executive Summary If you’ve followed our journey so far, you’ll know that Paiton is laser-focused on AMD-centric inference optimization. Our latest work takes DeepSeek R1 Distill Llama 8B to the next level, delivering 10–15% higher throughput, improved time-to-first-token (TTFT), and more stable performance at lower batch sizes, an area that previously needed a boost. In short, Paiton further cements its ability to exploit AMD hardware’s raw power, bridging ...","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[980,947,86,996,1370,1371,1194,1195,1260,86,954],"Deepseek","H100",{"path":1373,"title":1374,"description":1375,"date":1376,"slug":1377,"image":1378,"originalUrl":1379,"categories":1380},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","A First Look at Paiton in Action: Deepseek R1 Distill Llama 3.1 8B","Outperforming Stock Models on the AMD MI300X 1. Introduction We couldn’t wait to show what Paiton can really do. After detailing our AMD-centric approach and architecture-level optimizations in our previous blog post, we decided to test-drive Paiton on a hype-worthy model: Deepseek R1 Distill Llama 3.1 8B. By compiling the model into efficient libraries and fusing ...","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[980,947,86,996,1370,1371,1194,1195,1260,86,954],{"path":1382,"title":1383,"description":1384,"date":1385,"slug":1386,"image":1387,"originalUrl":1388,"categories":1389},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","In the fast-paced world of artificial intelligence, model efficiency and performance are paramount. At ElioVP, we’re redefining what’s possible by delivering unparalleled optimization solutions for AI models with Paiton. By compiling the model’s architecture and leveraging our custom-written kernels, Paiton enables faster inference and reduced resource consumption on AMD GPUs. Why Model Optimization is More ...","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[980,947,86,996,1371,1194,1195,1260,86,954],1788530427962]