[{"data":1,"prerenderedAt":1437},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":991},{"id":4,"title":5,"body":6,"categories":973,"date":979,"description":980,"extension":981,"heading":982,"image":983,"meta":984,"navigation":985,"originalUrl":986,"path":986,"seo":987,"slug":988,"stem":989,"__hash__":990},"blog\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700.md","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM",{"type":7,"value":8,"toc":961},"minimark",[9,13,21,44,55,61,67,72,75,140,166,173,180,187,196,200,203,209,258,269,280,284,287,290,296,350,356,359,366,370,385,391,394,397,403,465,468,479,483,486,525,532,543,546,550,553,575,588,594,601,605,612,615,741,744,751,755,776,779,797,801,804,844,864,867,920,932,936,948,957],[10,11,12],"p",{},"A product concept. A woodland photograph. A rainy street in watercolor. The useful part of local image generation is being able to change the prompt and try again, without sending the request to an external service.",[10,14,15,16,20],{},"Our free Paiton profile brings that workflow to ",[17,18,19],"strong",{},"FLUX.2 klein 4B on one AMD Radeon AI PRO R9700",", with a ready-to-run ComfyUI setup.",[10,22,23,24,27,28,31,32,35,36,39,40,43],{},"At ",[17,25,26],{},"1024 × 1024, four steps and batch size one",", complete warm generation averages ",[17,29,30],{},"1.054 seconds per image",", against ",[17,33,34],{},"1.258 seconds"," for the strongest stock configuration we qualified. That is ",[17,37,38],{},"16.2% lower generation latency"," and ",[17,41,42],{},"19.4% more projected images per hour",".",[10,45,46,47,50,51,54],{},"Memory use falls too: peak Torch allocation drops from ",[17,48,49],{},"19.3 to 12.9 GiB",", a reported ",[17,52,53],{},"33.4% reduction",". The text encoder, transformer and VAE stay on the GPU, with no CPU offload.",[10,56,57,60],{},[17,58,59],{},"The headline timing is warm prompt-to-image generation, not startup or browser-click latency."," It includes text encoding, denoising, VAE decoding and conversion to a PIL image; it excludes PNG encoding, file writing and the interface.",[10,62,63],{},[64,65,66],"em",{},"The fox pictured above is an actual retained benchmark output from the Paiton FLUX.2 klein profile, seed 42.",[68,69,71],"h2",{"id":70},"start-creating-locally","Start creating locally",[10,73,74],{},"On a Linux R9700 workstation with Docker, the Compose plugin and a working AMD GPU driver, launch the included workflow with:",[76,77,82],"pre",{"className":78,"code":79,"language":80,"meta":81,"style":81},"language-bash shiki shiki-themes github-light github-dark","git clone --depth 1 \\\n  --branch paiton-flux2-klein-gfx1201-v1.0.1 \\\n  https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\ncd paiton-vllm-plugin\n.\u002Fmodels\u002FFLUX.2-klein\u002Flaunch.sh\n","bash","",[83,84,85,108,119,125,134],"code",{"__ignoreMap":81},[86,87,90,94,98,102,105],"span",{"class":88,"line":89},"line",1,[86,91,93],{"class":92},"sScJk","git",[86,95,97],{"class":96},"sZZnC"," clone",[86,99,101],{"class":100},"sj4cs"," --depth",[86,103,104],{"class":100}," 1",[86,106,107],{"class":100}," \\\n",[86,109,111,114,117],{"class":88,"line":110},2,[86,112,113],{"class":100},"  --branch",[86,115,116],{"class":96}," paiton-flux2-klein-gfx1201-v1.0.1",[86,118,107],{"class":100},[86,120,122],{"class":88,"line":121},3,[86,123,124],{"class":96},"  https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\n",[86,126,128,131],{"class":88,"line":127},4,[86,129,130],{"class":100},"cd",[86,132,133],{"class":96}," paiton-vllm-plugin\n",[86,135,137],{"class":88,"line":136},5,[86,138,139],{"class":92},".\u002Fmodels\u002FFLUX.2-klein\u002Flaunch.sh\n",[10,141,142,143,150,151,154,155,158,159,162,163,43],{},"Open ",[144,145,149],"a",{"href":146,"rel":147},"http:\u002F\u002F127.0.0.1:8188\u002F?paiton=1",[148],"nofollow","ComfyUI at localhost:8188",", edit the prompt and click ",[17,152,153],{},"Run",". The connected workflow includes an engine selector for ",[17,156,157],{},"Paiton"," or ",[17,160,161],{},"Stock (Diffusers)",", a seed control, an image preview and a Save Image node. Images are saved in ",[83,164,165],{},"paiton-images\u002F",[10,167,168],{},[169,170],"img",{"alt":171,"src":172},"The included ComfyUI workflow, showing the Paiton engine selector, prompt and image preview","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fcomfyui-workflow.webp",[10,174,175,176,179],{},"The first image loads and compiles the selected engine and can take several minutes. Later prompts reuse it. Switching engines unloads the previous model and incurs setup again; they are not held in GPU memory together. To compare outputs, keep the prompt and seed unchanged and select ",[17,177,178],{},"fixed"," in the seed control.",[10,181,182,183,186],{},"The tested host has 16 GB of RAM. We recommend ",[17,184,185],{},"24 GB of system RAM and 60 GB of free disk space"," for compilation headroom, containers, model files, caches and outputs. Downloads and caches persist between starts. Once setup is complete, generation stays local; no paid inference service or cloud GPU is required.",[10,188,189,190,195],{},"The ",[144,191,194],{"href":192,"rel":193},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Ftree\u002Fmain\u002Fmodels\u002FFLUX.2-klein",[148],"model guide"," also covers a lightweight prompt interface, terminal generation and adding the node to an existing ComfyUI installation.",[68,197,199],{"id":198},"what-the-speed-result-includes","What the speed result includes",[10,201,202],{},"Both engines use the same model checkpoint and generation settings. The comparison measures the whole warm generation loop rather than extrapolating from a faster individual kernel.",[10,204,205],{},[169,206],{"alt":207,"src":208},"Warm generation takes 1.054 seconds with Paiton versus 1.258 seconds with qualified stock, with projected hourly output of 3,416 versus 2,862 images","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fgeneration-performance.webp",[210,211,212,228],"table",{},[213,214,215],"thead",{},[216,217,218,222,226],"tr",{},[219,220,221],"th",{},"Generation metric",[219,223,225],{"align":224},"right","Qualified stock + weight cache",[219,227,157],{"align":224},[229,230,231,245],"tbody",{},[216,232,233,237,240],{},[234,235,236],"td",{},"Mean warm seconds per image",[234,238,239],{"align":224},"1.258",[234,241,242],{"align":224},[17,243,244],{},"1.054",[216,246,247,250,253],{},[234,248,249],{},"Projected images per hour",[234,251,252],{"align":224},"2,862",[234,254,255],{"align":224},[17,256,257],{},"3,416",[10,259,260,261,264,265,268],{},"Projected output is ",[83,262,263],{},"3600 \u002F mean generation seconds",". It is ",[17,266,267],{},"not an hour-long throughput measurement",", and it excludes PNG writing and time spent editing prompts. The 19.4% output increase and 16.2% latency reduction express the same improvement in different ways.",[10,270,271,272,275,276,279],{},"The interface has its own costs. In a separate ComfyUI validation, two warm workflow executions took ",[17,273,274],{},"1.408 and 1.404 seconds",", averaging ",[17,277,278],{},"1.406 seconds"," including local engine transport, image handling and saving. Those server-side timings exclude browser display and are not the headline prompt-to-PIL benchmark.",[68,281,283],{"id":282},"lower-memory-use-without-cpu-offload","Lower memory use, without CPU offload",[10,285,286],{},"A model's download size is not its runtime memory requirement. The full pipeline also needs its text encoder, decoder, activations, temporary buffers and compilation allocations.",[10,288,289],{},"Paiton reduces all three reported memory measures in this comparison:",[10,291,292],{},[169,293],{"alt":294,"src":295},"Three separate memory measurements for stock and Paiton: peak Torch allocation 19.3 versus 12.9 GiB, peak reservation 22.1 versus 14.1 GiB, and maximum sampled driver VRAM 23.0 versus 14.6 GiB","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fpipeline-memory.webp",[210,297,298,309],{},[213,299,300],{},[216,301,302,305,307],{},[219,303,304],{},"Memory measurement",[219,306,225],{"align":224},[219,308,157],{"align":224},[229,310,311,324,337],{},[216,312,313,316,319],{},[234,314,315],{},"Peak Torch allocation",[234,317,318],{"align":224},"19.3 GiB",[234,320,321],{"align":224},[17,322,323],{},"12.9 GiB",[216,325,326,329,332],{},[234,327,328],{},"Peak Torch reservation",[234,330,331],{"align":224},"22.1 GiB",[234,333,334],{"align":224},[17,335,336],{},"14.1 GiB",[216,338,339,342,345],{},[234,340,341],{},"Maximum sampled driver VRAM",[234,343,344],{"align":224},"23.0 GiB",[234,346,347],{"align":224},[17,348,349],{},"14.6 GiB",[10,351,352,355],{},[17,353,354],{},"12.9 GiB is not total VRAM consumption."," Allocation records tensor memory; reservation records memory held by the Torch allocator; driver samples include memory outside that allocator. These views overlap and must not be added together.",[10,357,358],{},"Allocation and reservation peaks include compilation, warmup and generation. Driver sampling also includes loading. Values are rounded and use binary GiB; the card is marketed as 32 GB.",[10,360,361,362,365],{},"The practical benefit is more memory headroom on the tested R9700 while retaining GPU-resident generation. ",[17,363,364],{},"These measurements do not establish support for a 16 GB GPU or any other card."," This release qualifies the R9700 only.",[68,367,369],{"id":368},"same-settings-inspectable-images-not-identical-pixels","Same settings, inspectable images, not identical pixels",[10,371,372,373,376,377,380,381,384],{},"The retained examples cover wildlife photography at seed ",[17,374,375],{},"42",", product photography at ",[17,378,379],{},"31415",", and a watercolor bookshop at ",[17,382,383],{},"2026",". All final latents were finite, and each fixed-seed result repeated exactly within its benchmark process.",[10,386,387],{},[169,388],{"alt":389,"src":390},"The retained stock and Paiton product images, with a teal mug, lemon and PAITON card; both use seed 31415","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fproduct-comparison.webp",[10,392,393],{},"Both product images retain the teal cup, yellow lemon, window light and readable “PAITON” card. The fox pair preserves the overall pose and woodland lighting. Fine textures, reflections and shadows differ.",[10,395,396],{},"The bookshop pair shows a larger difference in scene details, including masonry, shelving and lettering. Both retain the watercolor style, wet street, red raincoat and bicycle. The person stands beside the bicycle rather than visibly riding it.",[10,398,399],{},[169,400],{"alt":401,"src":402},"The retained stock and Paiton watercolor bookshop images at seed 2026, showing similar composition with differences in scene details","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fbookshop-comparison.webp",[210,404,405,421],{},[213,406,407],{},[216,408,409,412,415,418],{},[219,410,411],{},"Prompt",[219,413,414],{"align":224},"RGB SSIM",[219,416,417],{"align":224},"RGB PSNR",[219,419,420],{"align":224},"Final latent relative RMSE",[229,422,423,437,451],{},[216,424,425,428,431,434],{},[234,426,427],{},"Fox",[234,429,430],{"align":224},"0.975",[234,432,433],{"align":224},"32.45 dB",[234,435,436],{"align":224},"0.146",[216,438,439,442,445,448],{},[234,440,441],{},"Product",[234,443,444],{"align":224},"0.974",[234,446,447],{"align":224},"30.10 dB",[234,449,450],{"align":224},"0.242",[216,452,453,456,459,462],{},[234,454,455],{},"Bookshop",[234,457,458],{"align":224},"0.859",[234,460,461],{"align":224},"21.38 dB",[234,463,464],{"align":224},"0.289",[10,466,467],{},"These metrics measure similarity to the stock output, not aesthetic quality or complete prompt adherence. Small numerical differences can propagate through denoising, so operator checks are accompanied by final-image inspection.",[10,469,470,471,474,475,478],{},"Repeatability also has a boundary: a separately compiled, empty-cache Paiton run produced a fox image with SSIM ",[17,472,473],{},"0.948"," against the populated-cache Paiton image. The cause was not isolated to one operation. This three-prompt check does ",[17,476,477],{},"not"," establish unchanged quality or identical pixels across all prompts, styles or independently compiled processes.",[68,480,482],{"id":481},"what-to-expect-at-startup","What to expect at startup",[10,484,485],{},"Once the engine is loaded, repeated generation is the fast path. A new process still pays for loading and graph initialization, even with persistent caches.",[210,487,488,501],{},[213,489,490],{},[216,491,492,495,498],{},[219,493,494],{},"Paiton startup condition",[219,496,497],{"align":224},"Process loading",[219,499,500],{"align":224},"First generation",[229,502,503,514],{},[216,504,505,508,511],{},[234,506,507],{},"Compilation caches populated",[234,509,510],{"align":224},"44.7 s",[234,512,513],{"align":224},"19.4 s",[216,515,516,519,522],{},[234,517,518],{},"Compilation caches empty; weights already prepared",[234,520,521],{"align":224},"48.2 s",[234,523,524],{"align":224},"104.1 s",[10,526,527,528,531],{},"The first-generation timings include graph setup and any remaining compilation. The next two warm generations in the empty-cache test took ",[17,529,530],{},"1.050 and 1.052 seconds",". These startup examples are separate from the six-image headline comparison.",[10,533,534,535,538,539,542],{},"Initial provisioning adds the download and model-preparation stages. On our system, the source download took ",[17,536,537],{},"110.7 seconds"," and conversion took ",[17,540,541],{},"88.7 seconds",". Network speed and cache history will change those numbers.",[10,544,545],{},"For repeated use, keep the selected engine loaded and iterate on prompts rather than switching backends after every image.",[68,547,549],{"id":548},"more-images-per-euro","More images per euro",[10,551,552],{},"At the same cost per productive hour, lower generation time gives the workstation more output capacity.",[10,554,555,556,559,560,563,564,567,568,571,572,43],{},"For an illustrative ownership model, assume ",[17,557,558],{},"€2,000"," for the workstation, ",[17,561,562],{},"4,000 productive generation hours",", electricity at ",[17,565,566],{},"€0.30\u002FkWh",", and equal assumed wall power of ",[17,569,570],{},"350 W"," for both engines. Hardware contributes €0.50 per hour and electricity adds €0.105, giving ",[17,573,574],{},"€0.605 per productive hour",[10,576,577,578,39,581,584,585,43],{},"At the reported generation rates, that produces approximately ",[17,579,580],{},"4,730 images per euro for stock",[17,582,583],{},"5,647 for Paiton",", ",[17,586,587],{},"19.4% more modeled output per euro",[10,589,590],{},[169,591],{"alt":592,"src":593},"Modeled generation capacity at the same assumed hourly cost: 4,730 images per euro for stock and 5,647 for Paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fmodeled-cost.webp",[10,595,596,597,600],{},"This is a capacity and cost model, ",[17,598,599],{},"not measured energy savings",". It assumes sustained productive use and excludes idle time, taxes, labor, financing, cooling outside the workstation and PNG writing. It also does not measure how many generated images a user will choose to keep. Lower utilization raises capital cost per useful image.",[68,602,604],{"id":603},"how-the-comparison-was-run","How the comparison was run",[10,606,607,608,611],{},"The baseline was the strongest stock path retained during our qualification, not an untouched default installation. It used ",[17,609,610],{},"Diffusers 0.40.0 and SDNQ 0.2.6",", prepared immutable weights, compiled pipeline components and graph capture. Other tuning options were tested; not every option improved the full pipeline.",[10,613,614],{},"Both engines used the same pinned checkpoint, prompts, fixed seeds, equivalent runtime precision and full GPU residency. Paiton contributes a qualified compiled execution path for this model and Radeon profile. The scheduler, step count and guidance remain unchanged; the result is not obtained by generating a smaller image or running fewer steps.",[210,616,617,627],{},[213,618,619],{},[216,620,621,624],{},[219,622,623],{},"Tested setting",[219,625,626],{},"Value",[229,628,629,637,645,653,661,669,677,685,693,701,711,721,731],{},[216,630,631,634],{},[234,632,633],{},"GPU",[234,635,636],{},"AMD Radeon AI PRO R9700, 32 GB",[216,638,639,642],{},[234,640,641],{},"Model profile",[234,643,644],{},"FLUX.2 klein 4B, text-to-image",[216,646,647,650],{},[234,648,649],{},"Resolution \u002F steps \u002F batch",[234,651,652],{},"1024 × 1024 \u002F 4 \u002F 1",[216,654,655,658],{},[234,656,657],{},"Guidance \u002F text sequence",[234,659,660],{},"1.0 \u002F 512 tokens",[216,662,663,666],{},[234,664,665],{},"Measurement date",[234,667,668],{},"September 7, 2026",[216,670,671,674],{},[234,672,673],{},"Repetitions",[234,675,676],{},"Three fixed prompts; two warmups and two measured runs per prompt",[216,678,679,682],{},[234,680,681],{},"Measured sample",[234,683,684],{},"Six images per backend",[216,686,687,690],{},[234,688,689],{},"Timing boundary",[234,691,692],{},"GPU-synchronized wall clock, prompt to PIL image",[216,694,695,698],{},[234,696,697],{},"GPU policy",[234,699,700],{},"AUTO performance level, COMPUTE profile",[216,702,703,706],{},[234,704,705],{},"Common PyTorch build",[234,707,708],{},[83,709,710],{},"2.12.0+rocm7.14.0",[216,712,713,716],{},[234,714,715],{},"HIP",[234,717,718],{},[83,719,720],{},"7.14.60850",[216,722,723,726],{},[234,724,725],{},"Triton",[234,727,728],{},[83,729,730],{},"3.7.1+git0263a6a6.rocm7.14.0",[216,732,733,736],{},[234,734,735],{},"Transformers",[234,737,738],{},[83,739,740],{},"5.15.1",[10,742,743],{},"No clock, power, voltage or fan limits were changed. Automatic clocks and temperatures varied. This small sample describes the tested workstation and settings; it is not a universal speed guarantee.",[10,745,746,747,750],{},"The supported release is deliberately specific: ",[17,748,749],{},"Linux, R9700, text-to-image, 1024-square, four steps, batch one",". Image editing, adapters, other resolutions and other GPUs are outside its qualified scope.",[68,752,754],{"id":753},"model-provenance-and-release-notes","Model provenance and release notes",[10,756,757,758,763,764,767,768,771,772,775],{},"The package downloads ",[144,759,762],{"href":760,"rel":761},"https:\u002F\u002Fhuggingface.co\u002FDisty0\u002FFLUX.2-klein-4B-SDNQ-4bit-dynamic",[148],"Disty0's SDNQ checkpoint",", pinned to revision ",[83,765,766],{},"45e9cc76cb70f84473ce5c6c2e2282d0ef3c6ecd",". The download is about ",[17,769,770],{},"5.46 GB","; a separate preparation stage produces about ",[17,773,774],{},"12 GB"," of tensor files. Checkpoint size should not be confused with runtime precision or GPU memory usage.",[10,777,778],{},"The release notes identify the original FLUX.2 klein 4B weights as Apache 2.0. They also flag a contradictory non-commercial link in the community checkpoint's metadata and an unpinned pre-quantization source revision. Weights are downloaded separately; those provenance limitations should be reviewed before commercial redistribution.",[10,780,781,782,584,787,39,792,796],{},"The SDNQ conversion\u002Fstock tools and ComfyUI retain their upstream source and notices. The Paiton inference path does not import SDNQ. See the ",[144,783,786],{"href":784,"rel":785},"https:\u002F\u002Fhuggingface.co\u002Fblack-forest-labs\u002FFLUX.2-klein-4B",[148],"official model card",[144,788,791],{"href":789,"rel":790},"https:\u002F\u002Fgithub.com\u002FDisty0\u002Fsdnq",[148],"SDNQ source",[144,793,795],{"href":192,"rel":794},[148],"release guide"," for the associated documentation.",[68,798,800],{"id":799},"reproduce-the-workflow","Reproduce the workflow",[10,802,803],{},"From the cloned repository, enter the model directory and generate an image:",[76,805,807],{"className":78,"code":806,"language":80,"meta":81,"style":81},"cd models\u002FFLUX.2-klein\n.\u002Frun.sh generate \\\n  --prompt 'A teal ceramic coffee cup beside a lemon, soft window light, product photograph' \\\n  --seed 42\n",[83,808,809,816,826,836],{"__ignoreMap":81},[86,810,811,813],{"class":88,"line":89},[86,812,130],{"class":100},[86,814,815],{"class":96}," models\u002FFLUX.2-klein\n",[86,817,818,821,824],{"class":88,"line":110},[86,819,820],{"class":92},".\u002Frun.sh",[86,822,823],{"class":96}," generate",[86,825,107],{"class":100},[86,827,828,831,834],{"class":88,"line":121},[86,829,830],{"class":100},"  --prompt",[86,832,833],{"class":96}," 'A teal ceramic coffee cup beside a lemon, soft window light, product photograph'",[86,835,107],{"class":100},[86,837,838,841],{"class":88,"line":127},[86,839,840],{"class":100},"  --seed",[86,842,843],{"class":100}," 42\n",[10,845,846,847,850,851,854,855,858,859,43],{},"Use ",[83,848,849],{},"--count 4"," to try consecutive seeds in one loaded process. The default output is ",[83,852,853],{},"outputs\u002Fimage.png",". For the lightweight interface, run ",[83,856,857],{},".\u002Flaunch.sh --ui simple"," and open ",[144,860,863],{"href":861,"rel":862},"http:\u002F\u002F127.0.0.1:7860",[148],"localhost:7860",[10,865,866],{},"Stop other generation services before comparing both engines:",[76,868,870],{"className":78,"code":869,"language":80,"meta":81,"style":81},".\u002Flaunch.sh --stop\n.\u002Frun.sh benchmark --backend stock --suite --output \u002Foutputs\u002Fstock\n.\u002Frun.sh benchmark --backend paiton --suite --output \u002Foutputs\u002Fpaiton\n",[83,871,872,880,902],{"__ignoreMap":81},[86,873,874,877],{"class":88,"line":89},[86,875,876],{"class":92},".\u002Flaunch.sh",[86,878,879],{"class":100}," --stop\n",[86,881,882,884,887,890,893,896,899],{"class":88,"line":110},[86,883,820],{"class":92},[86,885,886],{"class":96}," benchmark",[86,888,889],{"class":100}," --backend",[86,891,892],{"class":96}," stock",[86,894,895],{"class":100}," --suite",[86,897,898],{"class":100}," --output",[86,900,901],{"class":96}," \u002Foutputs\u002Fstock\n",[86,903,904,906,908,910,913,915,917],{"class":88,"line":121},[86,905,820],{"class":92},[86,907,886],{"class":96},[86,909,889],{"class":100},[86,911,912],{"class":96}," paiton",[86,914,895],{"class":100},[86,916,898],{"class":100},[86,918,919],{"class":96}," \u002Foutputs\u002Fpaiton\n",[10,921,189,922,927,928,931],{},[144,923,926],{"href":924,"rel":925},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FFLUX.2-klein-4B-Paiton-RDNA4",[148],"compiled artifacts on Hugging Face"," are already included in the containers. The release guide covers the build recipes, runtime bindings, workflow, launchers and notices. Use ",[83,929,930],{},".\u002Flaunch.sh --build"," to build the containers locally.",[68,933,935],{"id":934},"more-useful-work-from-amd-hardware","More useful work from AMD hardware",[10,937,938,939,39,943,947],{},"Our recent ",[144,940,942],{"href":941},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Qwen3.8",[144,944,946],{"href":945},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5"," releases focused on local language-model serving. This profile brings the same practical focus to a visual workflow: faster iteration, less memory pressure and a ComfyUI setup people can use on their own machine.",[10,949,950,953,954,43],{},[17,951,952],{},"Building image-generation services or running inference at scale on AMD CDNA? Talk to us about your workload."," Paiton's wider work targets throughput, memory efficiency and cost per useful output across AMD deployments. Learn more about ",[144,955,157],{"href":956},"\u002Fproducts\u002Fpaiton",[958,959,960],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":81,"searchDepth":110,"depth":110,"links":962},[963,964,965,966,967,968,969,970,971,972],{"id":70,"depth":110,"text":71},{"id":198,"depth":110,"text":199},{"id":282,"depth":110,"text":283},{"id":368,"depth":110,"text":369},{"id":481,"depth":110,"text":482},{"id":548,"depth":110,"text":549},{"id":603,"depth":110,"text":604},{"id":753,"depth":110,"text":754},{"id":799,"depth":110,"text":800},{"id":934,"depth":110,"text":935},[157,974,975,976,977,978],"AMD Radeon","Local AI","Image Generation","FLUX","ComfyUI","2026-09-07T09:00:00","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","md","Faster local FLUX.2 klein on Radeon, with less memory","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",{},true,"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700",{"title":5,"description":980},"paiton-flux2-klein-radeon-ai-pro-r9700","blog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","VFCKKIJgMKaqYEwhx3x9idSM_yQbaHPi0FFeyy1FPZw",[992,994,1010,1018,1033,1064,1076,1099,1117,1136,1154,1172,1188,1200,1216,1231,1243,1252,1260,1275,1287,1298,1309,1319,1332,1342,1355,1366,1376,1387,1396,1408,1419,1428],{"path":986,"title":5,"description":980,"date":979,"slug":988,"image":983,"originalUrl":986,"categories":993},[157,974,975,976,977,978],{"path":945,"title":995,"description":996,"date":997,"slug":998,"image":999,"originalUrl":1000,"categories":1001},"Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[157,1002,974,1003,1004,1005,1006,1007,1008,1009],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Inference Optimization","Large Language Models","vLLM","Cost Efficiency",{"path":941,"title":1011,"description":1012,"date":1013,"slug":1014,"image":1015,"originalUrl":1016,"categories":1017},"Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[157,1002,974,1003,1004,1005,1006,1007,1008,1009],{"path":1019,"title":1020,"description":1021,"date":1022,"slug":1023,"image":1024,"originalUrl":1025,"categories":1026},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",null,[1027,1028,1029,1030,1031,1032],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1034,"title":1035,"description":1036,"date":1037,"slug":1038,"image":1039,"originalUrl":1040,"categories":1041},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Paiton Returns to Its Diffusion Roots: Optimizing Wan2.2-T2V-A14B on AMD MI355X","When we first started building Paiton, one of our earliest focus areas was optimizing diffusion models. Stable Diffusion XL was one of the first large models where we showed that fused operators, efficient execution, and hardware-aware kernels could make a real difference. Now we are returning to those origins.With the growing interest in text-to-video generation, ...","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1027,1002,157,1042,1043,1044,1045,1046,1047,1048,1049,1050,633,1051,1052,1053,1054,1055,1056,1057,157,1058,1059,1060,1061,1062,1063],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1065,"title":1066,"description":1067,"date":1068,"slug":1069,"image":1070,"originalUrl":1071,"categories":1072},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","From the Attic to the Front Page: ElioVP Recognized as a Pioneer in Chip Optimization & Data Center Infrastructure","It has been some incredible weeks for the team here at Eliovp. We are extremely proud to share that our company was recently featured on the front page of De Tijd, Belgium’s leading business newspaper. Seeing our story, from our founder’s early days tinkering with wires in an attic to generating €215 million in revenue, ...","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1027,1002,1073,1074,1043,1075,1073,1055],"Modular DC","Uncategorized","De Tijd",{"path":1077,"title":1078,"description":1079,"date":1080,"slug":1081,"image":1082,"originalUrl":1083,"categories":1084},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","Privacy is geen IT-probleem meer, het is een strategische prioriteit: ","Privacyrisico’s, Psychologische Valkuilen en de Operationele Realiteit van Generatieve AI in de Benelux De recente verschijning van een frontpage-artikel over ons bedrijf in het gerespecteerde dagblad “De Tijd” heeft onze zichtbaarheid aanzienlijk vergroot, wat de aanleiding is voor dit artikel. Deze mediabelangstelling, gecombineerd met de talrijke uitnodigingen voor spreekbeurten die we hebben ontvangen, fungeert als ...","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1027,1002,1085,1074,1086,1087,1088,1089,1090,1091,1092,1093,1049,1094,1095,1096,1097,1098],"Trending","AI Act","Antropomorfisme","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Generatieve AI","Microsoft Copilot","Privacy","Shadow AI",{"path":1100,"title":1101,"description":1102,"date":1103,"slug":1104,"image":1105,"originalUrl":1106,"categories":1107},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Itsme? Bij ons is het “it’s not me”, en dit is waarom.","1. Inleiding: De Strategische Noodzaak van Weigering In het hedendaagse digitale landschap wordt de keuze voor een identiteit leverancier (IdP) vaak gereduceerd tot een discussie over User Experience (UX) en conversieratio’s. Deze reductionistische benadering verhult echter de diepgaande strategische, juridische en operationele risico’s die gepaard gaan met het uitbesteden van de “Sleutels tot het Koninkrijk”, ...","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1027,1108,1109,1110,1091,1111,1112,1113,1094,1114,1115,1116,1097],"AWS","Belgian Mobile ID","Cloud Act","Data Soevereiniteit","Digitale Identiteit","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1118,"title":1119,"description":1120,"date":1121,"slug":1122,"image":1123,"originalUrl":1124,"categories":1125},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","From Hype to Sovereign Infrastructure Nederlandse Versie Summary The narrative surrounding “Agentic AI” in 2025 is defined by a sharp contrast between market expectations and engineering reality. While the general public, conditioned by the ease of ChatGPT, expects “miracles” and instant integration, the reality of building autonomous agents for enterprise workflows is a discipline of ...","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1027,1002,1126,1085,1127,1128,1129,1130,1131,1132,1133,1134,1058,1135],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1137,"title":1138,"description":1139,"date":1140,"slug":1141,"image":1142,"originalUrl":1143,"categories":1144},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","How AI Neocloud Infrastructure Turns Circular Cash Flows Into Fake Growth Executive Summary The interval between 2023 and 2025 has birthed a capital allocation phenomenon arguably without precedent: the “Synthetic Bubble.” Driven by the scramble for AI dominance, the venture capital apparatus has directed billions into the “Neocloud” ecosystem. However, a rigorous analysis suggests that ...","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1027,1002,1085,1028,1145,1146,1147,1148,1149,1150,1151,1152,1153],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1155,"title":1156,"description":1157,"date":1158,"slug":1159,"image":1160,"originalUrl":1161,"categories":1162},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","Building the Engine for the AI Race: The 4-Month Path to NVIDIA GB300 NVL72 Power","In artificial intelligence infrastructure, speed is the foundation of competitive differentiation. From model training velocity to inference latency, every millisecond matters. But before any workload executes, there is a critical prerequisite that often determines success or failure: time to market. Traditional builds typically require 18–24 months. In the current AI cycle, that is simply too ...","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1027,1073,1074,1163,1028,1164,1165,1166,1167,1168,1169,1170,1171],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1173,"title":1174,"description":1175,"date":1176,"slug":1177,"image":1178,"originalUrl":1179,"categories":1180},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Every few years, a new solution pops up promising the same dream: On paper, that sounds perfect. Take your existing CUDA applications, swap out the toolchain, and suddenly you’re “portable.” And to be fair: if you’re running research code or trying to get an internal tool to compile on a non-NVIDIA box, that can absolutely ...","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1027,1002,157,1074,1181,1002,1182,1183,1184,1185,715,1186,157,1187],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","Kernel Tuning","ROCm",{"path":1189,"title":1190,"description":1191,"date":1192,"slug":1193,"image":1194,"originalUrl":1195,"categories":1196},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Let’s be honest, we’re not the marketing type.We’ve never taken a cent of outside investment, never burned cash on ad campaigns, and never hired a sales army.We just build things that work. In today’s world, it seems the companies shouting the loudest often get the spotlight, while the ones doing the actual engineering quietly build ...","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1027,1002,157,1003,1197,1181,1009,1198,1006,1186,157,1199,1008],"AMD Instinct","High Throughput","SGLang",{"path":1201,"title":1202,"description":1203,"date":1204,"slug":1205,"image":1206,"originalUrl":1207,"categories":1208},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Stop Overpaying: Paiton MI300X MoE Beats H200\u002FB200 on $\u002F1M Tokens","Short summary: We benchmarked Paiton with our new MoE support on Qwen\u002FQwen3-30B-A3B-Instruct-2507 to compare inference performance across several setups. Each configuration was run five times per batch size and we report the mean across runs. Why this benchmark Most published numbers use synthetic prompts or toy datasets. We focused on realistic conversational workloads (we always ...","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1027,1002,157,1209,1181,1210,1006,1211,1212,1213,1214,157,1215],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1217,"title":1218,"description":1219,"date":1220,"slug":1221,"image":1222,"originalUrl":1223,"categories":1224},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Agentic AI, But Make It Local: From Inbox to Insight to Action","(Nederlandse versie) We’ve built production-ready, local-first agentic AI that plugs into your existing email stack, auto-creates tickets, classifies messages, extracts multi-question threads, reads PDFs, spots invoices\u002Fquotes, analyzes images (yes, damage detection), and pushes structured reports into your systems, no dependency on OpenAI, Google, or Microsoft unless you want it. Tailor-made models trained on your data, ...","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1027,1002,1126,1074,1127,1225,1226,1227,1228,1132,1134,1058,1229,1230],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1232,"title":1233,"description":1234,"date":1235,"slug":1236,"image":1237,"originalUrl":1238,"categories":1239},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Data‑Parallel Benchmarks (8–64 GPUs): H200 Left Behind, B200 Within Reach","At ElioVP, we’re all about pushing AI inference past the limits, and packaging every squeeze of performance into a plug‑and‑play runtime.  Remember our last blog, where Paiton’s FP8 pipeline on AMD’s MI300X completely outclassed NVIDIA’s H200? Well, buckle up, because we’ve gone back to the drawing board. This time, we’re loading Llama-3.1-8B-Instruct-FP8-KV, the leaner, meaner ...","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[1027,1002,157,1074,1240,1043,1044,1241,1242,1054,1055,157,1008],"AI","H200","MI300X",{"path":1244,"title":1245,"description":1246,"date":1247,"slug":1248,"image":1249,"originalUrl":1250,"categories":1251},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Here at Eliovp, we continuously innovate when it comes to building practical solutions. If there’s one core strength, it’s our team’s ability to think outside the box. One key area of focus for us is developing applicable AI solutions, everyday usable AI implementations tailored specifically for our clients’ needs. In this blog, we’ll explore our ...","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1027,1002,1126,1127,1225,1226,1227,1228,1132,1134,1058,1229,1230],{"path":1253,"title":1254,"description":1255,"date":1256,"slug":1257,"image":81,"originalUrl":1258,"categories":1259},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Introduction AI is rapidly transforming every industry, but running large models efficiently remains a major technical and financial challenge. At ElioVP, we specialize in optimizing for AMD accelerators, helping organizations unlock the full potential of their hardware. Today, we’re excited to announce a new offering: free evaluation models that let you test our cutting-edge optimizations ...","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1027,1002,157],{"path":1261,"title":1262,"description":1263,"date":1264,"slug":1265,"image":1266,"originalUrl":1267,"categories":1268},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Paiton: Dramatically Faster Startup and Performance for Llama-3.1-405B","With Paiton, we’re not merely pursuing peak inference speeds, we’re fundamentally reshaping the entire lifecycle of large language model (LLM) deployment. Our latest endeavor pairs AMD’s cutting-edge MI300X GPUs with the colossal Llama-3.1-405B-Instruct-FP8-KV model, achieving groundbreaking milestones: Visual Demonstration: Startup Speed Showcase We’re excited to share a visual demonstration of Paiton’s revolutionary startup performance. Watch ...","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1027,1002,157,1074,1003,1181,1269,1183,1270,1271,1272,157,1273,1274],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1276,"title":1277,"description":1278,"date":1279,"slug":1280,"image":1281,"originalUrl":1282,"categories":1283},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","The world of AI is moving at an unprecedented pace, and efficient inference is key to deploying powerful models in real-world applications. At Eliovp, we’ve consistently pushed the boundaries of AI performance, as highlighted in our previous blogs showcasing significant inference speedups when benchmarking with fp16\u002Fbf16. Now, we’re thrilled to announce a further significant leap ...","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[1027,1002,157,1074,1181,1284,1131,1050,1004,1005,1007,1271,1285,1286],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1288,"title":1289,"description":1290,"date":1291,"slug":1292,"image":1293,"originalUrl":1294,"categories":1295},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","As large language models (LLMs) become a foundational part of modern applications, picking the right server for deployment is more important than ever. Whether you’re an enterprise scaling up inference, a startup optimizing for cost, or a researcher pushing throughput boundaries. This blog compares two high-profile server setups and two not so high-profile setups which ...","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1027,1002,157,1126,1074,1043,1242,1055,1296,1297],"RX7900XTX","tenstorrent",{"path":1299,"title":1300,"description":1301,"date":1302,"slug":1303,"image":1304,"originalUrl":1305,"categories":1306},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Empowering GPU Cluster Investors with Real-World Financial Insights","At Eliovp BV, we’ve spent years on the cutting edge of GPU cluster deployment and optimization across Europe. Our team supports leading organizations in AI, finance, and research, architecting, building, and scaling high-performance infrastructure. Over time, our customers, both newcomers and seasoned adopters, repeatedly asked the same question: “Can you help us build a P&L ...","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1027,1002,1073,1126,1044,1241,1307,1055,1308],"MI325x","pnl calculator",{"path":1310,"title":1311,"description":1312,"date":1313,"slug":1314,"image":1315,"originalUrl":1316,"categories":1317},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","Cranking Out Faster Tokens for Fewer Dollars: AMD MI300X vs. NVIDIA H200","Qwen3-32B on Paiton + AMD MI300x vs.NVIDIA H200 1. Introduction “While we’re actively training models for local customers, automating and streamlining critical business processes, we still found time to push our Paiton framework to the limit on Qwen3-32B.” In the competitive realm of LLMs, next-gen hardware like the NVIDIA H200 often steals the headlines. But ...","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1027,1002,157,1240,1043,1241,1318,1055,157,1008],"MI300",{"path":1320,"title":1321,"description":1322,"date":1323,"slug":1324,"image":1325,"originalUrl":1326,"categories":1327},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Power Meets Precision: High-Density Modular Data Center for NVIDIA NVL Deployments (1–2 MW)","Purpose-Built High-Density Infrastructure for Blackwell-Class AI Workloads At Eliovp, we’re engineering a new class of AI infrastructure. Our advanced modular platform is built to also support NVIDIA’s cutting-edge NVL architecture, from the efficient NVL4 to the ultra-scale NVL72, enabling deployments that range from distributed edge inference to full-stack model training at hyperscale. Designed to meet ...","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1027,1073,1328,1028,1165,1031,1166,1167,1329,1330,1170,1331],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1333,"title":1334,"description":1335,"date":1336,"slug":1337,"image":1338,"originalUrl":1339,"categories":1340},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","At Eliovp, we’re constantly keeping up with the newest AI trends. Consequently, we have been looking into AI agents and have created a medical agent designed to seamlessly interact with DICOM servers inside hospitals. This isn’t just another chatbot or AI tool. This is an intelligent assistant that understands the language of radiology and is ...","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1027,1002,1126,1074,1240,1043,1341],"Healthcare",{"path":1343,"title":1344,"description":1345,"date":1346,"slug":1347,"image":1348,"originalUrl":1349,"categories":1350},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","Eliovp BV: Your Trusted Partner for Supply Chain Resilience Amidst New U.S. Tariffs","In today’s rapidly evolving global trade landscape, businesses face unprecedented challenges in maintaining efficient and cost-effective IT infrastructure. The recent U.S. tariff adjustments have created waves of uncertainty across international markets, particularly for companies relying on high-performance computing and AI solutions. At Eliovp BV, we want to assure our valued clients that our comprehensive end-to-end ...","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1027,1085,1240,1043,1351,1352,1353,1354],"import","Taiwan","Tariffs","Trump",{"path":1356,"title":1357,"description":1358,"date":1359,"slug":1360,"image":1361,"originalUrl":1362,"categories":1363},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","1. Versatile IntegrationAI Agents are designed to integrate seamlessly with your existing software stack. This includes ERP, CRM, and marketing automation platforms. Instead of disrupting current systems, they complement and enhance them, all while learning from and adapting to your specific operational needs. 2. Intelligent Decision-MakingConventional automation scripts handle if-then scenarios, but they fall short ...","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1027,1002,1126,1240,1364,1365],"AI Agents","ERP",{"path":1367,"title":1368,"description":1369,"date":1370,"slug":1371,"image":1372,"originalUrl":1373,"categories":1374},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","In the rapidly evolving landscape of artificial intelligence (AI), open-source solutions are emerging as pivotal drivers of innovation and performance enhancement. These community-driven platforms democratize access to cutting-edge technologies, fostering collaboration and accelerating advancements in AI model optimization. The Open-Source Revolution in AI Open-source AI models have transformed the development and deployment of machine learning ...","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1027,1002,1085,1375,1043,633,1055],"AI news",{"path":1377,"title":1378,"description":1379,"date":1380,"slug":1381,"image":1382,"originalUrl":1383,"categories":1384},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","1. Introduction Benchmarking is an essential part of optimizing AI models and software applications. Whether you’re testing AI model inference speeds, profiling different hardware configurations, or ensuring system performance over time, having a reliable benchmarking tool is crucial. However, many existing tools suffer from issues like inconsistent environments, difficult configuration setups, and lack of automation. ...","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1027,1002,157,1240,1043,1385,1386,1242,157],"benchmark","LLM",{"path":1388,"title":1389,"description":1390,"date":1391,"slug":1392,"image":1393,"originalUrl":1394,"categories":1395},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","1. Introduction In the world of large language models (LLMs), most benchmarks center on Llama or DeepSeek derivatives. We decided to diversify by adding the Qwen2 architecture, using our Paiton framework. This 32-billion-parameter model pushes GPU resources to the limit, perfect for comparing NVIDIA’s new H200 to our AMD MI300X, which leverages Paiton for advanced ...","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1027,1002,157],{"path":1397,"title":1398,"description":1399,"date":1400,"slug":1401,"image":1402,"originalUrl":1403,"categories":1404},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","We’re excited to share that Eliovp was recently featured on AMD’s “Tech Talk” podcast! In this episode, our CEO, Elio Van Puyvelde sits down with Jim greene to talk about the origins of Eliovp, the passion and expertise that brought the company to life, and the innovative full end-to-end solutions we offer today. From our ...","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1027,1043,1405,1406,1407],"Jim Greene","Podcast","Tech Talk",{"path":1409,"title":1410,"description":1411,"date":1412,"slug":1413,"image":1414,"originalUrl":1415,"categories":1416},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Executive Summary If you’ve followed our journey so far, you’ll know that Paiton is laser-focused on AMD-centric inference optimization. Our latest work takes DeepSeek R1 Distill Llama 8B to the next level, delivering 10–15% higher throughput, improved time-to-first-token (TTFT), and more stable performance at lower batch sizes, an area that previously needed a boost. In short, Paiton further cements its ability to exploit AMD hardware’s raw power, bridging ...","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1027,1002,157,1043,1417,1418,1241,1242,1307,157,1008],"Deepseek","H100",{"path":1420,"title":1421,"description":1422,"date":1423,"slug":1424,"image":1425,"originalUrl":1426,"categories":1427},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","A First Look at Paiton in Action: Deepseek R1 Distill Llama 3.1 8B","Outperforming Stock Models on the AMD MI300X 1. Introduction We couldn’t wait to show what Paiton can really do. After detailing our AMD-centric approach and architecture-level optimizations in our previous blog post, we decided to test-drive Paiton on a hype-worthy model: Deepseek R1 Distill Llama 3.1 8B. By compiling the model into efficient libraries and fusing ...","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1027,1002,157,1043,1417,1418,1241,1242,1307,157,1008],{"path":1429,"title":1430,"description":1431,"date":1432,"slug":1433,"image":1434,"originalUrl":1435,"categories":1436},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","In the fast-paced world of artificial intelligence, model efficiency and performance are paramount. At ElioVP, we’re redefining what’s possible by delivering unparalleled optimization solutions for AI models with Paiton. By compiling the model’s architecture and leveraging our custom-written kernels, Paiton enables faster inference and reduced resource consumption on AMD GPUs. Why Model Optimization is More ...","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1027,1002,157,1043,1418,1241,1242,1307,157,1008],1788854468424]