[{"data":1,"prerenderedAt":1960},["ShallowReactive",2],{"blog-post-\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach":3,"blog-posts-sidebar":1549},{"id":4,"title":5,"body":6,"categories":1527,"date":1537,"description":1538,"extension":1539,"image":1540,"meta":1541,"navigation":1542,"originalUrl":1543,"path":1544,"seo":1545,"slug":1546,"stem":1547,"__hash__":1548},"blog\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach.md","MI300X FP8 Data‑Parallel Benchmarks (8–64 GPUs): H200 Left Behind, B200 Within Reach",{"type":7,"value":8,"toc":1515},"minimark",[9,13,24,27,30,36,82,85,90,101,108,111,250,253,259,262,270,278,281,292,300,303,309,312,320,323,329,335,338,347,350,357,363,370,376,379,382,388,391,397,403,405,411,417,420,426,433,436,439,444,447,452,455,461,466,776,782,786,976,979,993,998,1001,1004,1086,1131,1136,1172,1178,1181,1293,1298,1303,1325,1331,1334,1337,1340,1347,1364,1370,1396,1402,1405,1408,1411,1414,1420,1423,1432,1435,1443,1449],[10,11,12],"p",{},"At ElioVP, we’re all about pushing AI inference past the limits, and packaging every squeeze of performance into a plug‑and‑play runtime.",[10,14,15,16,23],{},"Remember ",[17,18,22],"a",{"href":19,"rel":20},"https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[21],"nofollow","our last blog",", where Paiton’s FP8 pipeline on AMD’s MI300X completely outclassed NVIDIA’s H200? Well, buckle up, because we’ve gone back to the drawing board.",[10,25,26],{},"This time, we’re loading Llama-3.1-8B-Instruct-FP8-KV, the leaner, meaner FP8‑quantized Llama variant, into not 8 GPUs but 64 virtual GPUs carved out of a single MI300X server.",[10,28,29],{},"Powered by vLLM and Paiton’s kernel magic, we expected modest gains in multi‑tenant scaling…what we got instead was an unexpected, yet amazing surprise and a near dead‑heat with Nvidia’s B200’s.",[10,31,32],{},[33,34,35],"strong",{},"Why did we do this?",[37,38,39,46,52,58,64,70,76],"ul",{},[40,41,42,45],"li",{},[33,43,44],{},"Maximize utilization",": Slice the silicon so every tenant only pays for, and uses, exactly the VRAM and compute they need.",[40,47,48,51],{},[33,49,50],{},"Elastic multi‑tenancy",": Spin up isolated vGPUs in seconds, eliminating noisy‑neighbor slowdowns and siloed resource contention.",[40,53,54,57],{},[33,55,56],{},"Granular SLAs",": Tailor QoS per slice, ultra‑low latency for chatbots, bulk throughput for batch jobs, without juggling hardware.",[40,59,60,63],{},[33,61,62],{},"Cost‑efficient scaling",": Right‑size your compute footprint (and your budget) by renting mini‑GPUs instead of the whole chip.",[40,65,66,69],{},[33,67,68],{},"Rapid CI\u002FCD provisioning",": Integrate GPU slices into your pipeline for instant A\u002FB tests, blue\u002Fgreen rollouts, and regression benchmarks.",[40,71,72,75],{},[33,73,74],{},"Fault isolation",": Contain OOMs and driver hiccups at the slice level, so one bad job doesn’t take down the entire server.",[40,77,78,81],{},[33,79,80],{},"Future‑proof flexibility",": Re‑slice on the fly to match new model footprints or quant formats, no forklift upgrades required.",[10,83,84],{},"With these building blocks in place, we set out to see how far Paiton could stretch inference on a partitioned MI300X, and the numbers? Let’s just say they’ll make you sit up and take notice.",[10,86,87],{},[33,88,89],{},"Goals",[37,91,92,95,98],{},[40,93,94],{},"Evaluate the inference scalability of Paiton on MI300X when using GPU partitioning.",[40,96,97],{},"Measure latency and throughput of Llama 3.1 8B in FP8 format using vLLM.",[40,99,100],{},"Validate memory efficiency and kernel fusion benefits of plug-and-play Paiton models.",[102,103,105],"h3",{"id":104},"benchmarking-testbed-methodology",[33,106,107],{},"Benchmarking Testbed & Methodology",[10,109,110],{},"Our benchmarking method follows a clear set of rules and steps. This makes sure our tests are open and reproducible.",[37,112,113,130,169,180,196,212,218],{},[40,114,115,118,119],{},[33,116,117],{},"Hardware Configuration",":\n",[37,120,121,124,127],{},[40,122,123],{},"8 x AMD MI300x",[40,125,126],{},"8 x Nvidia H200",[40,128,129],{},"8 x Nvidia B200",[40,131,132,118,135],{},[33,133,134],{},"Inference Library",[37,136,137,145,152,161],{},[40,138,139,140],{},"AMD MI300x (Paiton): ",[17,141,144],{"href":142,"rel":143},"https:\u002F\u002Fgithub.com\u002FROCm\u002Fvllm",[21],"vLLM v0.9.0",[40,146,147,148],{},"AMD MI300x (AITER): ",[17,149,151],{"href":142,"rel":150},[21],"v0.9.2",[40,153,154,155,160],{},"NVIDIA H200: ",[17,156,159],{"href":157,"rel":158},"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm",[21],"v0.10.0","  (V1 mode)",[40,162,163,164,168],{},"NVIDIA B200: ",[17,165,167],{"href":157,"rel":166},[21],"v0.10.1","  (V1 mode) (Had to build from source to support the B200 arch)",[40,170,171,174,175],{},[33,172,173],{},"Language Model:"," ",[17,176,179],{"href":177,"rel":178},"https:\u002F\u002Fhuggingface.co\u002Famd\u002FLlama-3.1-8B-Instruct-FP8-KV",[21],"Llama-3.1-8B-Instruct-FP8-KV",[40,181,182,118,185],{},[33,183,184],{},"Driver Stack",[37,186,187,190,193],{},[40,188,189],{},"AMD MI300x: ROCm 6.4.2",[40,191,192],{},"NVIDIA H200: CUDA 12.8.1",[40,194,195],{},"NVIDIA B200: CUDA 12.8.1",[40,197,198,201],{},[33,199,200],{},"Framework:",[37,202,203,206,209],{},[40,204,205],{},"AMD MI300x: Torch 2.7.1+rocm6.3",[40,207,208],{},"NVIDIA H200: Torch 2.7.1+cu128",[40,210,211],{},"NVIDIA B200: Torch 2.9.0.dev+cu128",[40,213,214,217],{},[33,215,216],{},"Batch Size",": 1024",[40,219,220,223,224],{},[33,221,222],{},"Measurement Protocol",": Each benchmark was run 10 times, and the numbers we report are overall averages. This helps reduce the effect of temporary system changes. Our careful measurement steps include:\n",[37,225,226,232,238,244],{},[40,227,228,231],{},[33,229,230],{},"Startup Times",": Important for checking how long it takes to load the model and get the system ready.",[40,233,234,237],{},[33,235,236],{},"Cold-Start TTFT (Time to First Token)",": Measures how long it takes from a new request until the first generated token appears. This is key for how quickly interactive applications respond.",[40,239,240,243],{},[33,241,242],{},"Steady-State TTFT",": Checks the TTFT after the system has been running steadily, showing typical performance under constant use.",[40,245,246,249],{},[33,247,248],{},"End-to-End Latency Metrics",": Gives a full picture of the time it takes for a complete inference request, from sending input to getting the final output.",[10,251,252],{},"This detailed method provides a strong way to check the specific performance details of Paiton in busy, partitioned GPU environments.",[102,254,256],{"id":255},"data-parallelism-without-partitioning",[33,257,258],{},"Data Parallelism Without Partitioning",[10,260,261],{},"Our first approach was to try to utilize vllm’s built in “–data-parallel-size” option, we quickly realized that this was not going to work out of the box and would require some serious modification. So instead, we took a different approach.",[10,263,264,265,269],{},"To run the benchmarks across 8 containers using vLLM, we first followed the official NGINX load balancing guide (",[17,266,267],{"href":267,"rel":268},"https:\u002F\u002Fdocs.vllm.ai\u002Fen\u002Fstable\u002Fdeployment\u002Fnginx.html",[21],")",[271,272,273],"ol",{},[40,274,275],{},[33,276,277],{},"NGINX Configuration",[10,279,280],{},"Here is the load balancing configuration we used in \u002Fetc\u002Fnginx\u002Fnginx.conf:",[282,283,289],"pre",{"className":284,"code":286,"language":287,"meta":288},[285],"language-text","upstream backend {\n    least_conn;\n    server vllm0:8000 max_fails=3 fail_timeout=10000s;\n    server vllm1:8000 max_fails=3 fail_timeout=10000s;\n    server vllm2:8000 max_fails=3 fail_timeout=10000s;\n    server vllm3:8000 max_fails=3 fail_timeout=10000s;\n    server vllm4:8000 max_fails=3 fail_timeout=10000s;\n    server vllm5:8000 max_fails=3 fail_timeout=10000s;\n    server vllm6:8000 max_fails=3 fail_timeout=10000s;\n    server vllm7:8000 max_fails=3 fail_timeout=10000s;\n}\n\nserver {\n    listen 80;\n    location \u002F {\n        proxy_pass http:\u002F\u002Fbackend;\n        proxy_set_header Host $host;\n        proxy_set_header X-Real-IP $remote_addr;\n        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;\n        proxy_set_header X-Forwarded-Proto $scheme;\n    }\n}\n","text","",[290,291,286],"code",{"__ignoreMap":288},[271,293,295],{"start":294},2,[40,296,297],{},[33,298,299],{},"Launching the Docker Containers",[10,301,302],{},"We used the following script to launch 8 containers using incremental device and port numbers:",[282,304,307],{"className":305,"code":306,"language":287,"meta":288},[285],"#!\u002Fbin\u002Fbash\n\necho \"Starting vLLM containers with incremental configuration…\"\n\nfor i in {0..7}; do\n    device_num=$((128 + (i * 8)))\n    device_path=\"\u002Fdev\u002Fdri\u002FrenderD${device_num}\"\n    port=$((8080 + i))\n    container_name=\"vllm${i}\"\n\n    echo \"Starting container ${container_name} on port ${port} with device ${device_path}…\"\n\n    docker run -itd \\\n        –ipc host \\\n        -v \u002Fdata:\u002Fdata \\\n        –network vllm_nginx \\\n        -e VLLM_ROCM_USE_AITER=True \\\n        -e HF_HOME=root\u002F.cache\u002Fhuggingface \\\n        -e HF_HUB_CACHE=\u002Froot\u002F.cache\u002Fhuggingface\u002Fhub \\\n        –device=\u002Fdev\u002Fkfd \\\n        –device=${device_path} \\\n        –group-add video \\\n        -p ${port}:8000 \\\n        –name ${container_name} \\\n        rocm\u002Fvllm:latest \\\n        vllm serve \\\n        amd\u002FLlama-3.1-8B-Instruct-FP8-KV \\\n        –num-scheduler-steps 10 \\\n        –kv-cache-dtype fp8 \\\n        –max-model-len 4096\n\n    if [ $? -eq 0 ]; then\n        echo \"✓ Container ${container_name} started successfully\"\n    else\n        echo \"✗ Failed to start container ${container_name}\"\n    fi\n\n    echo \"—\"\ndone\n\necho \"All containers started. Summary:\"\necho \"Containers: vllm0 through vllm7\"\necho \"Ports: 8081 through 8088\"\necho \"Devices: renderD128 through renderD184 (in steps of 8)\"\n",[290,308,306],{"__ignoreMap":288},[10,310,311],{},"This line, device_num=$((128 + (i * 8))), was necessary because of leftover render device entries in \u002Fdev\u002Fdri\u002F from previous GPU partitioning. Even after resetting the partitions, the device numbers did not reset to their original state. As a result, we had to offset each device path to correctly reference the available render nodes.",[271,313,315],{"start":314},3,[40,316,317],{},[33,318,319],{},"Benchmarking",[10,321,322],{},"Finally, we ran the following command to benchmark across all containers:",[282,324,327],{"className":325,"code":326,"language":287,"meta":288},[285],"for i in {1..10}; do\n    echo \"=== Running benchmark iteration $i\u002F10 ===\"\n    python3 ~\u002Fvllm\u002Fbenchmarks\u002Fbenchmark_serving.py \\\n      –backend vllm \\\n      –model amd\u002FLlama-3.1-8B-Instruct-FP8-KV \\\n      –dataset-name sharegpt \\\n      –dataset-path ~\u002Fvllm\u002FShareGPT_V3_unfiltered_cleaned_split.json \\\n      –num-prompts 1024 \\\n      –random-range-ratio 1.0 \\\n      –percentile-metrics ttft,tpot,itl,e2el \\\n      –sharegpt-output-len 256\n    echo \"=== Completed iteration $i\u002F10 ===\"\n    echo\ndone\n",[290,328,326],{"__ignoreMap":288},[102,330,332],{"id":331},"data-parallelism-with-partitioning",[33,333,334],{},"Data Parallelism With Partitioning",[10,336,337],{},"The most important first step was to partitionize our GPUs.",[10,339,340,341,346],{},"This was very straightforward and easy to do following ",[17,342,345],{"href":343,"rel":344},"https:\u002F\u002Finstinct.docs.amd.com\u002Fprojects\u002Famdgpu-docs\u002Fen\u002Flatest\u002Fgpu-partitioning\u002Findex.html",[21],"AMD’s official documentation",".",[10,348,349],{},"Steps:",[271,351,352],{},[40,353,354],{},[33,355,356],{},"Set the compute partitions.",[282,358,361],{"className":359,"code":360,"language":287,"meta":288},[285],"sudo amd-smi set –gpu all –compute-partition CPX\n",[290,362,360],{"__ignoreMap":288},[271,364,365],{"start":294},[40,366,367],{},[33,368,369],{},"Set the memory partitions.",[282,371,374],{"className":372,"code":373,"language":287,"meta":288},[285],"sudo amd-smi set –memory-partition NPS4\n",[290,375,373],{"__ignoreMap":288},[10,377,378],{},"Wait a few seconds and, done!",[10,380,381],{},"Result:",[10,383,384],{},[385,386],"img",{"alt":288,"src":387},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Frocmoverview.webp",[10,389,390],{},"Ready to go!",[10,392,393,394,269],{},"As mentioned in the previous section, to run the benchmarks across multiple containers using vLLM, we first followed the official NGINX load balancing guide (",[17,395,267],{"href":267,"rel":396},[21],[271,398,399],{},[40,400,401],{},[33,402,277],{},[10,404,280],{},[282,406,409],{"className":407,"code":408,"language":287,"meta":288},[285],"upstream backend {\n    least_conn;\n    server vllm0:8000 max_fails=3 fail_timeout=10000s;\n    .\n    .\n    .\n    server vllm63:8000 max_fails=3 fail_timeout=10000s;\n}\n\nserver {\n    listen 80;\n    location \u002F {\n        proxy_pass http:\u002F\u002Fbackend;\n        proxy_set_header Host $host;\n        proxy_set_header X-Real-IP $remote_addr;\n        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;\n        proxy_set_header X-Forwarded-Proto $scheme;\n    }\n}\n",[290,410,408],{"__ignoreMap":288},[271,412,413],{"start":314},[40,414,415],{},[33,416,299],{},[10,418,419],{},"We used the following script to launch 64 containers using incremental device and port numbers:",[282,421,424],{"className":422,"code":423,"language":287,"meta":288},[285],"#!\u002Fbin\u002Fbash\n\n# Script to run vLLM containers with incremental device, port, and name changes\n# Runs 64 containers with device=\u002Fdev\u002Fdri\u002FrenderD128 increasing in steps of 64\n# Port starting at 8081 and increasing by 1 each time\n# Container name starting at vllm0 and increasing by 1 each time\n\necho \"Starting vLLM containers with incremental configuration…\"\n\nfor i in {0..63}; do\n    # Calculate device number (renderD128, renderD192, renderD256, etc.)\n    device_num=$((128 + i))\n    device_path=\"\u002Fdev\u002Fdri\u002FrenderD${device_num}\"\n\n    # Calculate port (8081, 8082, 8083, etc.)\n    port=$((8081 + i))\n\n    # Calculate container name (vllm0, vllm1, vllm2, etc.)\n    container_name=\"vllm${i}\"\n\n    echo \"Starting container ${container_name} on port ${port} with device ${device_path}…\"\n    docker run -itd \\\n        –ipc host \\\n        -v \u002Fdata:\u002Fdata \\\n        –network vllm_nginx \\\n        -e VLLM_ROCM_USE_AITER=True \\\n        -e HF_HOME=root\u002F.cache\u002Fhuggingface \\\n        -e HF_HUB_CACHE=\u002Froot\u002F.cache\u002Fhuggingface\u002Fhub \\\n        –device=\u002Fdev\u002Fkfd \\\n        –device=${device_path} \\\n        –group-add video \\\n        -p ${port}:8000 \\\n        –name ${container_name} \\\n        rocm\u002Fvllm:latest \\\n        vllm serve \\\n        \u002Fdata\u002F.cache\u002Fhuggingface\u002Fhub\u002Fmodels–amd–Llama-3.1-8B-Instruct-FP8-KV\u002Fsnapshots\u002Ffa42f9a9105c545755fea25cf69f49ac8c8b40e1\u002F \\\n        –num-scheduler-steps 10 \\\n        –kv-cache-dtype fp8 \\\n        –max-model-len 4096\n\n    # Check if container started successfully\n    if [ $? -eq 0 ]; then\n        echo \"✓ Container ${container_name} started successfully\"\n    else\n        echo \"✗ Failed to start container ${container_name}\"\n    fi\n\n    echo \"—\"\ndone\n\necho \"All containers started. Summary:\"\necho \"Containers: vllm0 through vllm63\"\necho \"Ports: 8081 through 8144\"\necho \"Devices: renderD128 through renderD4160 (in steps of 64)\"\necho \"\"\necho \"To check container status: docker ps\"\necho \"To view logs: docker logs &lt;container_name&gt;\"\n",[290,425,423],{"__ignoreMap":288},[271,427,429],{"start":428},4,[40,430,431],{},[33,432,319],{},[10,434,435],{},"Same script as in the previous section.",[10,437,438],{},"Action view :)",[10,440,441],{},[385,442],{"alt":288,"src":443},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Flh7-rt.googleusercontent.com-dc3c57fbfc36.gif",[10,445,446],{},"Paiton MI300x",[10,448,449],{},[385,450],{"alt":288,"src":451},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Flh7-rt.googleusercontent.com-d1bd449add79.gif",[10,453,454],{},"Stock MI300x",[102,456,458],{"id":457},"benchmark-results",[33,459,460],{},"Benchmark Results",[462,463,465],"h4",{"id":464},"no-partitions-8-gpus","No Partitions – 8 GPUs",[467,468,469,516],"table",{},[470,471,472],"thead",{},[473,474,475,481,486,491,496,501,506,511],"tr",{},[476,477,478],"th",{},[33,479,480],{},"Metric",[476,482,483],{},[33,484,485],{},"Paiton ∆",[476,487,488],{},[33,489,490],{},"Stock",[476,492,493],{},[33,494,495],{},"∆ vs Stock",[476,497,498],{},[33,499,500],{},"H200",[476,502,503],{},[33,504,505],{},"∆ vs H200",[476,507,508],{},[33,509,510],{},"B200",[476,512,513],{},[33,514,515],{},"∆ vs B200",[517,518,519,552,584,616,648,680,712,744],"tbody",{},[473,520,521,527,530,533,538,541,546,549],{},[522,523,524],"td",{},[33,525,526],{},"Benchmark duration (s) ↓",[522,528,529],{},"4.812",[522,531,532],{},"11.029",[522,534,535],{},[33,536,537],{},"+129.20%",[522,539,540],{},"11.84",[522,542,543],{},[33,544,545],{},"+146.05%",[522,547,548],{},"4.59",[522,550,551],{},"-4.61%",[473,553,554,559,562,565,570,573,578,581],{},[522,555,556],{},[33,557,558],{},"Request throughput (req\u002Fs) ↑",[522,560,561],{},"213.55",[522,563,564],{},"94.308",[522,566,567],{},[33,568,569],{},"+126.44%",[522,571,572],{},"83.22",[522,574,575],{},[33,576,577],{},"+156.61%",[522,579,580],{},"225.99",[522,582,583],{},"–5.50%",[473,585,586,591,594,597,602,605,610,613],{},[522,587,588],{},[33,589,590],{},"Output token throughput (tok\u002Fs) ↑",[522,592,593],{},"53851.639",[522,595,596],{},"23809.63",[522,598,599],{},[33,600,601],{},"+126.18%",[522,603,604],{},"20940.86",[522,606,607],{},[33,608,609],{},"+157.16%",[522,611,612],{},"56989.26",[522,614,615],{},"-5.52%",[473,617,618,623,626,629,634,637,642,645],{},[522,619,620],{},[33,621,622],{},"Total Token throughput (tok\u002Fs) ↑",[522,624,625],{},"101941.667",[522,627,628],{},"45047.076",[522,630,631],{},[33,632,633],{},"+126.30%",[522,635,636],{},"39674.51",[522,638,639],{},[33,640,641],{},"+156.94%",[522,643,644],{},"107827.34",[522,646,647],{},"-5.46%",[473,649,650,655,658,661,666,669,674,677],{},[522,651,652],{},[33,653,654],{},"Mean TTFT (ms) ↓",[522,656,657],{},"543.799",[522,659,660],{},"4252.513",[522,662,663],{},[33,664,665],{},"+682.47%",[522,667,668],{},"3027.49",[522,670,671],{},[33,672,673],{},"+456.96%",[522,675,676],{},"1245.55",[522,678,679],{},"+129.05%",[473,681,682,687,690,693,698,701,706,709],{},[522,683,684],{},[33,685,686],{},"Mean TPOT (ms) ↓",[522,688,689],{},"15.075",[522,691,692],{},"16.872",[522,694,695],{},[33,696,697],{},"+11.92%",[522,699,700],{},"26.70",[522,702,703],{},[33,704,705],{},"+77.02%",[522,707,708],{},"10.27",[522,710,711],{},"-31.87%",[473,713,714,719,722,725,730,733,738,741],{},[522,715,716],{},[33,717,718],{},"Mean ITL (ms) ↓",[522,720,721],{},"15.025",[522,723,724],{},"16.509",[522,726,727],{},[33,728,729],{},"+9.88%",[522,731,732],{},"71.11",[522,734,735],{},[33,736,737],{},"+373.37%",[522,739,740],{},"32.62",[522,742,743],{},"+117.10%",[473,745,746,751,754,757,762,765,770,773],{},[522,747,748],{},[33,749,750],{},"Mean E2EL (ms) ↓",[522,752,753],{},"4317.43",[522,755,756],{},"8403.948",[522,758,759],{},[33,760,761],{},"+94.65%",[522,763,764],{},"9705.94",[522,766,767],{},[33,768,769],{},"+124.79%",[522,771,772],{},"3818.69",[522,774,775],{},"-11.51%",[10,777,778],{},[385,779],{"alt":288,"src":780,"title":781},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Faverage-throughput-2-2.webp","Chart",[462,783,785],{"id":784},"partitions-64-vgpus","Partitions – 64 vGPUs",[467,787,788,816],{},[470,789,790],{},[473,791,792,796,801,805,810,813],{},[476,793,794],{},[33,795,480],{},[476,797,798,800],{},[33,799,485],{},"*",[476,802,803],{},[33,804,490],{},[476,806,807],{},[33,808,809],{},"∆vs Stock(Ratio)",[476,811,812],{},"**H200",[476,814,815],{},"**∆vs H200",[517,817,818,838,858,878,897,917,937,956],{},[473,819,820,825,828,831,834,836],{},[522,821,822],{},[33,823,824],{},"Benchmark duration (s)",[522,826,827],{},"7.875",[522,829,830],{},"17.294",[522,832,833],{},"2.20",[522,835],{},[522,837],{},[473,839,840,845,848,851,854,856],{},[522,841,842],{},[33,843,844],{},"Request throughput (req\u002Fs)",[522,846,847],{},"130.234",[522,849,850],{},"59.727",[522,852,853],{},"2.18",[522,855],{},[522,857],{},[473,859,860,865,868,871,874,876],{},[522,861,862],{},[33,863,864],{},"Output token throughput (tok\u002Fs)",[522,866,867],{},"33339.931",[522,869,870],{},"15047.115",[522,872,873],{},"2.22",[522,875],{},[522,877],{},[473,879,880,885,888,891,893,895],{},[522,881,882],{},[33,883,884],{},"Total Token throughput (tok\u002Fs)",[522,886,887],{},"62667.914",[522,889,890],{},"28497.62",[522,892,833],{},[522,894],{},[522,896],{},[473,898,899,904,907,910,913,915],{},[522,900,901],{},[33,902,903],{},"Mean TTFT (ms)",[522,905,906],{},"1082.885",[522,908,909],{},"6255.879",[522,911,912],{},"5.78",[522,914],{},[522,916],{},[473,918,919,924,927,930,933,935],{},[522,920,921],{},[33,922,923],{},"Mean TPOT (ms)",[522,925,926],{},"20.99",[522,928,929],{},"31.289",[522,931,932],{},"1.49",[522,934],{},[522,936],{},[473,938,939,944,946,949,952,954],{},[522,940,941],{},[33,942,943],{},"Mean ITL (ms)",[522,945,926],{},[522,947,948],{},"31.13",[522,950,951],{},"1.48",[522,953],{},[522,955],{},[473,957,958,963,966,969,972,974],{},[522,959,960],{},[33,961,962],{},"Mean E2EL (ms)",[522,964,965],{},"6435.477",[522,967,968],{},"14067.724",[522,970,971],{},"2.19",[522,973],{},[522,975],{},[10,977,978],{},"**Note: We are working on improving these numbers even more.",[10,980,981,982,992],{},"**Note2: Not possible with Nvidia, or at least very difficult (",[983,984,985],"em",{},[17,986,989],{"href":987,"rel":988},"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\u002Fissues\u002F6551#issuecomment-2237624342",[21],[983,990,991],{},"complicated",")*",[10,994,995],{},[385,996],{"alt":288,"src":997,"title":781},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Faverage-throughput-2-2-1.webp",[10,999,1000],{},"Now let’s look at this from an ROI-driven perspective.",[10,1002,1003],{},"If we haven’t impressed you so far, we’re pretty sure this will. We use the MI300X server as a reference\u002Fbaseline to compare its cost factor and throughput against the H200 and B200.",[467,1005,1006,1026],{},[470,1007,1008],{},[473,1009,1010,1015,1020,1023],{},[476,1011,1012],{},[33,1013,1014],{},"Architecture",[476,1016,1017],{},[33,1018,1019],{},"Cost Factor vs Paiton",[476,1021,1022],{},"**Throughput Cost-Eff ***",[476,1024,1025],{},"**Latency Cost-Eff",[517,1027,1028,1042,1057,1072],{},[473,1029,1030,1035,1038,1040],{},[522,1031,1032],{},[33,1033,1034],{},"MI300X+Paiton",[522,1036,1037],{},"Ref",[522,1039,1037],{},[522,1041,1037],{},[473,1043,1044,1048,1051,1054],{},[522,1045,1046],{},[33,1047,454],{},[522,1049,1050],{},"1x",[522,1052,1053],{},"+126.31%",[522,1055,1056],{},"+94.57%",[473,1058,1059,1063,1066,1069],{},[522,1060,1061],{},[33,1062,500],{},[522,1064,1065],{},"1.375x",[522,1067,1068],{},"+253.30%",[522,1070,1071],{},"+209.07%",[473,1073,1074,1078,1081,1084],{},[522,1075,1076],{},[33,1077,510],{},[522,1079,1080],{},"2x",[522,1082,1083],{},"+89.18%",[522,1085,705],{},[1087,1088,1089,1112],"blockquote",{},[10,1090,1091,174,1094,174,1099,1102,1103,174,1106,1111],{},[983,1092,1093],{},"*Throughput Cost-Efficiency: = %",[983,1095,1096],{},[33,1097,1098],{},"more",[983,1100,1101],{},"total‑token","** ",[983,1104,1105],{},"throughput",[983,1107,1108],{},[33,1109,1110],{},"per dollar"," *vs each platform.",[10,1113,1114,1115,174,1120,174,1123,1102,1128],{},"**Latency Cost-Efficiency: = %* ",[983,1116,1117],{},[33,1118,1119],{},"better",[983,1121,1122],{},"end‑to‑end",[983,1124,1125,1110],{},[33,1126,1127],{},"latency",[983,1129,1130],{},"vs each platform.",[10,1132,1133],{},[33,1134,1135],{},"What this tells you",[37,1137,1138,1148,1162],{},[40,1139,1140,1143,1144,1147],{},[33,1141,1142],{},"Paiton delivers 2.5 × the throughput per $"," over an H200 and ",[33,1145,1146],{},"+126 %"," over stock.",[40,1149,1150,1153,1154,1157,1158,1161],{},[33,1151,1152],{},"Latency per $"," is ",[33,1155,1156],{},"3.1 × better than the H200"," and ",[33,1159,1160],{},"+94 % over stock",", solid ROI on every millisecond shaved.",[40,1163,1164,1165,1167,1168,1171],{},"The ",[33,1166,510],{}," gap is real, but remember it costs ",[33,1169,1170],{},"twice"," as much, Paiton still wins on cost‑efficiency across the board.",[102,1173,1175],{"id":1174},"cost-per-million-tokens",[33,1176,1177],{},"Cost per Million tokens",[10,1179,1180],{},"If we use available renting prices for the different systems, we could calculate the relative cost per 1M tokens:",[467,1182,1183,1216],{},[470,1184,1185],{},[473,1186,1187,1191,1196,1201,1206,1211],{},[476,1188,1189],{},[33,1190,1014],{},[476,1192,1193],{},[33,1194,1195],{},"Throughput (tok\u002Fs)",[476,1197,1198],{},[33,1199,1200],{},"GPU Count",[476,1202,1203],{},[33,1204,1205],{},"Approx. hourly cost",[476,1207,1208],{},[33,1209,1210],{},"Inference Cost \u002F 1M Tokens",[476,1212,1213],{},[33,1214,1215],{},"Relative Cost",[517,1217,1218,1238,1256,1275],{},[473,1219,1220,1224,1226,1229,1232,1235],{},[522,1221,1222],{},[33,1223,1034],{},[522,1225,625],{},[522,1227,1228],{},"~8",[522,1230,1231],{},"$20.50",[522,1233,1234],{},"$0.06",[522,1236,1237],{},"REF",[473,1239,1240,1244,1246,1248,1250,1253],{},[522,1241,1242],{},[33,1243,454],{},[522,1245,628],{},[522,1247,1228],{},[522,1249,1231],{},[522,1251,1252],{},"$0.13",[522,1254,1255],{},"2.26× ↑",[473,1257,1258,1262,1264,1266,1269,1272],{},[522,1259,1260],{},[33,1261,500],{},[522,1263,636],{},[522,1265,1228],{},[522,1267,1268],{},"$28.20",[522,1270,1271],{},"$0.20",[522,1273,1274],{},"3.54× ↑",[473,1276,1277,1281,1283,1285,1288,1290],{},[522,1278,1279],{},[33,1280,510],{},[522,1282,644],{},[522,1284,1228],{},[522,1286,1287],{},"$48.60",[522,1289,1252],{},[522,1291,1292],{},"2.24× ↑",[10,1294,1295],{},[385,1296],{"alt":288,"src":1297},"\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.jpg",[10,1299,1300],{},[33,1301,1302],{},"Insights:",[37,1304,1305,1311,1316],{},[40,1306,1307,1310],{},[33,1308,1309],{},"Paiton"," delivers 2.26× cost savings compared to unoptimized MI300X.",[40,1312,1313,1315],{},[33,1314,500],{}," costs 3.54× more than MI300X+Paiton per 1M tokens.",[40,1317,1318,1320,1321,1324],{},[33,1319,510],{}," is the most expensive, costing over ",[33,1322,1323],{},"2.24× more"," than the optimized MI300X setup.",[102,1326,1328],{"id":1327},"big-win-for-amd",[33,1329,1330],{},"Big win for AMD",[10,1332,1333],{},"Trying to figure out how MIG worked on Nvidia with vLLM was like trying to find a perfect gift for your spouse, it was exhausting. Eventually we ran into a NCCL error which seemed unsolvable and that was the last straw.",[10,1335,1336],{},"While MIG allows virtual partitioning on supported NVIDIA GPUs, as previously mentioned we encountered significant limitations when attempting to use it in conjunction with vLLM for data-parallel workloads. Specifically, vLLM was unable to properly leverage MIG slices for distributed inference.",[10,1338,1339],{},"In contrast, AMD’s architecture enabled straightforward partitioning and containerized deployment of vLLM instances without any issues. This streamlined setup, along with ROCm’s compatibility, made AMD far better suited for true multi-tenancy out of the box.",[10,1341,1342,1343,1346],{},"This represents a ",[33,1344,1345],{},"major win for AMD",", particularly for enterprises aiming to deploy isolated inference workloads across shared hardware without too much friction or compromise.",[1087,1348,1349],{},[1087,1350,1351],{},[1087,1352,1353,1358,1361],{},[10,1354,1355],{},[33,1356,1357],{},"Having methodically outpaced Intel in performance, AMD is now strategically poised to challenge NVIDIA’s leadership, an evolution we’re proud to drive.",[10,1359,1360],{},"***Kian Mohadjerin",[10,1362,1363],{},"Head of AI, Eliovp BV*",[102,1365,1367],{"id":1366},"key-results",[33,1368,1369],{},"Key Results",[37,1371,1372,1378,1384,1390],{},[40,1373,1374,1377],{},[33,1375,1376],{},"Throughput scaling"," was near-linear up to 64 partitions, thanks to Paiton’s minimized memory overhead and fast kernel dispatch.",[40,1379,1380,1383],{},[33,1381,1382],{},"Latency"," remained stable across parallel sessions, demonstrating the strength of Paiton’s per-GPU scheduling and shared memory optimizations.",[40,1385,1386,1389],{},[33,1387,1388],{},"Memory usage per partition"," was significantly lower compared to standard vLLM or runtimes, enabling high-density deployment.",[40,1391,1392,1395],{},[33,1393,1394],{},"Cost per Million Tokens"," was reduced by over 2× compared to high-end systems like the B200, showcasing Paiton’s ability to deliver industry-leading efficiency even on more affordable AMD hardware.",[102,1397,1399],{"id":1398},"conclusion",[33,1400,1401],{},"Conclusion",[10,1403,1404],{},"This experiment highlights Paiton’s ability to unlock the full potential of modern hardware like the MI300X through advanced packaging and optimization techniques. Running Llama 3.1 8B FP8 across 64 GPU partitions showcases how inference workloads can be massively parallelized without sacrificing too much performance or usability.",[10,1406,1407],{},"Imagine the potential of Paiton paired with AMD’s upcoming MI355X. With even more memory bandwidth, compute, and architectural improvements on the horizon, the synergy between next-gen hardware and the Paiton runtime could redefine the state of high-performance AI serving.",[10,1409,1410],{},"Stay tuned for future updates as we expand Paiton’s capabilities.",[10,1412,1413],{},"Don’t believe our results? Neither did we, so test Paiton for yourself and request an evaluation model.",[102,1415,1417],{"id":1416},"pricing",[33,1418,1419],{},"Pricing",[10,1421,1422],{},"If you’re curious about pricing with Paiton, our formula is quite simple:",[467,1424,1425],{},[470,1426,1427],{},[473,1428,1429],{},[476,1430,1431],{},"50% of x% costs saved per 1M tokens",[10,1433,1434],{},"The cost saved is measured by looking at the customer’s current throughput compared to the throughput using Paiton.",[10,1436,1437,1442],{},[17,1438,1441],{"href":1439,"rel":1440},"https:\u002F\u002Fai.eliovp.com\u002Fpaiton",[21],"Reach out and let’s talk"," :)",[102,1444,1446],{"id":1445},"references",[33,1447,1448],{},"References",[271,1450,1451,1458,1465,1471,1477,1483,1489,1495,1501,1508],{},[40,1452,1453],{},[17,1454,1457],{"href":1455,"rel":1456},"https:\u002F\u002Fwww.supermicro.com\u002Fen\u002Fproducts\u002Fsystem\u002Fgpu\u002F8u\u002Fas%20-8125gs-tnmr2",[21],"Supermicro GPU System AS-8125GS-TNMR2",[40,1459,1460],{},[17,1461,1464],{"href":1462,"rel":1463},"https:\u002F\u002Fwww.amd.com\u002Fen\u002Fproducts\u002Faccelerators\u002Finstinct\u002Fmi300\u002Fmi300x.html",[21],"AMD Instinct MI300X",[40,1466,1467],{},[17,1468,1470],{"href":19,"rel":1469},[21],"Paiton FP8 beats Nvidia’s H200 on AMD’s MI300X",[40,1472,1473],{},[17,1474,1476],{"href":142,"rel":1475},[21],"ROCm\u002Fvllm GitHub",[40,1478,1479],{},[17,1480,1482],{"href":157,"rel":1481},[21],"vllm-project\u002Fvllm GitHub",[40,1484,1485],{},[17,1486,1488],{"href":177,"rel":1487},[21],"Hugging Face AMD Llama-3.1-8B-Instruct-FP8-KV",[40,1490,1491],{},[17,1492,1494],{"href":267,"rel":1493},[21],"vLLM Nginx Deployment",[40,1496,1497],{},[17,1498,1500],{"href":343,"rel":1499},[21],"AMD GPU Partitioning Documentation",[40,1502,1503],{},[17,1504,1507],{"href":1505,"rel":1506},"https:\u002F\u002Frocm.blogs.amd.com\u002Fsoftware-tools-optimization\u002Fcompute-memory-modes\u002FREADME.html",[21],"ROCm Compute Memory Modes",[40,1509,1510],{},[17,1511,1514],{"href":1512,"rel":1513},"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\u002Fissues\u002F6551",[21],"vLLM GitHub Issue #6551",{"title":288,"searchDepth":294,"depth":294,"links":1516},[1517,1518,1519,1520,1521,1522,1523,1524,1525,1526],{"id":104,"depth":314,"text":107},{"id":255,"depth":314,"text":258},{"id":331,"depth":314,"text":334},{"id":457,"depth":314,"text":460},{"id":1174,"depth":314,"text":1177},{"id":1327,"depth":314,"text":1330},{"id":1366,"depth":314,"text":1369},{"id":1398,"depth":314,"text":1401},{"id":1416,"depth":314,"text":1419},{"id":1445,"depth":314,"text":1448},[1528,1529,1309,1530,1531,1532,510,500,1533,1534,1535,1309,1536],"All","Artificial Intelligence","Uncategorized","AI","AMD","MI300X","MI355x","NVidia","vLLM","2025-07-31T13:32:57","At ElioVP, we’re all about pushing AI inference past the limits, and packaging every squeeze of performance into a plug‑and‑play runtime.  Remember our last blog, where Paiton’s FP8 pipeline on AMD’s MI300X completely outclassed NVIDIA’s H200? Well, buckle up, because we’ve gone back to the drawing board. This time, we’re loading Llama-3.1-8B-Instruct-FP8-KV, the leaner, meaner ...","md","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp",{},true,"https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F","\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach",{"title":5,"description":1538},"mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","blog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","eLRmwCKJrL8zOFPJ-m5JwSPkhJ9f623SYZr1O8bd1vs",[1550,1564,1592,1603,1626,1644,1663,1681,1699,1716,1731,1747,1762,1764,1773,1781,1796,1810,1821,1832,1842,1855,1865,1878,1889,1899,1910,1919,1931,1942,1951],{"path":1551,"title":1552,"description":1553,"date":1554,"slug":1555,"image":1556,"originalUrl":1557,"categories":1558},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",null,[1528,1559,1560,1561,1562,1563],"AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1565,"title":1566,"description":1567,"date":1568,"slug":1569,"image":1570,"originalUrl":1571,"categories":1572},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Paiton Returns to Its Diffusion Roots: Optimizing Wan2.2-T2V-A14B on AMD MI355X","When we first started building Paiton, one of our earliest focus areas was optimizing diffusion models. Stable Diffusion XL was one of the first large models where we showed that fused operators, efficient execution, and hardware-aware kernels could make a real difference. Now we are returning to those origins.With the growing interest in text-to-video generation, ...","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1528,1529,1309,1573,1532,510,1574,1575,1576,1577,1578,1579,1580,1581,1582,1583,1534,1535,1584,1585,1309,1586,1587,1588,1589,1590,1591],"14B","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","GPU","Hardware","Inference","Instinct","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1593,"title":1594,"description":1595,"date":1596,"slug":1597,"image":1598,"originalUrl":1599,"categories":1600},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","From the Attic to the Front Page: ElioVP Recognized as a Pioneer in Chip Optimization & Data Center Infrastructure","It has been some incredible weeks for the team here at Eliovp. We are extremely proud to share that our company was recently featured on the front page of De Tijd, Belgium’s leading business newspaper. Seeing our story, from our founder’s early days tinkering with wires in an attic to generating €215 million in revenue, ...","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1528,1529,1601,1530,1532,1602,1601,1535],"Modular DC","De Tijd",{"path":1604,"title":1605,"description":1606,"date":1607,"slug":1608,"image":1609,"originalUrl":1610,"categories":1611},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","Privacy is geen IT-probleem meer, het is een strategische prioriteit: ","Privacyrisico’s, Psychologische Valkuilen en de Operationele Realiteit van Generatieve AI in de Benelux De recente verschijning van een frontpage-artikel over ons bedrijf in het gerespecteerde dagblad “De Tijd” heeft onze zichtbaarheid aanzienlijk vergroot, wat de aanleiding is voor dit artikel. Deze mediabelangstelling, gecombineerd met de talrijke uitnodigingen voor spreekbeurten die we hebben ontvangen, fungeert als ...","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1528,1529,1612,1530,1613,1614,1615,1616,1617,1618,1619,1620,1578,1621,1622,1623,1624,1625],"Trending","AI Act","Antropomorfisme","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Generatieve AI","Microsoft Copilot","Privacy","Shadow AI",{"path":1627,"title":1628,"description":1629,"date":1630,"slug":1631,"image":1632,"originalUrl":1633,"categories":1634},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Itsme? Bij ons is het “it’s not me”, en dit is waarom.","1. Inleiding: De Strategische Noodzaak van Weigering In het hedendaagse digitale landschap wordt de keuze voor een identiteit leverancier (IdP) vaak gereduceerd tot een discussie over User Experience (UX) en conversieratio’s. Deze reductionistische benadering verhult echter de diepgaande strategische, juridische en operationele risico’s die gepaard gaan met het uitbesteden van de “Sleutels tot het Koninkrijk”, ...","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1528,1635,1636,1637,1618,1638,1639,1640,1621,1641,1642,1643,1624],"AWS","Belgian Mobile ID","Cloud Act","Data Soevereiniteit","Digitale Identiteit","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1645,"title":1646,"description":1647,"date":1648,"slug":1649,"image":1650,"originalUrl":1651,"categories":1652},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","From Hype to Sovereign Infrastructure Nederlandse Versie Summary The narrative surrounding “Agentic AI” in 2025 is defined by a sharp contrast between market expectations and engineering reality. While the general public, conditioned by the ease of ChatGPT, expects “miracles” and instant integration, the reality of building autonomous agents for enterprise workflows is a discipline of ...","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1528,1529,1653,1612,1654,1655,1656,1657,1658,1659,1660,1661,1586,1662],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1664,"title":1665,"description":1666,"date":1667,"slug":1668,"image":1669,"originalUrl":1670,"categories":1671},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","How AI Neocloud Infrastructure Turns Circular Cash Flows Into Fake Growth Executive Summary The interval between 2023 and 2025 has birthed a capital allocation phenomenon arguably without precedent: the “Synthetic Bubble.” Driven by the scramble for AI dominance, the venture capital apparatus has directed billions into the “Neocloud” ecosystem. However, a rigorous analysis suggests that ...","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1528,1529,1612,1559,1672,1673,1674,1675,1676,1677,1678,1679,1680],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1682,"title":1683,"description":1684,"date":1685,"slug":1686,"image":1687,"originalUrl":1688,"categories":1689},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","Building the Engine for the AI Race: The 4-Month Path to NVIDIA GB300 NVL72 Power","In artificial intelligence infrastructure, speed is the foundation of competitive differentiation. From model training velocity to inference latency, every millisecond matters. But before any workload executes, there is a critical prerequisite that often determines success or failure: time to market. Traditional builds typically require 18–24 months. In the current AI cycle, that is simply too ...","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1528,1601,1530,1690,1559,1691,1692,1693,1694,1695,1696,1697,1698],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1700,"title":1701,"description":1702,"date":1703,"slug":1704,"image":1705,"originalUrl":1706,"categories":1707},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Every few years, a new solution pops up promising the same dream: On paper, that sounds perfect. Take your existing CUDA applications, swap out the toolchain, and suddenly you’re “portable.” And to be fair: if you’re running research code or trying to get an internal tool to compile on a non-NVIDIA box, that can absolutely ...","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1528,1529,1309,1530,1708,1529,1709,1710,1711,1712,1713,1714,1309,1715],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning","ROCm",{"path":1717,"title":1718,"description":1719,"date":1720,"slug":1721,"image":1722,"originalUrl":1723,"categories":1724},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Let’s be honest, we’re not the marketing type.We’ve never taken a cent of outside investment, never burned cash on ad campaigns, and never hired a sales army.We just build things that work. In today’s world, it seems the companies shouting the loudest often get the spotlight, while the ones doing the actual engineering quietly build ...","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1528,1529,1309,1725,1726,1708,1727,1728,1729,1714,1309,1730,1536],"AI Inference","AMD Instinct","Cost Efficiency","High Throughput","Inference Optimization","SGLang",{"path":1732,"title":1733,"description":1734,"date":1735,"slug":1736,"image":1737,"originalUrl":1738,"categories":1739},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Stop Overpaying: Paiton MI300X MoE Beats H200\u002FB200 on $\u002F1M Tokens","Short summary: We benchmarked Paiton with our new MoE support on Qwen\u002FQwen3-30B-A3B-Instruct-2507 to compare inference performance across several setups. Each configuration was run five times per batch size and we report the mean across runs. Why this benchmark Most published numbers use synthetic prompts or toy datasets. We focused on realistic conversational workloads (we always ...","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1528,1529,1309,1740,1708,1741,1729,1742,1743,1744,1745,1309,1746],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1748,"title":1749,"description":1750,"date":1751,"slug":1752,"image":1753,"originalUrl":1754,"categories":1755},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Agentic AI, But Make It Local: From Inbox to Insight to Action","(Nederlandse versie) We’ve built production-ready, local-first agentic AI that plugs into your existing email stack, auto-creates tickets, classifies messages, extracts multi-question threads, reads PDFs, spots invoices\u002Fquotes, analyzes images (yes, damage detection), and pushes structured reports into your systems, no dependency on OpenAI, Google, or Microsoft unless you want it. Tailor-made models trained on your data, ...","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1528,1529,1653,1530,1654,1756,1757,1758,1759,1659,1661,1586,1760,1761],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1544,"title":5,"description":1538,"date":1537,"slug":1546,"image":1540,"originalUrl":1543,"categories":1763},[1528,1529,1309,1530,1531,1532,510,500,1533,1534,1535,1309,1536],{"path":1765,"title":1766,"description":1767,"date":1768,"slug":1769,"image":1770,"originalUrl":1771,"categories":1772},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Here at Eliovp, we continuously innovate when it comes to building practical solutions. If there’s one core strength, it’s our team’s ability to think outside the box. One key area of focus for us is developing applicable AI solutions, everyday usable AI implementations tailored specifically for our clients’ needs. In this blog, we’ll explore our ...","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1528,1529,1653,1654,1756,1757,1758,1759,1659,1661,1586,1760,1761],{"path":1774,"title":1775,"description":1776,"date":1777,"slug":1778,"image":288,"originalUrl":1779,"categories":1780},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Introduction AI is rapidly transforming every industry, but running large models efficiently remains a major technical and financial challenge. At ElioVP, we specialize in optimizing for AMD accelerators, helping organizations unlock the full potential of their hardware. Today, we’re excited to announce a new offering: free evaluation models that let you test our cutting-edge optimizations ...","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1528,1529,1309],{"path":1782,"title":1783,"description":1784,"date":1785,"slug":1786,"image":1787,"originalUrl":1788,"categories":1789},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Paiton: Dramatically Faster Startup and Performance for Llama-3.1-405B","With Paiton, we’re not merely pursuing peak inference speeds, we’re fundamentally reshaping the entire lifecycle of large language model (LLM) deployment. Our latest endeavor pairs AMD’s cutting-edge MI300X GPUs with the colossal Llama-3.1-405B-Instruct-FP8-KV model, achieving groundbreaking milestones: Visual Demonstration: Startup Speed Showcase We’re excited to share a visual demonstration of Paiton’s revolutionary startup performance. Watch ...","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1528,1529,1309,1530,1725,1708,1790,1710,1791,1792,1793,1309,1794,1795],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1797,"title":1798,"description":1799,"date":1800,"slug":1801,"image":1802,"originalUrl":19,"categories":1803},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","The world of AI is moving at an unprecedented pace, and efficient inference is key to deploying powerful models in real-world applications. At Eliovp, we’ve consistently pushed the boundaries of AI performance, as highlighted in our previous blogs showcasing significant inference speedups when benchmarking with fp16\u002Fbf16. Now, we’re thrilled to announce a further significant leap ...","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp",[1528,1529,1309,1530,1708,1804,1658,1579,1805,1806,1807,1792,1808,1809],"Cold Start Optimization","GPU Performance","Inference Latency","Large Language Models","Model Serving","vLLM Optimization",{"path":1811,"title":1812,"description":1813,"date":1814,"slug":1815,"image":1816,"originalUrl":1817,"categories":1818},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","As large language models (LLMs) become a foundational part of modern applications, picking the right server for deployment is more important than ever. Whether you’re an enterprise scaling up inference, a startup optimizing for cost, or a researcher pushing throughput boundaries. This blog compares two high-profile server setups and two not so high-profile setups which ...","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1528,1529,1309,1653,1530,1532,1533,1535,1819,1820],"RX7900XTX","tenstorrent",{"path":1822,"title":1823,"description":1824,"date":1825,"slug":1826,"image":1827,"originalUrl":1828,"categories":1829},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Empowering GPU Cluster Investors with Real-World Financial Insights","At Eliovp BV, we’ve spent years on the cutting edge of GPU cluster deployment and optimization across Europe. Our team supports leading organizations in AI, finance, and research, architecting, building, and scaling high-performance infrastructure. Over time, our customers, both newcomers and seasoned adopters, repeatedly asked the same question: “Can you help us build a P&L ...","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1528,1529,1601,1653,510,500,1830,1535,1831],"MI325x","pnl calculator",{"path":1833,"title":1834,"description":1835,"date":1836,"slug":1837,"image":1838,"originalUrl":1839,"categories":1840},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","Cranking Out Faster Tokens for Fewer Dollars: AMD MI300X vs. NVIDIA H200","Qwen3-32B on Paiton + AMD MI300x vs.NVIDIA H200 1. Introduction “While we’re actively training models for local customers, automating and streamlining critical business processes, we still found time to push our Paiton framework to the limit on Qwen3-32B.” In the competitive realm of LLMs, next-gen hardware like the NVIDIA H200 often steals the headlines. But ...","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1528,1529,1309,1531,1532,500,1841,1535,1309,1536],"MI300",{"path":1843,"title":1844,"description":1845,"date":1846,"slug":1847,"image":1848,"originalUrl":1849,"categories":1850},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Power Meets Precision: High-Density Modular Data Center for NVIDIA NVL Deployments (1–2 MW)","Purpose-Built High-Density Infrastructure for Blackwell-Class AI Workloads At Eliovp, we’re engineering a new class of AI infrastructure. Our advanced modular platform is built to also support NVIDIA’s cutting-edge NVL architecture, from the efficient NVL4 to the ultra-scale NVL72, enabling deployments that range from distributed edge inference to full-stack model training at hyperscale. Designed to meet ...","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1528,1601,1851,1559,1692,1562,1693,1694,1852,1853,1697,1854],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1856,"title":1857,"description":1858,"date":1859,"slug":1860,"image":1861,"originalUrl":1862,"categories":1863},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","At Eliovp, we’re constantly keeping up with the newest AI trends. Consequently, we have been looking into AI agents and have created a medical agent designed to seamlessly interact with DICOM servers inside hospitals. This isn’t just another chatbot or AI tool. This is an intelligent assistant that understands the language of radiology and is ...","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1528,1529,1653,1530,1531,1532,1864],"Healthcare",{"path":1866,"title":1867,"description":1868,"date":1869,"slug":1870,"image":1871,"originalUrl":1872,"categories":1873},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","Eliovp BV: Your Trusted Partner for Supply Chain Resilience Amidst New U.S. Tariffs","In today’s rapidly evolving global trade landscape, businesses face unprecedented challenges in maintaining efficient and cost-effective IT infrastructure. The recent U.S. tariff adjustments have created waves of uncertainty across international markets, particularly for companies relying on high-performance computing and AI solutions. At Eliovp BV, we want to assure our valued clients that our comprehensive end-to-end ...","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1528,1612,1531,1532,1874,1875,1876,1877],"import","Taiwan","Tariffs","Trump",{"path":1879,"title":1880,"description":1881,"date":1882,"slug":1883,"image":1884,"originalUrl":1885,"categories":1886},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","1. Versatile IntegrationAI Agents are designed to integrate seamlessly with your existing software stack. This includes ERP, CRM, and marketing automation platforms. Instead of disrupting current systems, they complement and enhance them, all while learning from and adapting to your specific operational needs. 2. Intelligent Decision-MakingConventional automation scripts handle if-then scenarios, but they fall short ...","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1528,1529,1653,1531,1887,1888],"AI Agents","ERP",{"path":1890,"title":1891,"description":1892,"date":1893,"slug":1894,"image":1895,"originalUrl":1896,"categories":1897},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","In the rapidly evolving landscape of artificial intelligence (AI), open-source solutions are emerging as pivotal drivers of innovation and performance enhancement. These community-driven platforms democratize access to cutting-edge technologies, fostering collaboration and accelerating advancements in AI model optimization.​ The Open-Source Revolution in AI Open-source AI models have transformed the development and deployment of machine learning ...","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1528,1529,1612,1898,1532,1580,1535],"AI news",{"path":1900,"title":1901,"description":1902,"date":1903,"slug":1904,"image":1905,"originalUrl":1906,"categories":1907},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","1. Introduction Benchmarking is an essential part of optimizing AI models and software applications. Whether you’re testing AI model inference speeds, profiling different hardware configurations, or ensuring system performance over time, having a reliable benchmarking tool is crucial. However, many existing tools suffer from issues like inconsistent environments, difficult configuration setups, and lack of automation. ...","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1528,1529,1309,1531,1532,1908,1909,1533,1309],"benchmark","LLM",{"path":1911,"title":1912,"description":1913,"date":1914,"slug":1915,"image":1916,"originalUrl":1917,"categories":1918},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","1. Introduction In the world of large language models (LLMs), most benchmarks center on Llama or DeepSeek derivatives. We decided to diversify by adding the Qwen2 architecture, using our Paiton framework. This 32-billion-parameter model pushes GPU resources to the limit, perfect for comparing NVIDIA’s new H200 to our AMD MI300X, which leverages Paiton for advanced ...","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1528,1529,1309],{"path":1920,"title":1921,"description":1922,"date":1923,"slug":1924,"image":1925,"originalUrl":1926,"categories":1927},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","We’re excited to share that Eliovp was recently featured on AMD’s “Tech Talk” podcast! In this episode, our CEO, Elio Van Puyvelde sits down with Jim greene to talk about the origins of Eliovp, the passion and expertise that brought the company to life, and the innovative full end-to-end solutions we offer today. From our ...","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1528,1532,1928,1929,1930],"Jim Greene","Podcast","Tech Talk",{"path":1932,"title":1933,"description":1934,"date":1935,"slug":1936,"image":1937,"originalUrl":1938,"categories":1939},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Executive Summary If you’ve followed our journey so far, you’ll know that Paiton is laser-focused on AMD-centric inference optimization. Our latest work takes DeepSeek R1 Distill Llama 8B to the next level, delivering 10–15% higher throughput, improved time-to-first-token (TTFT), and more stable performance at lower batch sizes, an area that previously needed a boost. In short, Paiton further cements its ability to exploit AMD hardware’s raw power, bridging ...","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1528,1529,1309,1532,1940,1941,500,1533,1830,1309,1536],"Deepseek","H100",{"path":1943,"title":1944,"description":1945,"date":1946,"slug":1947,"image":1948,"originalUrl":1949,"categories":1950},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","A First Look at Paiton in Action: Deepseek R1 Distill Llama 3.1 8B","Outperforming Stock Models on the AMD MI300X 1. Introduction We couldn’t wait to show what Paiton can really do. After detailing our AMD-centric approach and architecture-level optimizations in our previous blog post, we decided to test-drive Paiton on a hype-worthy model: Deepseek R1 Distill Llama 3.1 8B. By compiling the model into efficient libraries and fusing ...","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1528,1529,1309,1532,1940,1941,500,1533,1830,1309,1536],{"path":1952,"title":1953,"description":1954,"date":1955,"slug":1956,"image":1957,"originalUrl":1958,"categories":1959},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","In the fast-paced world of artificial intelligence, model efficiency and performance are paramount. At ElioVP, we’re redefining what’s possible by delivering unparalleled optimization solutions for AI models with Paiton. By compiling the model’s architecture and leveraging our custom-written kernels, Paiton enables faster inference and reduced resource consumption on AMD GPUs. Why Model Optimization is More ...","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1528,1529,1309,1532,1941,500,1533,1830,1309,1536],1787908762610]