Skip to main content

[ 00 / 10 ]

[ PAITON | AMD GPU OPTIMIZATION | REAL WORKLOADS ]

MORE FROM AMD GPUS. PROVEN ON REAL MODELS.

Get more performance
from AMD GPUs.

Paiton finds what slows your model down, replaces the generic runtime path,
and turns AMD hardware into faster production inference without retraining.

[ 00 / 10 ]

[ LATEST BREAKTHROUGHS ]

MEASURED ON REAL GENERATION AND INFERENCE PATHS.

Benchmarks that make AMD
worth a second look

The gap is rarely the model. It is usually the runtime.
Paiton tunes the exact path your workload uses, then proves the gain with benchmarks.

Video diffusion: MI355X vs B200

17.6%

Wan2.2-T2V-A14B ran faster on AMD MI355X

5.501sAMD MI355X
6.672sNVIDIA B200

Paiton-Diffusers removed runtime waste from the generation path and let the hardware show up.

Read Wan2.2 Benchmark
Why it matters

You do not need a new model to improve unit economics. You need the runtime, kernels, and deployment path tuned around the workload you already run.

[ 00 / 10 ]

[ GPU ECONOMICS MODEL ]

TRANSLATE THROUGHPUT INTO REUSABLE CAPACITY.

EDITABLE PLANNING SCENARIO

Turn runtime gains
into a business case.

Faster inference is more than a benchmark number. It can reduce the GPU time required for the same workload,

or increase the output of hardware you already own.

Start with this illustrative scenario, then replace the assumptions with the uplift measured during your Paiton benchmark.

YOUR GPU SCENARIOAll four assumptions are editable
SAME WORKLOAD13.0% fewer GPU-hours per unit of output
16,000 current GPU-hours / month
Current runtime baselinePaiton-adjusted estimate: 13,913 GPU-hours

CAPACITY RELEASED

2,087GPU-hours / month

Equivalent to roughly 4.2 GPUs at this utilization level.

Potential compute-capacity value€6,261 / month€75,130 / year

SAME COMPUTE BUDGET

+15.0%potential output

Use the released runtime for more tokens, generations, jobs, or customer demand without expanding the GPU budget.

A published Radeon AI PRO R9700 test measured 39.77 versus 32.86 output tokens per second on the same Qwen3.8 workload. Gains remain workload-specific and that result is not applied automatically here.

View the R9700 benchmark
Benchmark my workload

Illustrative planning model. Actual gains depend on the model, batch size, sequence length, precision, runtime, hardware, scaling behaviour, and deployment. The calculation assumes throughput improvement translates proportionally into lower GPU runtime for the same workload. Released capacity may become a cash saving in usage-based infrastructure or reusable capacity on owned hardware. Figures exclude Paiton fees and do not guarantee savings or performance.

[ 00 / 10 ]

[ WHAT PAITON CHANGES ]

LLMS, VIDEO DIFFUSION, MOE, AND CUSTOM STACKS.

Paiton workload optimization visual

Turn AMD hardware
into production throughput.

Whether you run LLMs, image and video diffusion, MoE, or custom architectures,
Paiton starts where money is lost: latency, throughput, memory pressure, and cost per token.

Performance icon

More tokens, less waiting

Find the bottleneck that users and invoices feel. Qwen3.8 on one Radeon AI PRO R9700 delivered 21% more interactive output and 54.3% more coding output.

Weights unchanged icon

No retraining required

Keep the model and improve the path around it with AMD-aware operators and custom .so files.

Runtime icon

Runtime work that pays back

Local FLUX.2 klein image generation with ComfyUI on R9700, Wan2.2 video generation on Instinct, and vLLM support for language inference.

Language model icon

Language Models

Llama, Qwen, Deepseek, Gemma, Mistral, and CodeLlama with vLLM support, FP8 precision, and custom kernels when the gain is worth it.

Diffusion and video icon

Diffusion & Video

Wan2.2, Flux, Stable Diffusion, SDXL, ControlNet, and text-to-video workloads where scheduler, memory, and kernel choices move the number.

Advanced model icon

Advanced Models

MoE, MLA, multimodal systems, proprietary architectures, expert kernels, novel attention, and custom inference stacks.

[ 00 / 10 ]

[ ENGAGEMENT FLOW ]

MEASURE. TUNE. PROVE.

A practical optimization sprint

Bring the workload, the target GPU, and the performance goal.
We turn that into a focused path from measurement to production-ready gain.

01Measure

Run the real workload and identify where latency, memory pressure, or cost per token is leaking value.

02Tune

Build the AMD-specific path with custom .so files, FP8 precision, kernel fusion, and the same model weights.

03Prove

Deploy the optimized path on AMD Instinct or Radeon AI PRO hardware and compare the result against the baseline.

[ 00 / 10 ]

[ DEPLOYMENT FIT ]

ROCM, KERNELS, MULTI-GPU, AND RUNTIME WORK.

Serious AMD tuning,
for production inference.

Runtime work

vLLMPaiton-DiffusersComfyUI (FLUX.2 klein on R9700)ROCmHIPIndependent kernels

Performance levers

FP8 precisionTensor parallelismData parallelismKernel fusionCustom AMD operatorsMulti-GPU scaling

What gets tuned

FP8 can reduce memory pressure while preserving accuracy. Tensor parallelism spreads large models like Llama-3.1-405B across multiple AMD GPUs. Custom kernels target the bottlenecks generic runtimes leave behind.

CDNA 4.0

MI355X

288GB HBM3E

CDNA 3.0

MI325X / MI300X / MI300A

256GB HBM3E, 192GB HBM3, 128GB HBM3 APU

CDNA 2.0

MI250X / MI250 / MI210

Datacenter AMD GPU support

CDNA 1.0

MI100

First compute-optimized generation

RDNA 4.0

Radeon AI PRO R9700

Qualified Qwen3.8, Ornith 1.5 and FLUX.2 klein profiles on 32GB GDDR6

RDNA 3.0

RX 7900 XTX / RX 7900 XT

Supported for selected inference workloads

RDNA 2.0

RX 6800 XT / RX 6900 XT

Supported with workload-specific qualification

Paiton supports both AMD Instinct datacenter accelerators and Radeon GPUs. Every model, runtime, precision, and hardware combination is qualified against a defined deployment scope before delivery.

[ 00 / 10 ]

[ COMMUNITY ]

OUR CONTRIBUTION TO OPEN SOURCE.

Faster local AI.Shared with the community.

Better inference should reach beyond our customer projects. Our open-source Paiton vLLM plugin brings model-specific AMD GPU optimizations to developers running AI on their own hardware.

Explore the code, run the supported models locally, and share what you learn. It is our way of helping more people get more from the GPUs they already own.

Public runtime pluginApache-2.0

paiton-vllm-plugin

Code, containers and setup guides. Ready to explore on GitHub.

Explore the plugin on GitHub

Published profiles cover Qwen3.8 and Ornith 1.5 on Radeon AI PRO R9700, for single-user text inference. See the repository for exact configurations and requirements.

[ 00 / 10 ]

[ EVIDENCE ]

PUBLISHED BENCHMARKS AND CASE STUDIES.

Read the benchmarks
behind the claims

16.2% less generation time

FLUX.2 klein on Radeon AI PRO R9700

1.054 seconds per image with a warm pipeline, 33.4% lower peak Torch allocation, and a free local ComfyUI workflow. Tested at 1024 × 1024, four steps and batch one.

Read Benchmark

+27% output

Ornith 1.5 on one 32GB Radeon GPU

Paiton with DFlash delivers 44.63 output tokens per second on R9700, 27% ahead of tuned stock vLLM in the tested single-user workload.

Read Benchmark

+21% output

Qwen3.8 on Radeon AI PRO R9700

Higher interactive, coding, and long-context throughput from the same 32GB Radeon GPU.

Read Benchmark

17.6% faster

Wan2.2 on MI355X beats B200

Text-to-video diffusion speedup with Paiton-Diffusers.

Read Benchmark

405B model

Llama-3.1-405B acceleration

Faster startup and inference for a frontier-scale open model.

Read Case Study

$/1M tokens

MoE kernels beat H200/B200

MI300X plus Paiton runtime on long-context MoE economics.

Read Benchmark

[ 00 / 10 ]

[ NEXT STEP ]

START WITH THE MODEL AND TARGET GPU.
Paiton CTA background accent

Show us the workload.
We will show you the gain.

Share the model, target GPU, and production goal.
Paiton turns that into a benchmark-led optimization plan.