Benchmarking enterprise LLMs on AMD GPUs with Red Hat AI Inference and Podman

by , , | Sep 28, 2026 | AI

Deploying large language models (LLMs) on enterprise hardware requires a balance of reproducibility, security, and performance. Accurately benchmarking model performance under production-like workloads is both challenging and crucial for optimizing infrastructure investments and ensuring reliable service delivery.  Red Hat AI Inference provides an enterprise inference stack built on vLLM and llm-d for deploying and serving generative AI models across accelerators and clouds.

In this step-by-step guide, we’ll show you how to deploy an OpenAI-compatible LLM inference server on an AMD GPU instance using rootless Podman. We’ll verify the AMD Radeon Open Compute (ROCm) stack and container-level GPU access, test the deployed model, and benchmark its performance using GuideLLM.

This tutorial is intended for developers, performance engineers, data scientists, and system administrators who are familiar with Linux and containers but may be new to running and benchmarking LLM workloads on AMD GPUs.

Deployment architecture

This guide uses a single local machine equipped with an AMD GPU. 

vLLM is a high-performance inference and serving engine for LLMs that exposes an OpenAI-compatible API. GuideLLM is a benchmarking and load-generation tool that is part of vLLM, and regularly used to evaluate LLM inference endpoints under controlled workloads

All components run on the same host using Podman:

  • Red Hat AI Inference vLLM server: Serves the selected LLM through an OpenAI-compatible API.
  • GuideLLM: Generates benchmark workloads against the local inference endpoint.
  • Open WebUI: Provides an optional browser-based interface for interacting with the deployed model.

Because the inference server and benchmark clients run on the same machine, they communicate through the local host interface at 127.0.0.1.

1. Set up environment variables

Set up the model, container image, and benchmark parameters on the local AMD GPU system. Keep this shell open throughout the test.

# Model & Workload Settings  
export MODEL='Qwen/Qwen2.5-7B-Instruct'
export HF_TOKEN='hf_xxxxx' # Optional: set if using gated models
export GPU_VENDOR='amd'

# Red Hat AI Inference 3.5 AMD/ROCm and GuideLLM Container Images
export VLLM_IMAGE='registry.redhat.io/rhaii/vllm-rocm-rhel9:3.5.0-1786491157'
export GUIDELLM_IMAGE='registry.redhat.io/rhai/guidellm-rhel9:3.5.0-1787154406'

# Benchmarking Parameters (PSAP Methodology)
export TENSOR_PARALLEL_SIZE=1
export CONCURRENCY_LEVELS='1,50,100,200,300,500,650'
export BENCHMARK_SECONDS=600

2. Prepare and verify the AMD GPU stack

ROCm is AMD’s software platform for GPU computing. Before launching the inference server, we need to verify that the local system detects the AMD GPU and exposes the required ROCm device interfaces.

Install basic runtime packages

Install Podman and the required utilities on the local system:

sudo dnf install -y podman crun curl jq pciutils
sudo dnf module install -y container-tools:rhel9

 Note: Sudo is used here only to install host packages. All vLLM and GuideLLM container workloads will run entirely rootless without sudo.

Verify host hardware and ROCm drivers

Ensure that the host operating system properly detects the AMD GPUs and loads required kernel modules:

echo "=== AMD-SMI ===" && amd-smi list
echo -e "\n=== AMDGPU MODULE ===" && lsmod | grep amdgpu
echo -e "\n=== GPU DEVICES ===" && ls -l /dev/kfd /dev/dri /dev/dri/render*
echo -e "\n=== ROCMINFO ===" && rocminfo | grep -E "Name:|Marketing Name:|gfx" | head -100

 Verification checklist:

  • /dev/kfd is exposed (ROCm compute device interface).
  • /dev/dri and renderD* devices exist.
  • lsmod confirms amdgpu is loaded.
  • rocminfo lists your GPU architecture.

3. Pull Red Hat AI Inference and GuideLLM container images

Pull the Red Hat AI Inference and GuideLLM container images from registry.redhat.io:

# Log in to Red Hat Registry
podman login registry.redhat.io

# Pull AMD vLLM and GuideLLM images
podman pull "$VLLM_IMAGE"
podman pull "$GUIDELLM_IMAGE"

4. Validate rootless Podman and ROCm permissions

Check group access and OCI runtime

Ensure your user account belongs to the video and render groups for GPU access, and verify that Podman is configured to use the crun OCI runtime required for group preservation.

echo "User: $(id -un)"
echo "Groups: $(id -nG)"

for group in video render; do
  if ! id -nG | tr " " "\n" | grep -qx "$group"; then
    echo "ERROR: user $(id -un) is not a member of the $group group" >&2
    exit 1
  fi
done

OCI_RUNTIME=$(podman info --format "{{.Host.OCIRuntime.Name}}")
echo "Podman OCI runtime: $OCI_RUNTIME"

if [ "$OCI_RUNTIME" != "crun" ]; then
  echo "ERROR: --group-add=keep-groups requires the crun OCI runtime" >&2
  exit 1
fi

Validate container-level GPU detection

Confirm PyTorch inside the Red Hat AI Inference image recognizes the hardware:

export AVAILABLE_GPUS=$(podman run --rm \
  --device=/dev/kfd \
  --device=/dev/dri \
  --security-opt=label=disable \
  --group-add=keep-groups \
  --entrypoint=python3 \
  "$VLLM_IMAGE" \
  -c 'import torch; print(torch.cuda.device_count())')

echo "Detected GPUs: $AVAILABLE_GPUS"

5. Deploy the rootless Red Hat AI Inference instance 

Create a persistent Hugging Face cache directory, and start the inference server:

# Launch AMD vLLM Container
export LOCAL_CACHE="$HOME/rhaii-cache"

mkdir -p "$LOCAL_CACHE"
chmod g+rwX "$LOCAL_CACHE"

podman run -d \
  --name vllm-server \
  --network host \
  --device=/dev/kfd \
  --device=/dev/dri \
  --security-opt=label=disable \
  --group-add=keep-groups \
  --shm-size=4GB \
  -e HF_HUB_OFFLINE=0 \
  -e HF_TOKEN="$HF_TOKEN" \
  -v "$LOCAL_CACHE:/opt/app-root/src/.cache" \
  "$VLLM_IMAGE" \
  --model "$MODEL" \
  --host 127.0.0.1 \
  --port 8000 \
  --tensor-parallel-size "$TENSOR_PARALLEL_SIZE"

Wait for server health:

for attempt in $(seq 1 120); do
  if curl -fsS http://127.0.0.1:8000/health >/dev/null; then
    echo 'vLLM is healthy'
    break
  fi

  if [ "$attempt" -eq 120 ]; then
    echo 'ERROR: vLLM did not become healthy' >&2
    podman logs --tail 200 vllm-server >&2
    exit 1
  fi

  echo 'Waiting for vLLM...'
  sleep 5
done

6. Test inference and web UI

Execute an ad-hoc inference request

Test prompt execution using jq to build clean JSON payloads:

export PROMPT='This model runs on an AMD GPU stack'

PAYLOAD=$(jq -n \
  --arg model "$MODEL" \
  --arg prompt "$PROMPT" \
  '{model:$model,messages:[{role:"user",content:$prompt}]}')

curl -fsS http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  --data-binary "$PAYLOAD" | jq .

Connect Open WebUI

Launch Open WebUI on your local AMD GPU system:

# 1. Start Open WebUI container

podman run -d \
  --name open-webui \
  --network host \
  -v open-webui:/app/backend/data \
  -e OPENAI_API_BASE_URL=http://127.0.0.1:8000/v1 \
  -e OPENAI_API_KEY=none \
  -e 'DEFAULT_MODEL_METADATA={"capabilities":{"builtin_tools":false}}' \
  ghcr.io/open-webui/open-webui:main

7. Run GuideLLM load rest

Create result directories

Define a timestamped directory to organize your benchmark outputs, logs, and run metadata in a single location:

export RUN_ID="$(date +%Y%m%d-%H%M%S)-$(echo "$MODEL" | tr '/:' '__')"
export RESULTS_DIR="$HOME/vllm-gpu-bench-results/$RUN_ID"

mkdir -p "$RESULTS_DIR"
chmod g+rwX "$RESULTS_DIR"

printf '%s\n' \
  "run_id=$RUN_ID" \
  "model=$MODEL" \
  "gpu_vendor=$GPU_VENDOR" \
  "available_gpus=$AVAILABLE_GPUS" \
  "tensor_parallel_size=$TENSOR_PARALLEL_SIZE" \
  "vllm_image=$VLLM_IMAGE" \
  "guidellm_image=$GUIDELLM_IMAGE" \
  > "$RESULTS_DIR/run-info.txt"

echo "RESULTS_DIR=$RESULTS_DIR"

Execute GuideLLM using the standard balanced profile (1,000 prompt tokens / 1,000 output tokens):

This workload uses a 1,000-token input and requests a 1,000-token output. We run it across multiple concurrency levels to observe how throughput and latency change as the number of simultaneous requests increases.

podman run --rm \
  --network host \
  -e HF_TOKEN="$HF_TOKEN" \
  -v "$RESULTS_DIR:/results:Z" \
  "$GUIDELLM_IMAGE" \
  guidellm run \
  --backend kind=openai_http,target=http://127.0.0.1:8000,model=$MODEL,timeout=1000 \
  --data kind=synthetic_text,prompt_tokens=1000,output_tokens=1000 \
  --profile kind=concurrent \
  --override profile.streams "1,50,100,200,300,500,650" \
  --constraint kind=max_duration,seconds=600 \
  --constraint kind=max_error_rate,rate=0.05 \
  --output kind=json,path=/results/concurrent.json \
  --output kind=csv,path=/results/concurrent.csv \
  --output kind=html,path=/results/concurrent.html \
  | tee "$RESULTS_DIR/guidellm-concurrent-console.log"

Execute GuideLLM using multi-turn + prefix-cache

This scenario generates a four-turn conversation in which each new turn includes the previous conversation history. The repeated prefix creates an opportunity for vLLM’s automatic prefix caching to reuse previously processed tokens and reduce repeated prefill work across turns.

podman run --rm \
  --network host \
  -e HF_TOKEN="$HF_TOKEN" \
  -v "$RESULTS_DIR:/results:Z" \
  "$GUIDELLM_IMAGE" \
  guidellm run \
  --backend kind=openai_http,target=http://127.0.0.1:8000,model=$MODEL,timeout=1000,request_format=/v1/chat/completions \
  --data kind=synthetic_text,prompt_tokens=1000,output_tokens=1000,turns=4 \
  --profile kind=concurrent,streams=4 \
  --constraint kind=max_duration,seconds=600 \
  --constraint kind=max_error_rate,rate=0.05 \
  --output kind=json,path=/results/multiturn-prefix-cache.json \
  --output kind=csv,path=/results/multiturn-prefix-cache.csv \
  | tee "$RESULTS_DIR/guidellm-multiturn-prefix-cache-console.log"

Execute GuideLLM using Mooncake trace

Unlike the previous synthetic workloads, this scenario replays a Mooncake conversation trace containing request timing, input and output lengths, and shared-prefix information. This produces a more production-like, cache-sensitive workload for evaluating the inference server under realistic request patterns.

# Create directory for trace files
mkdir -p "$HOME/traces"

# Download trace file
curl -L \  https://raw.githubusercontent.com/kvcache-ai/Mooncake/main/FAST25-release/traces/conversation_trace.jsonl \
  -o "$HOME/traces/conversation_trace.jsonl"

# Verify downloaded trace file
ls -lh "$HOME/traces/conversation_trace.jsonl"
podman run --rm \
  --network host \
  -e HF_TOKEN="$HF_TOKEN" \
  -v "$RESULTS_DIR:/results:Z" \
  -v "$HOME/traces:/traces:ro,Z" \
  "$GUIDELLM_IMAGE" \
  guidellm run \
  --backend kind=openai_http,target=http://127.0.0.1:8000,model=$MODEL,timeout=1000 \
  --data kind=mooncake,path=/traces/conversation_trace.jsonl \
  --profile kind=replay,time_scale=0.001 \
  --constraint kind=max_error_rate,rate=0.05 \
  --output kind=json,path=/results/mooncake.json \
  --output kind=csv,path=/results/mooncake.csv \
  | tee "$RESULTS_DIR/guidellm-mooncake-console.log"

8. Clean up resources

After saving any benchmark results you want to keep, remove the containers, images, benchmark output, and model cache created during this walkthrough:

podman rm -f vllm-server 2>/dev/null || true
podman rmi -f "$VLLM_IMAGE" 2>/dev/null || true
podman rmi -f "$GUIDELLM_IMAGE" 2>/dev/null || true
rm -rf "$RESULTS_DIR"
rm -rf "$LOCAL_CACHE"

 

Compare performance in other scenarios

This guide walked you through deploying a vLLM-based inference server from Red Hat AI Inference on a local AMD GPU system using rootless Podman, allowing you to verify GPU access from the host and container and test the model through its OpenAI-compatible API. From there, you can use GuideLLM to run several benchmark workloads: a balanced workload provides for evaluating throughput and latency across concurrency levels, and long-context, multi-turn, and trace-replay scenarios expose different aspects of model-serving behavior.

From here, you can repeat the workflow with other supported models, GPU configurations, tensor-parallel settings, or GuideLLM workloads and compare the resulting performance data.

What’s next?

Now you are ready to build and benchmark your own solutions on top of the vLLM Red Hat Inference Server. Explore more in the Red Hat AI Inference demo and quick-start sections

Interested in more demos with Red Hat AI Inference or exploring other Red Hat AI offerings? Visit AI Quickstarts: https://docs.redhat.com/en/learn/ai-quickstarts/