Engineering notes · DeepSeek on AMD

How ProsGrow Cut DeepSeek v4 Flash First-Token Latency by Up to 32%

Less waiting before the first answer. ProsGrow reduced DeepSeek V4.1 Flash’s cold time to first token by up to 32.0% on the same eight AMD MI325X GPUs, while retaining the model’s original checkpoint precision.

· ProsGrow AI Engineering · 4 min read

DeepSeek cold first-token latency: 0.564 to 0.384 seconds at approximately 8K input tokens, a 32% reduction on the same eight AMD MI325X GPUs.

The result: 30%+ less waiting

For an assistant working through a document or a coding request, the wait before generation starts is part of the user experience. In our real-serving tests, ProsGrow cut that wait from 0.564 to 0.384 seconds for an approximately 8,000-token prompt: a 32.0% reduction in time to first token (TTFT).

The comparison used DeepSeek V4.1 Flash on the same 8 × AMD MI325X host, against our local baseline with the same model revision and runtime versions. Cold requests had no prefix-cache hits. Improvements extended across all three tested prompt lengths.

Median cold time to first token · real serving · lower is better
Input lengthLocal baselineProsGrow optimizedLatency reduction
~8K tokens0.564 s0.384 s32.0%
~60K tokens3.931 s2.813 s28.4%
~300K tokens27.532 s21.872 s20.6%

At the longest tested prompt, that is about 5.66 seconds less waiting before the first content arrives. The headline describes first-token latency; complete-request throughput is measured separately below.

Where the improvement came from

We focused on how the serving stack executes the model on AMD hardware, particularly the work required to process an incoming prompt before generation begins. Profiling and targeted runtime optimization recovered performance from the existing deployment.

The result builds on DeepSeek’s model and the upstream vLLM and AMD AITER runtime capabilities. We retained the original checkpoint precision and real target-model verification for speculative generation. No additional lossy quantization was introduced.

For teams operating private AI infrastructure, this points to an opportunity: measure the serving bottleneck before expanding hardware. The right optimization depends on the model, prompt lengths, and response-time requirements of the application.

Throughput and quality checks

A separate post-deployment check with repeated, cached prompts showed 7.7–10.2% higher full-request output throughput across the three prompt lengths. This measure includes the wait for the first token and the time to receive the complete response. Cached first-token latency rose slightly, by 2.4–3.2%, so the benefit depends on which metric matters to the workload.

Before promotion, we compared the same 128 math questions under short- and long-prompt conditions. No baseline-correct answer became wrong in either condition, and all six post-deployment functional checks passed. These bounded checks support the change; broader coding, multilingual, and customer-task quality still need their own evaluation.

InferenceX reference vs. ProsGrow

In a separate one-hour AgentX benchmark, our optimized deployment achieved 261.780 tokens/s/user in median interactivity, compared with 239.808 in the saved InferenceX MI325X reference: a 9.16% increase. Output throughput per GPU rose from 15.83257 to 16.55379 tokens/s, a 4.56% increase.

InferenceX reference versus ProsGrow on eight AMD MI325X GPUs: median interactivity 239.808 versus 261.780 tokens/s/user, up 9.16%; output throughput 15.83257 versus 16.55379 tokens/s/GPU, up 4.56%. Both bar charts start at zero.
DeepSeek V4.1 Flash · one-hour AgentX · eight MI325X GPUs · tensor parallelism 8 · concurrency 1. Reference: InferenceX, row 442872 dated September 22, 2026 (published run). ProsGrow: local measurement on September 24, 2026. Select the image to view it full size.

These runs use the benchmark’s synthetic speculative acceptance length of 3.51. This is a performance test setting, separate from the real verification used in our serving and quality tests. The chart’s gains are separate from the 32% cold first-token latency reduction.

Median interactivity is the inverse of median inter-token latency, using the dashboard’s rounding convention. Throughput includes trace idle time. The historical reference does not fully identify its checkpoint and runtime image; this is a reference comparison with partial reproduction. The published and optimized runs completed 244 and 252 profiled requests, respectively, so their completed request mixes differ. There is one run per configuration and no confidence interval. The ProsGrow measurements shown here are local results.

What we measured

Tests ran on September 24, 2026, with the model distributed across all eight GPUs. Each real-serving request generated exactly 1,024 output tokens. TTFT measures the time from request start to the first nonempty streamed content. Corresponding prompts and output lengths were controlled, and cold requests were verified to have zero prefix-cache hits.

Each cold configuration used one warmup and four measured requests per prompt length. Runs were sequential, with different compilation and cache histories. These are local measurements against our reproduced baseline; results will vary with traffic, concurrency, and prompt mix. The 32% is the best observed reduction across the three tested lengths, with no confidence interval or production SLA implied.

Private DeepSeek hosting and AMD GPU compute

For teams evaluating private DeepSeek hosting or dedicated AMD GPU servers, this benchmark connects deployment planning to model precision, prompt length, and response time. Time to first token (TTFT) measures when an answer starts; full-request throughput measures a different part of the experience. ProsGrow can evaluate these trade-offs for your private inference workload.

Explore your deployment’s potential

Running DeepSeek or planning a private AI deployment? Contact us to discuss the results, our optimization approach, and an evaluation using your model and workload. We can help identify where your infrastructure has room to improve.

For access to AMD GPU compute, dedicated GPU servers, or private inference infrastructure, contact us to discuss your requirements.

Book a call with us

Prefer email? contact@prosgrow.ai