LLM Tech · LLM Tech platform · LLM

nvidia/Qwen3.8-27B-NVFP4

Capability changeCapability change announced Sep 18, 2026

Notice

Capability change

Capability change announced Sep 18, 2026

LLM Tech documents a capability change: “2026-09-18 Service is back after the maintenance window, on new hardware: one RTX PRO 6000 Blackwell with 96 GB, in Italy.”

Hosting route
LLM Tech platform
Affected scope
Not stated in the notice
Announced
Sep 18, 2026
First seen by ModelClock
Oct 8, 2026

What the provider published

llmtech.eu ↗

2026-09-18 Service is back after the maintenance window, on new hardware: one RTX PRO 6000 Blackwell with 96 GB, in Italy. The engine moved from a nightly vLLM build to the 0.29.0 stable release, which closes the risk named in the entry of 28 August, and a second cache tier of 64 GiB now lives in host RAM. Measured on the card in service that day, with nothing else running on it: 61.6 tok/s on a single stream, 50.7 at twelve concurrent, 32.5 at thirty-two; 0.17 s to first token; 13,190 tok/s of fresh prefill; a cached prompt comes back 19.3 times faster than the same prompt cold. The public model id is now nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps working as an alias, so nothing built against it breaks. The prompt cache is

Read from the provider's text; the quote is the provider's exact lines.

Revision history

Loading…
Raw history JSON