Capability change announced Sep 18, 2026
LLM Tech documents a capability change: “2026-09-18 Service is back after the maintenance window, on new hardware: one RTX PRO 6000 Blackwell with 96 GB, in Italy.”
- Hosting route
- LLM Tech platform
- Affected scope
- Not stated in the notice
- Announced
- Sep 18, 2026
- First seen by ModelClock
- Oct 8, 2026
What the provider published
llmtech.eu ↗2026-09-18 Service is back after the maintenance window, on new hardware: one RTX PRO 6000 Blackwell with 96 GB, in Italy. The engine moved from a nightly vLLM build to the 0.29.0 stable release, which closes the risk named in the entry of 28 August, and a second cache tier of 64 GiB now lives in host RAM. Measured on the card in service that day, with nothing else running on it: 61.6 tok/s on a single stream, 50.7 at twelve concurrent, 32.5 at thirty-two; 0.17 s to first token; 13,190 tok/s of fresh prefill; a cached prompt comes back 19.3 times faster than the same prompt cold. The public model id is now nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps working as an alias, so nothing built against it breaks. The prompt cache is
Read from the provider's text; the quote is the provider's exact lines.