Shrinking Apertus 1.5 8B for an 8 GB Laptop GPU

Introduction

Swiss AI released Apertus 1.5 8B as a multimodal model with NVFP4 MLP weights, FP8 attention, and BF16 embedding and output matrices. The officially distributed checkpoint is onprem-ai/Apertus-v1.5-8B-NVFP4. It ships with a patched vLLM image that adds model support absent from stock vLLM releases. The model also has a large untied vocabulary, which becomes important later.

This article describes how we deployed that model on a Lenovo P14s Gen 6 laptop with an NVIDIA RTX PRO 1000 Blackwell GPU having 8 GB of VRAM, connected it to the Pi coding agent, benchmarked it, and then quantized it further to eliminate CPU offloading.

The result is a model that fits entirely in GPU memory with a 32K context window, runs 3 to 5× faster than the CPU-offloaded original, and serves as a conversational coding agent through Pi.

Hardware and Software

The workstation was:

  • Lenovo P14s Gen 6, Intel Core Ultra 7 265H, 64 GB RAM
  • NVIDIA RTX PRO 1000 Blackwell Laptop GPU, 8,151 MiB VRAM (approximately 8 GB)
  • WSL2 Ubuntu 24.04 on Windows
  • Docker with NVIDIA Container Toolkit
  • Pi coding agent 0.84.1
  • Swiss AI vLLM image: ghcr.io/swiss-ai/vllm_apertus_1.5_release:latest-amd64
  • Model checkpoint: onprem-ai/Apertus-v1.5-8B-NVFP4

Phase 1: Initial Deployment on 8 GB VRAM

The first challenge was getting Apertus to run at all. The NVFP4 checkpoint is approximately 8.4 GB on disk. vLLM reported 5.75 GiB of model residency, leaving almost nothing for KV cache on an 8 GB GPU.

The 8K starting point

The initial configuration used an 8,192-token context window with FP8 KV cache and 5 GB of CPU offloading:

--max-model-len 8192
--gpu-memory-utilization 0.86
--cpu-offload-gb 5
--kv-cache-dtype fp8
--max-num-seqs 1
--max-num-batched-tokens 2048

This worked for basic tool calls, but vLLM reported only 9,104 tokens of KV cache capacity. One large web-search result from Pi consumed the entire context and produced a maximum-context-length error. Eight thousand tokens was not a practical coding-agent context.

Why 32K did not fit

Quadrupling the context from 8K to 32K roughly quadrupled the KV cache requirement from 0.56 GiB to 2.0 GiB. Only 0.56 GiB was available after model loading and activation memory reservation. The problem was not just weight residency. The KV cache scaled with context length, and there was no room.

Attempt 1: CPU KV-cache offload

vLLM supported native KV offload to system RAM, but its startup admission check required one maximum-length request to fit in the GPU KV pool before the offload mechanism could engage. It rejected the 32K configuration before the offload path could help.

Attempt 2: Four-bit KV cache

NVFP4 KV cache was unavailable because Apertus uses a 128-byte attention head size that no available backend supported. The compatible alternative was per-token/per-head INT4:

--kv-cache-dtype int4_per_token_head

This reduced the 32K KV requirement from 2.0 GiB to 1.06 GiB. Still too much for the remaining 0.56 GiB.

Reducing activation memory

The original configuration processed 2,048 tokens per prefill batch, reserving 0.5 GiB for peak activations. Reducing the chunk size to 128 tokens cut the activation reservation to 0.03 GiB, freeing roughly 0.47 GiB. KV capacity rose from 0.56 GiB to 1.02 GiB, covering 31,568 tokens. Only 49 MiB short of 32K.

Finding the final margin

The GPU memory utilization was 86.0%. Raising it to 86.6% added 49 MiB without triggering vLLM’s startup-free threshold. The final admission report was (all figures as reported by vLLM’s startup log):

Available KV cache: 1.07 GiB
KV cache capacity: 33,040 tokens
Peak activation memory: 0.03 GiB
Model memory: 5.75 GiB

The working launch command

vllm serve onprem-ai/Apertus-v1.5-8B-NVFP4 \
--served-model-name onprem-ai/Apertus-v1.5-8B-NVFP4 \
--host 0.0.0.0 --port 8000 \
--chat-template-content-format string \
--max-model-len 32768 \
--gpu-memory-utilization 0.866 \
--cpu-offload-gb 7 \
--kv-cache-dtype int4_per_token_head \
--enable-mm-embeds \
--limit-mm-per-prompt '{"image":0,"audio":0}' \
--enforce-eager \
--max-num-seqs 1 \
--max-num-batched-tokens 128 \
--enable-chunked-prefill \
--enable-auto-tool-choice \
--tool-call-parser apertus

The --enable-mm-embeds and --limit-mm-per-prompt flags disable the vision and audio modalities at inference time, since this deployment uses Apertus as a text-only coding agent. The apertus tool-call parser tells vLLM how to decode the model’s structured function-call output into OpenAI-compatible tool invocations.

A 30,003-token synthetic prompt verified the full context. The server processed it in 198.7 seconds and returned a completion. Pi then reported a 32.8K context window with 2K max output tokens. A tool-call test through Pi produced a structured bash invocation, executed it, and consumed the result.

Connecting to Pi

Pi supports custom OpenAI-compatible providers through its models.json configuration. We registered the local vLLM server as vllm-apertus at http://127.0.0.1:8000/v1 with a dummy API key and a 32,768-token context window. The apertus tool-call parser enabled structured tool calls compatible with the model.

Operational limitations of the CPU-offloaded configuration

The CPU offloading meant roughly 4.2 billion parameters, the BF16 embedding and output head matrices plus overflow from the quantized layers, sat in system RAM and were fetched over PCIe on every forward pass. Generation speed was 5 to 5.6 tokens per second. A 30K prompt took 199 seconds to prefill. Only one request could run at a time. Other GPU models had to be unloaded first.

Phase 2: Benchmarking with Local Agent Bench

With the deployment stable, we ran the full Local Agent Bench suite (https://github.com/ebelo/local-agent-bench). The benchmark exercises two paths: a controlled ReAct harness where the framework manages tool execution, and Pi-native mode where the model drives Pi’s own tools directly.

The original benchmark

The benchmark ran 50 processes covering 260 unique tasks across five suites, with 10 repeats each. The run took 4 hours and 24 minutes. The vLLM server was cold-loaded once and kept resident throughout.

SuiteMean scorePercentMean latency/task
Controlled smoke3.90/578%51.3 s
Controlled agentic4.35/587%48.5 s
Pi-native basic1.35/527%38.8 s
Pi-native reality0.00/40%65.0 s
Pi-native ladder0.10/71.4%80.9 s

Apertus was competitive in the controlled harness. Its 87% agentic score tied Qwen2.5-Coder 7B, suggesting strong reasoning when the framework controls execution. Native Pi performance was poor. The model refused bash/curl calls, hallucinated weather data without retrieval, and confused Pi’s workspace with the benchmark repository. Native reality scored zero across 40 tasks.

For context, we had previously benchmarked six other models in the 7B to 9B range on the same hardware using the same Pi-native suites. In Pi-native basic, Mistral 7B scored 5.00/5 at 7.5 seconds per task. Apertus scored 1.35/5 at 38.8 seconds. In the reality suite, Mistral, Granite 4.1, Qwen3.5, and Ornith all scored above 80%. Apertus scored zero.

Apertus 1.5 8B had strong controlled-agent reasoning but could not reliably drive Pi’s native tools. The 5 to 5.6 tokens per second generation speed, caused by CPU offloading, made the situation worse.

Phase 3: Targeted W4 Quantization

The CPU offloading was the obvious bottleneck. If we could fit the entire model in GPU memory, we would eliminate the PCIe bottleneck and improve both speed and reliability. The question was what to quantize.

What was still in BF16

The NVFP4 checkpoint had already quantized the MLP layers to NVFP4 (4-bit float, group size 16) and the attention projections to FP8. But the input embedding and the output language-model head remained in BF16.

Apertus 1.5 has a large untied vocabulary. The input embedding has 266,752 rows and the output head has 131,072 rows. In BF16, these two matrices alone consumed roughly 3 GiB. They were the largest unquantized tensors in the checkpoint and the primary reason CPU offloading was needed.

The quantization recipe

We used llm-compressor 0.13.0 to apply W4A16 INT4 group-64 quantization to the embedding and LM head, leaving the NVFP4 MLP and FP8 attention untouched:

targets = {
"embedding": {
"targets": ["Embedding"],
"weights": {
"num_bits": 4, "type": "int", "symmetric": True,
"strategy": "group", "group_size": 64,
},
},
"output_head": {
"targets": ["re:.*lm_head$"],
"weights": {
"num_bits": 4, "type": "int", "symmetric": True,
"strategy": "group", "group_size": 64,
},
},
}
recipe = QuantizationModifier(config_groups=targets)
oneshot(model=model, recipe=recipe)

The quantization was data-free: no calibration data was needed because the weights were quantized directly. Group-64 symmetric INT4 on an embedding table is unusual (embeddings are looked up, not matmul’d), but vLLM’s compressed-tensors loader treats the quantized embedding as a standard linear layer with dequantization on load, so the group quantization is applied at weight-access time rather than during the forward pass. The resulting checkpoint was 6.1 GB on disk, down from 8.4 GB. The model residency in vLLM fell from 5.75 GiB to 5.11 GiB. The disk size drops more than residency because the original checkpoint includes vision and audio tokenizer weights that vLLM loads but does not count toward language-model residency.

Shard repacking for vLLM

vLLM’s Apertus loader expects three safetensors files with specific names: a language shard, a vision tokenizer shard, and an audio tokenizer shard. The llm-compressor output used Transformers’ default naming. We wrote split_for_vllm.py to repack the shards:

for shard in sorted(source.glob("*.safetensors")):
with safe_open(shard, framework="pt", device="cpu") as handle:
for key in handle.keys():
if key.startswith("model.vision_tokenizer."):
target = "model-vision_tokenizer-model.safetensors"
elif key.startswith("model.audio_tokenizer."):
target = "model-wavtokenizer-model.safetensors"
else:
target = "model-apertus-model-00001-of-00001.safetensors"
groups[target][key] = handle.get_tensor(key)

The script also resolved a loader requirement for model-vision_tokenizer-model.safetensors and ensured safetensor files were world-readable inside the container.

The compressed checkpoint’s quantization profile

The final checkpoint had four quantization groups:

GroupTargetFormatBitsStrategy
0Input embeddingpack-quantized4 INTgroup, size 64
1Attention q/k/v/o projectionsfloat-quantized8 FPchannel
2MLP up/down projectionsnvfp4-pack-quantized4 FPtensor-group, size 16
3LM headpack-quantized4 INTgroup, size 64

The audio tokenizer’s ConvNeXt layers were excluded from quantization to preserve fidelity.

Server configuration for the compressed checkpoint

The compressed checkpoint loaded without CPU offloading. The 32K configuration became:

vllm serve /model \
--max-model-len 32768 \
--gpu-memory-utilization 0.85 \
--kv-cache-dtype int4_per_token_head \
--enforce-eager \
--max-num-seqs 1 \
--max-num-batched-tokens 128 \
--enable-chunked-prefill \
--enable-auto-tool-choice \
--tool-call-parser apertus

No --cpu-offload-gb. The model fit entirely in GPU memory. vLLM reported 5.11 GiB model residency, 1.6 GiB KV cache, 49,360 token capacity, and 7.1 GiB steady GPU usage with 759 MiB free. Cold start dropped from roughly 3 minutes to 139 seconds.

vLLM selected MarlinNvFp4LinearKernel for the NVFP4 GEMM layers and MarlinLinearKernel for the compressed W4A16 weights. Marlin kernels are vLLM’s optimized quantized GEMM implementations, generated at load time from the checkpoint’s quantization format. The CUDA-fused xIELU activation remained unavailable, falling back to a Python implementation. This was a minor performance cost, not a correctness issue.

Generation speed improvement

The compressed checkpoint generated 14 to 22 tokens per second, up from 5 to 5.6. The improvement came from eliminating the PCIe round-trips to system RAM for the offloaded embedding, output head, and overflow parameters on every forward pass. The compressed benchmark ran later the same day and completed in approximately 50 minutes.

Phase 4: Benchmarking the Compressed Checkpoint

We ran the same Local Agent Bench suites against the compressed checkpoint: 10 repeats each of controlled smoke, controlled agentic, Pi-native basic, and Pi-native reality, plus one native ladder run. That produced 197 scored task-repeats (fewer than the original 260 because the ladder was not repeated 10 times).

SuiteOriginal meanW4 meanOriginal latency/taskW4 latency/task
Controlled smoke3.90/53.68/551.3 s10.0 s
Controlled agentic4.35/53.88/548.5 s14.1 s
Pi-native basic1.35/50.85/538.8 s16.6 s
Pi-native reality0.00/40.10/465.0 s27.5 s
Pi-native ladder0.10/70.00/780.9 s52.1 s

The compressed checkpoint was 3 to 5× faster per task. Controlled quality dropped 6% on smoke and 11% on agentic tasks. Native basic dropped 37%. Native reality and the ladder remained near zero.

Reliability

During sustained testing, the vLLM server entered an inference deadlock after dozens of successful requests. The /health endpoint continued returning 200 OK while direct /v1/chat/completions calls hung indefinitely. A server restart was required. After the restart, the server completed 422 successful chat-completion HTTP calls without another deadlock.

The benchmark runner was hardened with 120-second runtime and task timeouts, an inference probe that tested the completion endpoint rather than /health, and automatic server restart on failure. The reliability log recorded three events: two Pi process hangs and one confirmed inference deadlock.

The compressed-tensors version issue

We quantized in a Docker image built from the Swiss vLLM image with llmcompressor==0.13.0 installed on top, which upgraded compressed-tensors from 0.17.0 to 0.18.0 and broke the version contract. The correct approach would have been to quantize in a separate builder image and serve in the untouched Swiss image. The inference deadlock may have been caused by the dependency mismatch rather than the quantized weights. The performance and quality measurements remain valid because vLLM loaded and generated text correctly, but the reliability verdict is confounded until the benchmark is repeated with matched dependencies.

The final launcher update corrected this: the production vllm-apertus script uses the original Swiss image without llm-compressor, and the compressed checkpoint is served by vLLM’s native compressed-tensors loader.

Manual Pi Session: A Hands-On Test

After the compressed checkpoint was deployed, we sat down with Pi for a 15-minute manual session to get a qualitative feel for the model as a coding assistant, using the vllm-apertus provider with the quantized model.

What worked

The model responded to “hi” in 5.6 seconds and correctly listed its available tools: read, bash, edit, write, web_search, web_fetch. A pwd command returned the correct working directory in 2.4 seconds. An ls listed the files in the working directory. The model could fetch web content through web_fetch and web_search.

When asked to write a Python script computing 2+2, the model eventually produced a correct script and saved it to disk. A German translation of a paragraph from this article was fluent and accurate.

What did not work

Weather retrieval confusion. Asked for the current weather in Paris, the model first claimed it could not access weather data, despite having web_fetch and web_search tools. When explicitly directed to call https://wttr.in/Paris through web_fetch, it fetched the data but then confused Fahrenheit and Celsius. The wttr.in response reported 34°C, which the model initially presented as 34°F. After we corrected it, the model acknowledged the error but still produced a garbled forecast mixing Fahrenheit temperatures from AccuWeather with Celsius values from wttr.in.

Tool-call schema errors. When we asked the model to edit the Python script it had written, the model produced three consecutive malformed edit tool calls. Each had the right intent but the wrong argument structure, nesting oldText and newText one level too deep:

{
"edits": [
{ "edits[].oldText": "result = 2 + 2\nprint(result)",
"edits[].newText": "# Compute the sum\ntotal = 2 + 2\nprint(total)" }
]
}

The validation error was clear: edits.0.oldText: must have required properties oldText, newText. The model eventually fell back to write and overwrote the file, which worked but lost the original content.

Curl refusal. When we asked the model to run curl, it refused, saying curl was not in the available tool list. After we clarified “bash > curl”, the model executed curl --help through the bash tool. This mirrors the benchmark finding: Apertus understands tool concepts but frequently refuses or misroutes native tool calls, then succeeds when the instruction is explicit enough.

What the session showed

The manual test was consistent with what the benchmark had measured. The quantized model is conversational and responsive: most replies arrived in 2 to 6 seconds, a stark contrast to the 30 to 50 second latencies of the CPU-offloaded original. Basic file operations and text generation work. But the model has consistent trouble with native tool semantics: it refuses available tools, misformats structured arguments, and confuses retrieved data. These are model-level behavioral issues, not quantization artifacts. The W4 compression did not introduce them, and it did not fix them.

Phase 5: The Launcher Update

The original vllm-apertus launcher loaded the NVFP4 checkpoint from the Hugging Face cache with 7 GB CPU offloading. The updated launcher points to the local compressed checkpoint and removes the --cpu-offload-gb flag. GPU memory utilization drops from 0.866 to 0.85, and KV cache capacity rises from 33,040 to 49,360 tokens. Cold start drops from roughly 3 minutes to 139 seconds, and generation speed rises from 5 to 5.6 tokens per second to 14 to 22.

The original launcher was preserved as vllm-apertus.original for rollback.

Conclusion

The work produced a functional 32K coding-agent inference server on an 8 GB laptop GPU, running 3 to 5× faster than the original CPU-offloaded configuration. The targeted W4 quantization of the embedding and output head eliminated the need for CPU offloading while preserving the model’s NVFP4 MLP and FP8 attention layers.

The quality trade-off is modest on controlled tasks: a 6 to 11% drop in smoke and agentic scores. On the Pi-native tool-use suites, Apertus scored 0.85/5 on basic and 0.10/4 on reality. This looks poor in isolation, but the existing clean-room corpus puts it in context. The models we have run through Pi on the same hardware are, in descending order of native performance:

  • Mistral 7B: 5.00/5 basic, 3.00/3 reality, 7.5 s/task. The strongest Pi-native model on this GPU.
  • IBM Granite 4.1 8B: 4.80/5 basic, 3.00/3 reality, 20.4 s/task.
  • Qwen3.5 9B: 4.80/5 basic, 2.80/3 reality, 26.7 s/task.
  • Ornith 9B: 4.40/5 basic, 2.80/3 reality, 31.9 s/task.
  • LFM 2.5: 4.65/5 basic, 2.40/3 reality, 23.2 s/task.
  • Qwen2.5-Coder 7B: 4.30/5 basic, 1.85/3 reality, 15.4 s/task.
  • Apertus 1.5 8B (W4): 0.85/5 basic, 0.10/4 reality, 16.6 s/task. (Apertus uses a 4-task reality suite; the other models use a 3-task version. Reliability and native scores pending a clean-room rerun with matched compressed-tensors versions.)

Apertus sits at the bottom of the Pi-native ranking. Its controlled ReAct scores (3.68/5 smoke, 3.88/5 agentic) are competitive with Qwen2.5-Coder and Ornith, which means the model can reason through tool-use sequences when a framework manages execution. The gap is specifically in native tool invocation: the model refuses available tools, misformats structured arguments, and sometimes hallucinates instead of retrieving. The manual Pi session confirmed this: basic chat, file operations, and translation worked well, but weather retrieval, curl, and the edit tool all required explicit user intervention.

As a self-hosted, offline chat assistant, the quantized Apertus works well. It is conversational, responsive (2 to 6 seconds per reply), handles file operations, and produces fluent translations. The 32K context window makes it practical for longer documents and code. The W4 compression did not introduce the tool-use issues, and it did not fix them. They are model-level behaviors that would require fine-tuning or a different tool-call parser to address.

The compressed checkpoint is now the default in the production launcher, with the original preserved for rollback. The open work is a clean-room rerun of the benchmark with matched compressed-tensors versions to produce an unconfounded reliability verdict.

Discover more from Digital Pathlines

Subscribe now to keep reading and get access to the full archive.

Continue reading