7.8 KiB
Context window sizing — reference
Quick reference for picking context_size and max_tokens in
configs/portfolio-bot.yaml based on the host's RAM budget. Math,
recommendations, and tips.
1. context_size vs max_tokens
Two different budgets in the provider config:
context_size: 4096 # total window (input + output)
max_tokens: 2048 # generation cap per response
context_size (4096) = system prompt + RAG + history + respuesta
└────────── input ──────────┘ └output┘
max_tokens
context_size→ total tokens the model can see + produce. Maps to llama-server's--ctx-size.max_tokens→ cap on generation per response. Doesn't affect how much input fits, only how long the answer can be.
Rule of thumb for a Q&A bot: 512–1024 max_tokens is plenty. Bigger just
steals budget from the input side, where the auto-compactor then has to
fire sooner.
2. KV cache math
The constraint on context size is KV cache RAM, not the model's advertised window. KV cache grows linearly with context and is held in RAM per active stream:
KV cache (bytes) ≈ 2 × num_layers × num_kv_heads × head_dim × bytes × context_size
Reference values for the models the bot is usually paired with:
| Model | num_layers | num_kv_heads | head_dim | KB/token |
|---|---|---|---|---|
| qwen2.5-1.5b | 28 | 2 | 128 | ~18 |
| qwen2.5-3b | 36 | 4 | 128 | ~72 |
| gemma-3-1b | 18 | 1 | 256 | ~36 |
| gemma-3-4b | 34 | 4 | 256 | ~272 |
Q4_K_M model weights (also RAM-resident):
| Model | Size |
|---|---|
| qwen2.5-1.5b | ~1.0 GB |
| qwen2.5-3b | ~2.0 GB |
| gemma-3-1b | ~0.8 GB |
| gemma-3-4b | ~2.5 GB |
3. RAM budget on a single-model VPS
Fixed cost before we pick a context size:
Sistema + Go binary + SQLite ~0.5 GB
Modelo Q4_K_M weights (ver tabla arriba)
Buffer para picos y tmpfs ~0.5 GB
Available for KV cache + headroom = RAM_total − 0.5 GB − modelo.
4. Recommendations by RAM
4 GB VPS (bare-bones dev)
Solo viable con modelo chico y contexto bajo.
| Model | context_size | KV cache | RAM usada |
|---|---|---|---|
| gemma-3-1b Q4 | 8192 | ~290 MB | ~1.8 GB |
| qwen2.5-1.5b Q4 | 4096 | ~72 MB | ~1.6 GB |
8 GB VPS (típico)
| Model | context_size | KV cache | RAM usada | Veredicto |
|---|---|---|---|---|
| qwen2.5-1.5b Q4 | 16384 | ~290 MB | ~2.0 GB | muy cómodo |
| qwen2.5-3b Q4 | 8192 | ~580 MB | ~3.0 GB | sweet spot |
| qwen2.5-3b Q4 | 16384 | ~1.1 GB | ~3.6 GB | recomendado |
| qwen2.5-3b Q4 | 32768 | ~2.3 GB | ~4.8 GB | máximo útil |
| gemma-3-1b Q4 | 32768 | ~1.1 GB | ~2.4 GB | recomendado |
| gemma-3-4b Q4 | 8192 | ~2.2 GB | ~5.2 GB | ajustado |
16 GB VPS
| Model | context_size | KV cache | RAM usada |
|---|---|---|---|
| qwen2.5-3b Q4 | 32768 | ~2.3 GB | ~5.0 GB |
| gemma-3-4b Q4 | 16384 | ~4.5 GB | ~7.5 GB |
5. Worked examples
qwen2.5-3b en 8 GB
# configs/portfolio-bot.yaml
providers:
- name: llamacpp-local
type: llamacpp
model: qwen2.5-3b-instruct
endpoint: http://localhost:9100/v1
context_size: 16384 # ~1.1 GB KV, deja 4 GB libres
max_tokens: 1024 # respuestas moderadas
llama-server \
-m qwen2.5-3b-instruct-q4_k_m.gguf \
--ctx-size 16384 \
-ngl 0 -t 2 \
--mlock
gemma-3-1b en 8 GB
providers:
- name: llamacpp-local
type: llamacpp
model: gemma-3-1b-it
endpoint: http://localhost:9100/v1
context_size: 32768 # sobra RAM, contexto largo
max_tokens: 1024
llama-server \
-m gemma-3-1b-it-Q4_K_M.gguf \
--ctx-size 32768 \
-ngl 0 -t 2 \
--mlock
Gemma con chat template custom (sin system role líder)
Gemma 3 rechaza mensajes system antes del primer user. Dos opciones:
Opción A — template custom en ~/.llama/gemma3.jinja:
{% if messages[0]['role'] != 'system' and messages[0]['role'] != 'user' %}
{{ raise_exception('First message must be system or user') }}
{% endif %}
{% for message in messages %}
{% if message['role'] == 'system' %}
{{ message['content'] | trim + '\n\n' -}}
{% elif message['role'] == 'user' %}
{{- '<start_of_turn>user\n' + message['content'] | trim + '<end_of_turn>\n' -}}
{% elif message['role'] == 'assistant' or message['role'] == 'model' %}
{{- '<start_of_turn>model\n' + message['content'] | trim + '<end_of_turn>\n' -}}
{% endif %}
{% endfor %}
{% if add_generation_prompt %}
{{- '<start_of_turn>model\n' -}}
{% endif %}
llama-server \
-m gemma-3-1b-it-Q4_K_M.gguf \
--ctx-size 32768 \
-ngl 0 -t 2 \
--mlock \
--chat-template-file ~/.llama/gemma3.jinja
Opción B — usar Ollama, que mapea system → prefix del primer user automáticamente.
6. Tuning with auto-compaction
The bot has built-in auto-compaction (configs/portfolio-bot.yaml →
compaction: block). When the previous turn's input tokens exceed
threshold_ratio × MaxContextWindow, the older portion of the chat gets
summarized into a single system note. This means:
- A smaller
context_sizestill works for long conversations — the compactor frees up room by folding old turns. - Bigger
max_tokensmeans the compactor fires sooner (less budget left for input). - Default
threshold_ratio: 0.75triggers compaction at ~75% of the window. Lower it (e.g.0.5) for headroom on slow CPU where each request is expensive; raise it (e.g.0.9) when you want to keep more verbatim history.
For a portfolio bot with context_size: 16384 and max_tokens: 1024,
compaction fires when input exceeds ~12k tokens — leaving ~5k for the
fresh history, which is ~10-15 recent user turns. More than enough.
7. Tips
--mlockes oro en VPS. Bloquea el modelo en RAM y evita swaps cuando hay picos de memoria. Cuesta ~modelo_size de locked RAM.- Monitoreá con
htopofree -hla primera semana. Si ves swap, bajá el context. - Más contexto ≠ más rápido. El prefill (procesar el input) escala
lineal con la cantidad de tokens. Generación (output) no se ve
afectada. Con
--ctx-size 32768y un input de 500 tokens, el TTFT apenas cambia; con 20k tokens de input sí. - Streams concurrentes. Cada stream activo reserva su propio KV cache. En 8 GB no hagas más de 1-2 streams simultáneos — el rate limiter del bot (default 30 req/min) ya te protege.
max_tokensbajo ayuda. 512 es suficiente para Q&A. Bajarlo deja más presupuesto para input y retrasa la compactación.
8. Quick-pick table
Copy-paste según tu setup:
| Setup | context_size |
max_tokens |
|---|---|---|
| 4 GB + gemma-3-1b | 8192 | 512 |
| 4 GB + qwen2.5-1.5b | 4096 | 512 |
| 8 GB + qwen2.5-1.5b | 16384 | 768 |
| 8 GB + qwen2.5-3b | 16384 | 1024 |
| 8 GB + gemma-3-1b | 32768 | 1024 |
| 16 GB + qwen2.5-3b | 32768 | 1024 |
| 16 GB + gemma-3-4b | 16384 | 1024 |