diff --git a/docs/vps-context-sizing.md b/docs/vps-context-sizing.md new file mode 100644 index 0000000..91c1f6c --- /dev/null +++ b/docs/vps-context-sizing.md @@ -0,0 +1,243 @@ +# Context window sizing — reference + +Quick reference for picking `context_size` and `max_tokens` in +`configs/portfolio-bot.yaml` based on the host's RAM budget. Math, +recommendations, and tips. + +--- + +## 1. `context_size` vs `max_tokens` + +Two different budgets in the provider config: + +```yaml +context_size: 4096 # total window (input + output) +max_tokens: 2048 # generation cap per response +``` + +``` +context_size (4096) = system prompt + RAG + history + respuesta + └────────── input ──────────┘ └output┘ + max_tokens +``` + +- **`context_size`** → total tokens the model can see + produce. Maps to + llama-server's `--ctx-size`. +- **`max_tokens`** → cap on **generation** per response. Doesn't affect + how much input fits, only how long the answer can be. + +Rule of thumb for a Q&A bot: 512–1024 `max_tokens` is plenty. Bigger just +steals budget from the input side, where the auto-compactor then has to +fire sooner. + +--- + +## 2. KV cache math + +The constraint on context size is **KV cache RAM**, not the model's +advertised window. KV cache grows linearly with context and is held in +RAM per active stream: + +``` +KV cache (bytes) ≈ 2 × num_layers × num_kv_heads × head_dim × bytes × context_size +``` + +Reference values for the models the bot is usually paired with: + +| Model | num_layers | num_kv_heads | head_dim | KB/token | +|-------------------|-----------:|-------------:|---------:|---------:| +| qwen2.5-1.5b | 28 | 2 | 128 | ~18 | +| qwen2.5-3b | 36 | 4 | 128 | ~72 | +| gemma-3-1b | 18 | 1 | 256 | ~36 | +| gemma-3-4b | 34 | 4 | 256 | ~272 | + +Q4_K_M model weights (also RAM-resident): + +| Model | Size | +|-------------------|--------:| +| qwen2.5-1.5b | ~1.0 GB | +| qwen2.5-3b | ~2.0 GB | +| gemma-3-1b | ~0.8 GB | +| gemma-3-4b | ~2.5 GB | + +--- + +## 3. RAM budget on a single-model VPS + +Fixed cost before we pick a context size: + +``` +Sistema + Go binary + SQLite ~0.5 GB +Modelo Q4_K_M weights (ver tabla arriba) +Buffer para picos y tmpfs ~0.5 GB +``` + +Available for **KV cache + headroom** = `RAM_total − 0.5 GB − modelo`. + +--- + +## 4. Recommendations by RAM + +### 4 GB VPS (bare-bones dev) + +Solo viable con modelo chico y contexto bajo. + +| Model | context_size | KV cache | RAM usada | +|-------------------|-------------:|---------:|----------:| +| gemma-3-1b Q4 | 8192 | ~290 MB | ~1.8 GB | +| qwen2.5-1.5b Q4 | 4096 | ~72 MB | ~1.6 GB | + +### 8 GB VPS (típico) + +| Model | context_size | KV cache | RAM usada | Veredicto | +|-------------------|-------------:|---------:|----------:|---------------------| +| qwen2.5-1.5b Q4 | 16384 | ~290 MB | ~2.0 GB | muy cómodo | +| qwen2.5-3b Q4 | 8192 | ~580 MB | ~3.0 GB | **sweet spot** | +| qwen2.5-3b Q4 | 16384 | ~1.1 GB | ~3.6 GB | **recomendado** | +| qwen2.5-3b Q4 | 32768 | ~2.3 GB | ~4.8 GB | máximo útil | +| gemma-3-1b Q4 | 32768 | ~1.1 GB | ~2.4 GB | **recomendado** | +| gemma-3-4b Q4 | 8192 | ~2.2 GB | ~5.2 GB | ajustado | + +### 16 GB VPS + +| Model | context_size | KV cache | RAM usada | +|-------------------|-------------:|---------:|----------:| +| qwen2.5-3b Q4 | 32768 | ~2.3 GB | ~5.0 GB | +| gemma-3-4b Q4 | 16384 | ~4.5 GB | ~7.5 GB | + +--- + +## 5. Worked examples + +### qwen2.5-3b en 8 GB + +```yaml +# configs/portfolio-bot.yaml +providers: + - name: llamacpp-local + type: llamacpp + model: qwen2.5-3b-instruct + endpoint: http://localhost:9100/v1 + context_size: 16384 # ~1.1 GB KV, deja 4 GB libres + max_tokens: 1024 # respuestas moderadas +``` + +```bash +llama-server \ + -m qwen2.5-3b-instruct-q4_k_m.gguf \ + --ctx-size 16384 \ + -ngl 0 -t 2 \ + --mlock +``` + +### gemma-3-1b en 8 GB + +```yaml +providers: + - name: llamacpp-local + type: llamacpp + model: gemma-3-1b-it + endpoint: http://localhost:9100/v1 + context_size: 32768 # sobra RAM, contexto largo + max_tokens: 1024 +``` + +```bash +llama-server \ + -m gemma-3-1b-it-Q4_K_M.gguf \ + --ctx-size 32768 \ + -ngl 0 -t 2 \ + --mlock +``` + +### Gemma con chat template custom (sin system role líder) + +Gemma 3 rechaza mensajes `system` antes del primer `user`. Dos opciones: + +**Opción A** — template custom en `~/.llama/gemma3.jinja`: + +```jinja +{% if messages[0]['role'] != 'system' and messages[0]['role'] != 'user' %} +{{ raise_exception('First message must be system or user') }} +{% endif %} +{% for message in messages %} +{% if message['role'] == 'system' %} +{{ message['content'] | trim + '\n\n' -}} +{% elif message['role'] == 'user' %} +{{- 'user\n' + message['content'] | trim + '\n' -}} +{% elif message['role'] == 'assistant' or message['role'] == 'model' %} +{{- 'model\n' + message['content'] | trim + '\n' -}} +{% endif %} +{% endfor %} +{% if add_generation_prompt %} +{{- 'model\n' -}} +{% endif %} +``` + +```bash +llama-server \ + -m gemma-3-1b-it-Q4_K_M.gguf \ + --ctx-size 32768 \ + -ngl 0 -t 2 \ + --mlock \ + --chat-template-file ~/.llama/gemma3.jinja +``` + +**Opción B** — usar Ollama, que mapea system → prefix del primer user +automáticamente. + +--- + +## 6. Tuning with auto-compaction + +The bot has built-in auto-compaction (`configs/portfolio-bot.yaml` → +`compaction:` block). When the previous turn's input tokens exceed +`threshold_ratio × MaxContextWindow`, the older portion of the chat gets +summarized into a single system note. This means: + +- A **smaller `context_size`** still works for long conversations — the + compactor frees up room by folding old turns. +- **Bigger `max_tokens`** means the compactor fires sooner (less budget + left for input). +- Default `threshold_ratio: 0.75` triggers compaction at ~75% of the + window. Lower it (e.g. `0.5`) for headroom on slow CPU where each + request is expensive; raise it (e.g. `0.9`) when you want to keep + more verbatim history. + +For a portfolio bot with `context_size: 16384` and `max_tokens: 1024`, +compaction fires when input exceeds ~12k tokens — leaving ~5k for the +fresh history, which is ~10-15 recent user turns. More than enough. + +--- + +## 7. Tips + +1. **`--mlock`** es oro en VPS. Bloquea el modelo en RAM y evita swaps + cuando hay picos de memoria. Cuesta ~modelo_size de locked RAM. +2. **Monitoreá con `htop`** o `free -h` la primera semana. Si ves swap, + bajá el context. +3. **Más contexto ≠ más rápido.** El prefill (procesar el input) escala + lineal con la cantidad de tokens. Generación (output) no se ve + afectada. Con `--ctx-size 32768` y un input de 500 tokens, el TTFT + apenas cambia; con 20k tokens de input sí. +4. **Streams concurrentes.** Cada stream activo reserva su propio KV + cache. En 8 GB no hagas más de 1-2 streams simultáneos — el rate + limiter del bot (default 30 req/min) ya te protege. +5. **`max_tokens` bajo ayuda.** 512 es suficiente para Q&A. Bajarlo + deja más presupuesto para input y retrasa la compactación. + +--- + +## 8. Quick-pick table + +Copy-paste según tu setup: + +| Setup | `context_size` | `max_tokens` | +|-----------------------------|---------------:|-------------:| +| 4 GB + gemma-3-1b | 8192 | 512 | +| 4 GB + qwen2.5-1.5b | 4096 | 512 | +| 8 GB + qwen2.5-1.5b | 16384 | 768 | +| 8 GB + qwen2.5-3b | 16384 | 1024 | +| 8 GB + gemma-3-1b | 32768 | 1024 | +| 16 GB + qwen2.5-3b | 32768 | 1024 | +| 16 GB + gemma-3-4b | 16384 | 1024 |