rony-chat-bot/docs/architecture.md
Victor Hugo Vargas ab510d8b31 docs: document auto-compaction behavior in architecture guide
Adds the §4.5 Auto-compaction section to architecture.md and
architecture.es.md describing the trigger, fallback, persistence and
the new SSE 'compaction' event so consumers know how to react.

Includes the three placeholder projects used while exercising the
feature end to end (bot-onboarding, dashboard-metricas, tienda-ropa)
so /api/reindex picks them up without further setup.
2026-07-18 00:08:37 -07:00

1398 lines
No EOL
45 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 📋 Rony Chat Bot — Technical Design Document
**Version:** 1.0
**Author:** Victor Hugo Vargas
**Date:** 2026-06-28
**Status:** Complete specification for implementation
**Path:** `rony-chat-bot/docs/architecture.md`
> 🌐 **Language:** [English](./architecture.md) | [Español](./architecture.es.md)
>
> 📚 **Workspace:** This project is part of the `Rony/` workspace. See [`../README.md`](../../README.md).
>
> 🔑 **Depends on:** [`rony-llm-agent`](https://github.com/VictorVargas/rony-llm-agent) — core library that provides agent loop, LLM clients, RAG, persona system.
>
> 📐 **Methodology:** This project follows the **SDD + DDD + Hexagonal Architecture** approach. Functional Requirements are numbered as `CRF-XXX`. See [`../../METHODOLOGY.md`](../../METHODOLOGY.md).
---
## 🎯 1. Project Vision
### 1.1 What is Chat-Bot?
An **HTTP chatbot** that answers questions about Victor Hugo Vargas and his projects. Uses **RAG (Retrieval-Augmented Generation)** over markdown files describing each project, and a local LLM (or cloud) to generate responses.
### 1.2 Primary use case
Victor has a portfolio website (Astro + React). On the site there's a chat widget where visitors can ask:
- "What projects has Victor done?"
- "What's his experience with Go?"
- "How does Rony Harness work?"
- "Has Victor worked with PostgreSQL?"
The bot responds with accurate information extracted from the projects' markdown files + bio + skills.
### 1.3 Secondary use cases (future)
- **Client adaptation:** The same bot, with other data and another persona, serves car dealerships, restaurants, etc.
- **Standalone CLI:** `./chat-bot ask "what do you know about X?"` for terminal use.
- **Slack/Discord bot:** Wrapper that consumes the HTTP API.
### 1.4 Philosophy
- **Self-hosted by default** — works 100% local with Ollama + 1-3B models
- **Cloud optional** — if you need more quality, swap to Anthropic API
- **Portable** — easy to fork/customize for other contexts
- **Streaming** — token-by-token responses with SSE (no waiting for complete response)
- **Reuses `rony-llm-agent`** — doesn't reinvent the agent loop
---
## 🏗️ 2. Architecture
### 2.1 Overview
```
┌─────────────────────────────────────────────────────────────────┐
│ Browser (Astro site) │
│ ↓ HTTP POST /api/chat │
│ Astro SSR (proxy) ←────────── Serves portfolio + proxy chat │
│ ↓ HTTP POST /api/chat │
│ Chat-Bot HTTP server (:7331) │
│ ↓ │
│ Agent loop (rony-llm-agent) │
│ ↓ │
│ RAG retrieval → SQLite FTS5 over data/projects/*.md │
│ ↓ │
│ LLM (llama.cpp local default / Ollama or Anthropic optional) │
└─────────────────────────────────────────────────────────────────┘
```
### 2.2 Main components
| Component | Path | Responsibility |
|---|---|---|
| **HTTP server** | `internal/server/` | Gin/chi handlers, SSE streaming |
| **Agent runner** | `internal/agent/` | Wrapper over `rony-llm-agent` with specific config |
| **Portfolio loader** | `internal/portfolio/` | Reads `data/projects/*.md`, indexes in SQLite FTS5 |
| **Persona** | `internal/persona/` | Loads persona from `configs/portfolio-bot.yaml` |
| **CLI** | `cmd/chat-bot/` | Commands: `serve`, `reindex`, `ask`, `version` |
### 2.3 Tech stack
| Layer | Technology | Reason |
|---|---|---|
| **Language** | Go 1.26+ | Same as rony-harness, leverage `os.Root`, `iter.Seq` |
| **HTTP router** | `net/http` + `chi` | Stdlib + chi for middleware (CORS, logging) |
| **SSE** | `net/http` Flusher | Stdlib is enough, no external library needed |
| **Config** | `gopkg.in/yaml.v3` | Same as harness |
| **RAG backend** | SQLite + FTS5 (BM25) | Zero external deps, single file, fast |
| **LLM** | llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama as alt | Self-hosted by default |
| **Tests** | stdlib + testify | Consistency with the rest |
---
## 🔌 3. HTTP API
### 3.1 Endpoints
#### `POST /api/chat` — Chat with SSE streaming
**Request:**
```json
{
"messages": [
{"role": "user", "content": "What projects does Victor have?"}
],
"stream": true,
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e"
}
```
| Field | Required | Notes |
|---|---|---|
| `messages` | yes | At least one user message; alternation is not enforced. |
| `stream` | no, default `true` | `false` returns a single JSON body instead of SSE. |
| `conversation_id` | no | Hex string. If omitted, the server mints a new one and returns it (see below). Pass an existing ID to keep the thread. |
**Response (SSE):**
```
data: {"type":"start","conversation_id":"57f4aa3c7fab466bc4de9c43b296903e"}
data: {"type":"chunk","content":"Victor"}
data: {"type":"chunk","content":" has"}
data: {"type":"chunk","content":" several"}
data: {"type":"chunk","content":" projects"}
data: {"type":"sources","documents":["rony-harness.md","rony-llm-agent.md"]}
data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}
```
The `conversation_id` in the `start` event is what the client should store
(see §3.4 — *Conversation persistence*). When the client passed an
existing ID the server echoes it back; otherwise it's freshly minted.
**Without streaming** (`"stream": false`):
```json
{
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e",
"content": "Victor has several projects...",
"sources": ["rony-harness.md", "rony-llm-agent.md"],
"usage": {"input_tokens": 245, "output_tokens": 38}
}
```
#### `GET /api/conversations` — List recent conversations
Returns the most recent conversation summaries, newest first. Useful for a
"show my chats" sidebar in a custom UI.
**Query params:**
- `limit` (1200, default 50)
**Response:**
```json
{
"count": 2,
"conversations": [
{
"id": "57f4aa3c7fab466bc4de9c43b296903e",
"created_at": "2026-07-17T05:02:07Z",
"updated_at": "2026-07-17T05:04:31Z",
"preview": "What projects does Victor have?"
}
]
}
```
#### `GET /api/conversations/{id}` — Fetch one conversation
Returns the full history of a conversation with all messages in
chronological order.
**Response (200):**
```json
{
"id": "57f4aa3c7fab466bc4de9c43b296903e",
"created_at": "2026-07-17T05:02:07Z",
"updated_at": "2026-07-17T05:04:31Z",
"messages": [
{"id": 1, "role": "user", "content": "What projects does Victor have?", "created_at": "..."},
{"id": 2, "role": "assistant", "content": "Victor has several projects...", "sources": ["..."], "created_at": "..."}
]
}
```
**Response (404):** when the ID is unknown (e.g. server DB was wiped or the
client lost sync). The widget treats this as "start fresh".
> ⚠️ **Auth note:** the conversation ID is the only access token. For a
> public bot this is fine; for private contexts add auth at the proxy layer
> (e.g. require a session cookie before forwarding to this endpoint).
#### `DELETE /api/conversations/{id}` — Delete a conversation
Removes the conversation and all its messages (cascade). Returns 204 on
success, 404 if the ID doesn't exist.
#### `POST /api/reindex` — Re-index portfolio
Useful when files in `data/projects/` are modified.
**Request:** empty
**Response:**
```json
{
"indexed_files": 12,
"total_chunks": 87,
"duration_ms": 4321
}
```
#### `GET /api/health` — Health check (real)
Probes the LLM provider and the SQLite store in parallel and returns their
states. Designed for monitoring/load balancers. **Returns 200 when healthy
or degraded, 503 when unhealthy.**
- `?deep=true` adds a chunk count to the store probe (same latency budget).
**Status taxonomy:**
| `status` | HTTP | Meaning |
|---|---|---|
| `healthy` | 200 | LLM up, store up |
| `degraded` | 200 | LLM up, store down — bot still answers, just without RAG |
| `unhealthy` | 503 | LLM down — bot cannot answer, no point routing traffic here |
**Probe details:**
| Component | Probe | Latency |
|---|---|---|
| `llm` | `GET {provider}/health` (llamacpp, ollama) or `/models` (openai) | ~1ms for local llama-server |
| `store` | `SELECT 1` on the SQLite handle | ~100µs |
Each probe has a 2s timeout; the whole call returns within ~2.5s even if a
dependency hangs.
**Response shape (healthy):**
```json
{
"status": "healthy",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {
"status": "up",
"latency": "1.028ms",
"details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
},
"store": {
"status": "up",
"latency": "107µs"
}
}
}
```
**Response shape (degraded, with `?deep=true`):**
```json
{
"status": "degraded",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {"status": "up", "latency": "0.8ms", "details": {...}},
"store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
}
}
```
**Response shape (unhealthy):** HTTP 503, same JSON with `"status": "unhealthy"` and the failed component reporting `"status": "down"` plus an `error` field.
#### `GET /api/info` — Bot metadata
```json
{
"name": "Rony Chat Bot",
"model": "qwen2.5:1.5b",
"persona": "...",
"topics": ["projects", "experience", "technical skills"]
}
```
### 3.2 SSE Implementation
```go
// internal/server/chat.go
package server
import (
"encoding/json"
"fmt"
"net/http"
"github.com/VictorVargas/rony-llm-agent/pkg/agent"
)
func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
// SSE headers
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
w.Header().Set("Connection", "keep-alive")
w.Header().Set("X-Accel-Buffering", "no")
flusher, ok := w.(http.Flusher)
if !ok {
http.Error(w, "SSE not supported", http.StatusInternalServerError)
return
}
// Parse request
var req ChatRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
writeError(w, flusher, "invalid request", err)
return
}
// Start event
writeSSE(w, flusher, "start", map[string]string{
"conversation_id": generateConvID(),
})
// Run agent with streaming
sources := []string{}
for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
if err != nil {
writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
return
}
if chunk.Type == "source" {
sources = append(sources, chunk.Source)
}
writeSSE(w, flusher, chunk.Type, chunk.Data)
}
// Done event
writeSSE(w, flusher, "done", map[string]any{
"usage": map[string]int{
"input_tokens": 245,
"output_tokens": 38,
},
})
}
func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
payload, _ := json.Marshal(data)
fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
flusher.Flush()
}
```
### 3.3 Middleware
```go
// internal/server/middleware.go
package server
func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
// Wrap response writer to capture status
rw := &statusRecorder{ResponseWriter: w, status: 200}
next.ServeHTTP(rw, r)
slog.Info("http.request",
"method", r.Method,
"path", r.URL.Path,
"status", rw.status,
"duration_ms", time.Since(start).Milliseconds(),
"ip", r.RemoteAddr,
)
})
}
func (s *Server) corsMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
origin := r.Header.Get("Origin")
for _, allowed := range s.config.Server.CORSOrigins {
if origin == allowed {
w.Header().Set("Access-Control-Allow-Origin", origin)
w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
break
}
}
if r.Method == "OPTIONS" {
w.WriteHeader(204)
return
}
next.ServeHTTP(w, r)
})
}
func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if !limiter.Allow() {
http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
return
}
next.ServeHTTP(w, r)
})
}
```
### 3.4 Conversation persistence
The bot persists conversation threads in the same SQLite database as the
RAG index (`./data/portfolio.db`). Schema lives in `internal/portfolio/conversations.go`.
```sql
CREATE TABLE conversations (
id TEXT PRIMARY KEY, -- 16-byte random hex (32 chars)
created_at INTEGER NOT NULL,
updated_at INTEGER NOT NULL
);
CREATE TABLE messages (
id INTEGER PRIMARY KEY AUTOINCREMENT,
conversation_id TEXT NOT NULL,
role TEXT NOT NULL, -- user | assistant | system
content TEXT NOT NULL,
sources TEXT, -- JSON array, nullable
created_at INTEGER NOT NULL,
FOREIGN KEY (conversation_id) REFERENCES conversations(id) ON DELETE CASCADE
);
CREATE INDEX idx_messages_conv ON messages(conversation_id, id);
```
**Lifecycle:**
| When | What |
|---|---|
| `POST /api/chat` (no `conversation_id`) | Server mints a new hex ID, returns it in the `start` SSE event (or `conversation_id` field of the JSON response) |
| `POST /api/chat` (with `conversation_id`) | Server reuses the existing row; both user message and assistant reply are appended |
| User message | Persisted **before** the LLM runs, so it survives a model failure |
| Assistant message | Persisted **after** the stream completes, with the RAG sources attached |
| `GET /api/conversations/{id}` | Returns the full thread; 404 if unknown |
| `DELETE /api/conversations/{id}` | Cascade-deletes messages |
**Client responsibilities:**
1. On the first message, omit `conversation_id`. Capture the one the server
returns in the `start` SSE event.
2. Store it client-side (`localStorage["rony-chat-conv"]` in the widget).
3. On every subsequent message, send the ID back.
4. On page load, if you have a stored ID, call `GET /api/conversations/{id}`
to restore the thread. If 404, clear the stored ID and start fresh.
The widget (`web/chat-widget.js`) implements all four steps. Any other
client (a custom React component, an Astro endpoint, a CLI replay tool)
follows the same protocol.
**Auth model:**
The conversation ID is the only access token for `GET /api/conversations/{id}`.
It is 128 bits of random entropy, so guessing one is infeasible. For a
public portfolio bot this is the right trade-off — anyone who knows the
URL can read its history. For private contexts, add an auth layer in front
of the bot (proxy) that gates the conversation endpoints.
---
## 🧠 4. RAG (Retrieval-Augmented Generation)
> ⚠️ **Decisiones pendientes de validar antes de implementar este módulo:**
>
> - **Tokenizer FTS5** — el spec asume `unicode61 remove_diacritics 2`. Confirmar con datos reales si conviene cambiar a `porter` (stemming EN), `trigram` (sub-string matching) o un tokenizer custom para español. **Validar:** ejecutar queries representativas contra `data/projects/` y comparar recall antes de cerrar esta elección.
> - **Driver SQLite** — ✅ **DECIDIDO: `modernc.org/sqlite`** (puro Go, sin CGO). Ver benchmark abajo.
> - **Chunking** — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. **Validar:** medir recall con chunks por sección markdown (split por `#`/`##`) vs por tamaño.
> - **Sin similitud semántica** — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. **Validar:** tamaño del corpus y types of questions esperadas; si el corpus crece o las queries se vuelven abstractas, considerar agregar embeddings como capa secundaria.
### 4.0 Driver decision: benchmark results
Reproducible con `CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/`. Datos: 4 markdowns → 11 chunks.
| Operación | mattn (CGO) | modernc (puro Go) | Diferencia |
|---|---|---|---|
| **Insert** (11 chunks) | 2,802,843 ns/op | **1,465,646 ns/op** | modernc 1.9× más rápido |
| Insert alloc | 2,124,299 B/op | **9,770 B/op** | modernc usa 217× menos memoria |
| **Query** (8 queries BM25) | **244,047 ns/op** | 555,162 ns/op | mattn 2.3× más rápido |
| **Round-trip** (insert + 8 queries) | 3,543,417 ns/op | **2,267,669 ns/op** | modernc 1.6× más rápido |
| Binary size | 11 MB | 11 MB | igual |
| Build deps | gcc, CGO=1 | nada | modernc gana |
| CI/CD portable | requiere toolchain C | `go build` puro | modernc gana |
**Decisión: `modernc.org/sqlite`**.
Justificación:
1. Ambas latencias de query (~250µs vs ~550µs) son **2 órdenes de magnitud por debajo** del target de 50ms — imperceptible vs el LLM (varios segundos).
2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.
### 4.1 Indexing pipeline
```
data/projects/*.md
↓ (read all files)
Raw markdown content
↓ (split into chunks, ~500 chars, 50 overlap)
Chunks []
↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
Indexed corpus
```
**When it runs:**
- On bot startup (if `--reindex-on-start` flag)
- Manually: `./chat-bot reindex`
- Via HTTP: `POST /api/reindex`
### 4.2 Retrieval pipeline
```
User query "what projects does Victor have?"
↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
Top 5 relevant chunks
↓ (format as context block)
System prompt += relevant chunks
↓ (send to LLM)
LLM generates answer
```
### 4.3 Implementation
```go
// internal/portfolio/indexer.go
package portfolio
import (
"context"
"database/sql"
"fmt"
"log/slog"
"os"
"path/filepath"
"strings"
)
type Indexer struct {
dataPath string
db *sql.DB
chunkSize int
chunkOverlap int
}
func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
if err != nil {
return 0, err
}
// Rebuild FTS5 index from scratch (delete + insert is faster than diff for small corpora)
if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
return 0, fmt.Errorf("clear index: %w", err)
}
totalChunks := 0
for _, file := range files {
chunks, err := i.indexFile(ctx, file)
if err != nil {
slog.Warn("failed to index file", "file", file, "err", err)
continue
}
totalChunks += chunks
}
return totalChunks, nil
}
func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
content, err := os.ReadFile(path)
if err != nil {
return 0, err
}
projectID := strings.TrimSuffix(filepath.Base(path), ".md")
chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)
tx, err := i.db.BeginTx(ctx, nil)
if err != nil {
return 0, err
}
defer tx.Rollback()
stmt, err := tx.PrepareContext(ctx, `
INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
VALUES (?, ?, ?, ?, ?)
`)
if err != nil {
return 0, err
}
defer stmt.Close()
for idx, chunk := range chunks {
id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
return idx, err
}
}
if err := tx.Commit(); err != nil {
return 0, err
}
return len(chunks), nil
}
// schema.go — applied at startup
const schema = `
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
id UNINDEXED,
project_id UNINDEXED,
source_file UNINDEXED,
chunk_index UNINDEXED,
content,
tokenize = 'unicode61 remove_diacritics 2'
);
`
func splitIntoChunks(text string, size, overlap int) []string {
// Simple implementation: split by size with overlap
// Production version uses tokenizer-aware chunking
var chunks []string
for i := 0; i < len(text); i += size - overlap {
end := i + size
if end > len(text) {
end = len(text)
}
chunks = append(chunks, text[i:end])
}
return chunks
}
```
### 4.4 Retrieval in the agent loop
```go
// internal/portfolio/search.go
package portfolio
type Hit struct {
ProjectID string
SourceFile string
ChunkIndex int
Content string
Score float64 // BM25 score from FTS5
}
func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
// Escape user input: FTS5 syntax can break with special chars
ftsQuery := sanitizeFTS5(query)
rows, err := s.db.QueryContext(ctx, `
SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
FROM portfolio_chunks
WHERE portfolio_chunks MATCH ?
ORDER BY score
LIMIT ?
`, ftsQuery, topK)
if err != nil {
return nil, err
}
defer rows.Close()
var hits []Hit
for rows.Next() {
var h Hit
if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
return nil, err
}
hits = append(hits, h)
}
return hits, rows.Err()
}
// sanitizeFTS5 wraps the user query so reserved chars and unquoted strings don't crash FTS5.
// A pragmatic choice for a Q&A bot: append prefix-match wildcard to each token.
func sanitizeFTS5(q string) string {
tokens := strings.FieldsFunc(q, func(r rune) bool {
return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
(r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
r > 0x7F) // keep accented chars
})
if len(tokens) == 0 {
return `""`
}
for i, t := range tokens {
tokens[i] = `"` + strings.ToLower(t) + `"*`
}
return strings.Join(tokens, " ")
}
```
```go
// internal/agent/runner.go
package agent
func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
basePrompt := r.persona.SystemPrompt
hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
if err != nil {
return "", err
}
if len(hits) == 0 {
return basePrompt, nil
}
var contextBlock strings.Builder
contextBlock.WriteString(basePrompt)
contextBlock.WriteString("\n\n## Relevant context\n\n")
for _, h := range hits {
contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
h.SourceFile, h.Content))
}
return contextBlock.String(), nil
}
func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
return func(yield func(Chunk, error) bool) {
lastUserMsg := getLastUserMessage(messages)
systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
if err != nil {
yield(Chunk{}, err)
return
}
messages = prependSystem(messages, systemPrompt)
for chunk, err := range r.loop.RunStream(ctx, messages) {
if !yield(chunk, err) {
return
}
}
}
}
```
**Why this is simpler than embeddings:**
- No embedding model to download or run (saves ~270MB of RAM and ~200ms per query)
- One file (`data/portfolio.db`), one driver, no extra process
- BM25 ranking is excellent for keyword-based retrieval over structured docs like project READMEs
- Trade-off: no semantic similarity ("projects about AI" won't match "machine learning" without the literal words). Mitigation: `trigram` tokenizer handles morphology well for English/Spanish.
---
## 🗜️ 4.5 Auto-compaction
Long conversations eventually run out of context — at qwen2.5-3b's 4k window, the ~3k-token system prompt + RAG block leaves only room for 23 user turns. Auto-compaction solves this by folding the older portion of the conversation into a single summary system message when the previous turn's input tokens cross a configurable threshold.
### When it fires
`agent.Runner.Compact` runs once per `/api/chat` request, before the RAG search. It compares the runner's most recent `Usage.InputTokens` (reported by the provider in the previous streamed chunk) against `client.Capabilities().MaxContextWindow × threshold_ratio`.
| Setting | Default | What it controls |
|---|---|---|
| `compaction.enabled` | `false` | Master switch. |
| `compaction.threshold_ratio` | `0.75` | Trigger when used tokens ≥ window × ratio. |
| `compaction.keep_recent_turns` | `4` | How many of the latest user turns are kept verbatim after compaction. |
| `compaction.summary_system_prompt` | *(built-in bilingual)* | Override the instruction sent to the LLM when summarizing. |
Short-circuits silently when compaction is disabled, the provider doesn't report a window (`Capabilities().MaxContextWindow == 0`), the history is shorter than `keep_recent_turns`, or usage is still unknown (first turn).
### How the summary is made
1. `splitByTurns(history, keep_recent_turns)` divides messages into `(older, recent)` on user-role boundaries so a kept turn's user/assistant pair always stays together.
2. `renderTranscript(older)` flattens older messages into a `User:` / `Assistant:` transcript (skipping tool messages and empty assistant placeholders).
3. The runner calls `client.Generate(...)` with the summary prompt + transcript and a 512-token cap so the compaction step itself stays cheap.
4. The returned text is prepended as a system message (`"Earlier conversation summary:\n…"`), followed by the recent tail.
5. `LastCompaction()` returns `CompactionStats` so the SSE handler can emit a `compaction` event right before the streamed chunks.
### Failure mode
If `Generate` errors or returns an empty summary, compaction falls back to `truncateToBudget`: drop oldest user-turns one at a time until the remaining slice fits `threshold` tokens (heuristic: `len(s) / 4 + 1`). The current user turn is always preserved. The fallback is logged at WARN and the request still proceeds — a flaky summarize call never fails the user's request.
### Wire protocol
Streaming responses gain an optional `compaction` event:
```
data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500}
```
Emitted after `start` (when applicable) and before `sources` / `chunk`. The widget can render this as a subtle "Context compacted" hint or ignore it — both are valid.
### Persistence
Compaction is **per-request**. The full transcript is still saved to `messages` in `data/portfolio.db` verbatim, so `GET /api/conversations/{id}` always returns the original history. Only what we send to the LLM is reduced — the next session can re-read the full thread from the DB.
---
## 🌐 5. Embedding the widget
The bot ships with a drop-in vanilla-JS widget. Add two files to your site and it works.
### 5.1 The widget (any site)
```html
<link rel="stylesheet" href="/path/to/chat-widget.css">
<script src="/path/to/chat-widget.js"
data-api-url="https://chat.example.com"
data-title="Ask me anything"
data-greeting="Hi! Ask me about the projects."
data-position="bottom-right"
data-theme="auto"
defer></script>
```
A bubble appears bottom-right, opens a panel, talks SSE to `/api/chat`, streams the response, and cites sources. No build step, no React/Vue, no framework lock-in.
**Browser→bot options:**
| Topology | Trade-offs |
|---|---|
| **Direct** (browser → bot, same domain or CORS) | Simplest. Add the bot's origin to `cors_origins` in YAML. |
| **Reverse proxy** (nginx/Caddy in front) | Bot stays on private network, single public domain, no CORS to manage. |
| **Site proxies the bot** (Astro/Next API route) | Adds a hop and a bit of code, but gives you auth/session hooks in your site. |
The widget works the same in all three. Pick the topology that matches your infra.
> **Default dev setup is direct + CORS.** `cors_origins` in `configs/portfolio-bot.yaml` controls which sites can call the bot. Add your site's origin there.
### 5.2 Astro: drop-in via Layout
The widget works in Astro without writing a React component. Add this to your shared layout:
```astro
---
// src/layouts/BaseLayout.astro
import "../path/to/chat-widget.css";
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
---
<html>
<body>
<slot />
<script src="/path/to/chat-widget.js"
data-api-url={apiUrl}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
defer is:inline></script>
</body>
</html>
```
`is:inline` keeps Astro from hashing/transforming the script tag, so the `data-*` attributes survive.
### 5.3 React / Next.js: same script tag
```tsx
// app/layout.tsx
import Script from "next/script";
export default function RootLayout({ children }) {
return (
<html>
<head>
<link rel="stylesheet" href="/chat-widget.css" />
<Script src="/chat-widget.js"
data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
strategy="afterInteractive" />
</head>
<body>{children}</body>
</html>
);
}
```
### 5.4 If you want a server proxy (Astro/Next API route)
The widget can also call a same-origin endpoint that forwards to the bot. This is the right call when you need:
- Auth on `/api/chat` (logged-in users only)
- Centralized rate limiting at the site level
- Hiding the bot's origin from the browser
```typescript
// src/pages/api/chat.ts (Astro) or app/api/chat/route.ts (Next)
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";
export const POST = async ({ request }) => {
const body = await request.json();
// (optional) auth check, rate limit, session lookup here
const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
});
return new Response(resp.body, {
status: resp.status,
headers: {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"Connection": "keep-alive",
},
});
};
```
Then point the widget at `/api/chat` (same origin) instead of the bot's URL.
### 5.5 Widget configuration reference
All options are `data-*` attributes on the `<script>` tag:
| Attribute | Default | Notes |
|---|---|---|
| `data-api-url` | *(required)* | Base URL of the bot. No trailing slash. |
| `data-title` | `"Chat"` | Header text. |
| `data-greeting` | `""` | First assistant message when the panel opens. |
| `data-position` | `"bottom-right"` | `"bottom-right"` or `"bottom-left"`. |
| `data-theme` | `"auto"` | `"auto"` (follows OS), `"light"`, `"dark"`. |
Theming is via CSS custom properties on `.rony-chat-widget-root` (see `web/chat-widget.css`):
```css
.rony-chat-widget-root {
--rony-accent: #ff6b35;
--rony-radius: 4px;
--rony-font: "Inter", sans-serif;
}
```
### 5.6 What the widget doesn't do (yet)
- **Richer markdown** (tables, images) — the built-in renderer handles the common cases; for full CommonMark, swap `renderMarkdown` in `chat-widget.js` for `marked` or `markdown-it`.
- **Mobile swipe-to-dismiss** — panel goes full-screen on phones.
- **Conversation history sidebar** — only the active conversation is shown (the backend exposes `GET /api/conversations` for a future sidebar).
---
## 🤖 6. Self-hosting with llama.cpp (default)
### 6.1 Setup
llama-server is a separate process that the bot connects to over HTTP. **Both ports (the bot's and llama-server's) are configurable** — pick what fits your environment.
```bash
# 1. Make sure you have a GGUF model available
# Download from Hugging Face, e.g.:
# https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
export RONY_MODELS_PATH=/path/to/models
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf
# 2. Start llama-server (port is configurable; default llama.cpp is 8080)
llama-server \
-m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
--port 9100 \
--host 127.0.0.1 \
--ctx-size 4096 \
--mlock # prevents swap, critical on shared VPS
# 3. Make sure configs/portfolio-bot.yaml points to the same port
# providers[0].endpoint: http://localhost:9100/v1
# 4. Start the bot (default port 7331, also configurable)
./bin/chat-bot serve
# → Serves on http://localhost:7331
# → Override with: ./bin/chat-bot serve --port 9101 --host 127.0.0.1
```
**Port reference:**
| What | Default | How to change |
|---|---|---|
| `llama-server` HTTP port | 8080 (llama.cpp convention) | `--port N` flag when starting `llama-server` |
| chat-bot HTTP port | 7331 | `--port N` flag on `serve`, or `server.port` in YAML |
| chat-bot → llama-server URL | `http://localhost:8080/v1` | `endpoint` field on the provider in YAML |
The `llamacpp` provider is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
### 6.2 Alternative: Ollama (easier for development)
If you don't want to manage GGUF files manually, Ollama provides the same models with a simpler workflow:
```bash
# 1. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# 2. Download chat model
ollama pull qwen2.5:1.5b
# 3. Verify
ollama list
# 4. Edit configs/portfolio-bot.yaml to mark ollama-local as default:
# providers[0].default: true (and remove default from llamacpp-local)
# Ollama exposes an OpenAI-compatible API on :11434/v1
# 5. Start the bot
ollama serve &
./bin/chat-bot serve
```
### 6.3 Alternative: llama.cpp direct (advanced)
For more control or if Ollama doesn't work in your setup:
```yaml
providers:
- name: llamacpp-local
type: llamacpp
model: qwen2.5-3b-instruct
endpoint: http://localhost:9100/v1 # configurable, see §6.1
context_size: 4096
max_tokens: 2048
default: true
```
The `llamacpp` adapter is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
---
## 📦 7. Bot CLI
### 7.1 Commands
```bash
# Start HTTP server
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]
# Re-index portfolio (reads data/projects/*.md → SQLite FTS5)
chat-bot reindex
# Single question (no server, useful for tests)
chat-bot ask "What projects does Victor have?" [--no-rag]
# Validate config
chat-bot config validate
# Health check (useful for monitoring)
chat-bot health
# Version
chat-bot version
```
### 7.2 Implementation with Cobra
```go
// cmd/chat-bot/main.go
package main
import (
"github.com/spf13/cobra"
)
func main() {
root := &cobra.Command{
Use: "chat-bot",
Short: "Portfolio chatbot HTTP server",
}
root.AddCommand(serveCmd())
root.AddCommand(reindexCmd())
root.AddCommand(askCmd())
root.AddCommand(configCmd())
root.AddCommand(healthCmd())
root.AddCommand(versionCmd())
if err := root.Execute(); err != nil {
os.Exit(1)
}
}
func serveCmd() *cobra.Command {
var port int
var host string
var reindexOnStart bool
cmd := &cobra.Command{
Use: "serve",
Short: "Start HTTP server",
RunE: func(cmd *cobra.Command, args []string) error {
return server.Serve(server.Config{
Port: port,
Host: host,
ReindexOnStart: reindexOnStart,
})
},
}
cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
return cmd
}
```
---
## 🚀 8. Deployment
### 8.1 Recommendation: Self-hosted on VPS
```bash
# 1. Install dependencies
sudo apt install golang-go ollama
ollama pull qwen2.5:1.5b
# 2. Build
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot
# 3. systemd service
cat > /etc/systemd/system/chat-bot.service <<EOF
[Unit]
Description=Portfolio Chat Bot
After=network.target ollama.service
[Service]
Type=simple
User=chatbot
WorkingDirectory=/opt/chat-bot
ExecStart=/usr/local/bin/chat-bot serve
Restart=on-failure
Environment=RONY_MODELS_PATH=/opt/models
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl enable --now chat-bot
```
### 8.2 Reverse proxy (Caddy)
```
# /etc/caddy/Caddyfile
chat.victorvargas.dev {
reverse_proxy localhost:7331
}
```
### 8.3 Monitoring
```bash
# Health check periodic
curl -s http://localhost:7331/api/health | jq
# Logs
journalctl -u chat-bot -f
```
---
## 🧪 9. Testing
### 9.1 Unit tests
```go
// internal/server/chat_test.go
package server
func TestHandleChat_ValidRequest(t *testing.T) {
s := newTestServer(t)
req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
"messages": [{"role": "user", "content": "hello"}]
}`))
req.Header.Set("Content-Type", "application/json")
w := httptest.NewRecorder()
s.handleChat(w, req)
assert.Equal(t, 200, w.Code)
assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
}
func TestHandleChat_RateLimit(t *testing.T) {
s := newTestServerWithConfig(t, server.Config{
RateLimit: 1, // 1 request per minute
})
// First request OK
req1 := newChatRequest("hello")
w1 := httptest.NewRecorder()
s.handleChat(w1, req1)
assert.Equal(t, 200, w1.Code)
// Second request denied
req2 := newChatRequest("hello again")
w2 := httptest.NewRecorder()
s.handleChat(w2, req2)
assert.Equal(t, 429, w2.Code)
}
```
### 9.2 Integration tests with mock LLM
```go
// internal/agent/runner_test.go
func TestRunner_RAGContextIsInjected(t *testing.T) {
mockLLM := mock.New(mock.Responses{
{Match: "projects", Response: "Victor has several projects..."},
})
memory := newMockMemoryWithDocs(t, []rag.Fragment{
{Content: "Rony Harness: AI agent harness...", ProjectID: "rony-harness"},
{Content: "rony-llm-agent: Go library...", ProjectID: "rony-llm-agent"},
})
runner := agent.NewRunner(agent.Config{
LLM: mockLLM,
Memory: memory,
Persona: testPersona,
})
resp, _ := runner.Run(context.Background(), []llm.Message{
{Role: llm.RoleUser, Content: "what projects does Victor have?"},
})
// Verify LLM received context chunks in system prompt
lastReq := mockLLM.LastRequest()
assert.Contains(t, lastReq.Messages[0].Content, "Rony Harness")
assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
}
```
### 9.3 E2E test with Astro
```bash
# 1. Start chat-bot on :7331
./bin/chat-bot serve &
# 2. Start Astro on :4321
cd ../portfolio && npm run dev &
# 3. Make request to Astro's proxy
curl -X POST http://localhost:4321/api/chat \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"hello"}]}'
# 4. Verify SSE stream
```
---
## 📂 10. Project Structure
```
rony-chat-bot/
├── cmd/
│ └── chat-bot/
│ └── main.go # CLI entrypoint
├── internal/
│ ├── server/ # HTTP handlers
│ │ ├── server.go # chi router + middleware
│ │ ├── handlers.go # /api/chat, /api/health, /api/info, /api/reindex, /api/conversations
│ │ ├── conversations_test.go # round-trip, continue, list, 404, delete, streaming
│ │ └── middleware.go # RequestID, Logging, CORS, RateLimit
│ │
│ ├── agent/ # LLM client + RAG runner
│ │ ├── runner.go # Stream wrapper, RAG injection into system prompt
│ │ └── client.go # NewClient factory: llamacpp / ollama / openai / anthropic
│ │
│ ├── portfolio/ # RAG: markdown → SQLite FTS5 + conversation persistence
│ │ ├── chunker.go # Heading-based splitter
│ │ ├── indexer.go # Store: schema, Reindex, Search (BM25)
│ │ ├── conversations.go # Conversation + Message CRUD, persisted alongside RAG
│ │ └── chunker_test.go / store_test.go
│ │
│ ├── persona/ # Persona bridge to rony-llm-agent
│ │ └── persona.go # FromConfig, BuildSystemPrompt (with RAG context)
│ │
│ ├── streaming/ # SSE protocol helpers
│ │ └── sse.go # WriteStart/Chunk/Sources/Done/Error
│ │
│ ├── i18n/ # Language detection (ES/EN) for the response
│ │
│ └── config/ # YAML loader + validation
├── web/ # ← DROP-IN CHAT WIDGET
│ ├── chat-widget.js # Vanilla JS, ~12 KB
│ ├── chat-widget.css # Scoped styles, CSS-custom-prop themable
│ ├── example.html # Local demo (python -m http.server)
│ └── README.md # Integration guide (HTML, Astro, Next.js)
├── data/
│ └── projects/ # ← Markdown per project (one .md per project)
│ ├── rony-harness.md
│ ├── rony-llm-agent.md
│ └── example-project.md
├── configs/
│ └── portfolio-bot.yaml # Provider + RAG + persona config
├── docs/
│ ├── architecture.md # ← THIS FILE
│ └── architecture.es.md
├── bench/ # Reproducible SQLite driver benchmark
├── go.mod # require rony-llm-agent, modernc.org/sqlite
└── README.md
```
---
## 📅 11. Roadmap
### Phase 1: MVP (2-3 weeks)
- [ ] Project setup (`go mod init`, structure)
- [ ] Basic HTTP server with `/api/chat` endpoint
- [ ] Functional SSE streaming
- [ ] RAG indexer (reads `data/projects/*.md` → SQLite FTS5)
- [ ] RAG retriever (query → top-k chunks)
- [ ] Persona loader from YAML
- [ ] llama.cpp integration (qwen2.5:1.5b GGUF)
- [ ] CLI: `serve`, `reindex`, `ask`
- [ ] Basic tests
### Phase 2: Integration with Astro (1 week)
- [ ] Astro API route of the proxy
- [ ] React component of the chat widget
- [ ] E2E test: Astro → chat-bot → response
- [ ] Widget styling (TailwindCSS)
### Phase 3: Polish (1 week)
- [ ] Robust rate limiting
- [ ] Structured logging (JSON)
- [ ] Health checks for monitoring
- [ ] systemd service file
- [ ] README + deployment docs
### Phase 4: Optionals
- [ ] Support for multiple conversations (session ID)
- [ ] Persisted chat history
- [ ] Analysis of frequent questions
- [ ] Multi-language (EN/ES switch)
- [ ] More polished standalone CLI version (`chat-bot ask`)
---
## 📐 12. Quality Specifications
### 12.1 Performance metrics
| Metric | Target |
|---|---|
| TTFT (Time-to-first-token) | <500ms with llama.cpp local |
| End-to-end (question complete response) | <3s for typical responses |
| Memory at rest | <150MB |
| RAG indexing speed | ~100 docs/second |
| Retrieval latency | <50ms for top-5 |
### 12.2 Required tests
- Unit tests: coverage 70%
- Integration tests: with mock LLM + in-memory SQLite FTS5
- E2E: at least one complete Astro chat-bot flow
---
## 🔒 13. Security
### 13.1 Implemented
- **Rate limiting** per IP (default 30 req/min)
- **Restrictive CORS** only configured origins
- **Input validation** JSON schema validation on requests
- **No PII storage** we don't save conversations by default
- **Local-only by default** no calls to cloud APIs
### 13.2 Deferred / Optional
- Auth with API key (for private use)
- Query logging for analytics
- IP anonymization in logs
- HTTPS via reverse proxy (Caddy/nginx)
---
## 📚 14. References
- **SSE Spec:** https://html.spec.whatwg.org/multipage/server-sent-events.html
- **Ollama API:** https://github.com/ollama/ollama/blob/main/docs/api.md
- **SQLite FTS5:** https://www.sqlite.org/fts5.html
- **Go SQLite driver:** https://github.com/mattn/go-sqlite3 (CGO) or https://modernc.org/sqlite (pure Go)
- **qwen2.5:** https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
- **Astro API routes:** https://docs.astro.build/en/guides/endpoints/
- **rony-llm-agent:** https://github.com/VictorVargas/rony-llm-agent
---
**Document ready for implementation. 🚀**