rony-chat-bot/docs/architecture.md
Victor Hugo Vargas ab510d8b31 docs: document auto-compaction behavior in architecture guide
Adds the §4.5 Auto-compaction section to architecture.md and
architecture.es.md describing the trigger, fallback, persistence and
the new SSE 'compaction' event so consumers know how to react.

Includes the three placeholder projects used while exercising the
feature end to end (bot-onboarding, dashboard-metricas, tienda-ropa)
so /api/reindex picks them up without further setup.
2026-07-18 00:08:37 -07:00

45 KiB
Raw Blame History

📋 Rony Chat Bot — Technical Design Document

Version: 1.0 Author: Victor Hugo Vargas Date: 2026-06-28 Status: Complete specification for implementation Path: rony-chat-bot/docs/architecture.md

🌐 Language: English | Español

📚 Workspace: This project is part of the Rony/ workspace. See ../README.md.

🔑 Depends on: rony-llm-agent — core library that provides agent loop, LLM clients, RAG, persona system.

📐 Methodology: This project follows the SDD + DDD + Hexagonal Architecture approach. Functional Requirements are numbered as CRF-XXX. See ../../METHODOLOGY.md.


🎯 1. Project Vision

1.1 What is Chat-Bot?

An HTTP chatbot that answers questions about Victor Hugo Vargas and his projects. Uses RAG (Retrieval-Augmented Generation) over markdown files describing each project, and a local LLM (or cloud) to generate responses.

1.2 Primary use case

Victor has a portfolio website (Astro + React). On the site there's a chat widget where visitors can ask:

  • "What projects has Victor done?"
  • "What's his experience with Go?"
  • "How does Rony Harness work?"
  • "Has Victor worked with PostgreSQL?"

The bot responds with accurate information extracted from the projects' markdown files + bio + skills.

1.3 Secondary use cases (future)

  • Client adaptation: The same bot, with other data and another persona, serves car dealerships, restaurants, etc.
  • Standalone CLI: ./chat-bot ask "what do you know about X?" for terminal use.
  • Slack/Discord bot: Wrapper that consumes the HTTP API.

1.4 Philosophy

  • Self-hosted by default — works 100% local with Ollama + 1-3B models
  • Cloud optional — if you need more quality, swap to Anthropic API
  • Portable — easy to fork/customize for other contexts
  • Streaming — token-by-token responses with SSE (no waiting for complete response)
  • Reuses rony-llm-agent — doesn't reinvent the agent loop

🏗️ 2. Architecture

2.1 Overview

┌─────────────────────────────────────────────────────────────────┐
│  Browser (Astro site)                                            │
│      ↓ HTTP POST /api/chat                                       │
│  Astro SSR (proxy)  ←────────── Serves portfolio + proxy chat    │
│      ↓ HTTP POST /api/chat                                       │
│  Chat-Bot HTTP server (:7331)                                    │
│      ↓                                                           │
│  Agent loop (rony-llm-agent)                                       │
│      ↓                                                           │
│  RAG retrieval → SQLite FTS5 over data/projects/*.md              │
│      ↓                                                           │
│  LLM (llama.cpp local default / Ollama or Anthropic optional)   │
└─────────────────────────────────────────────────────────────────┘

2.2 Main components

Component Path Responsibility
HTTP server internal/server/ Gin/chi handlers, SSE streaming
Agent runner internal/agent/ Wrapper over rony-llm-agent with specific config
Portfolio loader internal/portfolio/ Reads data/projects/*.md, indexes in SQLite FTS5
Persona internal/persona/ Loads persona from configs/portfolio-bot.yaml
CLI cmd/chat-bot/ Commands: serve, reindex, ask, version

2.3 Tech stack

Layer Technology Reason
Language Go 1.26+ Same as rony-harness, leverage os.Root, iter.Seq
HTTP router net/http + chi Stdlib + chi for middleware (CORS, logging)
SSE net/http Flusher Stdlib is enough, no external library needed
Config gopkg.in/yaml.v3 Same as harness
RAG backend SQLite + FTS5 (BM25) Zero external deps, single file, fast
LLM llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama as alt Self-hosted by default
Tests stdlib + testify Consistency with the rest

🔌 3. HTTP API

3.1 Endpoints

POST /api/chat — Chat with SSE streaming

Request:

{
  "messages": [
    {"role": "user", "content": "What projects does Victor have?"}
  ],
  "stream": true,
  "conversation_id": "57f4aa3c7fab466bc4de9c43b296903e"
}
Field Required Notes
messages yes At least one user message; alternation is not enforced.
stream no, default true false returns a single JSON body instead of SSE.
conversation_id no Hex string. If omitted, the server mints a new one and returns it (see below). Pass an existing ID to keep the thread.

Response (SSE):

data: {"type":"start","conversation_id":"57f4aa3c7fab466bc4de9c43b296903e"}

data: {"type":"chunk","content":"Victor"}
data: {"type":"chunk","content":" has"}
data: {"type":"chunk","content":" several"}
data: {"type":"chunk","content":" projects"}

data: {"type":"sources","documents":["rony-harness.md","rony-llm-agent.md"]}

data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}

The conversation_id in the start event is what the client should store (see §3.4 — Conversation persistence). When the client passed an existing ID the server echoes it back; otherwise it's freshly minted.

Without streaming ("stream": false):

{
  "conversation_id": "57f4aa3c7fab466bc4de9c43b296903e",
  "content": "Victor has several projects...",
  "sources": ["rony-harness.md", "rony-llm-agent.md"],
  "usage": {"input_tokens": 245, "output_tokens": 38}
}

GET /api/conversations — List recent conversations

Returns the most recent conversation summaries, newest first. Useful for a "show my chats" sidebar in a custom UI.

Query params:

  • limit (1200, default 50)

Response:

{
  "count": 2,
  "conversations": [
    {
      "id": "57f4aa3c7fab466bc4de9c43b296903e",
      "created_at": "2026-07-17T05:02:07Z",
      "updated_at": "2026-07-17T05:04:31Z",
      "preview": "What projects does Victor have?"
    }
  ]
}

GET /api/conversations/{id} — Fetch one conversation

Returns the full history of a conversation with all messages in chronological order.

Response (200):

{
  "id": "57f4aa3c7fab466bc4de9c43b296903e",
  "created_at": "2026-07-17T05:02:07Z",
  "updated_at": "2026-07-17T05:04:31Z",
  "messages": [
    {"id": 1, "role": "user",      "content": "What projects does Victor have?", "created_at": "..."},
    {"id": 2, "role": "assistant", "content": "Victor has several projects...", "sources": ["..."], "created_at": "..."}
  ]
}

Response (404): when the ID is unknown (e.g. server DB was wiped or the client lost sync). The widget treats this as "start fresh".

⚠️ Auth note: the conversation ID is the only access token. For a public bot this is fine; for private contexts add auth at the proxy layer (e.g. require a session cookie before forwarding to this endpoint).

DELETE /api/conversations/{id} — Delete a conversation

Removes the conversation and all its messages (cascade). Returns 204 on success, 404 if the ID doesn't exist.

POST /api/reindex — Re-index portfolio

Useful when files in data/projects/ are modified.

Request: empty Response:

{
  "indexed_files": 12,
  "total_chunks": 87,
  "duration_ms": 4321
}

GET /api/health — Health check (real)

Probes the LLM provider and the SQLite store in parallel and returns their states. Designed for monitoring/load balancers. Returns 200 when healthy or degraded, 503 when unhealthy.

  • ?deep=true adds a chunk count to the store probe (same latency budget).

Status taxonomy:

status HTTP Meaning
healthy 200 LLM up, store up
degraded 200 LLM up, store down — bot still answers, just without RAG
unhealthy 503 LLM down — bot cannot answer, no point routing traffic here

Probe details:

Component Probe Latency
llm GET {provider}/health (llamacpp, ollama) or /models (openai) ~1ms for local llama-server
store SELECT 1 on the SQLite handle ~100µs

Each probe has a 2s timeout; the whole call returns within ~2.5s even if a dependency hangs.

Response shape (healthy):

{
  "status": "healthy",
  "version": "0.2.0-dev",
  "checked_at": "2026-07-17T05:02:07Z",
  "components": {
    "llm": {
      "status": "up",
      "latency": "1.028ms",
      "details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
    },
    "store": {
      "status": "up",
      "latency": "107µs"
    }
  }
}

Response shape (degraded, with ?deep=true):

{
  "status": "degraded",
  "version": "0.2.0-dev",
  "checked_at": "2026-07-17T05:02:07Z",
  "components": {
    "llm": {"status": "up", "latency": "0.8ms", "details": {...}},
    "store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
  }
}

Response shape (unhealthy): HTTP 503, same JSON with "status": "unhealthy" and the failed component reporting "status": "down" plus an error field.

GET /api/info — Bot metadata

{
  "name": "Rony Chat Bot",
  "model": "qwen2.5:1.5b",
  "persona": "...",
  "topics": ["projects", "experience", "technical skills"]
}

3.2 SSE Implementation

// internal/server/chat.go
package server

import (
    "encoding/json"
    "fmt"
    "net/http"
    "github.com/VictorVargas/rony-llm-agent/pkg/agent"
)

func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
    // SSE headers
    w.Header().Set("Content-Type", "text/event-stream")
    w.Header().Set("Cache-Control", "no-cache")
    w.Header().Set("Connection", "keep-alive")
    w.Header().Set("X-Accel-Buffering", "no")
    
    flusher, ok := w.(http.Flusher)
    if !ok {
        http.Error(w, "SSE not supported", http.StatusInternalServerError)
        return
    }
    
    // Parse request
    var req ChatRequest
    if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
        writeError(w, flusher, "invalid request", err)
        return
    }
    
    // Start event
    writeSSE(w, flusher, "start", map[string]string{
        "conversation_id": generateConvID(),
    })
    
    // Run agent with streaming
    sources := []string{}
    for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
        if err != nil {
            writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
            return
        }
        if chunk.Type == "source" {
            sources = append(sources, chunk.Source)
        }
        writeSSE(w, flusher, chunk.Type, chunk.Data)
    }
    
    // Done event
    writeSSE(w, flusher, "done", map[string]any{
        "usage": map[string]int{
            "input_tokens":  245,
            "output_tokens": 38,
        },
    })
}

func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
    payload, _ := json.Marshal(data)
    fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
    flusher.Flush()
}

3.3 Middleware

// internal/server/middleware.go
package server

func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        start := time.Now()
        // Wrap response writer to capture status
        rw := &statusRecorder{ResponseWriter: w, status: 200}
        next.ServeHTTP(rw, r)
        
        slog.Info("http.request",
            "method", r.Method,
            "path", r.URL.Path,
            "status", rw.status,
            "duration_ms", time.Since(start).Milliseconds(),
            "ip", r.RemoteAddr,
        )
    })
}

func (s *Server) corsMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        origin := r.Header.Get("Origin")
        for _, allowed := range s.config.Server.CORSOrigins {
            if origin == allowed {
                w.Header().Set("Access-Control-Allow-Origin", origin)
                w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
                w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
                break
            }
        }
        if r.Method == "OPTIONS" {
            w.WriteHeader(204)
            return
        }
        next.ServeHTTP(w, r)
    })
}

func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
    limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        if !limiter.Allow() {
            http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
            return
        }
        next.ServeHTTP(w, r)
    })
}

3.4 Conversation persistence

The bot persists conversation threads in the same SQLite database as the RAG index (./data/portfolio.db). Schema lives in internal/portfolio/conversations.go.

CREATE TABLE conversations (
    id         TEXT PRIMARY KEY,         -- 16-byte random hex (32 chars)
    created_at INTEGER NOT NULL,
    updated_at INTEGER NOT NULL
);

CREATE TABLE messages (
    id              INTEGER PRIMARY KEY AUTOINCREMENT,
    conversation_id TEXT    NOT NULL,
    role            TEXT    NOT NULL,    -- user | assistant | system
    content         TEXT    NOT NULL,
    sources         TEXT,                 -- JSON array, nullable
    created_at      INTEGER NOT NULL,
    FOREIGN KEY (conversation_id) REFERENCES conversations(id) ON DELETE CASCADE
);
CREATE INDEX idx_messages_conv ON messages(conversation_id, id);

Lifecycle:

When What
POST /api/chat (no conversation_id) Server mints a new hex ID, returns it in the start SSE event (or conversation_id field of the JSON response)
POST /api/chat (with conversation_id) Server reuses the existing row; both user message and assistant reply are appended
User message Persisted before the LLM runs, so it survives a model failure
Assistant message Persisted after the stream completes, with the RAG sources attached
GET /api/conversations/{id} Returns the full thread; 404 if unknown
DELETE /api/conversations/{id} Cascade-deletes messages

Client responsibilities:

  1. On the first message, omit conversation_id. Capture the one the server returns in the start SSE event.
  2. Store it client-side (localStorage["rony-chat-conv"] in the widget).
  3. On every subsequent message, send the ID back.
  4. On page load, if you have a stored ID, call GET /api/conversations/{id} to restore the thread. If 404, clear the stored ID and start fresh.

The widget (web/chat-widget.js) implements all four steps. Any other client (a custom React component, an Astro endpoint, a CLI replay tool) follows the same protocol.

Auth model:

The conversation ID is the only access token for GET /api/conversations/{id}. It is 128 bits of random entropy, so guessing one is infeasible. For a public portfolio bot this is the right trade-off — anyone who knows the URL can read its history. For private contexts, add an auth layer in front of the bot (proxy) that gates the conversation endpoints.


🧠 4. RAG (Retrieval-Augmented Generation)

⚠️ Decisiones pendientes de validar antes de implementar este módulo:

  • Tokenizer FTS5 — el spec asume unicode61 remove_diacritics 2. Confirmar con datos reales si conviene cambiar a porter (stemming EN), trigram (sub-string matching) o un tokenizer custom para español. Validar: ejecutar queries representativas contra data/projects/ y comparar recall antes de cerrar esta elección.
  • Driver SQLite DECIDIDO: modernc.org/sqlite (puro Go, sin CGO). Ver benchmark abajo.
  • Chunking — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. Validar: medir recall con chunks por sección markdown (split por #/##) vs por tamaño.
  • Sin similitud semántica — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. Validar: tamaño del corpus y types of questions esperadas; si el corpus crece o las queries se vuelven abstractas, considerar agregar embeddings como capa secundaria.

4.0 Driver decision: benchmark results

Reproducible con CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/. Datos: 4 markdowns → 11 chunks.

Operación mattn (CGO) modernc (puro Go) Diferencia
Insert (11 chunks) 2,802,843 ns/op 1,465,646 ns/op modernc 1.9× más rápido
Insert alloc 2,124,299 B/op 9,770 B/op modernc usa 217× menos memoria
Query (8 queries BM25) 244,047 ns/op 555,162 ns/op mattn 2.3× más rápido
Round-trip (insert + 8 queries) 3,543,417 ns/op 2,267,669 ns/op modernc 1.6× más rápido
Binary size 11 MB 11 MB igual
Build deps gcc, CGO=1 nada modernc gana
CI/CD portable requiere toolchain C go build puro modernc gana

Decisión: modernc.org/sqlite.

Justificación:

  1. Ambas latencias de query (~250µs vs ~550µs) son 2 órdenes de magnitud por debajo del target de 50ms — imperceptible vs el LLM (varios segundos).
  2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
  3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
  4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.

4.1 Indexing pipeline

data/projects/*.md
    ↓ (read all files)
Raw markdown content
    ↓ (split into chunks, ~500 chars, 50 overlap)
Chunks []
    ↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
Indexed corpus

When it runs:

  • On bot startup (if --reindex-on-start flag)
  • Manually: ./chat-bot reindex
  • Via HTTP: POST /api/reindex

4.2 Retrieval pipeline

User query "what projects does Victor have?"
    ↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
Top 5 relevant chunks
    ↓ (format as context block)
System prompt += relevant chunks
    ↓ (send to LLM)
LLM generates answer

4.3 Implementation

// internal/portfolio/indexer.go
package portfolio

import (
    "context"
    "database/sql"
    "fmt"
    "log/slog"
    "os"
    "path/filepath"
    "strings"
)

type Indexer struct {
    dataPath     string
    db           *sql.DB
    chunkSize    int
    chunkOverlap int
}

func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
    files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
    if err != nil {
        return 0, err
    }

    // Rebuild FTS5 index from scratch (delete + insert is faster than diff for small corpora)
    if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
        return 0, fmt.Errorf("clear index: %w", err)
    }

    totalChunks := 0
    for _, file := range files {
        chunks, err := i.indexFile(ctx, file)
        if err != nil {
            slog.Warn("failed to index file", "file", file, "err", err)
            continue
        }
        totalChunks += chunks
    }

    return totalChunks, nil
}

func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
    content, err := os.ReadFile(path)
    if err != nil {
        return 0, err
    }

    projectID := strings.TrimSuffix(filepath.Base(path), ".md")
    chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)

    tx, err := i.db.BeginTx(ctx, nil)
    if err != nil {
        return 0, err
    }
    defer tx.Rollback()

    stmt, err := tx.PrepareContext(ctx, `
        INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
        VALUES (?, ?, ?, ?, ?)
    `)
    if err != nil {
        return 0, err
    }
    defer stmt.Close()

    for idx, chunk := range chunks {
        id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
        if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
            return idx, err
        }
    }

    if err := tx.Commit(); err != nil {
        return 0, err
    }
    return len(chunks), nil
}

// schema.go — applied at startup
const schema = `
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
    id UNINDEXED,
    project_id UNINDEXED,
    source_file UNINDEXED,
    chunk_index UNINDEXED,
    content,
    tokenize = 'unicode61 remove_diacritics 2'
);
`

func splitIntoChunks(text string, size, overlap int) []string {
    // Simple implementation: split by size with overlap
    // Production version uses tokenizer-aware chunking
    var chunks []string
    for i := 0; i < len(text); i += size - overlap {
        end := i + size
        if end > len(text) {
            end = len(text)
        }
        chunks = append(chunks, text[i:end])
    }
    return chunks
}

4.4 Retrieval in the agent loop

// internal/portfolio/search.go
package portfolio

type Hit struct {
    ProjectID  string
    SourceFile string
    ChunkIndex int
    Content    string
    Score      float64 // BM25 score from FTS5
}

func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
    // Escape user input: FTS5 syntax can break with special chars
    ftsQuery := sanitizeFTS5(query)

    rows, err := s.db.QueryContext(ctx, `
        SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
        FROM portfolio_chunks
        WHERE portfolio_chunks MATCH ?
        ORDER BY score
        LIMIT ?
    `, ftsQuery, topK)
    if err != nil {
        return nil, err
    }
    defer rows.Close()

    var hits []Hit
    for rows.Next() {
        var h Hit
        if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
            return nil, err
        }
        hits = append(hits, h)
    }
    return hits, rows.Err()
}

// sanitizeFTS5 wraps the user query so reserved chars and unquoted strings don't crash FTS5.
// A pragmatic choice for a Q&A bot: append prefix-match wildcard to each token.
func sanitizeFTS5(q string) string {
    tokens := strings.FieldsFunc(q, func(r rune) bool {
        return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
            (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
            r > 0x7F) // keep accented chars
    })
    if len(tokens) == 0 {
        return `""`
    }
    for i, t := range tokens {
        tokens[i] = `"` + strings.ToLower(t) + `"*`
    }
    return strings.Join(tokens, " ")
}
// internal/agent/runner.go
package agent

func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
    basePrompt := r.persona.SystemPrompt

    hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
    if err != nil {
        return "", err
    }
    if len(hits) == 0 {
        return basePrompt, nil
    }

    var contextBlock strings.Builder
    contextBlock.WriteString(basePrompt)
    contextBlock.WriteString("\n\n## Relevant context\n\n")
    for _, h := range hits {
        contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
            h.SourceFile, h.Content))
    }
    return contextBlock.String(), nil
}

func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
    return func(yield func(Chunk, error) bool) {
        lastUserMsg := getLastUserMessage(messages)
        systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
        if err != nil {
            yield(Chunk{}, err)
            return
        }

        messages = prependSystem(messages, systemPrompt)

        for chunk, err := range r.loop.RunStream(ctx, messages) {
            if !yield(chunk, err) {
                return
            }
        }
    }
}

Why this is simpler than embeddings:

  • No embedding model to download or run (saves ~270MB of RAM and ~200ms per query)
  • One file (data/portfolio.db), one driver, no extra process
  • BM25 ranking is excellent for keyword-based retrieval over structured docs like project READMEs
  • Trade-off: no semantic similarity ("projects about AI" won't match "machine learning" without the literal words). Mitigation: trigram tokenizer handles morphology well for English/Spanish.

🗜️ 4.5 Auto-compaction

Long conversations eventually run out of context — at qwen2.5-3b's 4k window, the ~3k-token system prompt + RAG block leaves only room for 23 user turns. Auto-compaction solves this by folding the older portion of the conversation into a single summary system message when the previous turn's input tokens cross a configurable threshold.

When it fires

agent.Runner.Compact runs once per /api/chat request, before the RAG search. It compares the runner's most recent Usage.InputTokens (reported by the provider in the previous streamed chunk) against client.Capabilities().MaxContextWindow × threshold_ratio.

Setting Default What it controls
compaction.enabled false Master switch.
compaction.threshold_ratio 0.75 Trigger when used tokens ≥ window × ratio.
compaction.keep_recent_turns 4 How many of the latest user turns are kept verbatim after compaction.
compaction.summary_system_prompt (built-in bilingual) Override the instruction sent to the LLM when summarizing.

Short-circuits silently when compaction is disabled, the provider doesn't report a window (Capabilities().MaxContextWindow == 0), the history is shorter than keep_recent_turns, or usage is still unknown (first turn).

How the summary is made

  1. splitByTurns(history, keep_recent_turns) divides messages into (older, recent) on user-role boundaries so a kept turn's user/assistant pair always stays together.
  2. renderTranscript(older) flattens older messages into a User: / Assistant: transcript (skipping tool messages and empty assistant placeholders).
  3. The runner calls client.Generate(...) with the summary prompt + transcript and a 512-token cap so the compaction step itself stays cheap.
  4. The returned text is prepended as a system message ("Earlier conversation summary:\n…"), followed by the recent tail.
  5. LastCompaction() returns CompactionStats so the SSE handler can emit a compaction event right before the streamed chunks.

Failure mode

If Generate errors or returns an empty summary, compaction falls back to truncateToBudget: drop oldest user-turns one at a time until the remaining slice fits threshold tokens (heuristic: len(s) / 4 + 1). The current user turn is always preserved. The fallback is logged at WARN and the request still proceeds — a flaky summarize call never fails the user's request.

Wire protocol

Streaming responses gain an optional compaction event:

data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500}

Emitted after start (when applicable) and before sources / chunk. The widget can render this as a subtle "Context compacted" hint or ignore it — both are valid.

Persistence

Compaction is per-request. The full transcript is still saved to messages in data/portfolio.db verbatim, so GET /api/conversations/{id} always returns the original history. Only what we send to the LLM is reduced — the next session can re-read the full thread from the DB.


🌐 5. Embedding the widget

The bot ships with a drop-in vanilla-JS widget. Add two files to your site and it works.

5.1 The widget (any site)

<link rel="stylesheet" href="/path/to/chat-widget.css">
<script src="/path/to/chat-widget.js"
        data-api-url="https://chat.example.com"
        data-title="Ask me anything"
        data-greeting="Hi! Ask me about the projects."
        data-position="bottom-right"
        data-theme="auto"
        defer></script>

A bubble appears bottom-right, opens a panel, talks SSE to /api/chat, streams the response, and cites sources. No build step, no React/Vue, no framework lock-in.

Browser→bot options:

Topology Trade-offs
Direct (browser → bot, same domain or CORS) Simplest. Add the bot's origin to cors_origins in YAML.
Reverse proxy (nginx/Caddy in front) Bot stays on private network, single public domain, no CORS to manage.
Site proxies the bot (Astro/Next API route) Adds a hop and a bit of code, but gives you auth/session hooks in your site.

The widget works the same in all three. Pick the topology that matches your infra.

Default dev setup is direct + CORS. cors_origins in configs/portfolio-bot.yaml controls which sites can call the bot. Add your site's origin there.

5.2 Astro: drop-in via Layout

The widget works in Astro without writing a React component. Add this to your shared layout:

---
// src/layouts/BaseLayout.astro
import "../path/to/chat-widget.css";
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
---
<html>
  <body>
    <slot />
    <script src="/path/to/chat-widget.js"
            data-api-url={apiUrl}
            data-title="Ask me anything"
            data-position="bottom-right"
            data-theme="auto"
            defer is:inline></script>
  </body>
</html>

is:inline keeps Astro from hashing/transforming the script tag, so the data-* attributes survive.

5.3 React / Next.js: same script tag

// app/layout.tsx
import Script from "next/script";

export default function RootLayout({ children }) {
  return (
    <html>
      <head>
        <link rel="stylesheet" href="/chat-widget.css" />
        <Script src="/chat-widget.js"
                data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
                data-title="Ask me anything"
                data-position="bottom-right"
                data-theme="auto"
                strategy="afterInteractive" />
      </head>
      <body>{children}</body>
    </html>
  );
}

5.4 If you want a server proxy (Astro/Next API route)

The widget can also call a same-origin endpoint that forwards to the bot. This is the right call when you need:

  • Auth on /api/chat (logged-in users only)
  • Centralized rate limiting at the site level
  • Hiding the bot's origin from the browser
// src/pages/api/chat.ts (Astro) or app/api/chat/route.ts (Next)
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";

export const POST = async ({ request }) => {
    const body = await request.json();
    // (optional) auth check, rate limit, session lookup here

    const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify(body),
    });

    return new Response(resp.body, {
        status: resp.status,
        headers: {
            "Content-Type": "text/event-stream",
            "Cache-Control": "no-cache",
            "Connection": "keep-alive",
        },
    });
};

Then point the widget at /api/chat (same origin) instead of the bot's URL.

5.5 Widget configuration reference

All options are data-* attributes on the <script> tag:

Attribute Default Notes
data-api-url (required) Base URL of the bot. No trailing slash.
data-title "Chat" Header text.
data-greeting "" First assistant message when the panel opens.
data-position "bottom-right" "bottom-right" or "bottom-left".
data-theme "auto" "auto" (follows OS), "light", "dark".

Theming is via CSS custom properties on .rony-chat-widget-root (see web/chat-widget.css):

.rony-chat-widget-root {
  --rony-accent: #ff6b35;
  --rony-radius: 4px;
  --rony-font: "Inter", sans-serif;
}

5.6 What the widget doesn't do (yet)

  • Richer markdown (tables, images) — the built-in renderer handles the common cases; for full CommonMark, swap renderMarkdown in chat-widget.js for marked or markdown-it.
  • Mobile swipe-to-dismiss — panel goes full-screen on phones.
  • Conversation history sidebar — only the active conversation is shown (the backend exposes GET /api/conversations for a future sidebar).

🤖 6. Self-hosting with llama.cpp (default)

6.1 Setup

llama-server is a separate process that the bot connects to over HTTP. Both ports (the bot's and llama-server's) are configurable — pick what fits your environment.

# 1. Make sure you have a GGUF model available
# Download from Hugging Face, e.g.:
#   https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
export RONY_MODELS_PATH=/path/to/models
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf

# 2. Start llama-server (port is configurable; default llama.cpp is 8080)
llama-server \
  -m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
  --port 9100 \
  --host 127.0.0.1 \
  --ctx-size 4096 \
  --mlock            # prevents swap, critical on shared VPS

# 3. Make sure configs/portfolio-bot.yaml points to the same port
#    providers[0].endpoint: http://localhost:9100/v1

# 4. Start the bot (default port 7331, also configurable)
./bin/chat-bot serve
# → Serves on http://localhost:7331
# → Override with: ./bin/chat-bot serve --port 9101 --host 127.0.0.1

Port reference:

What Default How to change
llama-server HTTP port 8080 (llama.cpp convention) --port N flag when starting llama-server
chat-bot HTTP port 7331 --port N flag on serve, or server.port in YAML
chat-bot → llama-server URL http://localhost:8080/v1 endpoint field on the provider in YAML

The llamacpp provider is imported from rony-llm-agent/pkg/llm/providers/llamacpp and is compiled against llama.cpp via CGO or external binary.

6.2 Alternative: Ollama (easier for development)

If you don't want to manage GGUF files manually, Ollama provides the same models with a simpler workflow:

# 1. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# 2. Download chat model
ollama pull qwen2.5:1.5b

# 3. Verify
ollama list

# 4. Edit configs/portfolio-bot.yaml to mark ollama-local as default:
#    providers[0].default: true (and remove default from llamacpp-local)
#    Ollama exposes an OpenAI-compatible API on :11434/v1

# 5. Start the bot
ollama serve &
./bin/chat-bot serve

6.3 Alternative: llama.cpp direct (advanced)

For more control or if Ollama doesn't work in your setup:

providers:
  - name: llamacpp-local
    type: llamacpp
    model: qwen2.5-3b-instruct
    endpoint: http://localhost:9100/v1   # configurable, see §6.1
    context_size: 4096
    max_tokens: 2048
    default: true

The llamacpp adapter is imported from rony-llm-agent/pkg/llm/providers/llamacpp and is compiled against llama.cpp via CGO or external binary.


📦 7. Bot CLI

7.1 Commands

# Start HTTP server
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]

# Re-index portfolio (reads data/projects/*.md → SQLite FTS5)
chat-bot reindex

# Single question (no server, useful for tests)
chat-bot ask "What projects does Victor have?" [--no-rag]

# Validate config
chat-bot config validate

# Health check (useful for monitoring)
chat-bot health

# Version
chat-bot version

7.2 Implementation with Cobra

// cmd/chat-bot/main.go
package main

import (
    "github.com/spf13/cobra"
)

func main() {
    root := &cobra.Command{
        Use:   "chat-bot",
        Short: "Portfolio chatbot HTTP server",
    }
    
    root.AddCommand(serveCmd())
    root.AddCommand(reindexCmd())
    root.AddCommand(askCmd())
    root.AddCommand(configCmd())
    root.AddCommand(healthCmd())
    root.AddCommand(versionCmd())
    
    if err := root.Execute(); err != nil {
        os.Exit(1)
    }
}

func serveCmd() *cobra.Command {
    var port int
    var host string
    var reindexOnStart bool
    
    cmd := &cobra.Command{
        Use:   "serve",
        Short: "Start HTTP server",
        RunE: func(cmd *cobra.Command, args []string) error {
            return server.Serve(server.Config{
                Port:           port,
                Host:           host,
                ReindexOnStart: reindexOnStart,
            })
        },
    }
    
    cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
    cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
    cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
    
    return cmd
}

🚀 8. Deployment

8.1 Recommendation: Self-hosted on VPS

# 1. Install dependencies
sudo apt install golang-go ollama
ollama pull qwen2.5:1.5b

# 2. Build
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot

# 3. systemd service
cat > /etc/systemd/system/chat-bot.service <<EOF
[Unit]
Description=Portfolio Chat Bot
After=network.target ollama.service

[Service]
Type=simple
User=chatbot
WorkingDirectory=/opt/chat-bot
ExecStart=/usr/local/bin/chat-bot serve
Restart=on-failure
Environment=RONY_MODELS_PATH=/opt/models

[Install]
WantedBy=multi-user.target
EOF

sudo systemctl enable --now chat-bot

8.2 Reverse proxy (Caddy)

# /etc/caddy/Caddyfile
chat.victorvargas.dev {
    reverse_proxy localhost:7331
}

8.3 Monitoring

# Health check periodic
curl -s http://localhost:7331/api/health | jq

# Logs
journalctl -u chat-bot -f

🧪 9. Testing

9.1 Unit tests

// internal/server/chat_test.go
package server

func TestHandleChat_ValidRequest(t *testing.T) {
    s := newTestServer(t)
    
    req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
        "messages": [{"role": "user", "content": "hello"}]
    }`))
    req.Header.Set("Content-Type", "application/json")
    
    w := httptest.NewRecorder()
    s.handleChat(w, req)
    
    assert.Equal(t, 200, w.Code)
    assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
}

func TestHandleChat_RateLimit(t *testing.T) {
    s := newTestServerWithConfig(t, server.Config{
        RateLimit: 1, // 1 request per minute
    })
    
    // First request OK
    req1 := newChatRequest("hello")
    w1 := httptest.NewRecorder()
    s.handleChat(w1, req1)
    assert.Equal(t, 200, w1.Code)
    
    // Second request denied
    req2 := newChatRequest("hello again")
    w2 := httptest.NewRecorder()
    s.handleChat(w2, req2)
    assert.Equal(t, 429, w2.Code)
}

9.2 Integration tests with mock LLM

// internal/agent/runner_test.go
func TestRunner_RAGContextIsInjected(t *testing.T) {
    mockLLM := mock.New(mock.Responses{
        {Match: "projects", Response: "Victor has several projects..."},
    })
    
    memory := newMockMemoryWithDocs(t, []rag.Fragment{
        {Content: "Rony Harness: AI agent harness...", ProjectID: "rony-harness"},
        {Content: "rony-llm-agent: Go library...", ProjectID: "rony-llm-agent"},
    })
    
    runner := agent.NewRunner(agent.Config{
        LLM:    mockLLM,
        Memory: memory,
        Persona: testPersona,
    })
    
    resp, _ := runner.Run(context.Background(), []llm.Message{
        {Role: llm.RoleUser, Content: "what projects does Victor have?"},
    })
    
    // Verify LLM received context chunks in system prompt
    lastReq := mockLLM.LastRequest()
    assert.Contains(t, lastReq.Messages[0].Content, "Rony Harness")
    assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
}

9.3 E2E test with Astro

# 1. Start chat-bot on :7331
./bin/chat-bot serve &

# 2. Start Astro on :4321
cd ../portfolio && npm run dev &

# 3. Make request to Astro's proxy
curl -X POST http://localhost:4321/api/chat \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"hello"}]}'

# 4. Verify SSE stream

📂 10. Project Structure

rony-chat-bot/
├── cmd/
│   └── chat-bot/
│       └── main.go                 # CLI entrypoint
│
├── internal/
│   ├── server/                     # HTTP handlers
│   │   ├── server.go               # chi router + middleware
│   │   ├── handlers.go             # /api/chat, /api/health, /api/info, /api/reindex, /api/conversations
│   │   ├── conversations_test.go   # round-trip, continue, list, 404, delete, streaming
│   │   └── middleware.go           # RequestID, Logging, CORS, RateLimit
│   │
│   ├── agent/                      # LLM client + RAG runner
│   │   ├── runner.go               # Stream wrapper, RAG injection into system prompt
│   │   └── client.go               # NewClient factory: llamacpp / ollama / openai / anthropic
│   │
│   ├── portfolio/                  # RAG: markdown → SQLite FTS5 + conversation persistence
│   │   ├── chunker.go              # Heading-based splitter
│   │   ├── indexer.go              # Store: schema, Reindex, Search (BM25)
│   │   ├── conversations.go        # Conversation + Message CRUD, persisted alongside RAG
│   │   └── chunker_test.go / store_test.go
│   │
│   ├── persona/                    # Persona bridge to rony-llm-agent
│   │   └── persona.go              # FromConfig, BuildSystemPrompt (with RAG context)
│   │
│   ├── streaming/                  # SSE protocol helpers
│   │   └── sse.go                  # WriteStart/Chunk/Sources/Done/Error
│   │
│   ├── i18n/                       # Language detection (ES/EN) for the response
│   │
│   └── config/                     # YAML loader + validation
│
├── web/                            # ← DROP-IN CHAT WIDGET
│   ├── chat-widget.js              # Vanilla JS, ~12 KB
│   ├── chat-widget.css             # Scoped styles, CSS-custom-prop themable
│   ├── example.html                # Local demo (python -m http.server)
│   └── README.md                   # Integration guide (HTML, Astro, Next.js)
│
├── data/
│   └── projects/                   # ← Markdown per project (one .md per project)
│       ├── rony-harness.md
│       ├── rony-llm-agent.md
│       └── example-project.md
│
├── configs/
│   └── portfolio-bot.yaml          # Provider + RAG + persona config
│
├── docs/
│   ├── architecture.md             # ← THIS FILE
│   └── architecture.es.md
│
├── bench/                          # Reproducible SQLite driver benchmark
│
├── go.mod                          # require rony-llm-agent, modernc.org/sqlite
└── README.md

📅 11. Roadmap

Phase 1: MVP (2-3 weeks)

  • Project setup (go mod init, structure)
  • Basic HTTP server with /api/chat endpoint
  • Functional SSE streaming
  • RAG indexer (reads data/projects/*.md → SQLite FTS5)
  • RAG retriever (query → top-k chunks)
  • Persona loader from YAML
  • llama.cpp integration (qwen2.5:1.5b GGUF)
  • CLI: serve, reindex, ask
  • Basic tests

Phase 2: Integration with Astro (1 week)

  • Astro API route of the proxy
  • React component of the chat widget
  • E2E test: Astro → chat-bot → response
  • Widget styling (TailwindCSS)

Phase 3: Polish (1 week)

  • Robust rate limiting
  • Structured logging (JSON)
  • Health checks for monitoring
  • systemd service file
  • README + deployment docs

Phase 4: Optionals

  • Support for multiple conversations (session ID)
  • Persisted chat history
  • Analysis of frequent questions
  • Multi-language (EN/ES switch)
  • More polished standalone CLI version (chat-bot ask)

📐 12. Quality Specifications

12.1 Performance metrics

Metric Target
TTFT (Time-to-first-token) <500ms with llama.cpp local
End-to-end (question → complete response) <3s for typical responses
Memory at rest <150MB
RAG indexing speed ~100 docs/second
Retrieval latency <50ms for top-5

12.2 Required tests

  • Unit tests: coverage ≥70%
  • Integration tests: with mock LLM + in-memory SQLite FTS5
  • E2E: at least one complete Astro → chat-bot flow

🔒 13. Security

13.1 Implemented

  • Rate limiting per IP (default 30 req/min)
  • Restrictive CORS — only configured origins
  • Input validation — JSON schema validation on requests
  • No PII storage — we don't save conversations by default
  • Local-only by default — no calls to cloud APIs

13.2 Deferred / Optional

  • Auth with API key (for private use)
  • Query logging for analytics
  • IP anonymization in logs
  • HTTPS via reverse proxy (Caddy/nginx)

📚 14. References


Document ready for implementation. 🚀