Adds the §4.5 Auto-compaction section to architecture.md and architecture.es.md describing the trigger, fallback, persistence and the new SSE 'compaction' event so consumers know how to react. Includes the three placeholder projects used while exercising the feature end to end (bot-onboarding, dashboard-metricas, tienda-ropa) so /api/reindex picks them up without further setup.
45 KiB
📋 Rony Chat Bot — Technical Design Document
Version: 1.0
Author: Victor Hugo Vargas
Date: 2026-06-28
Status: Complete specification for implementation
Path: rony-chat-bot/docs/architecture.md
📚 Workspace: This project is part of the
Rony/workspace. See../README.md.🔑 Depends on:
rony-llm-agent— core library that provides agent loop, LLM clients, RAG, persona system.📐 Methodology: This project follows the SDD + DDD + Hexagonal Architecture approach. Functional Requirements are numbered as
CRF-XXX. See../../METHODOLOGY.md.
🎯 1. Project Vision
1.1 What is Chat-Bot?
An HTTP chatbot that answers questions about Victor Hugo Vargas and his projects. Uses RAG (Retrieval-Augmented Generation) over markdown files describing each project, and a local LLM (or cloud) to generate responses.
1.2 Primary use case
Victor has a portfolio website (Astro + React). On the site there's a chat widget where visitors can ask:
- "What projects has Victor done?"
- "What's his experience with Go?"
- "How does Rony Harness work?"
- "Has Victor worked with PostgreSQL?"
The bot responds with accurate information extracted from the projects' markdown files + bio + skills.
1.3 Secondary use cases (future)
- Client adaptation: The same bot, with other data and another persona, serves car dealerships, restaurants, etc.
- Standalone CLI:
./chat-bot ask "what do you know about X?"for terminal use. - Slack/Discord bot: Wrapper that consumes the HTTP API.
1.4 Philosophy
- Self-hosted by default — works 100% local with Ollama + 1-3B models
- Cloud optional — if you need more quality, swap to Anthropic API
- Portable — easy to fork/customize for other contexts
- Streaming — token-by-token responses with SSE (no waiting for complete response)
- Reuses
rony-llm-agent— doesn't reinvent the agent loop
🏗️ 2. Architecture
2.1 Overview
┌─────────────────────────────────────────────────────────────────┐
│ Browser (Astro site) │
│ ↓ HTTP POST /api/chat │
│ Astro SSR (proxy) ←────────── Serves portfolio + proxy chat │
│ ↓ HTTP POST /api/chat │
│ Chat-Bot HTTP server (:7331) │
│ ↓ │
│ Agent loop (rony-llm-agent) │
│ ↓ │
│ RAG retrieval → SQLite FTS5 over data/projects/*.md │
│ ↓ │
│ LLM (llama.cpp local default / Ollama or Anthropic optional) │
└─────────────────────────────────────────────────────────────────┘
2.2 Main components
| Component | Path | Responsibility |
|---|---|---|
| HTTP server | internal/server/ |
Gin/chi handlers, SSE streaming |
| Agent runner | internal/agent/ |
Wrapper over rony-llm-agent with specific config |
| Portfolio loader | internal/portfolio/ |
Reads data/projects/*.md, indexes in SQLite FTS5 |
| Persona | internal/persona/ |
Loads persona from configs/portfolio-bot.yaml |
| CLI | cmd/chat-bot/ |
Commands: serve, reindex, ask, version |
2.3 Tech stack
| Layer | Technology | Reason |
|---|---|---|
| Language | Go 1.26+ | Same as rony-harness, leverage os.Root, iter.Seq |
| HTTP router | net/http + chi |
Stdlib + chi for middleware (CORS, logging) |
| SSE | net/http Flusher |
Stdlib is enough, no external library needed |
| Config | gopkg.in/yaml.v3 |
Same as harness |
| RAG backend | SQLite + FTS5 (BM25) | Zero external deps, single file, fast |
| LLM | llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama as alt | Self-hosted by default |
| Tests | stdlib + testify | Consistency with the rest |
🔌 3. HTTP API
3.1 Endpoints
POST /api/chat — Chat with SSE streaming
Request:
{
"messages": [
{"role": "user", "content": "What projects does Victor have?"}
],
"stream": true,
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e"
}
| Field | Required | Notes |
|---|---|---|
messages |
yes | At least one user message; alternation is not enforced. |
stream |
no, default true |
false returns a single JSON body instead of SSE. |
conversation_id |
no | Hex string. If omitted, the server mints a new one and returns it (see below). Pass an existing ID to keep the thread. |
Response (SSE):
data: {"type":"start","conversation_id":"57f4aa3c7fab466bc4de9c43b296903e"}
data: {"type":"chunk","content":"Victor"}
data: {"type":"chunk","content":" has"}
data: {"type":"chunk","content":" several"}
data: {"type":"chunk","content":" projects"}
data: {"type":"sources","documents":["rony-harness.md","rony-llm-agent.md"]}
data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}
The conversation_id in the start event is what the client should store
(see §3.4 — Conversation persistence). When the client passed an
existing ID the server echoes it back; otherwise it's freshly minted.
Without streaming ("stream": false):
{
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e",
"content": "Victor has several projects...",
"sources": ["rony-harness.md", "rony-llm-agent.md"],
"usage": {"input_tokens": 245, "output_tokens": 38}
}
GET /api/conversations — List recent conversations
Returns the most recent conversation summaries, newest first. Useful for a "show my chats" sidebar in a custom UI.
Query params:
limit(1–200, default 50)
Response:
{
"count": 2,
"conversations": [
{
"id": "57f4aa3c7fab466bc4de9c43b296903e",
"created_at": "2026-07-17T05:02:07Z",
"updated_at": "2026-07-17T05:04:31Z",
"preview": "What projects does Victor have?"
}
]
}
GET /api/conversations/{id} — Fetch one conversation
Returns the full history of a conversation with all messages in chronological order.
Response (200):
{
"id": "57f4aa3c7fab466bc4de9c43b296903e",
"created_at": "2026-07-17T05:02:07Z",
"updated_at": "2026-07-17T05:04:31Z",
"messages": [
{"id": 1, "role": "user", "content": "What projects does Victor have?", "created_at": "..."},
{"id": 2, "role": "assistant", "content": "Victor has several projects...", "sources": ["..."], "created_at": "..."}
]
}
Response (404): when the ID is unknown (e.g. server DB was wiped or the client lost sync). The widget treats this as "start fresh".
⚠️ Auth note: the conversation ID is the only access token. For a public bot this is fine; for private contexts add auth at the proxy layer (e.g. require a session cookie before forwarding to this endpoint).
DELETE /api/conversations/{id} — Delete a conversation
Removes the conversation and all its messages (cascade). Returns 204 on success, 404 if the ID doesn't exist.
POST /api/reindex — Re-index portfolio
Useful when files in data/projects/ are modified.
Request: empty Response:
{
"indexed_files": 12,
"total_chunks": 87,
"duration_ms": 4321
}
GET /api/health — Health check (real)
Probes the LLM provider and the SQLite store in parallel and returns their states. Designed for monitoring/load balancers. Returns 200 when healthy or degraded, 503 when unhealthy.
?deep=trueadds a chunk count to the store probe (same latency budget).
Status taxonomy:
status |
HTTP | Meaning |
|---|---|---|
healthy |
200 | LLM up, store up |
degraded |
200 | LLM up, store down — bot still answers, just without RAG |
unhealthy |
503 | LLM down — bot cannot answer, no point routing traffic here |
Probe details:
| Component | Probe | Latency |
|---|---|---|
llm |
GET {provider}/health (llamacpp, ollama) or /models (openai) |
~1ms for local llama-server |
store |
SELECT 1 on the SQLite handle |
~100µs |
Each probe has a 2s timeout; the whole call returns within ~2.5s even if a dependency hangs.
Response shape (healthy):
{
"status": "healthy",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {
"status": "up",
"latency": "1.028ms",
"details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
},
"store": {
"status": "up",
"latency": "107µs"
}
}
}
Response shape (degraded, with ?deep=true):
{
"status": "degraded",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {"status": "up", "latency": "0.8ms", "details": {...}},
"store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
}
}
Response shape (unhealthy): HTTP 503, same JSON with "status": "unhealthy" and the failed component reporting "status": "down" plus an error field.
GET /api/info — Bot metadata
{
"name": "Rony Chat Bot",
"model": "qwen2.5:1.5b",
"persona": "...",
"topics": ["projects", "experience", "technical skills"]
}
3.2 SSE Implementation
// internal/server/chat.go
package server
import (
"encoding/json"
"fmt"
"net/http"
"github.com/VictorVargas/rony-llm-agent/pkg/agent"
)
func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
// SSE headers
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
w.Header().Set("Connection", "keep-alive")
w.Header().Set("X-Accel-Buffering", "no")
flusher, ok := w.(http.Flusher)
if !ok {
http.Error(w, "SSE not supported", http.StatusInternalServerError)
return
}
// Parse request
var req ChatRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
writeError(w, flusher, "invalid request", err)
return
}
// Start event
writeSSE(w, flusher, "start", map[string]string{
"conversation_id": generateConvID(),
})
// Run agent with streaming
sources := []string{}
for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
if err != nil {
writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
return
}
if chunk.Type == "source" {
sources = append(sources, chunk.Source)
}
writeSSE(w, flusher, chunk.Type, chunk.Data)
}
// Done event
writeSSE(w, flusher, "done", map[string]any{
"usage": map[string]int{
"input_tokens": 245,
"output_tokens": 38,
},
})
}
func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
payload, _ := json.Marshal(data)
fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
flusher.Flush()
}
3.3 Middleware
// internal/server/middleware.go
package server
func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
// Wrap response writer to capture status
rw := &statusRecorder{ResponseWriter: w, status: 200}
next.ServeHTTP(rw, r)
slog.Info("http.request",
"method", r.Method,
"path", r.URL.Path,
"status", rw.status,
"duration_ms", time.Since(start).Milliseconds(),
"ip", r.RemoteAddr,
)
})
}
func (s *Server) corsMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
origin := r.Header.Get("Origin")
for _, allowed := range s.config.Server.CORSOrigins {
if origin == allowed {
w.Header().Set("Access-Control-Allow-Origin", origin)
w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
break
}
}
if r.Method == "OPTIONS" {
w.WriteHeader(204)
return
}
next.ServeHTTP(w, r)
})
}
func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if !limiter.Allow() {
http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
return
}
next.ServeHTTP(w, r)
})
}
3.4 Conversation persistence
The bot persists conversation threads in the same SQLite database as the
RAG index (./data/portfolio.db). Schema lives in internal/portfolio/conversations.go.
CREATE TABLE conversations (
id TEXT PRIMARY KEY, -- 16-byte random hex (32 chars)
created_at INTEGER NOT NULL,
updated_at INTEGER NOT NULL
);
CREATE TABLE messages (
id INTEGER PRIMARY KEY AUTOINCREMENT,
conversation_id TEXT NOT NULL,
role TEXT NOT NULL, -- user | assistant | system
content TEXT NOT NULL,
sources TEXT, -- JSON array, nullable
created_at INTEGER NOT NULL,
FOREIGN KEY (conversation_id) REFERENCES conversations(id) ON DELETE CASCADE
);
CREATE INDEX idx_messages_conv ON messages(conversation_id, id);
Lifecycle:
| When | What |
|---|---|
POST /api/chat (no conversation_id) |
Server mints a new hex ID, returns it in the start SSE event (or conversation_id field of the JSON response) |
POST /api/chat (with conversation_id) |
Server reuses the existing row; both user message and assistant reply are appended |
| User message | Persisted before the LLM runs, so it survives a model failure |
| Assistant message | Persisted after the stream completes, with the RAG sources attached |
GET /api/conversations/{id} |
Returns the full thread; 404 if unknown |
DELETE /api/conversations/{id} |
Cascade-deletes messages |
Client responsibilities:
- On the first message, omit
conversation_id. Capture the one the server returns in thestartSSE event. - Store it client-side (
localStorage["rony-chat-conv"]in the widget). - On every subsequent message, send the ID back.
- On page load, if you have a stored ID, call
GET /api/conversations/{id}to restore the thread. If 404, clear the stored ID and start fresh.
The widget (web/chat-widget.js) implements all four steps. Any other
client (a custom React component, an Astro endpoint, a CLI replay tool)
follows the same protocol.
Auth model:
The conversation ID is the only access token for GET /api/conversations/{id}.
It is 128 bits of random entropy, so guessing one is infeasible. For a
public portfolio bot this is the right trade-off — anyone who knows the
URL can read its history. For private contexts, add an auth layer in front
of the bot (proxy) that gates the conversation endpoints.
🧠 4. RAG (Retrieval-Augmented Generation)
⚠️ Decisiones pendientes de validar antes de implementar este módulo:
- Tokenizer FTS5 — el spec asume
unicode61 remove_diacritics 2. Confirmar con datos reales si conviene cambiar aporter(stemming EN),trigram(sub-string matching) o un tokenizer custom para español. Validar: ejecutar queries representativas contradata/projects/y comparar recall antes de cerrar esta elección.- Driver SQLite — ✅ DECIDIDO:
modernc.org/sqlite(puro Go, sin CGO). Ver benchmark abajo.- Chunking — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. Validar: medir recall con chunks por sección markdown (split por
#/##) vs por tamaño.- Sin similitud semántica — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. Validar: tamaño del corpus y types of questions esperadas; si el corpus crece o las queries se vuelven abstractas, considerar agregar embeddings como capa secundaria.
4.0 Driver decision: benchmark results
Reproducible con CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/. Datos: 4 markdowns → 11 chunks.
| Operación | mattn (CGO) | modernc (puro Go) | Diferencia |
|---|---|---|---|
| Insert (11 chunks) | 2,802,843 ns/op | 1,465,646 ns/op | modernc 1.9× más rápido |
| Insert alloc | 2,124,299 B/op | 9,770 B/op | modernc usa 217× menos memoria |
| Query (8 queries BM25) | 244,047 ns/op | 555,162 ns/op | mattn 2.3× más rápido |
| Round-trip (insert + 8 queries) | 3,543,417 ns/op | 2,267,669 ns/op | modernc 1.6× más rápido |
| Binary size | 11 MB | 11 MB | igual |
| Build deps | gcc, CGO=1 | nada | modernc gana |
| CI/CD portable | requiere toolchain C | go build puro |
modernc gana |
Decisión: modernc.org/sqlite.
Justificación:
- Ambas latencias de query (~250µs vs ~550µs) son 2 órdenes de magnitud por debajo del target de 50ms — imperceptible vs el LLM (varios segundos).
- modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
- Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
- Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.
4.1 Indexing pipeline
data/projects/*.md
↓ (read all files)
Raw markdown content
↓ (split into chunks, ~500 chars, 50 overlap)
Chunks []
↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
Indexed corpus
When it runs:
- On bot startup (if
--reindex-on-startflag) - Manually:
./chat-bot reindex - Via HTTP:
POST /api/reindex
4.2 Retrieval pipeline
User query "what projects does Victor have?"
↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
Top 5 relevant chunks
↓ (format as context block)
System prompt += relevant chunks
↓ (send to LLM)
LLM generates answer
4.3 Implementation
// internal/portfolio/indexer.go
package portfolio
import (
"context"
"database/sql"
"fmt"
"log/slog"
"os"
"path/filepath"
"strings"
)
type Indexer struct {
dataPath string
db *sql.DB
chunkSize int
chunkOverlap int
}
func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
if err != nil {
return 0, err
}
// Rebuild FTS5 index from scratch (delete + insert is faster than diff for small corpora)
if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
return 0, fmt.Errorf("clear index: %w", err)
}
totalChunks := 0
for _, file := range files {
chunks, err := i.indexFile(ctx, file)
if err != nil {
slog.Warn("failed to index file", "file", file, "err", err)
continue
}
totalChunks += chunks
}
return totalChunks, nil
}
func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
content, err := os.ReadFile(path)
if err != nil {
return 0, err
}
projectID := strings.TrimSuffix(filepath.Base(path), ".md")
chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)
tx, err := i.db.BeginTx(ctx, nil)
if err != nil {
return 0, err
}
defer tx.Rollback()
stmt, err := tx.PrepareContext(ctx, `
INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
VALUES (?, ?, ?, ?, ?)
`)
if err != nil {
return 0, err
}
defer stmt.Close()
for idx, chunk := range chunks {
id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
return idx, err
}
}
if err := tx.Commit(); err != nil {
return 0, err
}
return len(chunks), nil
}
// schema.go — applied at startup
const schema = `
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
id UNINDEXED,
project_id UNINDEXED,
source_file UNINDEXED,
chunk_index UNINDEXED,
content,
tokenize = 'unicode61 remove_diacritics 2'
);
`
func splitIntoChunks(text string, size, overlap int) []string {
// Simple implementation: split by size with overlap
// Production version uses tokenizer-aware chunking
var chunks []string
for i := 0; i < len(text); i += size - overlap {
end := i + size
if end > len(text) {
end = len(text)
}
chunks = append(chunks, text[i:end])
}
return chunks
}
4.4 Retrieval in the agent loop
// internal/portfolio/search.go
package portfolio
type Hit struct {
ProjectID string
SourceFile string
ChunkIndex int
Content string
Score float64 // BM25 score from FTS5
}
func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
// Escape user input: FTS5 syntax can break with special chars
ftsQuery := sanitizeFTS5(query)
rows, err := s.db.QueryContext(ctx, `
SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
FROM portfolio_chunks
WHERE portfolio_chunks MATCH ?
ORDER BY score
LIMIT ?
`, ftsQuery, topK)
if err != nil {
return nil, err
}
defer rows.Close()
var hits []Hit
for rows.Next() {
var h Hit
if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
return nil, err
}
hits = append(hits, h)
}
return hits, rows.Err()
}
// sanitizeFTS5 wraps the user query so reserved chars and unquoted strings don't crash FTS5.
// A pragmatic choice for a Q&A bot: append prefix-match wildcard to each token.
func sanitizeFTS5(q string) string {
tokens := strings.FieldsFunc(q, func(r rune) bool {
return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
(r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
r > 0x7F) // keep accented chars
})
if len(tokens) == 0 {
return `""`
}
for i, t := range tokens {
tokens[i] = `"` + strings.ToLower(t) + `"*`
}
return strings.Join(tokens, " ")
}
// internal/agent/runner.go
package agent
func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
basePrompt := r.persona.SystemPrompt
hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
if err != nil {
return "", err
}
if len(hits) == 0 {
return basePrompt, nil
}
var contextBlock strings.Builder
contextBlock.WriteString(basePrompt)
contextBlock.WriteString("\n\n## Relevant context\n\n")
for _, h := range hits {
contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
h.SourceFile, h.Content))
}
return contextBlock.String(), nil
}
func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
return func(yield func(Chunk, error) bool) {
lastUserMsg := getLastUserMessage(messages)
systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
if err != nil {
yield(Chunk{}, err)
return
}
messages = prependSystem(messages, systemPrompt)
for chunk, err := range r.loop.RunStream(ctx, messages) {
if !yield(chunk, err) {
return
}
}
}
}
Why this is simpler than embeddings:
- No embedding model to download or run (saves ~270MB of RAM and ~200ms per query)
- One file (
data/portfolio.db), one driver, no extra process - BM25 ranking is excellent for keyword-based retrieval over structured docs like project READMEs
- Trade-off: no semantic similarity ("projects about AI" won't match "machine learning" without the literal words). Mitigation:
trigramtokenizer handles morphology well for English/Spanish.
🗜️ 4.5 Auto-compaction
Long conversations eventually run out of context — at qwen2.5-3b's 4k window, the ~3k-token system prompt + RAG block leaves only room for 2–3 user turns. Auto-compaction solves this by folding the older portion of the conversation into a single summary system message when the previous turn's input tokens cross a configurable threshold.
When it fires
agent.Runner.Compact runs once per /api/chat request, before the RAG search. It compares the runner's most recent Usage.InputTokens (reported by the provider in the previous streamed chunk) against client.Capabilities().MaxContextWindow × threshold_ratio.
| Setting | Default | What it controls |
|---|---|---|
compaction.enabled |
false |
Master switch. |
compaction.threshold_ratio |
0.75 |
Trigger when used tokens ≥ window × ratio. |
compaction.keep_recent_turns |
4 |
How many of the latest user turns are kept verbatim after compaction. |
compaction.summary_system_prompt |
(built-in bilingual) | Override the instruction sent to the LLM when summarizing. |
Short-circuits silently when compaction is disabled, the provider doesn't report a window (Capabilities().MaxContextWindow == 0), the history is shorter than keep_recent_turns, or usage is still unknown (first turn).
How the summary is made
splitByTurns(history, keep_recent_turns)divides messages into(older, recent)on user-role boundaries so a kept turn's user/assistant pair always stays together.renderTranscript(older)flattens older messages into aUser:/Assistant:transcript (skipping tool messages and empty assistant placeholders).- The runner calls
client.Generate(...)with the summary prompt + transcript and a 512-token cap so the compaction step itself stays cheap. - The returned text is prepended as a system message (
"Earlier conversation summary:\n…"), followed by the recent tail. LastCompaction()returnsCompactionStatsso the SSE handler can emit acompactionevent right before the streamed chunks.
Failure mode
If Generate errors or returns an empty summary, compaction falls back to truncateToBudget: drop oldest user-turns one at a time until the remaining slice fits threshold tokens (heuristic: len(s) / 4 + 1). The current user turn is always preserved. The fallback is logged at WARN and the request still proceeds — a flaky summarize call never fails the user's request.
Wire protocol
Streaming responses gain an optional compaction event:
data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500}
Emitted after start (when applicable) and before sources / chunk. The widget can render this as a subtle "Context compacted" hint or ignore it — both are valid.
Persistence
Compaction is per-request. The full transcript is still saved to messages in data/portfolio.db verbatim, so GET /api/conversations/{id} always returns the original history. Only what we send to the LLM is reduced — the next session can re-read the full thread from the DB.
🌐 5. Embedding the widget
The bot ships with a drop-in vanilla-JS widget. Add two files to your site and it works.
5.1 The widget (any site)
<link rel="stylesheet" href="/path/to/chat-widget.css">
<script src="/path/to/chat-widget.js"
data-api-url="https://chat.example.com"
data-title="Ask me anything"
data-greeting="Hi! Ask me about the projects."
data-position="bottom-right"
data-theme="auto"
defer></script>
A bubble appears bottom-right, opens a panel, talks SSE to /api/chat, streams the response, and cites sources. No build step, no React/Vue, no framework lock-in.
Browser→bot options:
| Topology | Trade-offs |
|---|---|
| Direct (browser → bot, same domain or CORS) | Simplest. Add the bot's origin to cors_origins in YAML. |
| Reverse proxy (nginx/Caddy in front) | Bot stays on private network, single public domain, no CORS to manage. |
| Site proxies the bot (Astro/Next API route) | Adds a hop and a bit of code, but gives you auth/session hooks in your site. |
The widget works the same in all three. Pick the topology that matches your infra.
Default dev setup is direct + CORS.
cors_originsinconfigs/portfolio-bot.yamlcontrols which sites can call the bot. Add your site's origin there.
5.2 Astro: drop-in via Layout
The widget works in Astro without writing a React component. Add this to your shared layout:
---
// src/layouts/BaseLayout.astro
import "../path/to/chat-widget.css";
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
---
<html>
<body>
<slot />
<script src="/path/to/chat-widget.js"
data-api-url={apiUrl}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
defer is:inline></script>
</body>
</html>
is:inline keeps Astro from hashing/transforming the script tag, so the data-* attributes survive.
5.3 React / Next.js: same script tag
// app/layout.tsx
import Script from "next/script";
export default function RootLayout({ children }) {
return (
<html>
<head>
<link rel="stylesheet" href="/chat-widget.css" />
<Script src="/chat-widget.js"
data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
strategy="afterInteractive" />
</head>
<body>{children}</body>
</html>
);
}
5.4 If you want a server proxy (Astro/Next API route)
The widget can also call a same-origin endpoint that forwards to the bot. This is the right call when you need:
- Auth on
/api/chat(logged-in users only) - Centralized rate limiting at the site level
- Hiding the bot's origin from the browser
// src/pages/api/chat.ts (Astro) or app/api/chat/route.ts (Next)
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";
export const POST = async ({ request }) => {
const body = await request.json();
// (optional) auth check, rate limit, session lookup here
const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
});
return new Response(resp.body, {
status: resp.status,
headers: {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"Connection": "keep-alive",
},
});
};
Then point the widget at /api/chat (same origin) instead of the bot's URL.
5.5 Widget configuration reference
All options are data-* attributes on the <script> tag:
| Attribute | Default | Notes |
|---|---|---|
data-api-url |
(required) | Base URL of the bot. No trailing slash. |
data-title |
"Chat" |
Header text. |
data-greeting |
"" |
First assistant message when the panel opens. |
data-position |
"bottom-right" |
"bottom-right" or "bottom-left". |
data-theme |
"auto" |
"auto" (follows OS), "light", "dark". |
Theming is via CSS custom properties on .rony-chat-widget-root (see web/chat-widget.css):
.rony-chat-widget-root {
--rony-accent: #ff6b35;
--rony-radius: 4px;
--rony-font: "Inter", sans-serif;
}
5.6 What the widget doesn't do (yet)
- Richer markdown (tables, images) — the built-in renderer handles the common cases; for full CommonMark, swap
renderMarkdowninchat-widget.jsformarkedormarkdown-it. - Mobile swipe-to-dismiss — panel goes full-screen on phones.
- Conversation history sidebar — only the active conversation is shown (the backend exposes
GET /api/conversationsfor a future sidebar).
🤖 6. Self-hosting with llama.cpp (default)
6.1 Setup
llama-server is a separate process that the bot connects to over HTTP. Both ports (the bot's and llama-server's) are configurable — pick what fits your environment.
# 1. Make sure you have a GGUF model available
# Download from Hugging Face, e.g.:
# https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
export RONY_MODELS_PATH=/path/to/models
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf
# 2. Start llama-server (port is configurable; default llama.cpp is 8080)
llama-server \
-m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
--port 9100 \
--host 127.0.0.1 \
--ctx-size 4096 \
--mlock # prevents swap, critical on shared VPS
# 3. Make sure configs/portfolio-bot.yaml points to the same port
# providers[0].endpoint: http://localhost:9100/v1
# 4. Start the bot (default port 7331, also configurable)
./bin/chat-bot serve
# → Serves on http://localhost:7331
# → Override with: ./bin/chat-bot serve --port 9101 --host 127.0.0.1
Port reference:
| What | Default | How to change |
|---|---|---|
llama-server HTTP port |
8080 (llama.cpp convention) | --port N flag when starting llama-server |
| chat-bot HTTP port | 7331 | --port N flag on serve, or server.port in YAML |
| chat-bot → llama-server URL | http://localhost:8080/v1 |
endpoint field on the provider in YAML |
The llamacpp provider is imported from rony-llm-agent/pkg/llm/providers/llamacpp and is compiled against llama.cpp via CGO or external binary.
6.2 Alternative: Ollama (easier for development)
If you don't want to manage GGUF files manually, Ollama provides the same models with a simpler workflow:
# 1. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# 2. Download chat model
ollama pull qwen2.5:1.5b
# 3. Verify
ollama list
# 4. Edit configs/portfolio-bot.yaml to mark ollama-local as default:
# providers[0].default: true (and remove default from llamacpp-local)
# Ollama exposes an OpenAI-compatible API on :11434/v1
# 5. Start the bot
ollama serve &
./bin/chat-bot serve
6.3 Alternative: llama.cpp direct (advanced)
For more control or if Ollama doesn't work in your setup:
providers:
- name: llamacpp-local
type: llamacpp
model: qwen2.5-3b-instruct
endpoint: http://localhost:9100/v1 # configurable, see §6.1
context_size: 4096
max_tokens: 2048
default: true
The llamacpp adapter is imported from rony-llm-agent/pkg/llm/providers/llamacpp and is compiled against llama.cpp via CGO or external binary.
📦 7. Bot CLI
7.1 Commands
# Start HTTP server
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]
# Re-index portfolio (reads data/projects/*.md → SQLite FTS5)
chat-bot reindex
# Single question (no server, useful for tests)
chat-bot ask "What projects does Victor have?" [--no-rag]
# Validate config
chat-bot config validate
# Health check (useful for monitoring)
chat-bot health
# Version
chat-bot version
7.2 Implementation with Cobra
// cmd/chat-bot/main.go
package main
import (
"github.com/spf13/cobra"
)
func main() {
root := &cobra.Command{
Use: "chat-bot",
Short: "Portfolio chatbot HTTP server",
}
root.AddCommand(serveCmd())
root.AddCommand(reindexCmd())
root.AddCommand(askCmd())
root.AddCommand(configCmd())
root.AddCommand(healthCmd())
root.AddCommand(versionCmd())
if err := root.Execute(); err != nil {
os.Exit(1)
}
}
func serveCmd() *cobra.Command {
var port int
var host string
var reindexOnStart bool
cmd := &cobra.Command{
Use: "serve",
Short: "Start HTTP server",
RunE: func(cmd *cobra.Command, args []string) error {
return server.Serve(server.Config{
Port: port,
Host: host,
ReindexOnStart: reindexOnStart,
})
},
}
cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
return cmd
}
🚀 8. Deployment
8.1 Recommendation: Self-hosted on VPS
# 1. Install dependencies
sudo apt install golang-go ollama
ollama pull qwen2.5:1.5b
# 2. Build
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot
# 3. systemd service
cat > /etc/systemd/system/chat-bot.service <<EOF
[Unit]
Description=Portfolio Chat Bot
After=network.target ollama.service
[Service]
Type=simple
User=chatbot
WorkingDirectory=/opt/chat-bot
ExecStart=/usr/local/bin/chat-bot serve
Restart=on-failure
Environment=RONY_MODELS_PATH=/opt/models
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl enable --now chat-bot
8.2 Reverse proxy (Caddy)
# /etc/caddy/Caddyfile
chat.victorvargas.dev {
reverse_proxy localhost:7331
}
8.3 Monitoring
# Health check periodic
curl -s http://localhost:7331/api/health | jq
# Logs
journalctl -u chat-bot -f
🧪 9. Testing
9.1 Unit tests
// internal/server/chat_test.go
package server
func TestHandleChat_ValidRequest(t *testing.T) {
s := newTestServer(t)
req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
"messages": [{"role": "user", "content": "hello"}]
}`))
req.Header.Set("Content-Type", "application/json")
w := httptest.NewRecorder()
s.handleChat(w, req)
assert.Equal(t, 200, w.Code)
assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
}
func TestHandleChat_RateLimit(t *testing.T) {
s := newTestServerWithConfig(t, server.Config{
RateLimit: 1, // 1 request per minute
})
// First request OK
req1 := newChatRequest("hello")
w1 := httptest.NewRecorder()
s.handleChat(w1, req1)
assert.Equal(t, 200, w1.Code)
// Second request denied
req2 := newChatRequest("hello again")
w2 := httptest.NewRecorder()
s.handleChat(w2, req2)
assert.Equal(t, 429, w2.Code)
}
9.2 Integration tests with mock LLM
// internal/agent/runner_test.go
func TestRunner_RAGContextIsInjected(t *testing.T) {
mockLLM := mock.New(mock.Responses{
{Match: "projects", Response: "Victor has several projects..."},
})
memory := newMockMemoryWithDocs(t, []rag.Fragment{
{Content: "Rony Harness: AI agent harness...", ProjectID: "rony-harness"},
{Content: "rony-llm-agent: Go library...", ProjectID: "rony-llm-agent"},
})
runner := agent.NewRunner(agent.Config{
LLM: mockLLM,
Memory: memory,
Persona: testPersona,
})
resp, _ := runner.Run(context.Background(), []llm.Message{
{Role: llm.RoleUser, Content: "what projects does Victor have?"},
})
// Verify LLM received context chunks in system prompt
lastReq := mockLLM.LastRequest()
assert.Contains(t, lastReq.Messages[0].Content, "Rony Harness")
assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
}
9.3 E2E test with Astro
# 1. Start chat-bot on :7331
./bin/chat-bot serve &
# 2. Start Astro on :4321
cd ../portfolio && npm run dev &
# 3. Make request to Astro's proxy
curl -X POST http://localhost:4321/api/chat \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"hello"}]}'
# 4. Verify SSE stream
📂 10. Project Structure
rony-chat-bot/
├── cmd/
│ └── chat-bot/
│ └── main.go # CLI entrypoint
│
├── internal/
│ ├── server/ # HTTP handlers
│ │ ├── server.go # chi router + middleware
│ │ ├── handlers.go # /api/chat, /api/health, /api/info, /api/reindex, /api/conversations
│ │ ├── conversations_test.go # round-trip, continue, list, 404, delete, streaming
│ │ └── middleware.go # RequestID, Logging, CORS, RateLimit
│ │
│ ├── agent/ # LLM client + RAG runner
│ │ ├── runner.go # Stream wrapper, RAG injection into system prompt
│ │ └── client.go # NewClient factory: llamacpp / ollama / openai / anthropic
│ │
│ ├── portfolio/ # RAG: markdown → SQLite FTS5 + conversation persistence
│ │ ├── chunker.go # Heading-based splitter
│ │ ├── indexer.go # Store: schema, Reindex, Search (BM25)
│ │ ├── conversations.go # Conversation + Message CRUD, persisted alongside RAG
│ │ └── chunker_test.go / store_test.go
│ │
│ ├── persona/ # Persona bridge to rony-llm-agent
│ │ └── persona.go # FromConfig, BuildSystemPrompt (with RAG context)
│ │
│ ├── streaming/ # SSE protocol helpers
│ │ └── sse.go # WriteStart/Chunk/Sources/Done/Error
│ │
│ ├── i18n/ # Language detection (ES/EN) for the response
│ │
│ └── config/ # YAML loader + validation
│
├── web/ # ← DROP-IN CHAT WIDGET
│ ├── chat-widget.js # Vanilla JS, ~12 KB
│ ├── chat-widget.css # Scoped styles, CSS-custom-prop themable
│ ├── example.html # Local demo (python -m http.server)
│ └── README.md # Integration guide (HTML, Astro, Next.js)
│
├── data/
│ └── projects/ # ← Markdown per project (one .md per project)
│ ├── rony-harness.md
│ ├── rony-llm-agent.md
│ └── example-project.md
│
├── configs/
│ └── portfolio-bot.yaml # Provider + RAG + persona config
│
├── docs/
│ ├── architecture.md # ← THIS FILE
│ └── architecture.es.md
│
├── bench/ # Reproducible SQLite driver benchmark
│
├── go.mod # require rony-llm-agent, modernc.org/sqlite
└── README.md
📅 11. Roadmap
Phase 1: MVP (2-3 weeks)
- Project setup (
go mod init, structure) - Basic HTTP server with
/api/chatendpoint - Functional SSE streaming
- RAG indexer (reads
data/projects/*.md→ SQLite FTS5) - RAG retriever (query → top-k chunks)
- Persona loader from YAML
- llama.cpp integration (qwen2.5:1.5b GGUF)
- CLI:
serve,reindex,ask - Basic tests
Phase 2: Integration with Astro (1 week)
- Astro API route of the proxy
- React component of the chat widget
- E2E test: Astro → chat-bot → response
- Widget styling (TailwindCSS)
Phase 3: Polish (1 week)
- Robust rate limiting
- Structured logging (JSON)
- Health checks for monitoring
- systemd service file
- README + deployment docs
Phase 4: Optionals
- Support for multiple conversations (session ID)
- Persisted chat history
- Analysis of frequent questions
- Multi-language (EN/ES switch)
- More polished standalone CLI version (
chat-bot ask)
📐 12. Quality Specifications
12.1 Performance metrics
| Metric | Target |
|---|---|
| TTFT (Time-to-first-token) | <500ms with llama.cpp local |
| End-to-end (question → complete response) | <3s for typical responses |
| Memory at rest | <150MB |
| RAG indexing speed | ~100 docs/second |
| Retrieval latency | <50ms for top-5 |
12.2 Required tests
- Unit tests: coverage ≥70%
- Integration tests: with mock LLM + in-memory SQLite FTS5
- E2E: at least one complete Astro → chat-bot flow
🔒 13. Security
13.1 Implemented
- Rate limiting per IP (default 30 req/min)
- Restrictive CORS — only configured origins
- Input validation — JSON schema validation on requests
- No PII storage — we don't save conversations by default
- Local-only by default — no calls to cloud APIs
13.2 Deferred / Optional
- Auth with API key (for private use)
- Query logging for analytics
- IP anonymization in logs
- HTTPS via reverse proxy (Caddy/nginx)
📚 14. References
- SSE Spec: https://html.spec.whatwg.org/multipage/server-sent-events.html
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
- SQLite FTS5: https://www.sqlite.org/fts5.html
- Go SQLite driver: https://github.com/mattn/go-sqlite3 (CGO) or https://modernc.org/sqlite (pure Go)
- qwen2.5: https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
- Astro API routes: https://docs.astro.build/en/guides/endpoints/
- rony-llm-agent: https://github.com/VictorVargas/rony-llm-agent
Document ready for implementation. 🚀