Adds the §4.5 Auto-compaction section to architecture.md and architecture.es.md describing the trigger, fallback, persistence and the new SSE 'compaction' event so consumers know how to react. Includes the three placeholder projects used while exercising the feature end to end (bot-onboarding, dashboard-metricas, tienda-ropa) so /api/reindex picks them up without further setup.
1398 lines
No EOL
45 KiB
Markdown
1398 lines
No EOL
45 KiB
Markdown
# 📋 Rony Chat Bot — Technical Design Document
|
||
|
||
**Version:** 1.0
|
||
**Author:** Victor Hugo Vargas
|
||
**Date:** 2026-06-28
|
||
**Status:** Complete specification for implementation
|
||
**Path:** `rony-chat-bot/docs/architecture.md`
|
||
|
||
> 🌐 **Language:** [English](./architecture.md) | [Español](./architecture.es.md)
|
||
>
|
||
> 📚 **Workspace:** This project is part of the `Rony/` workspace. See [`../README.md`](../../README.md).
|
||
>
|
||
> 🔑 **Depends on:** [`rony-llm-agent`](https://github.com/VictorVargas/rony-llm-agent) — core library that provides agent loop, LLM clients, RAG, persona system.
|
||
>
|
||
> 📐 **Methodology:** This project follows the **SDD + DDD + Hexagonal Architecture** approach. Functional Requirements are numbered as `CRF-XXX`. See [`../../METHODOLOGY.md`](../../METHODOLOGY.md).
|
||
|
||
---
|
||
|
||
## 🎯 1. Project Vision
|
||
|
||
### 1.1 What is Chat-Bot?
|
||
|
||
An **HTTP chatbot** that answers questions about Victor Hugo Vargas and his projects. Uses **RAG (Retrieval-Augmented Generation)** over markdown files describing each project, and a local LLM (or cloud) to generate responses.
|
||
|
||
### 1.2 Primary use case
|
||
|
||
Victor has a portfolio website (Astro + React). On the site there's a chat widget where visitors can ask:
|
||
|
||
- "What projects has Victor done?"
|
||
- "What's his experience with Go?"
|
||
- "How does Rony Harness work?"
|
||
- "Has Victor worked with PostgreSQL?"
|
||
|
||
The bot responds with accurate information extracted from the projects' markdown files + bio + skills.
|
||
|
||
### 1.3 Secondary use cases (future)
|
||
|
||
- **Client adaptation:** The same bot, with other data and another persona, serves car dealerships, restaurants, etc.
|
||
- **Standalone CLI:** `./chat-bot ask "what do you know about X?"` for terminal use.
|
||
- **Slack/Discord bot:** Wrapper that consumes the HTTP API.
|
||
|
||
### 1.4 Philosophy
|
||
|
||
- **Self-hosted by default** — works 100% local with Ollama + 1-3B models
|
||
- **Cloud optional** — if you need more quality, swap to Anthropic API
|
||
- **Portable** — easy to fork/customize for other contexts
|
||
- **Streaming** — token-by-token responses with SSE (no waiting for complete response)
|
||
- **Reuses `rony-llm-agent`** — doesn't reinvent the agent loop
|
||
|
||
---
|
||
|
||
## 🏗️ 2. Architecture
|
||
|
||
### 2.1 Overview
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────────────────┐
|
||
│ Browser (Astro site) │
|
||
│ ↓ HTTP POST /api/chat │
|
||
│ Astro SSR (proxy) ←────────── Serves portfolio + proxy chat │
|
||
│ ↓ HTTP POST /api/chat │
|
||
│ Chat-Bot HTTP server (:7331) │
|
||
│ ↓ │
|
||
│ Agent loop (rony-llm-agent) │
|
||
│ ↓ │
|
||
│ RAG retrieval → SQLite FTS5 over data/projects/*.md │
|
||
│ ↓ │
|
||
│ LLM (llama.cpp local default / Ollama or Anthropic optional) │
|
||
└─────────────────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
### 2.2 Main components
|
||
|
||
| Component | Path | Responsibility |
|
||
|---|---|---|
|
||
| **HTTP server** | `internal/server/` | Gin/chi handlers, SSE streaming |
|
||
| **Agent runner** | `internal/agent/` | Wrapper over `rony-llm-agent` with specific config |
|
||
| **Portfolio loader** | `internal/portfolio/` | Reads `data/projects/*.md`, indexes in SQLite FTS5 |
|
||
| **Persona** | `internal/persona/` | Loads persona from `configs/portfolio-bot.yaml` |
|
||
| **CLI** | `cmd/chat-bot/` | Commands: `serve`, `reindex`, `ask`, `version` |
|
||
|
||
### 2.3 Tech stack
|
||
|
||
| Layer | Technology | Reason |
|
||
|---|---|---|
|
||
| **Language** | Go 1.26+ | Same as rony-harness, leverage `os.Root`, `iter.Seq` |
|
||
| **HTTP router** | `net/http` + `chi` | Stdlib + chi for middleware (CORS, logging) |
|
||
| **SSE** | `net/http` Flusher | Stdlib is enough, no external library needed |
|
||
| **Config** | `gopkg.in/yaml.v3` | Same as harness |
|
||
| **RAG backend** | SQLite + FTS5 (BM25) | Zero external deps, single file, fast |
|
||
| **LLM** | llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama as alt | Self-hosted by default |
|
||
| **Tests** | stdlib + testify | Consistency with the rest |
|
||
|
||
---
|
||
|
||
## 🔌 3. HTTP API
|
||
|
||
### 3.1 Endpoints
|
||
|
||
#### `POST /api/chat` — Chat with SSE streaming
|
||
|
||
**Request:**
|
||
```json
|
||
{
|
||
"messages": [
|
||
{"role": "user", "content": "What projects does Victor have?"}
|
||
],
|
||
"stream": true,
|
||
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e"
|
||
}
|
||
```
|
||
|
||
| Field | Required | Notes |
|
||
|---|---|---|
|
||
| `messages` | yes | At least one user message; alternation is not enforced. |
|
||
| `stream` | no, default `true` | `false` returns a single JSON body instead of SSE. |
|
||
| `conversation_id` | no | Hex string. If omitted, the server mints a new one and returns it (see below). Pass an existing ID to keep the thread. |
|
||
|
||
**Response (SSE):**
|
||
```
|
||
data: {"type":"start","conversation_id":"57f4aa3c7fab466bc4de9c43b296903e"}
|
||
|
||
data: {"type":"chunk","content":"Victor"}
|
||
data: {"type":"chunk","content":" has"}
|
||
data: {"type":"chunk","content":" several"}
|
||
data: {"type":"chunk","content":" projects"}
|
||
|
||
data: {"type":"sources","documents":["rony-harness.md","rony-llm-agent.md"]}
|
||
|
||
data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}
|
||
```
|
||
|
||
The `conversation_id` in the `start` event is what the client should store
|
||
(see §3.4 — *Conversation persistence*). When the client passed an
|
||
existing ID the server echoes it back; otherwise it's freshly minted.
|
||
|
||
**Without streaming** (`"stream": false`):
|
||
```json
|
||
{
|
||
"conversation_id": "57f4aa3c7fab466bc4de9c43b296903e",
|
||
"content": "Victor has several projects...",
|
||
"sources": ["rony-harness.md", "rony-llm-agent.md"],
|
||
"usage": {"input_tokens": 245, "output_tokens": 38}
|
||
}
|
||
```
|
||
|
||
#### `GET /api/conversations` — List recent conversations
|
||
|
||
Returns the most recent conversation summaries, newest first. Useful for a
|
||
"show my chats" sidebar in a custom UI.
|
||
|
||
**Query params:**
|
||
- `limit` (1–200, default 50)
|
||
|
||
**Response:**
|
||
```json
|
||
{
|
||
"count": 2,
|
||
"conversations": [
|
||
{
|
||
"id": "57f4aa3c7fab466bc4de9c43b296903e",
|
||
"created_at": "2026-07-17T05:02:07Z",
|
||
"updated_at": "2026-07-17T05:04:31Z",
|
||
"preview": "What projects does Victor have?"
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
#### `GET /api/conversations/{id}` — Fetch one conversation
|
||
|
||
Returns the full history of a conversation with all messages in
|
||
chronological order.
|
||
|
||
**Response (200):**
|
||
```json
|
||
{
|
||
"id": "57f4aa3c7fab466bc4de9c43b296903e",
|
||
"created_at": "2026-07-17T05:02:07Z",
|
||
"updated_at": "2026-07-17T05:04:31Z",
|
||
"messages": [
|
||
{"id": 1, "role": "user", "content": "What projects does Victor have?", "created_at": "..."},
|
||
{"id": 2, "role": "assistant", "content": "Victor has several projects...", "sources": ["..."], "created_at": "..."}
|
||
]
|
||
}
|
||
```
|
||
|
||
**Response (404):** when the ID is unknown (e.g. server DB was wiped or the
|
||
client lost sync). The widget treats this as "start fresh".
|
||
|
||
> ⚠️ **Auth note:** the conversation ID is the only access token. For a
|
||
> public bot this is fine; for private contexts add auth at the proxy layer
|
||
> (e.g. require a session cookie before forwarding to this endpoint).
|
||
|
||
#### `DELETE /api/conversations/{id}` — Delete a conversation
|
||
|
||
Removes the conversation and all its messages (cascade). Returns 204 on
|
||
success, 404 if the ID doesn't exist.
|
||
|
||
#### `POST /api/reindex` — Re-index portfolio
|
||
|
||
Useful when files in `data/projects/` are modified.
|
||
|
||
**Request:** empty
|
||
**Response:**
|
||
```json
|
||
{
|
||
"indexed_files": 12,
|
||
"total_chunks": 87,
|
||
"duration_ms": 4321
|
||
}
|
||
```
|
||
|
||
#### `GET /api/health` — Health check (real)
|
||
|
||
Probes the LLM provider and the SQLite store in parallel and returns their
|
||
states. Designed for monitoring/load balancers. **Returns 200 when healthy
|
||
or degraded, 503 when unhealthy.**
|
||
|
||
- `?deep=true` adds a chunk count to the store probe (same latency budget).
|
||
|
||
**Status taxonomy:**
|
||
|
||
| `status` | HTTP | Meaning |
|
||
|---|---|---|
|
||
| `healthy` | 200 | LLM up, store up |
|
||
| `degraded` | 200 | LLM up, store down — bot still answers, just without RAG |
|
||
| `unhealthy` | 503 | LLM down — bot cannot answer, no point routing traffic here |
|
||
|
||
**Probe details:**
|
||
|
||
| Component | Probe | Latency |
|
||
|---|---|---|
|
||
| `llm` | `GET {provider}/health` (llamacpp, ollama) or `/models` (openai) | ~1ms for local llama-server |
|
||
| `store` | `SELECT 1` on the SQLite handle | ~100µs |
|
||
|
||
Each probe has a 2s timeout; the whole call returns within ~2.5s even if a
|
||
dependency hangs.
|
||
|
||
**Response shape (healthy):**
|
||
|
||
```json
|
||
{
|
||
"status": "healthy",
|
||
"version": "0.2.0-dev",
|
||
"checked_at": "2026-07-17T05:02:07Z",
|
||
"components": {
|
||
"llm": {
|
||
"status": "up",
|
||
"latency": "1.028ms",
|
||
"details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
|
||
},
|
||
"store": {
|
||
"status": "up",
|
||
"latency": "107µs"
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
**Response shape (degraded, with `?deep=true`):**
|
||
|
||
```json
|
||
{
|
||
"status": "degraded",
|
||
"version": "0.2.0-dev",
|
||
"checked_at": "2026-07-17T05:02:07Z",
|
||
"components": {
|
||
"llm": {"status": "up", "latency": "0.8ms", "details": {...}},
|
||
"store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
|
||
}
|
||
}
|
||
```
|
||
|
||
**Response shape (unhealthy):** HTTP 503, same JSON with `"status": "unhealthy"` and the failed component reporting `"status": "down"` plus an `error` field.
|
||
|
||
#### `GET /api/info` — Bot metadata
|
||
|
||
```json
|
||
{
|
||
"name": "Rony Chat Bot",
|
||
"model": "qwen2.5:1.5b",
|
||
"persona": "...",
|
||
"topics": ["projects", "experience", "technical skills"]
|
||
}
|
||
```
|
||
|
||
### 3.2 SSE Implementation
|
||
|
||
```go
|
||
// internal/server/chat.go
|
||
package server
|
||
|
||
import (
|
||
"encoding/json"
|
||
"fmt"
|
||
"net/http"
|
||
"github.com/VictorVargas/rony-llm-agent/pkg/agent"
|
||
)
|
||
|
||
func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
|
||
// SSE headers
|
||
w.Header().Set("Content-Type", "text/event-stream")
|
||
w.Header().Set("Cache-Control", "no-cache")
|
||
w.Header().Set("Connection", "keep-alive")
|
||
w.Header().Set("X-Accel-Buffering", "no")
|
||
|
||
flusher, ok := w.(http.Flusher)
|
||
if !ok {
|
||
http.Error(w, "SSE not supported", http.StatusInternalServerError)
|
||
return
|
||
}
|
||
|
||
// Parse request
|
||
var req ChatRequest
|
||
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||
writeError(w, flusher, "invalid request", err)
|
||
return
|
||
}
|
||
|
||
// Start event
|
||
writeSSE(w, flusher, "start", map[string]string{
|
||
"conversation_id": generateConvID(),
|
||
})
|
||
|
||
// Run agent with streaming
|
||
sources := []string{}
|
||
for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
|
||
if err != nil {
|
||
writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
|
||
return
|
||
}
|
||
if chunk.Type == "source" {
|
||
sources = append(sources, chunk.Source)
|
||
}
|
||
writeSSE(w, flusher, chunk.Type, chunk.Data)
|
||
}
|
||
|
||
// Done event
|
||
writeSSE(w, flusher, "done", map[string]any{
|
||
"usage": map[string]int{
|
||
"input_tokens": 245,
|
||
"output_tokens": 38,
|
||
},
|
||
})
|
||
}
|
||
|
||
func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
|
||
payload, _ := json.Marshal(data)
|
||
fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
|
||
flusher.Flush()
|
||
}
|
||
```
|
||
|
||
### 3.3 Middleware
|
||
|
||
```go
|
||
// internal/server/middleware.go
|
||
package server
|
||
|
||
func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
start := time.Now()
|
||
// Wrap response writer to capture status
|
||
rw := &statusRecorder{ResponseWriter: w, status: 200}
|
||
next.ServeHTTP(rw, r)
|
||
|
||
slog.Info("http.request",
|
||
"method", r.Method,
|
||
"path", r.URL.Path,
|
||
"status", rw.status,
|
||
"duration_ms", time.Since(start).Milliseconds(),
|
||
"ip", r.RemoteAddr,
|
||
)
|
||
})
|
||
}
|
||
|
||
func (s *Server) corsMiddleware(next http.Handler) http.Handler {
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
origin := r.Header.Get("Origin")
|
||
for _, allowed := range s.config.Server.CORSOrigins {
|
||
if origin == allowed {
|
||
w.Header().Set("Access-Control-Allow-Origin", origin)
|
||
w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
|
||
w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
|
||
break
|
||
}
|
||
}
|
||
if r.Method == "OPTIONS" {
|
||
w.WriteHeader(204)
|
||
return
|
||
}
|
||
next.ServeHTTP(w, r)
|
||
})
|
||
}
|
||
|
||
func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
|
||
limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
if !limiter.Allow() {
|
||
http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
|
||
return
|
||
}
|
||
next.ServeHTTP(w, r)
|
||
})
|
||
}
|
||
```
|
||
|
||
### 3.4 Conversation persistence
|
||
|
||
The bot persists conversation threads in the same SQLite database as the
|
||
RAG index (`./data/portfolio.db`). Schema lives in `internal/portfolio/conversations.go`.
|
||
|
||
```sql
|
||
CREATE TABLE conversations (
|
||
id TEXT PRIMARY KEY, -- 16-byte random hex (32 chars)
|
||
created_at INTEGER NOT NULL,
|
||
updated_at INTEGER NOT NULL
|
||
);
|
||
|
||
CREATE TABLE messages (
|
||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||
conversation_id TEXT NOT NULL,
|
||
role TEXT NOT NULL, -- user | assistant | system
|
||
content TEXT NOT NULL,
|
||
sources TEXT, -- JSON array, nullable
|
||
created_at INTEGER NOT NULL,
|
||
FOREIGN KEY (conversation_id) REFERENCES conversations(id) ON DELETE CASCADE
|
||
);
|
||
CREATE INDEX idx_messages_conv ON messages(conversation_id, id);
|
||
```
|
||
|
||
**Lifecycle:**
|
||
|
||
| When | What |
|
||
|---|---|
|
||
| `POST /api/chat` (no `conversation_id`) | Server mints a new hex ID, returns it in the `start` SSE event (or `conversation_id` field of the JSON response) |
|
||
| `POST /api/chat` (with `conversation_id`) | Server reuses the existing row; both user message and assistant reply are appended |
|
||
| User message | Persisted **before** the LLM runs, so it survives a model failure |
|
||
| Assistant message | Persisted **after** the stream completes, with the RAG sources attached |
|
||
| `GET /api/conversations/{id}` | Returns the full thread; 404 if unknown |
|
||
| `DELETE /api/conversations/{id}` | Cascade-deletes messages |
|
||
|
||
**Client responsibilities:**
|
||
|
||
1. On the first message, omit `conversation_id`. Capture the one the server
|
||
returns in the `start` SSE event.
|
||
2. Store it client-side (`localStorage["rony-chat-conv"]` in the widget).
|
||
3. On every subsequent message, send the ID back.
|
||
4. On page load, if you have a stored ID, call `GET /api/conversations/{id}`
|
||
to restore the thread. If 404, clear the stored ID and start fresh.
|
||
|
||
The widget (`web/chat-widget.js`) implements all four steps. Any other
|
||
client (a custom React component, an Astro endpoint, a CLI replay tool)
|
||
follows the same protocol.
|
||
|
||
**Auth model:**
|
||
|
||
The conversation ID is the only access token for `GET /api/conversations/{id}`.
|
||
It is 128 bits of random entropy, so guessing one is infeasible. For a
|
||
public portfolio bot this is the right trade-off — anyone who knows the
|
||
URL can read its history. For private contexts, add an auth layer in front
|
||
of the bot (proxy) that gates the conversation endpoints.
|
||
|
||
---
|
||
|
||
## 🧠 4. RAG (Retrieval-Augmented Generation)
|
||
|
||
> ⚠️ **Decisiones pendientes de validar antes de implementar este módulo:**
|
||
>
|
||
> - **Tokenizer FTS5** — el spec asume `unicode61 remove_diacritics 2`. Confirmar con datos reales si conviene cambiar a `porter` (stemming EN), `trigram` (sub-string matching) o un tokenizer custom para español. **Validar:** ejecutar queries representativas contra `data/projects/` y comparar recall antes de cerrar esta elección.
|
||
> - **Driver SQLite** — ✅ **DECIDIDO: `modernc.org/sqlite`** (puro Go, sin CGO). Ver benchmark abajo.
|
||
> - **Chunking** — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. **Validar:** medir recall con chunks por sección markdown (split por `#`/`##`) vs por tamaño.
|
||
> - **Sin similitud semántica** — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. **Validar:** tamaño del corpus y types of questions esperadas; si el corpus crece o las queries se vuelven abstractas, considerar agregar embeddings como capa secundaria.
|
||
|
||
### 4.0 Driver decision: benchmark results
|
||
|
||
Reproducible con `CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/`. Datos: 4 markdowns → 11 chunks.
|
||
|
||
| Operación | mattn (CGO) | modernc (puro Go) | Diferencia |
|
||
|---|---|---|---|
|
||
| **Insert** (11 chunks) | 2,802,843 ns/op | **1,465,646 ns/op** | modernc 1.9× más rápido |
|
||
| Insert alloc | 2,124,299 B/op | **9,770 B/op** | modernc usa 217× menos memoria |
|
||
| **Query** (8 queries BM25) | **244,047 ns/op** | 555,162 ns/op | mattn 2.3× más rápido |
|
||
| **Round-trip** (insert + 8 queries) | 3,543,417 ns/op | **2,267,669 ns/op** | modernc 1.6× más rápido |
|
||
| Binary size | 11 MB | 11 MB | igual |
|
||
| Build deps | gcc, CGO=1 | nada | modernc gana |
|
||
| CI/CD portable | requiere toolchain C | `go build` puro | modernc gana |
|
||
|
||
**Decisión: `modernc.org/sqlite`**.
|
||
|
||
Justificación:
|
||
1. Ambas latencias de query (~250µs vs ~550µs) son **2 órdenes de magnitud por debajo** del target de 50ms — imperceptible vs el LLM (varios segundos).
|
||
2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
|
||
3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
|
||
4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.
|
||
|
||
### 4.1 Indexing pipeline
|
||
|
||
```
|
||
data/projects/*.md
|
||
↓ (read all files)
|
||
Raw markdown content
|
||
↓ (split into chunks, ~500 chars, 50 overlap)
|
||
Chunks []
|
||
↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
|
||
Indexed corpus
|
||
```
|
||
|
||
**When it runs:**
|
||
- On bot startup (if `--reindex-on-start` flag)
|
||
- Manually: `./chat-bot reindex`
|
||
- Via HTTP: `POST /api/reindex`
|
||
|
||
### 4.2 Retrieval pipeline
|
||
|
||
```
|
||
User query "what projects does Victor have?"
|
||
↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
|
||
Top 5 relevant chunks
|
||
↓ (format as context block)
|
||
System prompt += relevant chunks
|
||
↓ (send to LLM)
|
||
LLM generates answer
|
||
```
|
||
|
||
### 4.3 Implementation
|
||
|
||
```go
|
||
// internal/portfolio/indexer.go
|
||
package portfolio
|
||
|
||
import (
|
||
"context"
|
||
"database/sql"
|
||
"fmt"
|
||
"log/slog"
|
||
"os"
|
||
"path/filepath"
|
||
"strings"
|
||
)
|
||
|
||
type Indexer struct {
|
||
dataPath string
|
||
db *sql.DB
|
||
chunkSize int
|
||
chunkOverlap int
|
||
}
|
||
|
||
func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
|
||
files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
|
||
if err != nil {
|
||
return 0, err
|
||
}
|
||
|
||
// Rebuild FTS5 index from scratch (delete + insert is faster than diff for small corpora)
|
||
if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
|
||
return 0, fmt.Errorf("clear index: %w", err)
|
||
}
|
||
|
||
totalChunks := 0
|
||
for _, file := range files {
|
||
chunks, err := i.indexFile(ctx, file)
|
||
if err != nil {
|
||
slog.Warn("failed to index file", "file", file, "err", err)
|
||
continue
|
||
}
|
||
totalChunks += chunks
|
||
}
|
||
|
||
return totalChunks, nil
|
||
}
|
||
|
||
func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
|
||
content, err := os.ReadFile(path)
|
||
if err != nil {
|
||
return 0, err
|
||
}
|
||
|
||
projectID := strings.TrimSuffix(filepath.Base(path), ".md")
|
||
chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)
|
||
|
||
tx, err := i.db.BeginTx(ctx, nil)
|
||
if err != nil {
|
||
return 0, err
|
||
}
|
||
defer tx.Rollback()
|
||
|
||
stmt, err := tx.PrepareContext(ctx, `
|
||
INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
|
||
VALUES (?, ?, ?, ?, ?)
|
||
`)
|
||
if err != nil {
|
||
return 0, err
|
||
}
|
||
defer stmt.Close()
|
||
|
||
for idx, chunk := range chunks {
|
||
id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
|
||
if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
|
||
return idx, err
|
||
}
|
||
}
|
||
|
||
if err := tx.Commit(); err != nil {
|
||
return 0, err
|
||
}
|
||
return len(chunks), nil
|
||
}
|
||
|
||
// schema.go — applied at startup
|
||
const schema = `
|
||
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
|
||
id UNINDEXED,
|
||
project_id UNINDEXED,
|
||
source_file UNINDEXED,
|
||
chunk_index UNINDEXED,
|
||
content,
|
||
tokenize = 'unicode61 remove_diacritics 2'
|
||
);
|
||
`
|
||
|
||
func splitIntoChunks(text string, size, overlap int) []string {
|
||
// Simple implementation: split by size with overlap
|
||
// Production version uses tokenizer-aware chunking
|
||
var chunks []string
|
||
for i := 0; i < len(text); i += size - overlap {
|
||
end := i + size
|
||
if end > len(text) {
|
||
end = len(text)
|
||
}
|
||
chunks = append(chunks, text[i:end])
|
||
}
|
||
return chunks
|
||
}
|
||
```
|
||
|
||
### 4.4 Retrieval in the agent loop
|
||
|
||
```go
|
||
// internal/portfolio/search.go
|
||
package portfolio
|
||
|
||
type Hit struct {
|
||
ProjectID string
|
||
SourceFile string
|
||
ChunkIndex int
|
||
Content string
|
||
Score float64 // BM25 score from FTS5
|
||
}
|
||
|
||
func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
|
||
// Escape user input: FTS5 syntax can break with special chars
|
||
ftsQuery := sanitizeFTS5(query)
|
||
|
||
rows, err := s.db.QueryContext(ctx, `
|
||
SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
|
||
FROM portfolio_chunks
|
||
WHERE portfolio_chunks MATCH ?
|
||
ORDER BY score
|
||
LIMIT ?
|
||
`, ftsQuery, topK)
|
||
if err != nil {
|
||
return nil, err
|
||
}
|
||
defer rows.Close()
|
||
|
||
var hits []Hit
|
||
for rows.Next() {
|
||
var h Hit
|
||
if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
|
||
return nil, err
|
||
}
|
||
hits = append(hits, h)
|
||
}
|
||
return hits, rows.Err()
|
||
}
|
||
|
||
// sanitizeFTS5 wraps the user query so reserved chars and unquoted strings don't crash FTS5.
|
||
// A pragmatic choice for a Q&A bot: append prefix-match wildcard to each token.
|
||
func sanitizeFTS5(q string) string {
|
||
tokens := strings.FieldsFunc(q, func(r rune) bool {
|
||
return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
|
||
(r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
|
||
r > 0x7F) // keep accented chars
|
||
})
|
||
if len(tokens) == 0 {
|
||
return `""`
|
||
}
|
||
for i, t := range tokens {
|
||
tokens[i] = `"` + strings.ToLower(t) + `"*`
|
||
}
|
||
return strings.Join(tokens, " ")
|
||
}
|
||
```
|
||
|
||
```go
|
||
// internal/agent/runner.go
|
||
package agent
|
||
|
||
func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
|
||
basePrompt := r.persona.SystemPrompt
|
||
|
||
hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
|
||
if err != nil {
|
||
return "", err
|
||
}
|
||
if len(hits) == 0 {
|
||
return basePrompt, nil
|
||
}
|
||
|
||
var contextBlock strings.Builder
|
||
contextBlock.WriteString(basePrompt)
|
||
contextBlock.WriteString("\n\n## Relevant context\n\n")
|
||
for _, h := range hits {
|
||
contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
|
||
h.SourceFile, h.Content))
|
||
}
|
||
return contextBlock.String(), nil
|
||
}
|
||
|
||
func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
|
||
return func(yield func(Chunk, error) bool) {
|
||
lastUserMsg := getLastUserMessage(messages)
|
||
systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
|
||
if err != nil {
|
||
yield(Chunk{}, err)
|
||
return
|
||
}
|
||
|
||
messages = prependSystem(messages, systemPrompt)
|
||
|
||
for chunk, err := range r.loop.RunStream(ctx, messages) {
|
||
if !yield(chunk, err) {
|
||
return
|
||
}
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
**Why this is simpler than embeddings:**
|
||
- No embedding model to download or run (saves ~270MB of RAM and ~200ms per query)
|
||
- One file (`data/portfolio.db`), one driver, no extra process
|
||
- BM25 ranking is excellent for keyword-based retrieval over structured docs like project READMEs
|
||
- Trade-off: no semantic similarity ("projects about AI" won't match "machine learning" without the literal words). Mitigation: `trigram` tokenizer handles morphology well for English/Spanish.
|
||
|
||
---
|
||
|
||
## 🗜️ 4.5 Auto-compaction
|
||
|
||
Long conversations eventually run out of context — at qwen2.5-3b's 4k window, the ~3k-token system prompt + RAG block leaves only room for 2–3 user turns. Auto-compaction solves this by folding the older portion of the conversation into a single summary system message when the previous turn's input tokens cross a configurable threshold.
|
||
|
||
### When it fires
|
||
|
||
`agent.Runner.Compact` runs once per `/api/chat` request, before the RAG search. It compares the runner's most recent `Usage.InputTokens` (reported by the provider in the previous streamed chunk) against `client.Capabilities().MaxContextWindow × threshold_ratio`.
|
||
|
||
| Setting | Default | What it controls |
|
||
|---|---|---|
|
||
| `compaction.enabled` | `false` | Master switch. |
|
||
| `compaction.threshold_ratio` | `0.75` | Trigger when used tokens ≥ window × ratio. |
|
||
| `compaction.keep_recent_turns` | `4` | How many of the latest user turns are kept verbatim after compaction. |
|
||
| `compaction.summary_system_prompt` | *(built-in bilingual)* | Override the instruction sent to the LLM when summarizing. |
|
||
|
||
Short-circuits silently when compaction is disabled, the provider doesn't report a window (`Capabilities().MaxContextWindow == 0`), the history is shorter than `keep_recent_turns`, or usage is still unknown (first turn).
|
||
|
||
### How the summary is made
|
||
|
||
1. `splitByTurns(history, keep_recent_turns)` divides messages into `(older, recent)` on user-role boundaries so a kept turn's user/assistant pair always stays together.
|
||
2. `renderTranscript(older)` flattens older messages into a `User:` / `Assistant:` transcript (skipping tool messages and empty assistant placeholders).
|
||
3. The runner calls `client.Generate(...)` with the summary prompt + transcript and a 512-token cap so the compaction step itself stays cheap.
|
||
4. The returned text is prepended as a system message (`"Earlier conversation summary:\n…"`), followed by the recent tail.
|
||
5. `LastCompaction()` returns `CompactionStats` so the SSE handler can emit a `compaction` event right before the streamed chunks.
|
||
|
||
### Failure mode
|
||
|
||
If `Generate` errors or returns an empty summary, compaction falls back to `truncateToBudget`: drop oldest user-turns one at a time until the remaining slice fits `threshold` tokens (heuristic: `len(s) / 4 + 1`). The current user turn is always preserved. The fallback is logged at WARN and the request still proceeds — a flaky summarize call never fails the user's request.
|
||
|
||
### Wire protocol
|
||
|
||
Streaming responses gain an optional `compaction` event:
|
||
|
||
```
|
||
data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500}
|
||
```
|
||
|
||
Emitted after `start` (when applicable) and before `sources` / `chunk`. The widget can render this as a subtle "Context compacted" hint or ignore it — both are valid.
|
||
|
||
### Persistence
|
||
|
||
Compaction is **per-request**. The full transcript is still saved to `messages` in `data/portfolio.db` verbatim, so `GET /api/conversations/{id}` always returns the original history. Only what we send to the LLM is reduced — the next session can re-read the full thread from the DB.
|
||
|
||
---
|
||
|
||
## 🌐 5. Embedding the widget
|
||
|
||
The bot ships with a drop-in vanilla-JS widget. Add two files to your site and it works.
|
||
|
||
### 5.1 The widget (any site)
|
||
|
||
```html
|
||
<link rel="stylesheet" href="/path/to/chat-widget.css">
|
||
<script src="/path/to/chat-widget.js"
|
||
data-api-url="https://chat.example.com"
|
||
data-title="Ask me anything"
|
||
data-greeting="Hi! Ask me about the projects."
|
||
data-position="bottom-right"
|
||
data-theme="auto"
|
||
defer></script>
|
||
```
|
||
|
||
A bubble appears bottom-right, opens a panel, talks SSE to `/api/chat`, streams the response, and cites sources. No build step, no React/Vue, no framework lock-in.
|
||
|
||
**Browser→bot options:**
|
||
|
||
| Topology | Trade-offs |
|
||
|---|---|
|
||
| **Direct** (browser → bot, same domain or CORS) | Simplest. Add the bot's origin to `cors_origins` in YAML. |
|
||
| **Reverse proxy** (nginx/Caddy in front) | Bot stays on private network, single public domain, no CORS to manage. |
|
||
| **Site proxies the bot** (Astro/Next API route) | Adds a hop and a bit of code, but gives you auth/session hooks in your site. |
|
||
|
||
The widget works the same in all three. Pick the topology that matches your infra.
|
||
|
||
> **Default dev setup is direct + CORS.** `cors_origins` in `configs/portfolio-bot.yaml` controls which sites can call the bot. Add your site's origin there.
|
||
|
||
### 5.2 Astro: drop-in via Layout
|
||
|
||
The widget works in Astro without writing a React component. Add this to your shared layout:
|
||
|
||
```astro
|
||
---
|
||
// src/layouts/BaseLayout.astro
|
||
import "../path/to/chat-widget.css";
|
||
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
|
||
---
|
||
<html>
|
||
<body>
|
||
<slot />
|
||
<script src="/path/to/chat-widget.js"
|
||
data-api-url={apiUrl}
|
||
data-title="Ask me anything"
|
||
data-position="bottom-right"
|
||
data-theme="auto"
|
||
defer is:inline></script>
|
||
</body>
|
||
</html>
|
||
```
|
||
|
||
`is:inline` keeps Astro from hashing/transforming the script tag, so the `data-*` attributes survive.
|
||
|
||
### 5.3 React / Next.js: same script tag
|
||
|
||
```tsx
|
||
// app/layout.tsx
|
||
import Script from "next/script";
|
||
|
||
export default function RootLayout({ children }) {
|
||
return (
|
||
<html>
|
||
<head>
|
||
<link rel="stylesheet" href="/chat-widget.css" />
|
||
<Script src="/chat-widget.js"
|
||
data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
|
||
data-title="Ask me anything"
|
||
data-position="bottom-right"
|
||
data-theme="auto"
|
||
strategy="afterInteractive" />
|
||
</head>
|
||
<body>{children}</body>
|
||
</html>
|
||
);
|
||
}
|
||
```
|
||
|
||
### 5.4 If you want a server proxy (Astro/Next API route)
|
||
|
||
The widget can also call a same-origin endpoint that forwards to the bot. This is the right call when you need:
|
||
- Auth on `/api/chat` (logged-in users only)
|
||
- Centralized rate limiting at the site level
|
||
- Hiding the bot's origin from the browser
|
||
|
||
```typescript
|
||
// src/pages/api/chat.ts (Astro) or app/api/chat/route.ts (Next)
|
||
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";
|
||
|
||
export const POST = async ({ request }) => {
|
||
const body = await request.json();
|
||
// (optional) auth check, rate limit, session lookup here
|
||
|
||
const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
|
||
method: "POST",
|
||
headers: { "Content-Type": "application/json" },
|
||
body: JSON.stringify(body),
|
||
});
|
||
|
||
return new Response(resp.body, {
|
||
status: resp.status,
|
||
headers: {
|
||
"Content-Type": "text/event-stream",
|
||
"Cache-Control": "no-cache",
|
||
"Connection": "keep-alive",
|
||
},
|
||
});
|
||
};
|
||
```
|
||
|
||
Then point the widget at `/api/chat` (same origin) instead of the bot's URL.
|
||
|
||
### 5.5 Widget configuration reference
|
||
|
||
All options are `data-*` attributes on the `<script>` tag:
|
||
|
||
| Attribute | Default | Notes |
|
||
|---|---|---|
|
||
| `data-api-url` | *(required)* | Base URL of the bot. No trailing slash. |
|
||
| `data-title` | `"Chat"` | Header text. |
|
||
| `data-greeting` | `""` | First assistant message when the panel opens. |
|
||
| `data-position` | `"bottom-right"` | `"bottom-right"` or `"bottom-left"`. |
|
||
| `data-theme` | `"auto"` | `"auto"` (follows OS), `"light"`, `"dark"`. |
|
||
|
||
Theming is via CSS custom properties on `.rony-chat-widget-root` (see `web/chat-widget.css`):
|
||
|
||
```css
|
||
.rony-chat-widget-root {
|
||
--rony-accent: #ff6b35;
|
||
--rony-radius: 4px;
|
||
--rony-font: "Inter", sans-serif;
|
||
}
|
||
```
|
||
|
||
### 5.6 What the widget doesn't do (yet)
|
||
|
||
- **Richer markdown** (tables, images) — the built-in renderer handles the common cases; for full CommonMark, swap `renderMarkdown` in `chat-widget.js` for `marked` or `markdown-it`.
|
||
- **Mobile swipe-to-dismiss** — panel goes full-screen on phones.
|
||
- **Conversation history sidebar** — only the active conversation is shown (the backend exposes `GET /api/conversations` for a future sidebar).
|
||
|
||
---
|
||
|
||
## 🤖 6. Self-hosting with llama.cpp (default)
|
||
|
||
### 6.1 Setup
|
||
|
||
llama-server is a separate process that the bot connects to over HTTP. **Both ports (the bot's and llama-server's) are configurable** — pick what fits your environment.
|
||
|
||
```bash
|
||
# 1. Make sure you have a GGUF model available
|
||
# Download from Hugging Face, e.g.:
|
||
# https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
|
||
export RONY_MODELS_PATH=/path/to/models
|
||
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf
|
||
|
||
# 2. Start llama-server (port is configurable; default llama.cpp is 8080)
|
||
llama-server \
|
||
-m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
|
||
--port 9100 \
|
||
--host 127.0.0.1 \
|
||
--ctx-size 4096 \
|
||
--mlock # prevents swap, critical on shared VPS
|
||
|
||
# 3. Make sure configs/portfolio-bot.yaml points to the same port
|
||
# providers[0].endpoint: http://localhost:9100/v1
|
||
|
||
# 4. Start the bot (default port 7331, also configurable)
|
||
./bin/chat-bot serve
|
||
# → Serves on http://localhost:7331
|
||
# → Override with: ./bin/chat-bot serve --port 9101 --host 127.0.0.1
|
||
```
|
||
|
||
**Port reference:**
|
||
|
||
| What | Default | How to change |
|
||
|---|---|---|
|
||
| `llama-server` HTTP port | 8080 (llama.cpp convention) | `--port N` flag when starting `llama-server` |
|
||
| chat-bot HTTP port | 7331 | `--port N` flag on `serve`, or `server.port` in YAML |
|
||
| chat-bot → llama-server URL | `http://localhost:8080/v1` | `endpoint` field on the provider in YAML |
|
||
|
||
The `llamacpp` provider is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
|
||
|
||
### 6.2 Alternative: Ollama (easier for development)
|
||
|
||
If you don't want to manage GGUF files manually, Ollama provides the same models with a simpler workflow:
|
||
|
||
```bash
|
||
# 1. Install Ollama
|
||
curl -fsSL https://ollama.com/install.sh | sh
|
||
|
||
# 2. Download chat model
|
||
ollama pull qwen2.5:1.5b
|
||
|
||
# 3. Verify
|
||
ollama list
|
||
|
||
# 4. Edit configs/portfolio-bot.yaml to mark ollama-local as default:
|
||
# providers[0].default: true (and remove default from llamacpp-local)
|
||
# Ollama exposes an OpenAI-compatible API on :11434/v1
|
||
|
||
# 5. Start the bot
|
||
ollama serve &
|
||
./bin/chat-bot serve
|
||
```
|
||
|
||
### 6.3 Alternative: llama.cpp direct (advanced)
|
||
|
||
For more control or if Ollama doesn't work in your setup:
|
||
|
||
```yaml
|
||
providers:
|
||
- name: llamacpp-local
|
||
type: llamacpp
|
||
model: qwen2.5-3b-instruct
|
||
endpoint: http://localhost:9100/v1 # configurable, see §6.1
|
||
context_size: 4096
|
||
max_tokens: 2048
|
||
default: true
|
||
```
|
||
|
||
The `llamacpp` adapter is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
|
||
|
||
---
|
||
|
||
## 📦 7. Bot CLI
|
||
|
||
### 7.1 Commands
|
||
|
||
```bash
|
||
# Start HTTP server
|
||
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]
|
||
|
||
# Re-index portfolio (reads data/projects/*.md → SQLite FTS5)
|
||
chat-bot reindex
|
||
|
||
# Single question (no server, useful for tests)
|
||
chat-bot ask "What projects does Victor have?" [--no-rag]
|
||
|
||
# Validate config
|
||
chat-bot config validate
|
||
|
||
# Health check (useful for monitoring)
|
||
chat-bot health
|
||
|
||
# Version
|
||
chat-bot version
|
||
```
|
||
|
||
### 7.2 Implementation with Cobra
|
||
|
||
```go
|
||
// cmd/chat-bot/main.go
|
||
package main
|
||
|
||
import (
|
||
"github.com/spf13/cobra"
|
||
)
|
||
|
||
func main() {
|
||
root := &cobra.Command{
|
||
Use: "chat-bot",
|
||
Short: "Portfolio chatbot HTTP server",
|
||
}
|
||
|
||
root.AddCommand(serveCmd())
|
||
root.AddCommand(reindexCmd())
|
||
root.AddCommand(askCmd())
|
||
root.AddCommand(configCmd())
|
||
root.AddCommand(healthCmd())
|
||
root.AddCommand(versionCmd())
|
||
|
||
if err := root.Execute(); err != nil {
|
||
os.Exit(1)
|
||
}
|
||
}
|
||
|
||
func serveCmd() *cobra.Command {
|
||
var port int
|
||
var host string
|
||
var reindexOnStart bool
|
||
|
||
cmd := &cobra.Command{
|
||
Use: "serve",
|
||
Short: "Start HTTP server",
|
||
RunE: func(cmd *cobra.Command, args []string) error {
|
||
return server.Serve(server.Config{
|
||
Port: port,
|
||
Host: host,
|
||
ReindexOnStart: reindexOnStart,
|
||
})
|
||
},
|
||
}
|
||
|
||
cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
|
||
cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
|
||
cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
|
||
|
||
return cmd
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## 🚀 8. Deployment
|
||
|
||
### 8.1 Recommendation: Self-hosted on VPS
|
||
|
||
```bash
|
||
# 1. Install dependencies
|
||
sudo apt install golang-go ollama
|
||
ollama pull qwen2.5:1.5b
|
||
|
||
# 2. Build
|
||
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot
|
||
|
||
# 3. systemd service
|
||
cat > /etc/systemd/system/chat-bot.service <<EOF
|
||
[Unit]
|
||
Description=Portfolio Chat Bot
|
||
After=network.target ollama.service
|
||
|
||
[Service]
|
||
Type=simple
|
||
User=chatbot
|
||
WorkingDirectory=/opt/chat-bot
|
||
ExecStart=/usr/local/bin/chat-bot serve
|
||
Restart=on-failure
|
||
Environment=RONY_MODELS_PATH=/opt/models
|
||
|
||
[Install]
|
||
WantedBy=multi-user.target
|
||
EOF
|
||
|
||
sudo systemctl enable --now chat-bot
|
||
```
|
||
|
||
### 8.2 Reverse proxy (Caddy)
|
||
|
||
```
|
||
# /etc/caddy/Caddyfile
|
||
chat.victorvargas.dev {
|
||
reverse_proxy localhost:7331
|
||
}
|
||
```
|
||
|
||
### 8.3 Monitoring
|
||
|
||
```bash
|
||
# Health check periodic
|
||
curl -s http://localhost:7331/api/health | jq
|
||
|
||
# Logs
|
||
journalctl -u chat-bot -f
|
||
```
|
||
|
||
---
|
||
|
||
## 🧪 9. Testing
|
||
|
||
### 9.1 Unit tests
|
||
|
||
```go
|
||
// internal/server/chat_test.go
|
||
package server
|
||
|
||
func TestHandleChat_ValidRequest(t *testing.T) {
|
||
s := newTestServer(t)
|
||
|
||
req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
|
||
"messages": [{"role": "user", "content": "hello"}]
|
||
}`))
|
||
req.Header.Set("Content-Type", "application/json")
|
||
|
||
w := httptest.NewRecorder()
|
||
s.handleChat(w, req)
|
||
|
||
assert.Equal(t, 200, w.Code)
|
||
assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
|
||
}
|
||
|
||
func TestHandleChat_RateLimit(t *testing.T) {
|
||
s := newTestServerWithConfig(t, server.Config{
|
||
RateLimit: 1, // 1 request per minute
|
||
})
|
||
|
||
// First request OK
|
||
req1 := newChatRequest("hello")
|
||
w1 := httptest.NewRecorder()
|
||
s.handleChat(w1, req1)
|
||
assert.Equal(t, 200, w1.Code)
|
||
|
||
// Second request denied
|
||
req2 := newChatRequest("hello again")
|
||
w2 := httptest.NewRecorder()
|
||
s.handleChat(w2, req2)
|
||
assert.Equal(t, 429, w2.Code)
|
||
}
|
||
```
|
||
|
||
### 9.2 Integration tests with mock LLM
|
||
|
||
```go
|
||
// internal/agent/runner_test.go
|
||
func TestRunner_RAGContextIsInjected(t *testing.T) {
|
||
mockLLM := mock.New(mock.Responses{
|
||
{Match: "projects", Response: "Victor has several projects..."},
|
||
})
|
||
|
||
memory := newMockMemoryWithDocs(t, []rag.Fragment{
|
||
{Content: "Rony Harness: AI agent harness...", ProjectID: "rony-harness"},
|
||
{Content: "rony-llm-agent: Go library...", ProjectID: "rony-llm-agent"},
|
||
})
|
||
|
||
runner := agent.NewRunner(agent.Config{
|
||
LLM: mockLLM,
|
||
Memory: memory,
|
||
Persona: testPersona,
|
||
})
|
||
|
||
resp, _ := runner.Run(context.Background(), []llm.Message{
|
||
{Role: llm.RoleUser, Content: "what projects does Victor have?"},
|
||
})
|
||
|
||
// Verify LLM received context chunks in system prompt
|
||
lastReq := mockLLM.LastRequest()
|
||
assert.Contains(t, lastReq.Messages[0].Content, "Rony Harness")
|
||
assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
|
||
}
|
||
```
|
||
|
||
### 9.3 E2E test with Astro
|
||
|
||
```bash
|
||
# 1. Start chat-bot on :7331
|
||
./bin/chat-bot serve &
|
||
|
||
# 2. Start Astro on :4321
|
||
cd ../portfolio && npm run dev &
|
||
|
||
# 3. Make request to Astro's proxy
|
||
curl -X POST http://localhost:4321/api/chat \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"messages":[{"role":"user","content":"hello"}]}'
|
||
|
||
# 4. Verify SSE stream
|
||
```
|
||
|
||
---
|
||
|
||
## 📂 10. Project Structure
|
||
|
||
```
|
||
rony-chat-bot/
|
||
├── cmd/
|
||
│ └── chat-bot/
|
||
│ └── main.go # CLI entrypoint
|
||
│
|
||
├── internal/
|
||
│ ├── server/ # HTTP handlers
|
||
│ │ ├── server.go # chi router + middleware
|
||
│ │ ├── handlers.go # /api/chat, /api/health, /api/info, /api/reindex, /api/conversations
|
||
│ │ ├── conversations_test.go # round-trip, continue, list, 404, delete, streaming
|
||
│ │ └── middleware.go # RequestID, Logging, CORS, RateLimit
|
||
│ │
|
||
│ ├── agent/ # LLM client + RAG runner
|
||
│ │ ├── runner.go # Stream wrapper, RAG injection into system prompt
|
||
│ │ └── client.go # NewClient factory: llamacpp / ollama / openai / anthropic
|
||
│ │
|
||
│ ├── portfolio/ # RAG: markdown → SQLite FTS5 + conversation persistence
|
||
│ │ ├── chunker.go # Heading-based splitter
|
||
│ │ ├── indexer.go # Store: schema, Reindex, Search (BM25)
|
||
│ │ ├── conversations.go # Conversation + Message CRUD, persisted alongside RAG
|
||
│ │ └── chunker_test.go / store_test.go
|
||
│ │
|
||
│ ├── persona/ # Persona bridge to rony-llm-agent
|
||
│ │ └── persona.go # FromConfig, BuildSystemPrompt (with RAG context)
|
||
│ │
|
||
│ ├── streaming/ # SSE protocol helpers
|
||
│ │ └── sse.go # WriteStart/Chunk/Sources/Done/Error
|
||
│ │
|
||
│ ├── i18n/ # Language detection (ES/EN) for the response
|
||
│ │
|
||
│ └── config/ # YAML loader + validation
|
||
│
|
||
├── web/ # ← DROP-IN CHAT WIDGET
|
||
│ ├── chat-widget.js # Vanilla JS, ~12 KB
|
||
│ ├── chat-widget.css # Scoped styles, CSS-custom-prop themable
|
||
│ ├── example.html # Local demo (python -m http.server)
|
||
│ └── README.md # Integration guide (HTML, Astro, Next.js)
|
||
│
|
||
├── data/
|
||
│ └── projects/ # ← Markdown per project (one .md per project)
|
||
│ ├── rony-harness.md
|
||
│ ├── rony-llm-agent.md
|
||
│ └── example-project.md
|
||
│
|
||
├── configs/
|
||
│ └── portfolio-bot.yaml # Provider + RAG + persona config
|
||
│
|
||
├── docs/
|
||
│ ├── architecture.md # ← THIS FILE
|
||
│ └── architecture.es.md
|
||
│
|
||
├── bench/ # Reproducible SQLite driver benchmark
|
||
│
|
||
├── go.mod # require rony-llm-agent, modernc.org/sqlite
|
||
└── README.md
|
||
```
|
||
|
||
---
|
||
|
||
## 📅 11. Roadmap
|
||
|
||
### Phase 1: MVP (2-3 weeks)
|
||
|
||
- [ ] Project setup (`go mod init`, structure)
|
||
- [ ] Basic HTTP server with `/api/chat` endpoint
|
||
- [ ] Functional SSE streaming
|
||
- [ ] RAG indexer (reads `data/projects/*.md` → SQLite FTS5)
|
||
- [ ] RAG retriever (query → top-k chunks)
|
||
- [ ] Persona loader from YAML
|
||
- [ ] llama.cpp integration (qwen2.5:1.5b GGUF)
|
||
- [ ] CLI: `serve`, `reindex`, `ask`
|
||
- [ ] Basic tests
|
||
|
||
### Phase 2: Integration with Astro (1 week)
|
||
|
||
- [ ] Astro API route of the proxy
|
||
- [ ] React component of the chat widget
|
||
- [ ] E2E test: Astro → chat-bot → response
|
||
- [ ] Widget styling (TailwindCSS)
|
||
|
||
### Phase 3: Polish (1 week)
|
||
|
||
- [ ] Robust rate limiting
|
||
- [ ] Structured logging (JSON)
|
||
- [ ] Health checks for monitoring
|
||
- [ ] systemd service file
|
||
- [ ] README + deployment docs
|
||
|
||
### Phase 4: Optionals
|
||
|
||
- [ ] Support for multiple conversations (session ID)
|
||
- [ ] Persisted chat history
|
||
- [ ] Analysis of frequent questions
|
||
- [ ] Multi-language (EN/ES switch)
|
||
- [ ] More polished standalone CLI version (`chat-bot ask`)
|
||
|
||
---
|
||
|
||
## 📐 12. Quality Specifications
|
||
|
||
### 12.1 Performance metrics
|
||
|
||
| Metric | Target |
|
||
|---|---|
|
||
| TTFT (Time-to-first-token) | <500ms with llama.cpp local |
|
||
| End-to-end (question → complete response) | <3s for typical responses |
|
||
| Memory at rest | <150MB |
|
||
| RAG indexing speed | ~100 docs/second |
|
||
| Retrieval latency | <50ms for top-5 |
|
||
|
||
### 12.2 Required tests
|
||
|
||
- Unit tests: coverage ≥70%
|
||
- Integration tests: with mock LLM + in-memory SQLite FTS5
|
||
- E2E: at least one complete Astro → chat-bot flow
|
||
|
||
---
|
||
|
||
## 🔒 13. Security
|
||
|
||
### 13.1 Implemented
|
||
|
||
- **Rate limiting** per IP (default 30 req/min)
|
||
- **Restrictive CORS** — only configured origins
|
||
- **Input validation** — JSON schema validation on requests
|
||
- **No PII storage** — we don't save conversations by default
|
||
- **Local-only by default** — no calls to cloud APIs
|
||
|
||
### 13.2 Deferred / Optional
|
||
|
||
- Auth with API key (for private use)
|
||
- Query logging for analytics
|
||
- IP anonymization in logs
|
||
- HTTPS via reverse proxy (Caddy/nginx)
|
||
|
||
---
|
||
|
||
## 📚 14. References
|
||
|
||
- **SSE Spec:** https://html.spec.whatwg.org/multipage/server-sent-events.html
|
||
- **Ollama API:** https://github.com/ollama/ollama/blob/main/docs/api.md
|
||
- **SQLite FTS5:** https://www.sqlite.org/fts5.html
|
||
- **Go SQLite driver:** https://github.com/mattn/go-sqlite3 (CGO) or https://modernc.org/sqlite (pure Go)
|
||
- **qwen2.5:** https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
|
||
- **Astro API routes:** https://docs.astro.build/en/guides/endpoints/
|
||
- **rony-llm-agent:** https://github.com/VictorVargas/rony-llm-agent
|
||
|
||
---
|
||
|
||
**Document ready for implementation. 🚀** |