rony-chat-bot/docs/architecture.md
Victor Hugo Vargas f33708534a feat: bootstrap rony-chat-bot Go module
Initial implementation of the bot:

- cmd/chat-bot: CLI entrypoint (serve, reindex, ask, version)
- internal/agent: LLM provider client + agent runner with RAG injection
- internal/config: YAML config loader (providers, RAG, persona, server)
- internal/i18n: response-language detection (EN/ES)
- internal/persona: persona system prompt assembly from YAML
- internal/portfolio: heading-based chunker + SQLite FTS5 indexer
- internal/server: chi router with /api/chat (SSE), /api/health, /api/info,
  /api/reindex, middleware (RequestID, Logging, CORS, RateLimit)
- internal/streaming: SSE protocol helpers (start, chunk, sources, done, error)
- web/: drop-in vanilla-JS chat widget (no build, no deps) + demo + README
- bench/: reproducible driver benchmark (modernc vs mattn SQLite)
- configs/portfolio-bot.yaml: llama.cpp default provider, SQLite RAG, canine persona
- docs/architecture.md / .es.md: aligned with SQLite FTS5 + llama.cpp decisions
- data/projects/README*.md: project data documentation
- README.md / .es.md: updated for current implementation

All tests pass (go test ./...). Bot is functional end-to-end with the
configured LLM provider.
2026-07-17 00:56:06 -07:00

1231 lines
No EOL
37 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 📋 Rony Chat Bot — Technical Design Document
**Version:** 1.0
**Author:** Victor Hugo Vargas
**Date:** 2026-06-28
**Status:** Complete specification for implementation
**Path:** `rony-chat-bot/docs/architecture.md`
> 🌐 **Language:** [English](./architecture.md) | [Español](./architecture.es.md)
>
> 📚 **Workspace:** This project is part of the `Rony/` workspace. See [`../README.md`](../../README.md).
>
> 🔑 **Depends on:** [`rony-llm-agent`](https://github.com/VictorVargas/rony-llm-agent) — core library that provides agent loop, LLM clients, RAG, persona system.
>
> 📐 **Methodology:** This project follows the **SDD + DDD + Hexagonal Architecture** approach. Functional Requirements are numbered as `CRF-XXX`. See [`../../METHODOLOGY.md`](../../METHODOLOGY.md).
---
## 🎯 1. Project Vision
### 1.1 What is Chat-Bot?
An **HTTP chatbot** that answers questions about Victor Hugo Vargas and his projects. Uses **RAG (Retrieval-Augmented Generation)** over markdown files describing each project, and a local LLM (or cloud) to generate responses.
### 1.2 Primary use case
Victor has a portfolio website (Astro + React). On the site there's a chat widget where visitors can ask:
- "What projects has Victor done?"
- "What's his experience with Go?"
- "How does Rony Harness work?"
- "Has Victor worked with PostgreSQL?"
The bot responds with accurate information extracted from the projects' markdown files + bio + skills.
### 1.3 Secondary use cases (future)
- **Client adaptation:** The same bot, with other data and another persona, serves car dealerships, restaurants, etc.
- **Standalone CLI:** `./chat-bot ask "what do you know about X?"` for terminal use.
- **Slack/Discord bot:** Wrapper that consumes the HTTP API.
### 1.4 Philosophy
- **Self-hosted by default** — works 100% local with Ollama + 1-3B models
- **Cloud optional** — if you need more quality, swap to Anthropic API
- **Portable** — easy to fork/customize for other contexts
- **Streaming** — token-by-token responses with SSE (no waiting for complete response)
- **Reuses `rony-llm-agent`** — doesn't reinvent the agent loop
---
## 🏗️ 2. Architecture
### 2.1 Overview
```
┌─────────────────────────────────────────────────────────────────┐
│ Browser (Astro site) │
│ ↓ HTTP POST /api/chat │
│ Astro SSR (proxy) ←────────── Serves portfolio + proxy chat │
│ ↓ HTTP POST /api/chat │
│ Chat-Bot HTTP server (:7331) │
│ ↓ │
│ Agent loop (rony-llm-agent) │
│ ↓ │
│ RAG retrieval → SQLite FTS5 over data/projects/*.md │
│ ↓ │
│ LLM (llama.cpp local default / Ollama or Anthropic optional) │
└─────────────────────────────────────────────────────────────────┘
```
### 2.2 Main components
| Component | Path | Responsibility |
|---|---|---|
| **HTTP server** | `internal/server/` | Gin/chi handlers, SSE streaming |
| **Agent runner** | `internal/agent/` | Wrapper over `rony-llm-agent` with specific config |
| **Portfolio loader** | `internal/portfolio/` | Reads `data/projects/*.md`, indexes in SQLite FTS5 |
| **Persona** | `internal/persona/` | Loads persona from `configs/portfolio-bot.yaml` |
| **CLI** | `cmd/chat-bot/` | Commands: `serve`, `reindex`, `ask`, `version` |
### 2.3 Tech stack
| Layer | Technology | Reason |
|---|---|---|
| **Language** | Go 1.26+ | Same as rony-harness, leverage `os.Root`, `iter.Seq` |
| **HTTP router** | `net/http` + `chi` | Stdlib + chi for middleware (CORS, logging) |
| **SSE** | `net/http` Flusher | Stdlib is enough, no external library needed |
| **Config** | `gopkg.in/yaml.v3` | Same as harness |
| **RAG backend** | SQLite + FTS5 (BM25) | Zero external deps, single file, fast |
| **LLM** | llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama as alt | Self-hosted by default |
| **Tests** | stdlib + testify | Consistency with the rest |
---
## 🔌 3. HTTP API
### 3.1 Endpoints
#### `POST /api/chat` — Chat with SSE streaming
**Request:**
```json
{
"messages": [
{"role": "user", "content": "What projects does Victor have?"}
],
"stream": true
}
```
**Response (SSE):**
```
data: {"type":"start","conversation_id":"abc123"}
data: {"type":"chunk","content":"Victor"}
data: {"type":"chunk","content":" has"}
data: {"type":"chunk","content":" several"}
data: {"type":"chunk","content":" projects"}
data: {"type":"sources","documents":["rony-harness.md","rony-llm-agent.md"]}
data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}
```
**Without streaming** (`"stream": false`):
```json
{
"content": "Victor has several projects...",
"sources": ["rony-harness.md", "rony-llm-agent.md"],
"usage": {"input_tokens": 245, "output_tokens": 38}
}
```
#### `POST /api/reindex` — Re-index portfolio
Useful when files in `data/projects/` are modified.
**Request:** empty
**Response:**
```json
{
"indexed_files": 12,
"total_chunks": 87,
"duration_ms": 4321
}
```
#### `GET /api/health` — Health check (real)
Probes the LLM provider and the SQLite store in parallel and returns their
states. Designed for monitoring/load balancers. **Returns 200 when healthy
or degraded, 503 when unhealthy.**
- `?deep=true` adds a chunk count to the store probe (same latency budget).
**Status taxonomy:**
| `status` | HTTP | Meaning |
|---|---|---|
| `healthy` | 200 | LLM up, store up |
| `degraded` | 200 | LLM up, store down — bot still answers, just without RAG |
| `unhealthy` | 503 | LLM down — bot cannot answer, no point routing traffic here |
**Probe details:**
| Component | Probe | Latency |
|---|---|---|
| `llm` | `GET {provider}/health` (llamacpp, ollama) or `/models` (openai) | ~1ms for local llama-server |
| `store` | `SELECT 1` on the SQLite handle | ~100µs |
Each probe has a 2s timeout; the whole call returns within ~2.5s even if a
dependency hangs.
**Response shape (healthy):**
```json
{
"status": "healthy",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {
"status": "up",
"latency": "1.028ms",
"details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
},
"store": {
"status": "up",
"latency": "107µs"
}
}
}
```
**Response shape (degraded, with `?deep=true`):**
```json
{
"status": "degraded",
"version": "0.2.0-dev",
"checked_at": "2026-07-17T05:02:07Z",
"components": {
"llm": {"status": "up", "latency": "0.8ms", "details": {...}},
"store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
}
}
```
**Response shape (unhealthy):** HTTP 503, same JSON with `"status": "unhealthy"` and the failed component reporting `"status": "down"` plus an `error` field.
#### `GET /api/info` — Bot metadata
```json
{
"name": "Rony Chat Bot",
"model": "qwen2.5:1.5b",
"persona": "...",
"topics": ["projects", "experience", "technical skills"]
}
```
### 3.2 SSE Implementation
```go
// internal/server/chat.go
package server
import (
"encoding/json"
"fmt"
"net/http"
"github.com/VictorVargas/rony-llm-agent/pkg/agent"
)
func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
// SSE headers
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
w.Header().Set("Connection", "keep-alive")
w.Header().Set("X-Accel-Buffering", "no")
flusher, ok := w.(http.Flusher)
if !ok {
http.Error(w, "SSE not supported", http.StatusInternalServerError)
return
}
// Parse request
var req ChatRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
writeError(w, flusher, "invalid request", err)
return
}
// Start event
writeSSE(w, flusher, "start", map[string]string{
"conversation_id": generateConvID(),
})
// Run agent with streaming
sources := []string{}
for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
if err != nil {
writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
return
}
if chunk.Type == "source" {
sources = append(sources, chunk.Source)
}
writeSSE(w, flusher, chunk.Type, chunk.Data)
}
// Done event
writeSSE(w, flusher, "done", map[string]any{
"usage": map[string]int{
"input_tokens": 245,
"output_tokens": 38,
},
})
}
func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
payload, _ := json.Marshal(data)
fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
flusher.Flush()
}
```
### 3.3 Middleware
```go
// internal/server/middleware.go
package server
func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
// Wrap response writer to capture status
rw := &statusRecorder{ResponseWriter: w, status: 200}
next.ServeHTTP(rw, r)
slog.Info("http.request",
"method", r.Method,
"path", r.URL.Path,
"status", rw.status,
"duration_ms", time.Since(start).Milliseconds(),
"ip", r.RemoteAddr,
)
})
}
func (s *Server) corsMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
origin := r.Header.Get("Origin")
for _, allowed := range s.config.Server.CORSOrigins {
if origin == allowed {
w.Header().Set("Access-Control-Allow-Origin", origin)
w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
break
}
}
if r.Method == "OPTIONS" {
w.WriteHeader(204)
return
}
next.ServeHTTP(w, r)
})
}
func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if !limiter.Allow() {
http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
return
}
next.ServeHTTP(w, r)
})
}
```
---
## 🧠 4. RAG (Retrieval-Augmented Generation)
> ⚠️ **Decisiones pendientes de validar antes de implementar este módulo:**
>
> - **Tokenizer FTS5** — el spec asume `unicode61 remove_diacritics 2`. Confirmar con datos reales si conviene cambiar a `porter` (stemming EN), `trigram` (sub-string matching) o un tokenizer custom para español. **Validar:** ejecutar queries representativas contra `data/projects/` y comparar recall antes de cerrar esta elección.
> - **Driver SQLite** — ✅ **DECIDIDO: `modernc.org/sqlite`** (puro Go, sin CGO). Ver benchmark abajo.
> - **Chunking** — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. **Validar:** medir recall con chunks por sección markdown (split por `#`/`##`) vs por tamaño.
> - **Sin similitud semántica** — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. **Validar:** tamaño del corpus y types of questions esperadas; si el corpus crece o las queries se vuelven abstractas, considerar agregar embeddings como capa secundaria.
### 4.0 Driver decision: benchmark results
Reproducible con `CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/`. Datos: 4 markdowns → 11 chunks.
| Operación | mattn (CGO) | modernc (puro Go) | Diferencia |
|---|---|---|---|
| **Insert** (11 chunks) | 2,802,843 ns/op | **1,465,646 ns/op** | modernc 1.9× más rápido |
| Insert alloc | 2,124,299 B/op | **9,770 B/op** | modernc usa 217× menos memoria |
| **Query** (8 queries BM25) | **244,047 ns/op** | 555,162 ns/op | mattn 2.3× más rápido |
| **Round-trip** (insert + 8 queries) | 3,543,417 ns/op | **2,267,669 ns/op** | modernc 1.6× más rápido |
| Binary size | 11 MB | 11 MB | igual |
| Build deps | gcc, CGO=1 | nada | modernc gana |
| CI/CD portable | requiere toolchain C | `go build` puro | modernc gana |
**Decisión: `modernc.org/sqlite`**.
Justificación:
1. Ambas latencias de query (~250µs vs ~550µs) son **2 órdenes de magnitud por debajo** del target de 50ms — imperceptible vs el LLM (varios segundos).
2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.
### 4.1 Indexing pipeline
```
data/projects/*.md
↓ (read all files)
Raw markdown content
↓ (split into chunks, ~500 chars, 50 overlap)
Chunks []
↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
Indexed corpus
```
**When it runs:**
- On bot startup (if `--reindex-on-start` flag)
- Manually: `./chat-bot reindex`
- Via HTTP: `POST /api/reindex`
### 4.2 Retrieval pipeline
```
User query "what projects does Victor have?"
↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
Top 5 relevant chunks
↓ (format as context block)
System prompt += relevant chunks
↓ (send to LLM)
LLM generates answer
```
### 4.3 Implementation
```go
// internal/portfolio/indexer.go
package portfolio
import (
"context"
"database/sql"
"fmt"
"log/slog"
"os"
"path/filepath"
"strings"
)
type Indexer struct {
dataPath string
db *sql.DB
chunkSize int
chunkOverlap int
}
func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
if err != nil {
return 0, err
}
// Rebuild FTS5 index from scratch (delete + insert is faster than diff for small corpora)
if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
return 0, fmt.Errorf("clear index: %w", err)
}
totalChunks := 0
for _, file := range files {
chunks, err := i.indexFile(ctx, file)
if err != nil {
slog.Warn("failed to index file", "file", file, "err", err)
continue
}
totalChunks += chunks
}
return totalChunks, nil
}
func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
content, err := os.ReadFile(path)
if err != nil {
return 0, err
}
projectID := strings.TrimSuffix(filepath.Base(path), ".md")
chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)
tx, err := i.db.BeginTx(ctx, nil)
if err != nil {
return 0, err
}
defer tx.Rollback()
stmt, err := tx.PrepareContext(ctx, `
INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
VALUES (?, ?, ?, ?, ?)
`)
if err != nil {
return 0, err
}
defer stmt.Close()
for idx, chunk := range chunks {
id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
return idx, err
}
}
if err := tx.Commit(); err != nil {
return 0, err
}
return len(chunks), nil
}
// schema.go — applied at startup
const schema = `
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
id UNINDEXED,
project_id UNINDEXED,
source_file UNINDEXED,
chunk_index UNINDEXED,
content,
tokenize = 'unicode61 remove_diacritics 2'
);
`
func splitIntoChunks(text string, size, overlap int) []string {
// Simple implementation: split by size with overlap
// Production version uses tokenizer-aware chunking
var chunks []string
for i := 0; i < len(text); i += size - overlap {
end := i + size
if end > len(text) {
end = len(text)
}
chunks = append(chunks, text[i:end])
}
return chunks
}
```
### 4.4 Retrieval in the agent loop
```go
// internal/portfolio/search.go
package portfolio
type Hit struct {
ProjectID string
SourceFile string
ChunkIndex int
Content string
Score float64 // BM25 score from FTS5
}
func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
// Escape user input: FTS5 syntax can break with special chars
ftsQuery := sanitizeFTS5(query)
rows, err := s.db.QueryContext(ctx, `
SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
FROM portfolio_chunks
WHERE portfolio_chunks MATCH ?
ORDER BY score
LIMIT ?
`, ftsQuery, topK)
if err != nil {
return nil, err
}
defer rows.Close()
var hits []Hit
for rows.Next() {
var h Hit
if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
return nil, err
}
hits = append(hits, h)
}
return hits, rows.Err()
}
// sanitizeFTS5 wraps the user query so reserved chars and unquoted strings don't crash FTS5.
// A pragmatic choice for a Q&A bot: append prefix-match wildcard to each token.
func sanitizeFTS5(q string) string {
tokens := strings.FieldsFunc(q, func(r rune) bool {
return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
(r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
r > 0x7F) // keep accented chars
})
if len(tokens) == 0 {
return `""`
}
for i, t := range tokens {
tokens[i] = `"` + strings.ToLower(t) + `"*`
}
return strings.Join(tokens, " ")
}
```
```go
// internal/agent/runner.go
package agent
func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
basePrompt := r.persona.SystemPrompt
hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
if err != nil {
return "", err
}
if len(hits) == 0 {
return basePrompt, nil
}
var contextBlock strings.Builder
contextBlock.WriteString(basePrompt)
contextBlock.WriteString("\n\n## Relevant context\n\n")
for _, h := range hits {
contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
h.SourceFile, h.Content))
}
return contextBlock.String(), nil
}
func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
return func(yield func(Chunk, error) bool) {
lastUserMsg := getLastUserMessage(messages)
systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
if err != nil {
yield(Chunk{}, err)
return
}
messages = prependSystem(messages, systemPrompt)
for chunk, err := range r.loop.RunStream(ctx, messages) {
if !yield(chunk, err) {
return
}
}
}
}
```
**Why this is simpler than embeddings:**
- No embedding model to download or run (saves ~270MB of RAM and ~200ms per query)
- One file (`data/portfolio.db`), one driver, no extra process
- BM25 ranking is excellent for keyword-based retrieval over structured docs like project READMEs
- Trade-off: no semantic similarity ("projects about AI" won't match "machine learning" without the literal words). Mitigation: `trigram` tokenizer handles morphology well for English/Spanish.
---
## 🌐 5. Embedding the widget
The bot ships with a drop-in vanilla-JS widget. Add two files to your site and it works.
### 5.1 The widget (any site)
```html
<link rel="stylesheet" href="/path/to/chat-widget.css">
<script src="/path/to/chat-widget.js"
data-api-url="https://chat.example.com"
data-title="Ask me anything"
data-greeting="Hi! Ask me about the projects."
data-position="bottom-right"
data-theme="auto"
defer></script>
```
A bubble appears bottom-right, opens a panel, talks SSE to `/api/chat`, streams the response, and cites sources. No build step, no React/Vue, no framework lock-in.
**Browser→bot options:**
| Topology | Trade-offs |
|---|---|
| **Direct** (browser → bot, same domain or CORS) | Simplest. Add the bot's origin to `cors_origins` in YAML. |
| **Reverse proxy** (nginx/Caddy in front) | Bot stays on private network, single public domain, no CORS to manage. |
| **Site proxies the bot** (Astro/Next API route) | Adds a hop and a bit of code, but gives you auth/session hooks in your site. |
The widget works the same in all three. Pick the topology that matches your infra.
> **Default dev setup is direct + CORS.** `cors_origins` in `configs/portfolio-bot.yaml` controls which sites can call the bot. Add your site's origin there.
### 5.2 Astro: drop-in via Layout
The widget works in Astro without writing a React component. Add this to your shared layout:
```astro
---
// src/layouts/BaseLayout.astro
import "../path/to/chat-widget.css";
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
---
<html>
<body>
<slot />
<script src="/path/to/chat-widget.js"
data-api-url={apiUrl}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
defer is:inline></script>
</body>
</html>
```
`is:inline` keeps Astro from hashing/transforming the script tag, so the `data-*` attributes survive.
### 5.3 React / Next.js: same script tag
```tsx
// app/layout.tsx
import Script from "next/script";
export default function RootLayout({ children }) {
return (
<html>
<head>
<link rel="stylesheet" href="/chat-widget.css" />
<Script src="/chat-widget.js"
data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
data-title="Ask me anything"
data-position="bottom-right"
data-theme="auto"
strategy="afterInteractive" />
</head>
<body>{children}</body>
</html>
);
}
```
### 5.4 If you want a server proxy (Astro/Next API route)
The widget can also call a same-origin endpoint that forwards to the bot. This is the right call when you need:
- Auth on `/api/chat` (logged-in users only)
- Centralized rate limiting at the site level
- Hiding the bot's origin from the browser
```typescript
// src/pages/api/chat.ts (Astro) or app/api/chat/route.ts (Next)
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";
export const POST = async ({ request }) => {
const body = await request.json();
// (optional) auth check, rate limit, session lookup here
const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
});
return new Response(resp.body, {
status: resp.status,
headers: {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"Connection": "keep-alive",
},
});
};
```
Then point the widget at `/api/chat` (same origin) instead of the bot's URL.
### 5.5 Widget configuration reference
All options are `data-*` attributes on the `<script>` tag:
| Attribute | Default | Notes |
|---|---|---|
| `data-api-url` | *(required)* | Base URL of the bot. No trailing slash. |
| `data-title` | `"Chat"` | Header text. |
| `data-greeting` | `""` | First assistant message when the panel opens. |
| `data-position` | `"bottom-right"` | `"bottom-right"` or `"bottom-left"`. |
| `data-theme` | `"auto"` | `"auto"` (follows OS), `"light"`, `"dark"`. |
Theming is via CSS custom properties on `.rony-chat-widget-root` (see `web/chat-widget.css`):
```css
.rony-chat-widget-root {
--rony-accent: #ff6b35;
--rony-radius: 4px;
--rony-font: "Inter", sans-serif;
}
```
### 5.6 What the widget doesn't do (yet)
- **Conversation persistence** — each visit is a fresh conversation. Bot is stateless.
- **Richer markdown** (tables, images) — the built-in renderer handles the common cases; for full CommonMark, swap `renderMarkdown` in `chat-widget.js` for `marked` or `markdown-it`.
- **Mobile swipe-to-dismiss** — panel goes full-screen on phones.
- **Conversation history sidebar** — only the active conversation is shown.
---
## 🤖 6. Self-hosting with llama.cpp (default)
### 6.1 Setup
llama-server is a separate process that the bot connects to over HTTP. **Both ports (the bot's and llama-server's) are configurable** — pick what fits your environment.
```bash
# 1. Make sure you have a GGUF model available
# Download from Hugging Face, e.g.:
# https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
export RONY_MODELS_PATH=/path/to/models
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf
# 2. Start llama-server (port is configurable; default llama.cpp is 8080)
llama-server \
-m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
--port 9100 \
--host 127.0.0.1 \
--ctx-size 4096 \
--mlock # prevents swap, critical on shared VPS
# 3. Make sure configs/portfolio-bot.yaml points to the same port
# providers[0].endpoint: http://localhost:9100/v1
# 4. Start the bot (default port 7331, also configurable)
./bin/chat-bot serve
# → Serves on http://localhost:7331
# → Override with: ./bin/chat-bot serve --port 9101 --host 127.0.0.1
```
**Port reference:**
| What | Default | How to change |
|---|---|---|
| `llama-server` HTTP port | 8080 (llama.cpp convention) | `--port N` flag when starting `llama-server` |
| chat-bot HTTP port | 7331 | `--port N` flag on `serve`, or `server.port` in YAML |
| chat-bot → llama-server URL | `http://localhost:8080/v1` | `endpoint` field on the provider in YAML |
The `llamacpp` provider is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
### 6.2 Alternative: Ollama (easier for development)
If you don't want to manage GGUF files manually, Ollama provides the same models with a simpler workflow:
```bash
# 1. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# 2. Download chat model
ollama pull qwen2.5:1.5b
# 3. Verify
ollama list
# 4. Edit configs/portfolio-bot.yaml to mark ollama-local as default:
# providers[0].default: true (and remove default from llamacpp-local)
# Ollama exposes an OpenAI-compatible API on :11434/v1
# 5. Start the bot
ollama serve &
./bin/chat-bot serve
```
### 6.3 Alternative: llama.cpp direct (advanced)
For more control or if Ollama doesn't work in your setup:
```yaml
providers:
- name: llamacpp-local
type: llamacpp
model: qwen2.5-3b-instruct
endpoint: http://localhost:9100/v1 # configurable, see §6.1
context_size: 4096
max_tokens: 2048
default: true
```
The `llamacpp` adapter is imported from `rony-llm-agent/pkg/llm/providers/llamacpp` and is compiled against `llama.cpp` via CGO or external binary.
---
## 📦 7. Bot CLI
### 7.1 Commands
```bash
# Start HTTP server
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]
# Re-index portfolio (reads data/projects/*.md → SQLite FTS5)
chat-bot reindex
# Single question (no server, useful for tests)
chat-bot ask "What projects does Victor have?" [--no-rag]
# Validate config
chat-bot config validate
# Health check (useful for monitoring)
chat-bot health
# Version
chat-bot version
```
### 7.2 Implementation with Cobra
```go
// cmd/chat-bot/main.go
package main
import (
"github.com/spf13/cobra"
)
func main() {
root := &cobra.Command{
Use: "chat-bot",
Short: "Portfolio chatbot HTTP server",
}
root.AddCommand(serveCmd())
root.AddCommand(reindexCmd())
root.AddCommand(askCmd())
root.AddCommand(configCmd())
root.AddCommand(healthCmd())
root.AddCommand(versionCmd())
if err := root.Execute(); err != nil {
os.Exit(1)
}
}
func serveCmd() *cobra.Command {
var port int
var host string
var reindexOnStart bool
cmd := &cobra.Command{
Use: "serve",
Short: "Start HTTP server",
RunE: func(cmd *cobra.Command, args []string) error {
return server.Serve(server.Config{
Port: port,
Host: host,
ReindexOnStart: reindexOnStart,
})
},
}
cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
return cmd
}
```
---
## 🚀 8. Deployment
### 8.1 Recommendation: Self-hosted on VPS
```bash
# 1. Install dependencies
sudo apt install golang-go ollama
ollama pull qwen2.5:1.5b
# 2. Build
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot
# 3. systemd service
cat > /etc/systemd/system/chat-bot.service <<EOF
[Unit]
Description=Portfolio Chat Bot
After=network.target ollama.service
[Service]
Type=simple
User=chatbot
WorkingDirectory=/opt/chat-bot
ExecStart=/usr/local/bin/chat-bot serve
Restart=on-failure
Environment=RONY_MODELS_PATH=/opt/models
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl enable --now chat-bot
```
### 8.2 Reverse proxy (Caddy)
```
# /etc/caddy/Caddyfile
chat.victorvargas.dev {
reverse_proxy localhost:7331
}
```
### 8.3 Monitoring
```bash
# Health check periodic
curl -s http://localhost:7331/api/health | jq
# Logs
journalctl -u chat-bot -f
```
---
## 🧪 9. Testing
### 9.1 Unit tests
```go
// internal/server/chat_test.go
package server
func TestHandleChat_ValidRequest(t *testing.T) {
s := newTestServer(t)
req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
"messages": [{"role": "user", "content": "hello"}]
}`))
req.Header.Set("Content-Type", "application/json")
w := httptest.NewRecorder()
s.handleChat(w, req)
assert.Equal(t, 200, w.Code)
assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
}
func TestHandleChat_RateLimit(t *testing.T) {
s := newTestServerWithConfig(t, server.Config{
RateLimit: 1, // 1 request per minute
})
// First request OK
req1 := newChatRequest("hello")
w1 := httptest.NewRecorder()
s.handleChat(w1, req1)
assert.Equal(t, 200, w1.Code)
// Second request denied
req2 := newChatRequest("hello again")
w2 := httptest.NewRecorder()
s.handleChat(w2, req2)
assert.Equal(t, 429, w2.Code)
}
```
### 9.2 Integration tests with mock LLM
```go
// internal/agent/runner_test.go
func TestRunner_RAGContextIsInjected(t *testing.T) {
mockLLM := mock.New(mock.Responses{
{Match: "projects", Response: "Victor has several projects..."},
})
memory := newMockMemoryWithDocs(t, []rag.Fragment{
{Content: "Rony Harness: AI agent harness...", ProjectID: "rony-harness"},
{Content: "rony-llm-agent: Go library...", ProjectID: "rony-llm-agent"},
})
runner := agent.NewRunner(agent.Config{
LLM: mockLLM,
Memory: memory,
Persona: testPersona,
})
resp, _ := runner.Run(context.Background(), []llm.Message{
{Role: llm.RoleUser, Content: "what projects does Victor have?"},
})
// Verify LLM received context chunks in system prompt
lastReq := mockLLM.LastRequest()
assert.Contains(t, lastReq.Messages[0].Content, "Rony Harness")
assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
}
```
### 9.3 E2E test with Astro
```bash
# 1. Start chat-bot on :7331
./bin/chat-bot serve &
# 2. Start Astro on :4321
cd ../portfolio && npm run dev &
# 3. Make request to Astro's proxy
curl -X POST http://localhost:4321/api/chat \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"hello"}]}'
# 4. Verify SSE stream
```
---
## 📂 10. Project Structure
```
rony-chat-bot/
├── cmd/
│ └── chat-bot/
│ └── main.go # CLI entrypoint
├── internal/
│ ├── server/ # HTTP handlers
│ │ ├── server.go # chi router + middleware
│ │ ├── handlers.go # /api/chat, /api/health, /api/info, /api/reindex
│ │ └── middleware.go # RequestID, Logging, CORS, RateLimit
│ │
│ ├── agent/ # LLM client + RAG runner
│ │ ├── runner.go # Stream wrapper, RAG injection into system prompt
│ │ └── client.go # NewClient factory: llamacpp / ollama / openai / anthropic
│ │
│ ├── portfolio/ # RAG: markdown → SQLite FTS5
│ │ ├── chunker.go # Heading-based splitter
│ │ ├── indexer.go # Store: schema, Reindex, Search (BM25)
│ │ └── chunker_test.go / store_test.go
│ │
│ ├── persona/ # Persona bridge to rony-llm-agent
│ │ └── persona.go # FromConfig, BuildSystemPrompt (with RAG context)
│ │
│ ├── streaming/ # SSE protocol helpers
│ │ └── sse.go # WriteStart/Chunk/Sources/Done/Error
│ │
│ ├── i18n/ # Language detection (ES/EN) for the response
│ │
│ └── config/ # YAML loader + validation
├── web/ # ← DROP-IN CHAT WIDGET
│ ├── chat-widget.js # Vanilla JS, ~12 KB
│ ├── chat-widget.css # Scoped styles, CSS-custom-prop themable
│ ├── example.html # Local demo (python -m http.server)
│ └── README.md # Integration guide (HTML, Astro, Next.js)
├── data/
│ └── projects/ # ← Markdown per project (one .md per project)
│ ├── rony-harness.md
│ ├── rony-llm-agent.md
│ └── example-project.md
├── configs/
│ └── portfolio-bot.yaml # Provider + RAG + persona config
├── docs/
│ ├── architecture.md # ← THIS FILE
│ └── architecture.es.md
├── bench/ # Reproducible SQLite driver benchmark
├── go.mod # require rony-llm-agent, modernc.org/sqlite
└── README.md
```
---
## 📅 11. Roadmap
### Phase 1: MVP (2-3 weeks)
- [ ] Project setup (`go mod init`, structure)
- [ ] Basic HTTP server with `/api/chat` endpoint
- [ ] Functional SSE streaming
- [ ] RAG indexer (reads `data/projects/*.md` → SQLite FTS5)
- [ ] RAG retriever (query → top-k chunks)
- [ ] Persona loader from YAML
- [ ] llama.cpp integration (qwen2.5:1.5b GGUF)
- [ ] CLI: `serve`, `reindex`, `ask`
- [ ] Basic tests
### Phase 2: Integration with Astro (1 week)
- [ ] Astro API route of the proxy
- [ ] React component of the chat widget
- [ ] E2E test: Astro → chat-bot → response
- [ ] Widget styling (TailwindCSS)
### Phase 3: Polish (1 week)
- [ ] Robust rate limiting
- [ ] Structured logging (JSON)
- [ ] Health checks for monitoring
- [ ] systemd service file
- [ ] README + deployment docs
### Phase 4: Optionals
- [ ] Support for multiple conversations (session ID)
- [ ] Persisted chat history
- [ ] Analysis of frequent questions
- [ ] Multi-language (EN/ES switch)
- [ ] More polished standalone CLI version (`chat-bot ask`)
---
## 📐 12. Quality Specifications
### 12.1 Performance metrics
| Metric | Target |
|---|---|
| TTFT (Time-to-first-token) | <500ms with llama.cpp local |
| End-to-end (question complete response) | <3s for typical responses |
| Memory at rest | <150MB |
| RAG indexing speed | ~100 docs/second |
| Retrieval latency | <50ms for top-5 |
### 12.2 Required tests
- Unit tests: coverage 70%
- Integration tests: with mock LLM + in-memory SQLite FTS5
- E2E: at least one complete Astro chat-bot flow
---
## 🔒 13. Security
### 13.1 Implemented
- **Rate limiting** per IP (default 30 req/min)
- **Restrictive CORS** only configured origins
- **Input validation** JSON schema validation on requests
- **No PII storage** we don't save conversations by default
- **Local-only by default** no calls to cloud APIs
### 13.2 Deferred / Optional
- Auth with API key (for private use)
- Query logging for analytics
- IP anonymization in logs
- HTTPS via reverse proxy (Caddy/nginx)
---
## 📚 14. References
- **SSE Spec:** https://html.spec.whatwg.org/multipage/server-sent-events.html
- **Ollama API:** https://github.com/ollama/ollama/blob/main/docs/api.md
- **SQLite FTS5:** https://www.sqlite.org/fts5.html
- **Go SQLite driver:** https://github.com/mattn/go-sqlite3 (CGO) or https://modernc.org/sqlite (pure Go)
- **qwen2.5:** https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
- **Astro API routes:** https://docs.astro.build/en/guides/endpoints/
- **rony-llm-agent:** https://github.com/VictorVargas/rony-llm-agent
---
**Document ready for implementation. 🚀**