rony-chat-bot/docs/architecture.es.md
Victor Hugo Vargas ab510d8b31 docs: document auto-compaction behavior in architecture guide
Adds the §4.5 Auto-compaction section to architecture.md and
architecture.es.md describing the trigger, fallback, persistence and
the new SSE 'compaction' event so consumers know how to react.

Includes the three placeholder projects used while exercising the
feature end to end (bot-onboarding, dashboard-metricas, tienda-ropa)
so /api/reindex picks them up without further setup.
2026-07-18 00:08:37 -07:00

41 KiB
Raw Blame History

📋 Rony Chat Bot — Technical Design Document

🌐 Idioma: English | Español

Versión: 1.0
Autor: Victor Hugo Vargas
Fecha: 2026-06-28
Estado: Especificación completa para implementación Path: rony-chat-bot/docs/architecture.md

📚 Workspace: Este proyecto es parte del workspace Rony/. Ver ../README.md.

🔑 Depende de: rony-llm-agent — librería core que provee agent loop, LLM clients, RAG, persona system.

📐 Metodología: Este proyecto sigue el enfoque SDD + DDD + Hexagonal Architecture. Los Requisitos Funcionales se numeran como CRF-XXX. Ver ../../METHODOLOGY.md.


🎯 1. Visión del Proyecto

1.1 ¿Qué es Chat-Bot?

Un chatbot HTTP que responde preguntas sobre Victor Hugo Vargas y sus proyectos. Usa RAG (Retrieval-Augmented Generation) sobre archivos markdown que describen cada proyecto, y un LLM local (o cloud) para generar respuestas.

1.2 Caso de uso primario

Victor tiene un portfolio web (Astro + React). En el sitio hay un widget de chat donde visitantes pueden preguntar:

  • "¿Qué proyectos ha hecho Victor?"
  • "¿Cuál es su experiencia con Go?"
  • "¿Cómo funciona Rony TUI?"
  • "¿Victor ha trabajado con PostgreSQL?"

El bot responde con información precisa extraída de los archivos markdown de proyectos + bio + skills.

1.3 Casos de uso secundarios (futuro)

  • Adaptación a clientes: El mismo bot, con otra data y otra persona, sirve para concesionarios, restaurantes, etc.
  • Standalone CLI: ./chat-bot ask "¿qué sabes de X?" para uso desde terminal.
  • Slack/Discord bot: Wrapper que consume el HTTP API.

1.4 Filosofía

  • Self-hosted por defecto — funciona 100% local con Ollama + modelos 1-3B
  • Cloud opcional — si se necesita más calidad, swap a Anthropic API
  • Portable — fácil de fork/customizar para otros contextos
  • Streaming — respuestas token-por-token con SSE (no espera a respuesta completa)
  • Reutiliza rony-llm-agent — no reinventar el agent loop

🏗️ 2. Arquitectura

2.1 Vista general

┌─────────────────────────────────────────────────────────────────┐
│  Browser (Astro site)                                            │
│      ↓ HTTP POST /api/chat                                       │
│  Astro SSR (proxy)  ←────────── Sirve portfolio + proxy chat    │
│      ↓ HTTP POST /api/chat                                       │
│  Chat-Bot HTTP server (:7331)                                    │
│      ↓                                                           │
│  Agent loop (rony-llm-agent)                                       │
│      ↓                                                           │
│  RAG retrieval → SQLite FTS5 sobre data/projects/*.md             │
│      ↓                                                           │
│  LLM (llama.cpp local default / Ollama o Anthropic opcionales)   │
└─────────────────────────────────────────────────────────────────┘

2.2 Componentes principales

Componente Path Responsabilidad
HTTP server internal/server/ Gin/chi handlers, SSE streaming
Agent runner internal/agent/ Wrapper sobre rony-llm-agent con config específica
Portfolio loader internal/portfolio/ Lee data/projects/*.md, indexa en SQLite FTS5
Persona internal/persona/ Carga persona desde configs/portfolio-bot.yaml
CLI cm./rony-chat-bot/ Comandos: serve, reindex, ask, version

2.3 Stack tecnológico

Capa Tecnología Razón
Lenguaje Go 1.26+ Mismo que harness, aprovechar os.Root, iter.Seq
HTTP router net/http + chi Stdlib + chi para middleware (CORS, logging)
SSE net/http Flusher Stdlib es suficiente, no necesita librería externa
Config gopkg.in/yaml.v3 Mismo que harness
RAG backend SQLite + FTS5 (BM25) Sin dependencias externas, un solo archivo, rápido
LLM llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama como alternativa Self-hosted por defecto
Tests stdlib + testify Consistencia con el resto

🔌 3. HTTP API

3.1 Endpoints

POST /api/chat — Chat con streaming SSE

Request:

{
  "messages": [
    {"role": "user", "content": "¿Qué proyectos tiene Victor?"}
  ],
  "stream": true
}

Response (SSE):

data: {"type":"start","conversation_id":"abc123"}

data: {"type":"chunk","content":"Victor"}
data: {"type":"chunk","content":" tiene"}
data: {"type":"chunk","content":" varios"}
data: {"type":"chunk","content":" proyectos"}

data: {"type":"sources","documents":["rony-tui.md","rony-llm-agent.md"]}

data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}}

Sin streaming ("stream": false):

{
  "content": "Victor tiene varios proyectos...",
  "sources": ["rony-tui.md", "rony-llm-agent.md"],
  "usage": {"input_tokens": 245, "output_tokens": 38}
}

POST /api/reindex — Re-indexar portfolio

Útil cuando se modifican archivos en data/projects/.

Request: vacío Response:

{
  "indexed_files": 12,
  "total_chunks": 87,
  "duration_ms": 4321
}

GET /api/health — Health check (real)

Prueba el LLM provider y el store SQLite en paralelo y reporta su estado. Pensado para monitoring / load balancers. Devuelve 200 cuando está healthy o degraded, 503 cuando está unhealthy.

  • ?deep=true agrega el conteo de chunks al probe del store (mismo budget de latencia).

Taxonomía de status:

status HTTP Significado
healthy 200 LLM up, store up
degraded 200 LLM up, store down — el bot igual responde, sin RAG
unhealthy 503 LLM down — el bot no puede responder, no tiene sentido rutear tráfico acá

Probes:

Componente Probe Latencia típica
llm GET {provider}/health (llamacpp, ollama) o /models (openai) ~1ms para llama-server local
store SELECT 1 sobre el handle SQLite ~100µs

Cada probe tiene 2s de timeout; toda la llamada vuelve en ~2.5s aunque una dependencia esté colgada.

Shape de respuesta (healthy):

{
  "status": "healthy",
  "version": "0.2.0-dev",
  "checked_at": "2026-07-17T05:02:07Z",
  "components": {
    "llm": {
      "status": "up",
      "latency": "1.028ms",
      "details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"}
    },
    "store": {
      "status": "up",
      "latency": "107µs"
    }
  }
}

Shape (degraded, con ?deep=true):

{
  "status": "degraded",
  "version": "0.2.0-dev",
  "checked_at": "2026-07-17T05:02:07Z",
  "components": {
    "llm": {"status": "up", "latency": "0.8ms", "details": {...}},
    "store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}}
  }
}

Shape (unhealthy): HTTP 503, mismo JSON con "status": "unhealthy" y el componente fallido reportando "status": "down" más un campo error.

GET /api/info — Metadata del bot

{
  "name": "Asistente de Victor Hugo Vargas",
  "model": "qwen2.5:1.5b",
  "persona": "...",
  "topics": ["proyectos", "experiencia", "skills técnicas"]
}

3.2 SSE Implementation

// internal/server/chat.go
package server

import (
    "encoding/json"
    "fmt"
    "net/http"
    "github.com/VictorVargas/rony-llm-agent/pkg/agent"
)

func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) {
    // Headers SSE
    w.Header().Set("Content-Type", "text/event-stream")
    w.Header().Set("Cache-Control", "no-cache")
    w.Header().Set("Connection", "keep-alive")
    w.Header().Set("X-Accel-Buffering", "no")
    
    flusher, ok := w.(http.Flusher)
    if !ok {
        http.Error(w, "SSE no soportado", http.StatusInternalServerError)
        return
    }
    
    // Parse request
    var req ChatRequest
    if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
        writeError(w, flusher, "invalid request", err)
        return
    }
    
    // Start event
    writeSSE(w, flusher, "start", map[string]string{
        "conversation_id": generateConvID(),
    })
    
    // Run agent con streaming
    sources := []string{}
    for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) {
        if err != nil {
            writeSSE(w, flusher, "error", map[string]string{"message": err.Error()})
            return
        }
        if chunk.Type == "source" {
            sources = append(sources, chunk.Source)
        }
        writeSSE(w, flusher, chunk.Type, chunk.Data)
    }
    
    // Done event
    writeSSE(w, flusher, "done", map[string]any{
        "usage": map[string]int{
            "input_tokens":  245,
            "output_tokens": 38,
        },
    })
}

func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) {
    payload, _ := json.Marshal(data)
    fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload)
    flusher.Flush()
}

3.3 Middleware

// internal/server/middleware.go
package server

func (s *Server) loggingMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        start := time.Now()
        // Wrap response writer para capturar status
        rw := &statusRecorder{ResponseWriter: w, status: 200}
        next.ServeHTTP(rw, r)
        
        slog.Info("http.request",
            "method", r.Method,
            "path", r.URL.Path,
            "status", rw.status,
            "duration_ms", time.Since(start).Milliseconds(),
            "ip", r.RemoteAddr,
        )
    })
}

func (s *Server) corsMiddleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        origin := r.Header.Get("Origin")
        for _, allowed := range s.config.Server.CORSOrigins {
            if origin == allowed {
                w.Header().Set("Access-Control-Allow-Origin", origin)
                w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS")
                w.Header().Set("Access-Control-Allow-Headers", "Content-Type")
                break
            }
        }
        if r.Method == "OPTIONS" {
            w.WriteHeader(204)
            return
        }
        next.ServeHTTP(w, r)
    })
}

func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler {
    limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst)
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        if !limiter.Allow() {
            http.Error(w, "rate limit exceeded", http.StatusTooManyRequests)
            return
        }
        next.ServeHTTP(w, r)
    })
}

🧠 4. RAG (Retrieval-Augmented Generation)

⚠️ Decisiones pendientes de validar antes de implementar este módulo:

  • Tokenizer FTS5 — el spec asume unicode61 remove_diacritics 2. Confirmar con datos reales si conviene cambiar a porter (stemming EN), trigram (sub-string matching) o un tokenizer custom para español. Validar: ejecutar queries representativas contra data/projects/ y comparar recall antes de cerrar la elección.
  • Driver SQLite DECIDIDO: modernc.org/sqlite (puro Go, sin CGO). Ver benchmark abajo.
  • Chunking — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. Validar: medir recall con chunks por sección markdown (split por #/##) vs por tamaño.
  • Sin similitud semántica — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. Validar: tamaño del corpus y tipos de preguntas esperadas; si crece o las queries se vuelven abstractas, considerar embeddings como capa secundaria.

4.0 Decisión de driver: resultados del benchmark

Reproducible con CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/. Datos: 4 markdowns → 11 chunks.

Operación mattn (CGO) modernc (puro Go) Diferencia
Insert (11 chunks) 2,802,843 ns/op 1,465,646 ns/op modernc 1.9× más rápido
Insert alloc 2,124,299 B/op 9,770 B/op modernc usa 217× menos memoria
Query (8 queries BM25) 244,047 ns/op 555,162 ns/op mattn 2.3× más rápido
Round-trip (insert + 8 queries) 3,543,417 ns/op 2,267,669 ns/op modernc 1.6× más rápido
Tamaño binario 11 MB 11 MB igual
Dependencias build gcc, CGO=1 ninguna gana modernc
CI/CD portable requiere toolchain C go build puro gana modernc

Decisión: modernc.org/sqlite.

Justificación:

  1. Ambas latencias de query (~250µs vs ~550µs) son 2 órdenes de magnitud por debajo del target de 50ms — imperceptible vs el LLM (varios segundos).
  2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex.
  3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles).
  4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no.

4.1 Pipeline de indexación

data/projects/*.md
    ↓ (read all files)
Raw markdown content
    ↓ (split into chunks, ~500 chars, 50 overlap)
Chunks []
    ↓ (insert into SQLite FTS5 virtual table "portfolio_chunks")
Indexed corpus

Cuándo se ejecuta:

  • Al arrancar el bot (si --reindex-on-start flag)
  • Manualmente: ./chat-bot reindex
  • Vía HTTP: POST /api/reindex

4.2 Pipeline de retrieval

User query "¿qué proyectos tiene Victor?"
    ↓ (FTS5 MATCH query, BM25 ranking, top_k=5)
Top 5 chunks relevantes
    ↓ (format as context block)
System prompt += chunks relevantes
    ↓ (send to LLM)
LLM generates answer

4.3 Implementación

// internal/portfolio/indexer.go
package portfolio

import (
    "context"
    "database/sql"
    "fmt"
    "log/slog"
    "os"
    "path/filepath"
    "strings"
)

type Indexer struct {
    dataPath     string
    db           *sql.DB
    chunkSize    int
    chunkOverlap int
}

func (i *Indexer) IndexAll(ctx context.Context) (int, error) {
    files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md"))
    if err != nil {
        return 0, err
    }

    // Reconstruir el índice FTS5 desde cero (DELETE+INSERT es más rápido
    // que diff para corpus pequeños)
    if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil {
        return 0, fmt.Errorf("clear index: %w", err)
    }

    totalChunks := 0
    for _, file := range files {
        chunks, err := i.indexFile(ctx, file)
        if err != nil {
            slog.Warn("failed to index file", "file", file, "err", err)
            continue
        }
        totalChunks += chunks
    }

    return totalChunks, nil
}

func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) {
    content, err := os.ReadFile(path)
    if err != nil {
        return 0, err
    }

    projectID := strings.TrimSuffix(filepath.Base(path), ".md")
    chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap)

    tx, err := i.db.BeginTx(ctx, nil)
    if err != nil {
        return 0, err
    }
    defer tx.Rollback()

    stmt, err := tx.PrepareContext(ctx, `
        INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content)
        VALUES (?, ?, ?, ?, ?)
    `)
    if err != nil {
        return 0, err
    }
    defer stmt.Close()

    for idx, chunk := range chunks {
        id := fmt.Sprintf("%s-chunk-%d", projectID, idx)
        if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil {
            return idx, err
        }
    }

    if err := tx.Commit(); err != nil {
        return 0, err
    }
    return len(chunks), nil
}

// schema.go — aplicado al arrancar
const schema = `
CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5(
    id UNINDEXED,
    project_id UNINDEXED,
    source_file UNINDEXED,
    chunk_index UNINDEXED,
    content,
    tokenize = 'unicode61 remove_diacritics 2'
);
`

func splitIntoChunks(text string, size, overlap int) []string {
    // Implementación simple: split por tamaño con overlap
    // Versión production usa tokenizer-aware chunking
    var chunks []string
    for i := 0; i < len(text); i += size - overlap {
        end := i + size
        if end > len(text) {
            end = len(text)
        }
        chunks = append(chunks, text[i:end])
    }
    return chunks
}

4.4 Retrieval en el agent loop

// internal/portfolio/search.go
package portfolio

type Hit struct {
    ProjectID  string
    SourceFile string
    ChunkIndex int
    Content    string
    Score      float64 // BM25 score devuelto por FTS5
}

func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) {
    // Escapar input del usuario: la sintaxis FTS5 puede romperse con caracteres especiales
    ftsQuery := sanitizeFTS5(query)

    rows, err := s.db.QueryContext(ctx, `
        SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score
        FROM portfolio_chunks
        WHERE portfolio_chunks MATCH ?
        ORDER BY score
        LIMIT ?
    `, ftsQuery, topK)
    if err != nil {
        return nil, err
    }
    defer rows.Close()

    var hits []Hit
    for rows.Next() {
        var h Hit
        if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil {
            return nil, err
        }
        hits = append(hits, h)
    }
    return hits, rows.Err()
}

// sanitizeFTS5 envuelve la consulta para que chars reservados no rompan FTS5.
// Para un bot de Q&A: agrega wildcard prefix-match a cada token.
func sanitizeFTS5(q string) string {
    tokens := strings.FieldsFunc(q, func(r rune) bool {
        return !(r == '-' || r == '_' || (r >= '0' && r <= '9') ||
            (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') ||
            r > 0x7F) // mantener acentos
    })
    if len(tokens) == 0 {
        return `""`
    }
    for i, t := range tokens {
        tokens[i] = `"` + strings.ToLower(t) + `"*`
    }
    return strings.Join(tokens, " ")
}
// internal/agent/runner.go
package agent

func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) {
    basePrompt := r.persona.SystemPrompt

    hits, err := r.store.Search(ctx, query, r.config.RAG.TopK)
    if err != nil {
        return "", err
    }
    if len(hits) == 0 {
        return basePrompt, nil
    }

    var contextBlock strings.Builder
    contextBlock.WriteString(basePrompt)
    contextBlock.WriteString("\n\n## Relevant context\n\n")
    for _, h := range hits {
        contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n",
            h.SourceFile, h.Content))
    }
    return contextBlock.String(), nil
}

func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] {
    return func(yield func(Chunk, error) bool) {
        lastUserMsg := getLastUserMessage(messages)
        systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg)
        if err != nil {
            yield(Chunk{}, err)
            return
        }

        messages = prependSystem(messages, systemPrompt)

        for chunk, err := range r.loop.RunStream(ctx, messages) {
            if !yield(chunk, err) {
                return
            }
        }
    }
}

Por qué esto es más simple que embeddings:

  • Sin modelo de embeddings que descargar ni ejecutar (ahorra ~270MB de RAM y ~200ms por consulta)
  • Un archivo (data/portfolio.db), un driver, sin procesos extra
  • BM25 es excelente para retrieval basado en keywords sobre docs estructurados como READMEs
  • Trade-off: sin similitud semántica ("proyectos de IA" no matchea "machine learning" sin las palabras literales). Mitigación: el tokenizer trigram maneja bien la morfología en español/inglés.

🗜️ 4.5 Auto-compactación

Las conversaciones largas eventualmente agotan el contexto — con la ventana de 4k de qwen2.5-3b, el system prompt de ~3k tokens + el bloque RAG sólo deja espacio para 23 turnos del usuario. La auto-compactación resuelve esto plegando la parte más antigua de la conversación en un único mensaje-resumen del sistema cuando los tokens de entrada del turno anterior cruzan un umbral configurable.

Cuándo se dispara

agent.Runner.Compact corre una vez por request a /api/chat, antes de la búsqueda RAG. Compara los Usage.InputTokens más recientes del runner (reportados por el provider en el chunk streameado previo) contra client.Capabilities().MaxContextWindow × threshold_ratio.

Config Default Qué controla
compaction.enabled false Switch maestro.
compaction.threshold_ratio 0.75 Dispara cuando tokens usados ≥ ventana × ratio.
compaction.keep_recent_turns 4 Cuántos turnos recientes del usuario se preservan literales tras la compactación.
compaction.summary_system_prompt (bilingüe built-in) Override de la instrucción enviada al LLM al resumir.

Sale silenciosamente cuando la compactación está deshabilitada, el provider no reporta ventana (Capabilities().MaxContextWindow == 0), la historia es más corta que keep_recent_turns, o el usage aún es desconocido (primer turno).

Cómo se hace el resumen

  1. splitByTurns(history, keep_recent_turns) divide los mensajes en (older, recent) cortando en límites de rol user, así el par user/assistant de un turno preservado queda siempre junto.
  2. renderTranscript(older) aplana los mensajes antiguos en una transcripción User: / Assistant: (saltando mensajes tool y placeholders vacíos de assistant).
  3. El runner llama a client.Generate(...) con el prompt de resumen + la transcripción y un cap de 512 tokens para que la compactación en sí misma sea barata.
  4. El texto devuelto se antepone como mensaje de sistema ("Earlier conversation summary:\n…"), seguido por la cola reciente.
  5. LastCompaction() devuelve CompactionStats para que el handler SSE emita el evento compaction justo antes de los chunks streameados.

Modo de falla

Si Generate falla o devuelve un resumen vacío, la compactación cae a truncateToBudget: descarta turnos antiguos del usuario uno por uno hasta que el slice restante entre en threshold tokens (heurística: len(s) / 4 + 1). El turno actual del usuario siempre se preserva. El fallback se loggea a nivel WARN y el request sigue — un fallo del resumidor nunca rompe la request del usuario.

Protocolo de cable

Las respuestas streameadas ganan un evento opcional compaction:

data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500}

Se emite después del start (cuando aplica) y antes de sources / chunk. El widget puede renderizar esto como un hint sutil "Contexto compactado" o ignorarlo — ambas son válidas.

Persistencia

La compactación es por-request. La transcripción completa igual se guarda en messages en data/portfolio.db literal, así que GET /api/conversations/{id} siempre devuelve la historia original. Sólo se reduce lo que se le manda al LLM — la próxima sesión puede releer el thread completo desde la DB.


🌐 5. Embebiendo el widget

El bot viene con un widget vanilla-JS drop-in. Agrega dos archivos a tu sitio y funciona.

5.1 El widget (cualquier sitio)

<link rel="stylesheet" href="/path/to/chat-widget.css">
<script src="/path/to/chat-widget.js"
        data-api-url="https://chat.example.com"
        data-title="Pregúntame lo que sea"
        data-greeting="¡Hola! Pregúntame sobre los proyectos."
        data-position="bottom-right"
        data-theme="auto"
        defer></script>

Aparece una burbuja abajo a la derecha, abre un panel, habla SSE con /api/chat, streamea la respuesta y cita las fuentes. Sin build step, sin React/Vue, sin lock-in de framework.

Opciones browser→bot:

Topología Trade-offs
Directo (browser → bot, mismo dominio o CORS) Lo más simple. Agrega el origen del bot a cors_origins en YAML.
Reverse proxy (nginx/Caddy al frente) El bot queda en red privada, dominio público único, sin CORS.
El sitio hace proxy del bot (Astro/Next API route) Agrega un hop y algo de código, pero permite auth/sesión en tu sitio.

El widget funciona igual en las tres. Elige la que se ajuste a tu infra.

El setup dev default es directo + CORS. cors_origins en configs/portfolio-bot.yaml controla qué sitios pueden llamar al bot. Agregá el origen de tu sitio ahí.

5.2 Astro: drop-in vía Layout

El widget funciona en Astro sin escribir un componente React. Agregá esto a tu layout compartido:

---
// src/layouts/BaseLayout.astro
import "../path/to/chat-widget.css";
const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331";
---
<html>
  <body>
    <slot />
    <script src="/path/to/chat-widget.js"
            data-api-url={apiUrl}
            data-title="Pregúntame lo que sea"
            data-position="bottom-right"
            data-theme="auto"
            defer is:inline></script>
  </body>
</html>

is:inline evita que Astro transforme/hash el <script>, así los atributos data-* sobreviven.

5.3 React / Next.js: el mismo <script>

// app/layout.tsx
import Script from "next/script";

export default function RootLayout({ children }) {
  return (
    <html>
      <head>
        <link rel="stylesheet" href="/chat-widget.css" />
        <Script src="/chat-widget.js"
                data-api-url={process.env.NEXT_PUBLIC_CHAT_API_URL}
                data-title="Pregúntame lo que sea"
                data-position="bottom-right"
                data-theme="auto"
                strategy="afterInteractive" />
      </head>
      <body>{children}</body>
    </html>
  );
}

5.4 Si querés un proxy server-side (Astro/Next API route)

El widget también puede llamar a un endpoint same-origin que reenvía al bot. Esto tiene sentido cuando necesitás:

  • Auth en /api/chat (solo usuarios logueados)
  • Rate limiting centralizado a nivel sitio
  • Ocultar el origen del bot al browser
// src/pages/api/chat.ts (Astro) o app/api/chat/route.ts (Next)
const CHAT_BOT_URL = process.env.CHAT_BOT_URL || "http://localhost:7331";

export const POST = async ({ request }) => {
    const body = await request.json();
    // (opcional) auth check, rate limit, session lookup acá

    const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify(body),
    });

    return new Response(resp.body, {
        status: resp.status,
        headers: {
            "Content-Type": "text/event-stream",
            "Cache-Control": "no-cache",
            "Connection": "keep-alive",
        },
    });
};

Entonces apuntás el widget a /api/chat (mismo origen) en vez de la URL del bot.

5.5 Referencia de configuración del widget

Todas las opciones son atributos data-* en el <script>:

Atributo Default Notas
data-api-url (requerido) URL base del bot. Sin slash final.
data-title "Chat" Texto del header.
data-greeting "" Primer mensaje del asistente al abrir el panel.
data-position "bottom-right" "bottom-right" o "bottom-left".
data-theme "auto" "auto" (sigue el OS), "light", "dark".

El theming se hace vía CSS custom properties en .rony-chat-widget-root (ver web/chat-widget.css):

.rony-chat-widget-root {
  --rony-accent: #ff6b35;
  --rony-radius: 4px;
  --rony-font: "Inter", sans-serif;
}

5.6 Lo que el widget NO hace (aún)

  • Persistencia de conversación — cada visita es nueva. El bot es stateless.
  • Markdown enriquecido (tablas, imágenes) — el renderer built-in cubre los casos comunes; para CommonMark completo, cambiá renderMarkdown por marked o markdown-it.
  • Swipe-to-dismiss en mobile — el panel pasa a full-screen en mobile, sin gesto.
  • Historial de conversaciones — solo se ve la conversación activa.

🤖 6. Self-hosting con llama.cpp (default)

6.1 Setup

llama-server es un proceso separado al que el bot se conecta por HTTP. Ambos puertos (el del bot y el de llama-server) son configurables — elegí lo que se ajuste a tu entorno.

# 1. Asegúrate de tener un modelo GGUF disponible
# Descárgalo de Hugging Face, ej.:
#   https://huggingface.co/Qwen/Qwen2.5-3B-Instruct-GGUF
export RONY_MODELS_PATH=/path/to/models
ls $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf

# 2. Arrancar llama-server (puerto configurable; default de llama.cpp es 8080)
llama-server \
  -m $RONY_MODELS_PATH/qwen2.5-3b-instruct-q4_k_m.gguf \
  --port 9100 \
  --host 127.0.0.1 \
  --ctx-size 4096 \
  --mlock            # previene swap, crítico en VPS compartido

# 3. Verifica que configs/portfolio-bot.yaml apunte al mismo puerto
#    providers[0].endpoint: http://localhost:9100/v1

# 4. Arrancar el bot (puerto default 7331, también configurable)
./bin/chat-bot serve
# → Sirve en http://localhost:7331
# → Override: ./bin/chat-bot serve --port 9101 --host 127.0.0.1

Referencia de puertos:

Qué Default Cómo cambiarlo
Puerto HTTP de llama-server 8080 (convención de llama.cpp) flag --port N al arrancar llama-server
Puerto HTTP del chat-bot 7331 flag --port N en serve, o server.port en YAML
URL bot → llama-server http://localhost:8080/v1 campo endpoint del provider en YAML

El provider llamacpp se importa desde rony-llm-agent/pkg/llm/providers/llamacpp y se compila contra llama.cpp vía CGO o binario externo.

6.2 Alternativa: Ollama (más fácil para desarrollo)

Si prefieres no gestionar archivos GGUF manualmente, Ollama ofrece los mismos modelos con un flujo más simple:

# 1. Instalar Ollama
curl -fsSL https://ollama.com/install.sh | sh

# 2. Descargar modelo de chat
ollama pull qwen2.5:1.5b

# 3. Verificar
ollama list

# 4. Editar configs/portfolio-bot.yaml para marcar ollama-local como default:
#    providers[0].default: true (y quitar default de llamacpp-local)
#    Ollama expone una API OpenAI-compatible en :11434/v1

# 5. Arrancar el bot
ollama serve &
./bin/chat-bot serve

6.3 Alternativa: llama.cpp directo (avanzado)

Para más control o si Ollama no funciona en tu setup:

providers:
  - name: llamacpp-local
    type: llamacpp
    model: qwen2.5-3b-instruct
    endpoint: http://localhost:9100/v1   # configurable, ver §6.1
    context_size: 4096
    max_tokens: 2048
    default: true

El adapter llamacpp se importa desde rony-llm-agent/pkg/llm/providers/llamacpp y se compila contra llama.cpp vía CGO o binario externo.


📦 7. CLI del bot

7.1 Comandos

# Arrancar servidor HTTP
chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start]

# Re-indexar portfolio (lee data/projects/*.md → SQLite FTS5)
chat-bot reindex

# Pregunta única (sin servidor, útil para tests)
chat-bot ask "¿Qué proyectos tiene Victor?" [--no-rag]

# Validar config
chat-bot config validate

# Health check (útil para monitoring)
chat-bot health

# Versión
chat-bot version

7.2 Implementación con Cobra

// cm./rony-chat-bot/main.go
package main

import (
    "github.com/spf13/cobra"
)

func main() {
    root := &cobra.Command{
        Use:   "chat-bot",
        Short: "Portfolio chatbot HTTP server",
    }
    
    root.AddCommand(serveCmd())
    root.AddCommand(reindexCmd())
    root.AddCommand(askCmd())
    root.AddCommand(configCmd())
    root.AddCommand(healthCmd())
    root.AddCommand(versionCmd())
    
    if err := root.Execute(); err != nil {
        os.Exit(1)
    }
}

func serveCmd() *cobra.Command {
    var port int
    var host string
    var reindexOnStart bool
    
    cmd := &cobra.Command{
        Use:   "serve",
        Short: "Start HTTP server",
        RunE: func(cmd *cobra.Command, args []string) error {
            return server.Serve(server.Config{
                Port:           port,
                Host:           host,
                ReindexOnStart: reindexOnStart,
            })
        },
    }
    
    cmd.Flags().IntVar(&port, "port", 7331, "HTTP port")
    cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host")
    cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving")
    
    return cmd
}

🚀 8. Deployment

8.1 Recomendación: Self-hosted en VPS

# 1. Instalar dependencias
sudo apt install golang-go ollama
ollama pull qwen2.5:1.5b

# 2. Build
go build -o /usr/local/bin/chat-bot ./cmd/chat-bot

# 3. systemd service
cat > /etc/systemd/system/chat-bot.service <<EOF
[Unit]
Description=Portfolio Chat Bot
After=network.target ollama.service

[Service]
Type=simple
User=chatbot
WorkingDirectory=/opt/chat-bot
ExecStart=/usr/local/bin/chat-bot serve
Restart=on-failure
Environment=RONY_MODELS_PATH=/opt/models

[Install]
WantedBy=multi-user.target
EOF

sudo systemctl enable --now chat-bot

8.2 Reverse proxy (Caddy)

# /etc/caddy/Caddyfile
chat.victorvargas.dev {
    reverse_proxy localhost:7331
}

8.3 Monitoring

# Health check periódico
curl -s http://localhost:7331/api/health | jq

# Logs
journalctl -u chat-bot -f

🧪 9. Testing

9.1 Unit tests

// internal/server/chat_test.go
package server

func TestHandleChat_ValidRequest(t *testing.T) {
    s := newTestServer(t)
    
    req := httptest.NewRequest("POST", "/api/chat", strings.NewReader(`{
        "messages": [{"role": "user", "content": "hola"}]
    }`))
    req.Header.Set("Content-Type", "application/json")
    
    w := httptest.NewRecorder()
    s.handleChat(w, req)
    
    assert.Equal(t, 200, w.Code)
    assert.Equal(t, "text/event-stream", w.Header().Get("Content-Type"))
}

func TestHandleChat_RateLimit(t *testing.T) {
    s := newTestServerWithConfig(t, server.Config{
        RateLimit: 1, // 1 request per minute
    })
    
    // First request OK
    req1 := newChatRequest("hola")
    w1 := httptest.NewRecorder()
    s.handleChat(w1, req1)
    assert.Equal(t, 200, w1.Code)
    
    // Second request denied
    req2 := newChatRequest("hola de nuevo")
    w2 := httptest.NewRecorder()
    s.handleChat(w2, req2)
    assert.Equal(t, 429, w2.Code)
}

9.2 Integration tests con mock LLM

// internal/agent/runner_test.go
func TestRunner_RAGContextIsInjected(t *testing.T) {
    mockLLM := mock.New(mock.Responses{
        {Match: "proyectos", Response: "Victor tiene varios proyectos..."},
    })
    
    memory := newMockMemoryWithDocs(t, []rag.Fragment{
        {Content: "Rony TUI: AI agent harness...", ProjectID: "rony-tui"},
        {Content: "rony-llm-agent: librería Go...", ProjectID: "rony-llm-agent"},
    })
    
    runner := agent.NewRunner(agent.Config{
        LLM:    mockLLM,
        Memory: memory,
        Persona: testPersona,
    })
    
    resp, _ := runner.Run(context.Background(), []llm.Message{
        {Role: llm.RoleUser, Content: "¿qué proyectos tiene Victor?"},
    })
    
    // Verify LLM received context chunks in system prompt
    lastReq := mockLLM.LastRequest()
    assert.Contains(t, lastReq.Messages[0].Content, "Rony TUI")
    assert.Contains(t, lastReq.Messages[0].Content, "rony-llm-agent")
}

9.3 E2E test con Astro

# 1. Arrancar chat-bot en :7331
./bin/chat-bot serve &

# 2. Arrancar Astro en :4321
cd ../portfolio && npm run dev &

# 3. Hacer request al proxy de Astro
curl -X POST http://localhost:4321/api/chat \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"hola"}]}'

# 4. Verificar SSE stream

📂 10. Estructura del Proyecto

rony-chat-bot/
├── cmd/
│   └── chat-bot/
│       └── main.go                 # Entrypoint CLI
│
├── internal/
│   ├── server/                     # HTTP handlers
│   │   ├── server.go               # chi router + middleware
│   │   ├── handlers.go             # /api/chat, /api/health, /api/info, /api/reindex
│   │   └── middleware.go           # RequestID, Logging, CORS, RateLimit
│   │
│   ├── agent/                      # LLM client + RAG runner
│   │   ├── runner.go               # Wrapper Stream, inyección de RAG en system prompt
│   │   └── client.go               # Factory NewClient: llamacpp / ollama / openai / anthropic
│   │
│   ├── portfolio/                  # RAG: markdown → SQLite FTS5
│   │   ├── chunker.go              # Heading-based splitter
│   │   ├── indexer.go              # Store: schema, Reindex, Search (BM25)
│   │   └── chunker_test.go / store_test.go
│   │
│   ├── persona/                    # Bridge persona → rony-llm-agent
│   │   └── persona.go              # FromConfig, BuildSystemPrompt (con contexto RAG)
│   │
│   ├── streaming/                  # Helpers protocolo SSE
│   │   └── sse.go                  # WriteStart/Chunk/Sources/Done/Error
│   │
│   ├── i18n/                       # Detección de idioma (ES/EN) para la respuesta
│   │
│   └── config/                     # Loader YAML + validación
│
├── web/                            # ← WIDGET DE CHAT DROP-IN
│   ├── chat-widget.js              # Vanilla JS, ~12 KB
│   ├── chat-widget.css             # Estilos scoped, themable vía CSS custom props
│   ├── example.html                # Demo local (python -m http.server)
│   └── README.md                   # Guía de integración (HTML, Astro, Next.js)
│
├── data/
│   └── projects/                   # ← Markdown por proyecto (un .md por proyecto)
│       ├── rony-tui.md
│       ├── rony-llm-agent.md
│       └── example-project.md
│
├── configs/
│   └── portfolio-bot.yaml          # Provider + RAG + persona config
│
├── docs/
│   ├── architecture.md             # ← THIS FILE
│   └── architecture.es.md
│
├── bench/                          # Benchmark reproducible de drivers SQLite
│
├── go.mod                          # require rony-llm-agent, modernc.org/sqlite
└── README.md

📅 11. Roadmap

Fase 1: MVP (2-3 semanas)

  • Setup proyecto (go mod init, estructura)
  • HTTP server básico con un endpoint /api/chat
  • SSE streaming funcional
  • RAG indexer (lee data/projects/*.md → SQLite FTS5)
  • RAG retriever (query → top-k chunks)
  • Persona loader desde YAML
  • Integración con llama.cpp (qwen2.5:1.5b GGUF)
  • CLI: serve, reindex, ask
  • Tests básicos

Fase 2: Integración con Astro (1 semana)

  • Astro API route del proxy
  • React component del chat widget
  • E2E test: Astro → chat-bot → respuesta
  • Styling del widget (TailwindCSS)

Fase 3: Polish (1 semana)

  • Rate limiting robusto
  • Logging estructurado (JSON)
  • Health checks para monitoring
  • systemd service file
  • README + docs de deployment

Fase 4: Opcionales

  • Soporte para múltiples conversaciones (session ID)
  • Historial de chats persistido
  • Análisis de preguntas frecuentes
  • Multi-idioma (EN/ES switch)
  • Versión standalone CLI más pulida (chat-bot ask)

📐 12. Especificaciones de Calidad

12.1 Métricas de rendimiento

Métrica Target
TTFT (Time-to-first-token) <500ms con llama.cpp local
End-to-end (pregunta → respuesta completa) <3s para respuestas típicas
Memoria en reposo <150MB
RAG indexing speed ~100 docs/segundo
Retrieval latency <50ms para top-5

12.2 Pruebas requeridas

  • Unit tests: cobertura ≥70%
  • Integration tests: con mock LLM + SQLite FTS5 en memoria
  • E2E: al menos un flujo completo Astro → chat-bot

🔒 13. Seguridad

13.1 Implementado

  • Rate limiting por IP (default 30 req/min)
  • CORS restrictivo — solo origins configurados
  • Input validation — JSON schema validation en requests
  • No PII storage — no guardamos conversaciones por default
  • Local-only por default — sin llamadas a APIs cloud

13.2 Diferido / Opcional

  • Auth con API key (para uso privado)
  • Logging de queries para analytics
  • Anonymization de IPs en logs
  • HTTPS via reverse proxy (Caddy/nginx)

📚 14. Referencias


Documento listo para implementación. 🚀