# 📋 Rony Chat Bot — Technical Design Document **Versión:** 1.0 **Autor:** Victor Hugo Vargas **Fecha:** 2026-06-28 **Estado:** Especificación completa para implementación **Path:** `rony-chat-bot/docs/architecture.md` > 📚 **Workspace:** Este proyecto es parte del workspace `Rony/`. Ver [`../README.md`](../../README.md). > > 🔑 **Depende de:** [`rony-llm-agent`](https://github.com/VictorVargas/rony-llm-agent) — librería core que provee agent loop, LLM clients, RAG, persona system. > > 📐 **Metodología:** Este proyecto sigue el enfoque **SDD + DDD + Hexagonal Architecture**. Los Requisitos Funcionales se numeran como `CRF-XXX`. Ver [`../../METHODOLOGY.md`](../../METHODOLOGY.md). --- ## 🎯 1. Visión del Proyecto ### 1.1 ¿Qué es Chat-Bot? Un **chatbot HTTP** que responde preguntas sobre Victor Hugo Vargas y sus proyectos. Usa **RAG (Retrieval-Augmented Generation)** sobre archivos markdown que describen cada proyecto, y un LLM local (o cloud) para generar respuestas. ### 1.2 Caso de uso primario Victor tiene un portfolio web (Astro + React). En el sitio hay un widget de chat donde visitantes pueden preguntar: - "¿Qué proyectos ha hecho Victor?" - "¿Cuál es su experiencia con Go?" - "¿Cómo funciona Rony TUI?" - "¿Victor ha trabajado con PostgreSQL?" El bot responde con información precisa extraída de los archivos markdown de proyectos + bio + skills. ### 1.3 Casos de uso secundarios (futuro) - **Adaptación a clientes:** El mismo bot, con otra data y otra persona, sirve para concesionarios, restaurantes, etc. - **Standalone CLI:** `./chat-bot ask "¿qué sabes de X?"` para uso desde terminal. - **Slack/Discord bot:** Wrapper que consume el HTTP API. ### 1.4 Filosofía - **Self-hosted por defecto** — funciona 100% local con Ollama + modelos 1-3B - **Cloud opcional** — si se necesita más calidad, swap a Anthropic API - **Portable** — fácil de fork/customizar para otros contextos - **Streaming** — respuestas token-por-token con SSE (no espera a respuesta completa) - **Reutiliza `rony-llm-agent`** — no reinventar el agent loop --- ## 🏗️ 2. Arquitectura ### 2.1 Vista general ``` ┌─────────────────────────────────────────────────────────────────┐ │ Browser (Astro site) │ │ ↓ HTTP POST /api/chat │ │ Astro SSR (proxy) ←────────── Sirve portfolio + proxy chat │ │ ↓ HTTP POST /api/chat │ │ Chat-Bot HTTP server (:7331) │ │ ↓ │ │ Agent loop (rony-llm-agent) │ │ ↓ │ │ RAG retrieval → ChromaDB sobre data/projects/*.md │ │ ↓ │ │ LLM (Ollama local / Anthropic cloud) │ └─────────────────────────────────────────────────────────────────┘ ``` ### 2.2 Componentes principales | Componente | Path | Responsabilidad | |---|---|---| | **HTTP server** | `internal/server/` | Gin/chi handlers, SSE streaming | | **Agent runner** | `internal/agent/` | Wrapper sobre `rony-llm-agent` con config específica | | **Portfolio loader** | `internal/portfolio/` | Lee `data/projects/*.md`, indexa en ChromaDB | | **Persona** | `internal/persona/` | Carga persona desde `configs/portfolio-bot.yaml` | | **CLI** | `cm./rony-chat-bot/` | Comandos: `serve`, `reindex`, `ask`, `version` | ### 2.3 Stack tecnológico | Capa | Tecnología | Razón | |---|---|---| | **Lenguaje** | Go 1.26+ | Mismo que `harness`, aprovechar `os.Root`, `iter.Seq` | | **HTTP router** | `net/http` + `chi` | Stdlib + chi para middleware (CORS, logging) | | **SSE** | `net/http` Flusher | Stdlib es suficiente, no necesita librería externa | | **Config** | `gopkg.in/yaml.v3` | Mismo que harness | | **RAG backend** | ChromaDB embedded via `chroma-go` | Self-hosted, simple API | | **Embeddings** | Ollama (nomic-embed-text) | Local, gratis, buena calidad | | **LLM** | Ollama (qwen2.5:1.5b) o llama.cpp | Self-hosted por defecto | | **Tests** | stdlib + testify | Consistencia con el resto | --- ## 🔌 3. HTTP API ### 3.1 Endpoints #### `POST /api/chat` — Chat con streaming SSE **Request:** ```json { "messages": [ {"role": "user", "content": "¿Qué proyectos tiene Victor?"} ], "stream": true } ``` **Response (SSE):** ``` data: {"type":"start","conversation_id":"abc123"} data: {"type":"chunk","content":"Victor"} data: {"type":"chunk","content":" tiene"} data: {"type":"chunk","content":" varios"} data: {"type":"chunk","content":" proyectos"} data: {"type":"sources","documents":["rony-tui.md","rony-llm-agent.md"]} data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}} ``` **Sin streaming** (`"stream": false`): ```json { "content": "Victor tiene varios proyectos...", "sources": ["rony-tui.md", "rony-llm-agent.md"], "usage": {"input_tokens": 245, "output_tokens": 38} } ``` #### `POST /api/reindex` — Re-indexar portfolio Útil cuando se modifican archivos en `data/projects/`. **Request:** vacío **Response:** ```json { "indexed_files": 12, "total_chunks": 87, "duration_ms": 4321 } ``` #### `GET /api/health` — Health check ```json { "status": "ok", "version": "1.0.0", "providers": ["ollama-local"], "rag": { "documents": 12, "chunks": 87, "last_index": "2026-06-28T10:23:45Z" } } ``` #### `GET /api/info` — Metadata del bot ```json { "name": "Asistente de Victor Hugo Vargas", "model": "qwen2.5:1.5b", "persona": "...", "topics": ["proyectos", "experiencia", "skills técnicas"] } ``` ### 3.2 SSE Implementation ```go // internal/server/chat.go package server import ( "encoding/json" "fmt" "net/http" "github.com/VictorVargas/rony-llm-agent/pkg/agent" ) func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) { // Headers SSE w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") w.Header().Set("X-Accel-Buffering", "no") flusher, ok := w.(http.Flusher) if !ok { http.Error(w, "SSE no soportado", http.StatusInternalServerError) return } // Parse request var req ChatRequest if err := json.NewDecoder(r.Body).Decode(&req); err != nil { writeError(w, flusher, "invalid request", err) return } // Start event writeSSE(w, flusher, "start", map[string]string{ "conversation_id": generateConvID(), }) // Run agent con streaming sources := []string{} for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) { if err != nil { writeSSE(w, flusher, "error", map[string]string{"message": err.Error()}) return } if chunk.Type == "source" { sources = append(sources, chunk.Source) } writeSSE(w, flusher, chunk.Type, chunk.Data) } // Done event writeSSE(w, flusher, "done", map[string]any{ "usage": map[string]int{ "input_tokens": 245, "output_tokens": 38, }, }) } func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) { payload, _ := json.Marshal(data) fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload) flusher.Flush() } ``` ### 3.3 Middleware ```go // internal/server/middleware.go package server func (s *Server) loggingMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { start := time.Now() // Wrap response writer para capturar status rw := &statusRecorder{ResponseWriter: w, status: 200} next.ServeHTTP(rw, r) slog.Info("http.request", "method", r.Method, "path", r.URL.Path, "status", rw.status, "duration_ms", time.Since(start).Milliseconds(), "ip", r.RemoteAddr, ) }) } func (s *Server) corsMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { origin := r.Header.Get("Origin") for _, allowed := range s.config.Server.CORSOrigins { if origin == allowed { w.Header().Set("Access-Control-Allow-Origin", origin) w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS") w.Header().Set("Access-Control-Allow-Headers", "Content-Type") break } } if r.Method == "OPTIONS" { w.WriteHeader(204) return } next.ServeHTTP(w, r) }) } func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler { limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst) return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { if !limiter.Allow() { http.Error(w, "rate limit exceeded", http.StatusTooManyRequests) return } next.ServeHTTP(w, r) }) } ``` --- ## 🧠 4. RAG (Retrieval-Augmented Generation) ### 4.1 Pipeline de indexación ``` data/projects/*.md ↓ (read all files) Raw markdown content ↓ (split into chunks, ~500 chars, 50 overlap) Chunks [] ↓ (embed each chunk via Ollama nomic-embed-text) Vectors [][]float32 ↓ (store in ChromaDB collection "portfolio") Indexed corpus ``` **Cuándo se ejecuta:** - Al arrancar el bot (si `--reindex-on-start` flag) - Manualmente: `./chat-bot reindex` - Vía HTTP: `POST /api/reindex` ### 4.2 Pipeline de retrieval ``` User query "¿qué proyectos tiene Victor?" ↓ (embed query) Query vector ↓ (cosine similarity search en ChromaDB, top_k=5) Top 5 chunks relevantes ↓ (format as context block) System prompt += chunks relevantes ↓ (send to LLM) LLM generates answer ``` ### 4.3 Implementación ```go // internal/portfolio/indexer.go package portfolio import ( "context" "os" "path/filepath" "strings" "github.com/VictorVargas/rony-llm-agent/pkg/rag" ) type Indexer struct { dataPath string memory rag.Memory embedder rag.Embedder chunkSize int chunkOverlap int } func (i *Indexer) IndexAll(ctx context.Context) (int, error) { files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md")) if err != nil { return 0, err } totalChunks := 0 for _, file := range files { chunks, err := i.indexFile(ctx, file) if err != nil { slog.Warn("failed to index file", "file", file, "err", err) continue } totalChunks += chunks } return totalChunks, nil } func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) { content, err := os.ReadFile(path) if err != nil { return 0, err } projectID := strings.TrimSuffix(filepath.Base(path), ".md") chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap) for idx, chunk := range chunks { embedding, err := i.embedder.Embed(ctx, chunk) if err != nil { return idx, err } fragment := rag.Fragment{ ID: fmt.Sprintf("%s-chunk-%d", projectID, idx), Content: chunk, Vector: embedding, ProjectID: projectID, Metadata: map[string]string{ "source_file": path, "chunk_index": fmt.Sprint(idx), }, } if err := i.memory.Add(ctx, fragment); err != nil { return idx, err } } return len(chunks), nil } func splitIntoChunks(text string, size, overlap int) []string { // Implementación simple: split por tamaño con overlap // Versión production usa tokenizer-aware chunking var chunks []string for i := 0; i < len(text); i += size - overlap { end := i + size if end > len(text) { end = len(text) } chunks = append(chunks, text[i:end]) } return chunks } ``` ### 4.4 Retrieval en el agent loop ```go // internal/agent/runner.go package agent func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) { // 1. Base persona prompt basePrompt := r.persona.SystemPrompt // 2. Retrieve relevant chunks fragments, err := r.memory.Search(ctx, query, r.config.RAG.TopK) if err != nil { return "", err } // 3. Format as context var contextBlock strings.Builder contextBlock.WriteString(basePrompt) contextBlock.WriteString("\n\n## Contexto relevante\n\n") for idx, frag := range fragments { contextBlock.WriteString(fmt.Sprintf("### Fuente: %s\n%s\n\n", frag.Metadata["source_file"], frag.Content)) } return contextBlock.String(), nil } func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] { return func(yield func(Chunk, error) bool) { // Build prompt with RAG context lastUserMsg := getLastUserMessage(messages) systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg) if err != nil { yield(Chunk{}, err) return } // Inject system prompt messages = prependSystem(messages, systemPrompt) // Run agent loop for chunk, err := range r.loop.RunStream(ctx, messages) { if !yield(chunk, err) { return } } } } ``` --- ## 🌐 5. Integración con Astro (Portfolio) ### 5.1 Patrón recomendado: Astro proxy ``` [Browser] ←→ [Astro SSR :4321] ←→ [Chat-Bot :7331] ``` **Por qué proxy y no llamada directa del browser al chat-bot:** - ✅ Single domain (no CORS) - ✅ Astro maneja auth/sesión si se necesita - ✅ Puede haber rate limiting centralizado en Astro - ✅ El chat-bot queda en red privada (no expuesto a internet directamente) ### 5.2 Astro: API route del proxy ```typescript // portfolio/src/pages/api/chat.ts import type { APIRoute } from 'astro'; const CHAT_BOT_URL = process.env.CHAT_BOT_URL || 'http://localhost:7331'; export const POST: APIRoute = async ({ request }) => { const body = await request.json(); const resp = await fetch(`${CHAT_BOT_URL}/api/chat`, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify(body), }); if (!resp.ok) { return new Response('Chat bot error', { status: resp.status }); } // Stream SSE de vuelta al browser return new Response(resp.body, { status: 200, headers: { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache', 'Connection': 'keep-alive', }, }); }; ``` ### 5.3 React: Componente del chat ```tsx // portfolio/src/components/Chat.tsx import { useState, useRef } from 'react'; interface Message { role: 'user' | 'assistant'; content: string; } export default function Chat() { const [messages, setMessages] = useState([]); const [input, setInput] = useState(''); const [streaming, setStreaming] = useState(false); const abortRef = useRef(null); const send = async () => { if (!input.trim() || streaming) return; const userMsg: Message = { role: 'user', content: input }; setMessages(prev => [...prev, userMsg]); setInput(''); setStreaming(true); // Placeholder para streaming const assistantMsg: Message = { role: 'assistant', content: '' }; setMessages(prev => [...prev, assistantMsg]); abortRef.current = new AbortController(); try { const resp = await fetch('/api/chat', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ messages: [...messages, userMsg], stream: true, }), signal: abortRef.current.signal, }); const reader = resp.body!.getReader(); const decoder = new TextDecoder(); let buffer = ''; while (true) { const { done, value } = await reader.read(); if (done) break; buffer += decoder.decode(value, { stream: true }); const lines = buffer.split('\n\n'); buffer = lines.pop() || ''; for (const line of lines) { if (!line.startsWith('data: ')) continue; const event = JSON.parse(line.slice(6)); if (event.type === 'chunk') { setMessages(prev => { const updated = [...prev]; updated[updated.length - 1].content += event.data.content; return updated; }); } } } } catch (err) { if ((err as Error).name !== 'AbortError') { console.error(err); } } finally { setStreaming(false); abortRef.current = null; } }; const stop = () => abortRef.current?.abort(); return (
{messages.map((m, i) => (
{m.content || (streaming && i === messages.length - 1 ? '...' : '')}
))}
setInput(e.target.value)} onKeyDown={e => e.key === 'Enter' && send()} placeholder="Pregunta sobre Victor..." disabled={streaming} /> {streaming ? ( ) : ( )}
); } ``` --- ## 🤖 6. Self-hosting con Ollama ### 6.1 Setup ```bash # 1. Instalar Ollama curl -fsSL https://ollama.com/install.sh | sh # 2. Descargar modelo de chat ollama pull qwen2.5:1.5b # 3. Descargar modelo de embeddings ollama pull nomic-embed-text # 4. Verificar ollama list ``` ### 6.2 Configuración por defecto `configs/portfolio-bot.yaml` ya viene con Ollama como default. Solo necesitas: ```bash # Asegurar que Ollama está corriendo ollama serve # Arrancar el bot ./bin/chat-bot serve ``` ### 6.3 Alternativa: llama.cpp directo Para más control o si Ollama no funciona en tu setup: ```yaml providers: - name: llamacpp-local type: llamacpp model_path: ${RONY_MODELS_PATH}/qwen2.5-1.5b-instruct-q5_k_m.gguf context_size: 4096 n_gpu_layers: 999 # offload todo a GPU default: true ``` El adapter `llamacpp` se importa desde `rony-llm-agent/pkg/llm/providers/llamacpp` y se compila contra `llama.cpp` vía CGO o binario externo. --- ## 📦 7. CLI del bot ### 7.1 Comandos ```bash # Arrancar servidor HTTP chat-bot serve [--port 7331] [--host 0.0.0.0] [--reindex-on-start] # Re-indexar portfolio (lee data/projects/*.md → ChromaDB) chat-bot reindex # Pregunta única (sin servidor, útil para tests) chat-bot ask "¿Qué proyectos tiene Victor?" [--no-rag] # Validar config chat-bot config validate # Health check (útil para monitoring) chat-bot health # Versión chat-bot version ``` ### 7.2 Implementación con Cobra ```go // cm./rony-chat-bot/main.go package main import ( "github.com/spf13/cobra" ) func main() { root := &cobra.Command{ Use: "chat-bot", Short: "Portfolio chatbot HTTP server", } root.AddCommand(serveCmd()) root.AddCommand(reindexCmd()) root.AddCommand(askCmd()) root.AddCommand(configCmd()) root.AddCommand(healthCmd()) root.AddCommand(versionCmd()) if err := root.Execute(); err != nil { os.Exit(1) } } func serveCmd() *cobra.Command { var port int var host string var reindexOnStart bool cmd := &cobra.Command{ Use: "serve", Short: "Start HTTP server", RunE: func(cmd *cobra.Command, args []string) error { return server.Serve(server.Config{ Port: port, Host: host, ReindexOnStart: reindexOnStart, }) }, } cmd.Flags().IntVar(&port, "port", 7331, "HTTP port") cmd.Flags().StringVar(&host, "host", "0.0.0.0", "HTTP host") cmd.Flags().BoolVar(&reindexOnStart, "reindex-on-start", false, "Re-index RAG before serving") return cmd } ``` --- ## 🚀 8. Deployment ### 8.1 Recomendación: Self-hosted en VPS ```bash # 1. Instalar dependencias sudo apt install golang-go ollama ollama pull qwen2.5:1.5b ollama pull nomic-embed-text # 2. Build go build -o /usr/local/bin/chat-bot ./cmd/chat-bot # 3. systemd service cat > /etc/systemd/system/chat-bot.service <