# 📋 Rony Chat Bot — Technical Design Document > 🌐 **Idioma:** [English](architecture.md) | [Español](architecture.es.md) **Versión:** 1.0 **Autor:** Victor Hugo Vargas **Fecha:** 2026-06-28 **Estado:** Especificación completa para implementación **Path:** `rony-chat-bot/docs/architecture.md` > 📚 **Workspace:** Este proyecto es parte del workspace `Rony/`. Ver [`../README.md`](../../README.md). > > 🔑 **Depende de:** [`rony-llm-agent`](https://github.com/VictorVargas/rony-llm-agent) — librería core que provee agent loop, LLM clients, RAG, persona system. > > 📐 **Metodología:** Este proyecto sigue el enfoque **SDD + DDD + Hexagonal Architecture**. Los Requisitos Funcionales se numeran como `CRF-XXX`. Ver [`../../METHODOLOGY.md`](../../METHODOLOGY.md). --- ## 🎯 1. Visión del Proyecto ### 1.1 ¿Qué es Chat-Bot? Un **chatbot HTTP** que responde preguntas sobre Victor Hugo Vargas y sus proyectos. Usa **RAG (Retrieval-Augmented Generation)** sobre archivos markdown que describen cada proyecto, y un LLM local (o cloud) para generar respuestas. ### 1.2 Caso de uso primario Victor tiene un portfolio web (Astro + React). En el sitio hay un widget de chat donde visitantes pueden preguntar: - "¿Qué proyectos ha hecho Victor?" - "¿Cuál es su experiencia con Go?" - "¿Cómo funciona Rony TUI?" - "¿Victor ha trabajado con PostgreSQL?" El bot responde con información precisa extraída de los archivos markdown de proyectos + bio + skills. ### 1.3 Casos de uso secundarios (futuro) - **Adaptación a clientes:** El mismo bot, con otra data y otra persona, sirve para concesionarios, restaurantes, etc. - **Standalone CLI:** `./chat-bot ask "¿qué sabes de X?"` para uso desde terminal. - **Slack/Discord bot:** Wrapper que consume el HTTP API. ### 1.4 Filosofía - **Self-hosted por defecto** — funciona 100% local con Ollama + modelos 1-3B - **Cloud opcional** — si se necesita más calidad, swap a Anthropic API - **Portable** — fácil de fork/customizar para otros contextos - **Streaming** — respuestas token-por-token con SSE (no espera a respuesta completa) - **Reutiliza `rony-llm-agent`** — no reinventar el agent loop --- ## 🏗️ 2. Arquitectura ### 2.1 Vista general ``` ┌─────────────────────────────────────────────────────────────────┐ │ Browser (Astro site) │ │ ↓ HTTP POST /api/chat │ │ Astro SSR (proxy) ←────────── Sirve portfolio + proxy chat │ │ ↓ HTTP POST /api/chat │ │ Chat-Bot HTTP server (:7331) │ │ ↓ │ │ Agent loop (rony-llm-agent) │ │ ↓ │ │ RAG retrieval → SQLite FTS5 sobre data/projects/*.md │ │ ↓ │ │ LLM (llama.cpp local default / Ollama o Anthropic opcionales) │ └─────────────────────────────────────────────────────────────────┘ ``` ### 2.2 Componentes principales | Componente | Path | Responsabilidad | |---|---|---| | **HTTP server** | `internal/server/` | Gin/chi handlers, SSE streaming | | **Agent runner** | `internal/agent/` | Wrapper sobre `rony-llm-agent` con config específica | | **Portfolio loader** | `internal/portfolio/` | Lee `data/projects/*.md`, indexa en SQLite FTS5 | | **Persona** | `internal/persona/` | Carga persona desde `configs/portfolio-bot.yaml` | | **CLI** | `cm./rony-chat-bot/` | Comandos: `serve`, `reindex`, `ask`, `version` | ### 2.3 Stack tecnológico | Capa | Tecnología | Razón | |---|---|---| | **Lenguaje** | Go 1.26+ | Mismo que `harness`, aprovechar `os.Root`, `iter.Seq` | | **HTTP router** | `net/http` + `chi` | Stdlib + chi para middleware (CORS, logging) | | **SSE** | `net/http` Flusher | Stdlib es suficiente, no necesita librería externa | | **Config** | `gopkg.in/yaml.v3` | Mismo que harness | | **RAG backend** | SQLite + FTS5 (BM25) | Sin dependencias externas, un solo archivo, rápido | | **LLM** | llama.cpp (qwen2.5:1.5b GGUF) — default; Ollama como alternativa | Self-hosted por defecto | | **Tests** | stdlib + testify | Consistencia con el resto | --- ## 🔌 3. HTTP API ### 3.1 Endpoints #### `POST /api/chat` — Chat con streaming SSE **Request:** ```json { "messages": [ {"role": "user", "content": "¿Qué proyectos tiene Victor?"} ], "stream": true } ``` **Response (SSE):** ``` data: {"type":"start","conversation_id":"abc123"} data: {"type":"chunk","content":"Victor"} data: {"type":"chunk","content":" tiene"} data: {"type":"chunk","content":" varios"} data: {"type":"chunk","content":" proyectos"} data: {"type":"sources","documents":["rony-tui.md","rony-llm-agent.md"]} data: {"type":"done","usage":{"input_tokens":245,"output_tokens":38}} ``` **Sin streaming** (`"stream": false`): ```json { "content": "Victor tiene varios proyectos...", "sources": ["rony-tui.md", "rony-llm-agent.md"], "usage": {"input_tokens": 245, "output_tokens": 38} } ``` #### `POST /api/reindex` — Re-indexar portfolio Útil cuando se modifican archivos en `data/projects/`. **Request:** vacío **Response:** ```json { "indexed_files": 12, "total_chunks": 87, "duration_ms": 4321 } ``` #### `GET /api/health` — Health check (real) Prueba el LLM provider y el store SQLite en paralelo y reporta su estado. Pensado para monitoring / load balancers. **Devuelve 200 cuando está healthy o degraded, 503 cuando está unhealthy.** - `?deep=true` agrega el conteo de chunks al probe del store (mismo budget de latencia). **Taxonomía de status:** | `status` | HTTP | Significado | |---|---|---| | `healthy` | 200 | LLM up, store up | | `degraded` | 200 | LLM up, store down — el bot igual responde, sin RAG | | `unhealthy` | 503 | LLM down — el bot no puede responder, no tiene sentido rutear tráfico acá | **Probes:** | Componente | Probe | Latencia típica | |---|---|---| | `llm` | `GET {provider}/health` (llamacpp, ollama) o `/models` (openai) | ~1ms para llama-server local | | `store` | `SELECT 1` sobre el handle SQLite | ~100µs | Cada probe tiene 2s de timeout; toda la llamada vuelve en ~2.5s aunque una dependencia esté colgada. **Shape de respuesta (healthy):** ```json { "status": "healthy", "version": "0.2.0-dev", "checked_at": "2026-07-17T05:02:07Z", "components": { "llm": { "status": "up", "latency": "1.028ms", "details": {"provider": "llamacpp", "model": "qwen2.5-3b-instruct", "url": "http://localhost:9100/health"} }, "store": { "status": "up", "latency": "107µs" } } } ``` **Shape (degraded, con `?deep=true`):** ```json { "status": "degraded", "version": "0.2.0-dev", "checked_at": "2026-07-17T05:02:07Z", "components": { "llm": {"status": "up", "latency": "0.8ms", "details": {...}}, "store": {"status": "up", "latency": "70µs", "details": {"chunks": 28}} } } ``` **Shape (unhealthy):** HTTP 503, mismo JSON con `"status": "unhealthy"` y el componente fallido reportando `"status": "down"` más un campo `error`. #### `GET /api/info` — Metadata del bot ```json { "name": "Asistente de Victor Hugo Vargas", "model": "qwen2.5:1.5b", "persona": "...", "topics": ["proyectos", "experiencia", "skills técnicas"] } ``` ### 3.2 SSE Implementation ```go // internal/server/chat.go package server import ( "encoding/json" "fmt" "net/http" "github.com/VictorVargas/rony-llm-agent/pkg/agent" ) func (s *Server) handleChat(w http.ResponseWriter, r *http.Request) { // Headers SSE w.Header().Set("Content-Type", "text/event-stream") w.Header().Set("Cache-Control", "no-cache") w.Header().Set("Connection", "keep-alive") w.Header().Set("X-Accel-Buffering", "no") flusher, ok := w.(http.Flusher) if !ok { http.Error(w, "SSE no soportado", http.StatusInternalServerError) return } // Parse request var req ChatRequest if err := json.NewDecoder(r.Body).Decode(&req); err != nil { writeError(w, flusher, "invalid request", err) return } // Start event writeSSE(w, flusher, "start", map[string]string{ "conversation_id": generateConvID(), }) // Run agent con streaming sources := []string{} for chunk, err := range s.agent.RunStream(r.Context(), req.Messages) { if err != nil { writeSSE(w, flusher, "error", map[string]string{"message": err.Error()}) return } if chunk.Type == "source" { sources = append(sources, chunk.Source) } writeSSE(w, flusher, chunk.Type, chunk.Data) } // Done event writeSSE(w, flusher, "done", map[string]any{ "usage": map[string]int{ "input_tokens": 245, "output_tokens": 38, }, }) } func writeSSE(w http.ResponseWriter, flusher http.Flusher, eventType string, data any) { payload, _ := json.Marshal(data) fmt.Fprintf(w, "data: {\"type\":%q,\"data\":%s}\n\n", eventType, payload) flusher.Flush() } ``` ### 3.3 Middleware ```go // internal/server/middleware.go package server func (s *Server) loggingMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { start := time.Now() // Wrap response writer para capturar status rw := &statusRecorder{ResponseWriter: w, status: 200} next.ServeHTTP(rw, r) slog.Info("http.request", "method", r.Method, "path", r.URL.Path, "status", rw.status, "duration_ms", time.Since(start).Milliseconds(), "ip", r.RemoteAddr, ) }) } func (s *Server) corsMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { origin := r.Header.Get("Origin") for _, allowed := range s.config.Server.CORSOrigins { if origin == allowed { w.Header().Set("Access-Control-Allow-Origin", origin) w.Header().Set("Access-Control-Allow-Methods", "POST, GET, OPTIONS") w.Header().Set("Access-Control-Allow-Headers", "Content-Type") break } } if r.Method == "OPTIONS" { w.WriteHeader(204) return } next.ServeHTTP(w, r) }) } func (s *Server) rateLimitMiddleware(next http.Handler) http.Handler { limiter := rate.NewLimiter(rate.Every(time.Minute/time.Duration(s.config.Server.RateLimit.RequestsPerMinute)), s.config.Server.RateLimit.Burst) return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { if !limiter.Allow() { http.Error(w, "rate limit exceeded", http.StatusTooManyRequests) return } next.ServeHTTP(w, r) }) } ``` --- ## 🧠 4. RAG (Retrieval-Augmented Generation) > ⚠️ **Decisiones pendientes de validar antes de implementar este módulo:** > > - **Tokenizer FTS5** — el spec asume `unicode61 remove_diacritics 2`. Confirmar con datos reales si conviene cambiar a `porter` (stemming EN), `trigram` (sub-string matching) o un tokenizer custom para español. **Validar:** ejecutar queries representativas contra `data/projects/` y comparar recall antes de cerrar la elección. > - **Driver SQLite** — ✅ **DECIDIDO: `modernc.org/sqlite`** (puro Go, sin CGO). Ver benchmark abajo. > - **Chunking** — el split por tamaño fijo (500 chars / 50 overlap) corta headings y code blocks arbitrariamente. **Validar:** medir recall con chunks por sección markdown (split por `#`/`##`) vs por tamaño. > - **Sin similitud semántica** — BM25 no matchea "IA" con "machine learning" salvo que la palabra esté literal. **Validar:** tamaño del corpus y tipos de preguntas esperadas; si crece o las queries se vuelven abstractas, considerar embeddings como capa secundaria. ### 4.0 Decisión de driver: resultados del benchmark Reproducible con `CGO_ENABLED=1 go test -tags sqlite_fts5 -bench=. ./bench/`. Datos: 4 markdowns → 11 chunks. | Operación | mattn (CGO) | modernc (puro Go) | Diferencia | |---|---|---|---| | **Insert** (11 chunks) | 2,802,843 ns/op | **1,465,646 ns/op** | modernc 1.9× más rápido | | Insert alloc | 2,124,299 B/op | **9,770 B/op** | modernc usa 217× menos memoria | | **Query** (8 queries BM25) | **244,047 ns/op** | 555,162 ns/op | mattn 2.3× más rápido | | **Round-trip** (insert + 8 queries) | 3,543,417 ns/op | **2,267,669 ns/op** | modernc 1.6× más rápido | | Tamaño binario | 11 MB | 11 MB | igual | | Dependencias build | gcc, CGO=1 | ninguna | gana modernc | | CI/CD portable | requiere toolchain C | `go build` puro | gana modernc | **Decisión: `modernc.org/sqlite`**. Justificación: 1. Ambas latencias de query (~250µs vs ~550µs) son **2 órdenes de magnitud por debajo** del target de 50ms — imperceptible vs el LLM (varios segundos). 2. modernc gana en inserts (1.9×) y round-trip (1.6×), que es el path de reindex. 3. Sin CGO = CI/CD más simple (sin gcc, sin Alpine musl-dev, binarios reproducibles). 4. Si en el futuro el cuello de botella pasa a ser query latency (corpus >10k chunks), se puede reconsiderar. Hoy no. ### 4.1 Pipeline de indexación ``` data/projects/*.md ↓ (read all files) Raw markdown content ↓ (split into chunks, ~500 chars, 50 overlap) Chunks [] ↓ (insert into SQLite FTS5 virtual table "portfolio_chunks") Indexed corpus ``` **Cuándo se ejecuta:** - Al arrancar el bot (si `--reindex-on-start` flag) - Manualmente: `./chat-bot reindex` - Vía HTTP: `POST /api/reindex` ### 4.2 Pipeline de retrieval ``` User query "¿qué proyectos tiene Victor?" ↓ (FTS5 MATCH query, BM25 ranking, top_k=5) Top 5 chunks relevantes ↓ (format as context block) System prompt += chunks relevantes ↓ (send to LLM) LLM generates answer ``` ### 4.3 Implementación ```go // internal/portfolio/indexer.go package portfolio import ( "context" "database/sql" "fmt" "log/slog" "os" "path/filepath" "strings" ) type Indexer struct { dataPath string db *sql.DB chunkSize int chunkOverlap int } func (i *Indexer) IndexAll(ctx context.Context) (int, error) { files, err := filepath.Glob(filepath.Join(i.dataPath, "*.md")) if err != nil { return 0, err } // Reconstruir el índice FTS5 desde cero (DELETE+INSERT es más rápido // que diff para corpus pequeños) if _, err := i.db.ExecContext(ctx, `DELETE FROM portfolio_chunks`); err != nil { return 0, fmt.Errorf("clear index: %w", err) } totalChunks := 0 for _, file := range files { chunks, err := i.indexFile(ctx, file) if err != nil { slog.Warn("failed to index file", "file", file, "err", err) continue } totalChunks += chunks } return totalChunks, nil } func (i *Indexer) indexFile(ctx context.Context, path string) (int, error) { content, err := os.ReadFile(path) if err != nil { return 0, err } projectID := strings.TrimSuffix(filepath.Base(path), ".md") chunks := splitIntoChunks(string(content), i.chunkSize, i.chunkOverlap) tx, err := i.db.BeginTx(ctx, nil) if err != nil { return 0, err } defer tx.Rollback() stmt, err := tx.PrepareContext(ctx, ` INSERT INTO portfolio_chunks (id, project_id, source_file, chunk_index, content) VALUES (?, ?, ?, ?, ?) `) if err != nil { return 0, err } defer stmt.Close() for idx, chunk := range chunks { id := fmt.Sprintf("%s-chunk-%d", projectID, idx) if _, err := stmt.ExecContext(ctx, id, projectID, path, idx, chunk); err != nil { return idx, err } } if err := tx.Commit(); err != nil { return 0, err } return len(chunks), nil } // schema.go — aplicado al arrancar const schema = ` CREATE VIRTUAL TABLE IF NOT EXISTS portfolio_chunks USING fts5( id UNINDEXED, project_id UNINDEXED, source_file UNINDEXED, chunk_index UNINDEXED, content, tokenize = 'unicode61 remove_diacritics 2' ); ` func splitIntoChunks(text string, size, overlap int) []string { // Implementación simple: split por tamaño con overlap // Versión production usa tokenizer-aware chunking var chunks []string for i := 0; i < len(text); i += size - overlap { end := i + size if end > len(text) { end = len(text) } chunks = append(chunks, text[i:end]) } return chunks } ``` ### 4.4 Retrieval en el agent loop ```go // internal/portfolio/search.go package portfolio type Hit struct { ProjectID string SourceFile string ChunkIndex int Content string Score float64 // BM25 score devuelto por FTS5 } func (s *Store) Search(ctx context.Context, query string, topK int) ([]Hit, error) { // Escapar input del usuario: la sintaxis FTS5 puede romperse con caracteres especiales ftsQuery := sanitizeFTS5(query) rows, err := s.db.QueryContext(ctx, ` SELECT project_id, source_file, chunk_index, content, bm25(portfolio_chunks) AS score FROM portfolio_chunks WHERE portfolio_chunks MATCH ? ORDER BY score LIMIT ? `, ftsQuery, topK) if err != nil { return nil, err } defer rows.Close() var hits []Hit for rows.Next() { var h Hit if err := rows.Scan(&h.ProjectID, &h.SourceFile, &h.ChunkIndex, &h.Content, &h.Score); err != nil { return nil, err } hits = append(hits, h) } return hits, rows.Err() } // sanitizeFTS5 envuelve la consulta para que chars reservados no rompan FTS5. // Para un bot de Q&A: agrega wildcard prefix-match a cada token. func sanitizeFTS5(q string) string { tokens := strings.FieldsFunc(q, func(r rune) bool { return !(r == '-' || r == '_' || (r >= '0' && r <= '9') || (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || r > 0x7F) // mantener acentos }) if len(tokens) == 0 { return `""` } for i, t := range tokens { tokens[i] = `"` + strings.ToLower(t) + `"*` } return strings.Join(tokens, " ") } ``` ```go // internal/agent/runner.go package agent func (r *Runner) buildSystemPrompt(ctx context.Context, query string) (string, error) { basePrompt := r.persona.SystemPrompt hits, err := r.store.Search(ctx, query, r.config.RAG.TopK) if err != nil { return "", err } if len(hits) == 0 { return basePrompt, nil } var contextBlock strings.Builder contextBlock.WriteString(basePrompt) contextBlock.WriteString("\n\n## Relevant context\n\n") for _, h := range hits { contextBlock.WriteString(fmt.Sprintf("### Source: %s\n%s\n\n", h.SourceFile, h.Content)) } return contextBlock.String(), nil } func (r *Runner) RunStream(ctx context.Context, messages []llm.Message) iter.Seq2[Chunk, error] { return func(yield func(Chunk, error) bool) { lastUserMsg := getLastUserMessage(messages) systemPrompt, err := r.buildSystemPrompt(ctx, lastUserMsg) if err != nil { yield(Chunk{}, err) return } messages = prependSystem(messages, systemPrompt) for chunk, err := range r.loop.RunStream(ctx, messages) { if !yield(chunk, err) { return } } } } ``` **Por qué esto es más simple que embeddings:** - Sin modelo de embeddings que descargar ni ejecutar (ahorra ~270MB de RAM y ~200ms por consulta) - Un archivo (`data/portfolio.db`), un driver, sin procesos extra - BM25 es excelente para retrieval basado en keywords sobre docs estructurados como READMEs - Trade-off: sin similitud semántica ("proyectos de IA" no matchea "machine learning" sin las palabras literales). Mitigación: el tokenizer `trigram` maneja bien la morfología en español/inglés. --- ## 🗜️ 4.5 Auto-compactación Las conversaciones largas eventualmente agotan el contexto — con la ventana de 4k de qwen2.5-3b, el system prompt de ~3k tokens + el bloque RAG sólo deja espacio para 2–3 turnos del usuario. La auto-compactación resuelve esto plegando la parte más antigua de la conversación en un único mensaje-resumen del sistema cuando los tokens de entrada del turno anterior cruzan un umbral configurable. ### Cuándo se dispara `agent.Runner.Compact` corre una vez por request a `/api/chat`, antes de la búsqueda RAG. Compara los `Usage.InputTokens` más recientes del runner (reportados por el provider en el chunk streameado previo) contra `client.Capabilities().MaxContextWindow × threshold_ratio`. | Config | Default | Qué controla | |---|---|---| | `compaction.enabled` | `false` | Switch maestro. | | `compaction.threshold_ratio` | `0.75` | Dispara cuando tokens usados ≥ ventana × ratio. | | `compaction.keep_recent_turns` | `4` | Cuántos turnos recientes del usuario se preservan literales tras la compactación. | | `compaction.summary_system_prompt` | *(bilingüe built-in)* | Override de la instrucción enviada al LLM al resumir. | Sale silenciosamente cuando la compactación está deshabilitada, el provider no reporta ventana (`Capabilities().MaxContextWindow == 0`), la historia es más corta que `keep_recent_turns`, o el usage aún es desconocido (primer turno). ### Cómo se hace el resumen 1. `splitByTurns(history, keep_recent_turns)` divide los mensajes en `(older, recent)` cortando en límites de rol `user`, así el par user/assistant de un turno preservado queda siempre junto. 2. `renderTranscript(older)` aplana los mensajes antiguos en una transcripción `User:` / `Assistant:` (saltando mensajes tool y placeholders vacíos de assistant). 3. El runner llama a `client.Generate(...)` con el prompt de resumen + la transcripción y un cap de 512 tokens para que la compactación en sí misma sea barata. 4. El texto devuelto se antepone como mensaje de sistema (`"Earlier conversation summary:\n…"`), seguido por la cola reciente. 5. `LastCompaction()` devuelve `CompactionStats` para que el handler SSE emita el evento `compaction` justo antes de los chunks streameados. ### Modo de falla Si `Generate` falla o devuelve un resumen vacío, la compactación cae a `truncateToBudget`: descarta turnos antiguos del usuario uno por uno hasta que el slice restante entre en `threshold` tokens (heurística: `len(s) / 4 + 1`). El turno actual del usuario siempre se preserva. El fallback se loggea a nivel WARN y el request sigue — un fallo del resumidor nunca rompe la request del usuario. ### Protocolo de cable Las respuestas streameadas ganan un evento opcional `compaction`: ``` data: {"type":"compaction","older_turns":6,"kept_turns":2,"summary_tokens":120,"window_tokens":4096,"used_tokens":3500} ``` Se emite después del `start` (cuando aplica) y antes de `sources` / `chunk`. El widget puede renderizar esto como un hint sutil "Contexto compactado" o ignorarlo — ambas son válidas. ### Persistencia La compactación es **por-request**. La transcripción completa igual se guarda en `messages` en `data/portfolio.db` literal, así que `GET /api/conversations/{id}` siempre devuelve la historia original. Sólo se reduce lo que se le manda al LLM — la próxima sesión puede releer el thread completo desde la DB. --- ## 🌐 5. Embebiendo el widget El bot viene con un widget vanilla-JS drop-in. Agrega dos archivos a tu sitio y funciona. ### 5.1 El widget (cualquier sitio) ```html ``` Aparece una burbuja abajo a la derecha, abre un panel, habla SSE con `/api/chat`, streamea la respuesta y cita las fuentes. Sin build step, sin React/Vue, sin lock-in de framework. **Opciones browser→bot:** | Topología | Trade-offs | |---|---| | **Directo** (browser → bot, mismo dominio o CORS) | Lo más simple. Agrega el origen del bot a `cors_origins` en YAML. | | **Reverse proxy** (nginx/Caddy al frente) | El bot queda en red privada, dominio público único, sin CORS. | | **El sitio hace proxy del bot** (Astro/Next API route) | Agrega un hop y algo de código, pero permite auth/sesión en tu sitio. | El widget funciona igual en las tres. Elige la que se ajuste a tu infra. > **El setup dev default es directo + CORS.** `cors_origins` en `configs/portfolio-bot.yaml` controla qué sitios pueden llamar al bot. Agregá el origen de tu sitio ahí. ### 5.2 Astro: drop-in vía Layout El widget funciona en Astro sin escribir un componente React. Agregá esto a tu layout compartido: ```astro --- // src/layouts/BaseLayout.astro import "../path/to/chat-widget.css"; const apiUrl = import.meta.env.PUBLIC_CHAT_API_URL || "http://localhost:7331"; --- ``` `is:inline` evita que Astro transforme/hash el `