Message gains an optional Parts []ContentPart alongside the existing
plain-text Content, so a turn can carry text plus image/video
attachments. Content stays the single source of truth for every
existing text-only caller (sidebar.go, memory_tools.go, etc. are
untouched); Parts only matters to a provider client when non-empty.
openai and llamacpp (both OpenAI-compatible) serialize Parts into the
standard text/image_url content-array shape; llamacpp additionally
passes video through as a best-effort video_url part, since llama.cpp
itself has no video support but the whole point of this client is the
user's own OpenAI-compatible server sitting in front of a
video-capable model — the server decides whether it understands it,
not this client. anthropic converts image parts to its base64 image
content block, and rejects a video part outright with a clear error:
the Messages API has no video block type at all, so sending one would
just produce a confusing 400 instead.
ProviderCapabilities gains SupportsVideo, true only for llamacpp.
Two failure modes seen live with Qwen3.6 on llama.cpp ended turns silently
mid-task:
- The model writes its tool call as plain text inside its reasoning, the
server never parses it, and the round ends with nothing executed. The
loop now detects the markers and nudges the model to re-issue the call
for real (max 2 per turn).
- llama.cpp silently ignores the max_thinking_tokens field, so a model in
a reasoning spiral ran until max_tokens (seen live: 25k+ tokens of
nonstop thinking, ~20 min). The llamacpp client now enforces the budget
client-side during Stream: once exceeded while the round is still pure
reasoning, it cuts with FinishThinkingBudget and aborts the request
(freeing the server slot); the loop answers with its own corrective
nudge, on a separate counter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Formatting only (struct field alignment, import ordering) across the
files that didn't comply — no semantic changes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- llamacpp: 4MB SSE scanner buffer (the 64KB bufio.Scanner default killed
streams whose single line exceeded it, e.g. a write tool call carrying a
whole file) and an empty-choices guard in toResponse instead of a panic;
request payload now uses bytes.NewReader (drops a full string copy).
- anthropic: Capabilities() reported a 1M-token context window for any
non-haiku model. Callers use that number to decide when to compact, so
compaction would have fired far too late and requests overflowed the
real window. Default is now the standard 200k, configurable via
Config.ContextWindow for extended-window models/plans.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- pkg/llm/providers/anthropic: new client implementation (was previously
imported by rony-harness but never committed here, so a fresh clone
wouldn't build)
- pkg/llm/providers/llamacpp: Config/Client gain the full local-model
sampling surface (max_tokens, context_window, top_k/top_p/min_p,
presence/repetition penalty, max_thinking_tokens) to match the
llamacpp-local* entries added to configs/ai_providers.yaml
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- agent/loop.go: Record assistant message with ToolCalls before tool results,
and set ToolCallID on tool-result messages so follow-up requests have complete
context (prevents models from losing track of already-attempted tools).
- llm/providers/llamacpp/client.go: Buffer fragmented tool call deltas during
streaming, assemble them into complete calls when finish_reason arrives.
Add ToolCalls, ToolCallID, Name fields to request building.
- llm/providers/openai/client.go: Send ToolCalls, ToolCallID, Name when
building chat requests so messages are wire-format correct.
- llm/types.go: Add ToolCalls field to Message struct for serialization
back into conversation history.
- agent/integration_test.go: Move integration test skip from TestMain to a
per-test skipUnlessIntegration() so it doesn't hide other package tests.
- sandbox & tools: Add edge-case tests (relative traversal, array paths,
non-path strings, zero-value guards, sentinel errors).
Backend.Upsert never received the fragment's Content, so ChromaDB (and
any backend) stored the vector but silently dropped the actual text —
saved memories had nothing to retrieve later. Backend.Search now also
takes the raw query text, and a failed/missing embedding no longer
hard-fails Add/Search: it degrades to a nil vector so a lexical-capable
backend can still index/find the content (Chroma has no such fallback
and now says so explicitly instead of misbehaving).
Adds pkg/rag/backends/sqlitevec: a zero-dependency backend (pure-Go
SQLite, no external service) that does cosine similarity when a real
embedding vector is available and falls back to FTS5/BM25 full-text
search otherwise. Adds pkg/rag/embeddings.OpenAICompatible, covering
both a local llama.cpp server (`--embeddings` enabled) and real OpenAI
(or any OpenAI-shaped /embeddings endpoint) through the same client.
Also fixes token usage tracking for llama.cpp streaming: the client
never requested `stream_options.include_usage` nor parsed a usage-only
SSE event, and even when present, the agent loop's RunStream dropped
any chunk with no Delta/ReasoningDelta — silently discarding the only
chunk that carries usage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Add optional history parameter to Run/RunStream for conversation context
- Add Reasoning/ReasoningDelta fields to completion and stream types
- Update llama.cpp adapter to propagate reasoning content from responses
- Default persona language now adapts to the user's language dynamically
- Add ChatTemplateKwargs field to llm.CompletionRequest
- Propagate kwargs through agent loop in both Run() and RunStream()
- Pass kwargs to llama.cpp client chat request
- Fix tool schema marshaling to include type/function wrapper
- Fix stream indentation logic in RunStream with responseBuilder
- Remove indirect marker from uuid dependency