rony-llm-agent/pkg/llm
Victor Vargas e6830161bc fix(llamacpp,anthropic): SSE headroom, no-choices guard, real context window
- llamacpp: 4MB SSE scanner buffer (the 64KB bufio.Scanner default killed
  streams whose single line exceeded it, e.g. a write tool call carrying a
  whole file) and an empty-choices guard in toResponse instead of a panic;
  request payload now uses bytes.NewReader (drops a full string copy).
- anthropic: Capabilities() reported a 1M-token context window for any
  non-haiku model. Callers use that number to decide when to compact, so
  compaction would have fired far too late and requests overflowed the
  real window. Default is now the standard 200k, configurable via
  Config.ContextWindow for extended-window models/plans.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 16:14:15 -07:00
..
mock test(llm,embeddings): add unit tests for mock client, types, and Ollama embedder 2026-07-03 14:22:41 -07:00
providers fix(llamacpp,anthropic): SSE headroom, no-choices guard, real context window 2026-07-12 16:14:15 -07:00
README.es.md docs(i18n): translate all docs to English (with .es.md as Spanish alternative) 2026-06-30 13:40:37 -07:00
README.md docs(i18n): translate all docs to English (with .es.md as Spanish alternative) 2026-06-30 13:40:37 -07:00
types.go Fix tool call tracking and streaming assembly for all providers 2026-07-08 16:11:57 -07:00
types_test.go test(llm,embeddings): add unit tests for mock client, types, and Ollama embedder 2026-07-03 14:22:41 -07:00

pkg/llm

Multi-provider abstraction for language models.

Responsibility

Define a common interface (LLMClient) and adapters for the main providers.

Public API

type LLMClient interface {
    Generate(ctx context.Context, req CompletionRequest) (CompletionResponse, error)
    Stream(ctx context.Context, req CompletionRequest) iter.Seq2[StreamChunk, error]
    Name() string
    Capabilities() ProviderCapabilities
}

type CompletionRequest struct {
    Messages    []Message
    Tools       []tools.Tool
    ToolChoice  ToolChoice
    Model       string
    Temperature *float32
    MaxTokens   *int
}

type CompletionResponse struct {
    Content    string
    ToolCalls  []tools.Call
    Usage      TokenUsage
    StopReason string
}

type ProviderCapabilities struct {
    SupportsTools    bool
    SupportsVision   bool
    MaxContextWindow int
}

Included providers

Provider Package Tool support
OpenAI providers/openai
Anthropic providers/anthropic
Ollama providers/ollama (models that support it)
llama.cpp providers/llamacpp (with grammar)

Usage

import "github.com/VictorVargas/rony-llm-agent/pkg/llm/providers/anthropic"

client, err := anthropic.New(anthropic.Config{
    APIKey: os.Getenv("ANTHROPIC_API_KEY"),
    Model:  "claude-sonnet-4.5",
})

resp, err := client.Generate(ctx, llm.CompletionRequest{
    Messages: []llm.Message{
        {Role: llm.RoleUser, Content: "Hello"},
    },
})

Streaming

for chunk, err := range client.Stream(ctx, req) {
    if err != nil { return err }
    fmt.Print(chunk.Delta)
}

Mock for tests

import "github.com/VictorVargas/rony-llm-agent/pkg/llm/mock"

mockClient := mock.New(mock.Responses{
    {Match: "hello", Response: "Hi! How are you?"},
    {Match: "*",    Response: "default"},
})

See also