- llamacpp: 4MB SSE scanner buffer (the 64KB bufio.Scanner default killed streams whose single line exceeded it, e.g. a write tool call carrying a whole file) and an empty-choices guard in toResponse instead of a panic; request payload now uses bytes.NewReader (drops a full string copy). - anthropic: Capabilities() reported a 1M-token context window for any non-haiku model. Callers use that number to decide when to compact, so compaction would have fired far too late and requests overflowed the real window. Default is now the standard 200k, configurable via Config.ContextWindow for extended-window models/plans. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| mock | ||
| providers | ||
| README.es.md | ||
| README.md | ||
| types.go | ||
| types_test.go | ||
pkg/llm
Multi-provider abstraction for language models.
Responsibility
Define a common interface (LLMClient) and adapters for the main providers.
Public API
type LLMClient interface {
Generate(ctx context.Context, req CompletionRequest) (CompletionResponse, error)
Stream(ctx context.Context, req CompletionRequest) iter.Seq2[StreamChunk, error]
Name() string
Capabilities() ProviderCapabilities
}
type CompletionRequest struct {
Messages []Message
Tools []tools.Tool
ToolChoice ToolChoice
Model string
Temperature *float32
MaxTokens *int
}
type CompletionResponse struct {
Content string
ToolCalls []tools.Call
Usage TokenUsage
StopReason string
}
type ProviderCapabilities struct {
SupportsTools bool
SupportsVision bool
MaxContextWindow int
}
Included providers
| Provider | Package | Tool support |
|---|---|---|
| OpenAI | providers/openai |
✅ |
| Anthropic | providers/anthropic |
✅ |
| Ollama | providers/ollama |
✅ (models that support it) |
| llama.cpp | providers/llamacpp |
✅ (with grammar) |
Usage
import "github.com/VictorVargas/rony-llm-agent/pkg/llm/providers/anthropic"
client, err := anthropic.New(anthropic.Config{
APIKey: os.Getenv("ANTHROPIC_API_KEY"),
Model: "claude-sonnet-4.5",
})
resp, err := client.Generate(ctx, llm.CompletionRequest{
Messages: []llm.Message{
{Role: llm.RoleUser, Content: "Hello"},
},
})
Streaming
for chunk, err := range client.Stream(ctx, req) {
if err != nil { return err }
fmt.Print(chunk.Delta)
}
Mock for tests
import "github.com/VictorVargas/rony-llm-agent/pkg/llm/mock"
mockClient := mock.New(mock.Responses{
{Match: "hello", Response: "Hi! How are you?"},
{Match: "*", Response: "default"},
})