Two failure modes seen live with Qwen3.6 on llama.cpp ended turns silently mid-task: - The model writes its tool call as plain text inside its reasoning, the server never parses it, and the round ends with nothing executed. The loop now detects the markers and nudges the model to re-issue the call for real (max 2 per turn). - llama.cpp silently ignores the max_thinking_tokens field, so a model in a reasoning spiral ran until max_tokens (seen live: 25k+ tokens of nonstop thinking, ~20 min). The llamacpp client now enforces the budget client-side during Stream: once exceeded while the round is still pure reasoning, it cuts with FinishThinkingBudget and aborts the request (freeing the server slot); the loop answers with its own corrective nudge, on a separate counter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| mock | ||
| providers | ||
| README.es.md | ||
| README.md | ||
| types.go | ||
| types_test.go | ||
pkg/llm
Multi-provider abstraction for language models.
Responsibility
Define a common interface (LLMClient) and adapters for the main providers.
Public API
type LLMClient interface {
Generate(ctx context.Context, req CompletionRequest) (CompletionResponse, error)
Stream(ctx context.Context, req CompletionRequest) iter.Seq2[StreamChunk, error]
Name() string
Capabilities() ProviderCapabilities
}
type CompletionRequest struct {
Messages []Message
Tools []tools.Tool
ToolChoice ToolChoice
Model string
Temperature *float32
MaxTokens *int
}
type CompletionResponse struct {
Content string
ToolCalls []tools.Call
Usage TokenUsage
StopReason string
}
type ProviderCapabilities struct {
SupportsTools bool
SupportsVision bool
MaxContextWindow int
}
Included providers
| Provider | Package | Tool support |
|---|---|---|
| OpenAI | providers/openai |
✅ |
| Anthropic | providers/anthropic |
✅ |
| Ollama | providers/ollama |
✅ (models that support it) |
| llama.cpp | providers/llamacpp |
✅ (with grammar) |
Usage
import "github.com/VictorVargas/rony-llm-agent/pkg/llm/providers/anthropic"
client, err := anthropic.New(anthropic.Config{
APIKey: os.Getenv("ANTHROPIC_API_KEY"),
Model: "claude-sonnet-4.5",
})
resp, err := client.Generate(ctx, llm.CompletionRequest{
Messages: []llm.Message{
{Role: llm.RoleUser, Content: "Hello"},
},
})
Streaming
for chunk, err := range client.Stream(ctx, req) {
if err != nil { return err }
fmt.Print(chunk.Delta)
}
Mock for tests
import "github.com/VictorVargas/rony-llm-agent/pkg/llm/mock"
mockClient := mock.New(mock.Responses{
{Match: "hello", Response: "Hi! How are you?"},
{Match: "*", Response: "default"},
})