rony-llm-agent/pkg/llm
Victor Vargas 8e887c8c78 fix(agent,llamacpp): recover turns killed by unparsed tool calls and reasoning spirals
Two failure modes seen live with Qwen3.6 on llama.cpp ended turns silently
mid-task:

- The model writes its tool call as plain text inside its reasoning, the
  server never parses it, and the round ends with nothing executed. The
  loop now detects the markers and nudges the model to re-issue the call
  for real (max 2 per turn).

- llama.cpp silently ignores the max_thinking_tokens field, so a model in
  a reasoning spiral ran until max_tokens (seen live: 25k+ tokens of
  nonstop thinking, ~20 min). The llamacpp client now enforces the budget
  client-side during Stream: once exceeded while the round is still pure
  reasoning, it cuts with FinishThinkingBudget and aborts the request
  (freeing the server slot); the loop answers with its own corrective
  nudge, on a separate counter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 14:50:45 -07:00
..
mock style: gofmt 2026-07-12 16:14:15 -07:00
providers fix(agent,llamacpp): recover turns killed by unparsed tool calls and reasoning spirals 2026-07-15 14:50:45 -07:00
README.es.md docs(i18n): translate all docs to English (with .es.md as Spanish alternative) 2026-06-30 13:40:37 -07:00
README.md docs(i18n): translate all docs to English (with .es.md as Spanish alternative) 2026-06-30 13:40:37 -07:00
types.go fix(agent,llamacpp): recover turns killed by unparsed tool calls and reasoning spirals 2026-07-15 14:50:45 -07:00
types_test.go style: gofmt 2026-07-12 16:14:15 -07:00

pkg/llm

Multi-provider abstraction for language models.

Responsibility

Define a common interface (LLMClient) and adapters for the main providers.

Public API

type LLMClient interface {
    Generate(ctx context.Context, req CompletionRequest) (CompletionResponse, error)
    Stream(ctx context.Context, req CompletionRequest) iter.Seq2[StreamChunk, error]
    Name() string
    Capabilities() ProviderCapabilities
}

type CompletionRequest struct {
    Messages    []Message
    Tools       []tools.Tool
    ToolChoice  ToolChoice
    Model       string
    Temperature *float32
    MaxTokens   *int
}

type CompletionResponse struct {
    Content    string
    ToolCalls  []tools.Call
    Usage      TokenUsage
    StopReason string
}

type ProviderCapabilities struct {
    SupportsTools    bool
    SupportsVision   bool
    MaxContextWindow int
}

Included providers

Provider Package Tool support
OpenAI providers/openai
Anthropic providers/anthropic
Ollama providers/ollama (models that support it)
llama.cpp providers/llamacpp (with grammar)

Usage

import "github.com/VictorVargas/rony-llm-agent/pkg/llm/providers/anthropic"

client, err := anthropic.New(anthropic.Config{
    APIKey: os.Getenv("ANTHROPIC_API_KEY"),
    Model:  "claude-sonnet-4.5",
})

resp, err := client.Generate(ctx, llm.CompletionRequest{
    Messages: []llm.Message{
        {Role: llm.RoleUser, Content: "Hello"},
    },
})

Streaming

for chunk, err := range client.Stream(ctx, req) {
    if err != nil { return err }
    fmt.Print(chunk.Delta)
}

Mock for tests

import "github.com/VictorVargas/rony-llm-agent/pkg/llm/mock"

mockClient := mock.New(mock.Responses{
    {Match: "hello", Response: "Hi! How are you?"},
    {Match: "*",    Response: "default"},
})

See also