Runner.Compact folds the older portion of history into a single system-role summary when the previous turn's input tokens cross threshold_ratio × MaxContextWindow. When summarization fails, the runner falls back to truncateToBudget so a flaky summarize call never breaks the user's request. EstimatePromptTokens / totalPromptTokens give a conservative count (roughly 3 chars per token) used by both the compaction trigger and BuildMessages' new limitRAGContext / fitHistory helpers to cap the prompt inside the provider's reported window before the request goes out. Covers the first-turn case where no usage has been reported yet. Adds runner_compaction_test.go with table-driven coverage for the disabled, below-threshold, short-history, unknown-window, fallback and first-stream cases, plus a regression for the RAG-context trimmer. |
||
|---|---|---|
| .. | ||
| client.go | ||
| runner.go | ||
| runner_compaction_test.go | ||
| runner_test.go | ||
| testhelpers_test.go | ||