Run/RunStream took the new turn as a bare string, which had nowhere
to carry ContentPart attachments. Both now take an llm.Message
(Role is forced to RoleUser regardless of what the caller sets), so a
caller building a multimodal turn just fills in Content/Parts on it
instead of the loop needing a second, parallel parameter.
subagent.go and every test call site are updated to wrap their string
prompt as llm.Message{Role: llm.RoleUser, Content: ...} — SubAgent.Run
itself is untouched, it still takes a plain task string.
Two failure modes seen live with Qwen3.6 on llama.cpp ended turns silently
mid-task:
- The model writes its tool call as plain text inside its reasoning, the
server never parses it, and the round ends with nothing executed. The
loop now detects the markers and nudges the model to re-issue the call
for real (max 2 per turn).
- llama.cpp silently ignores the max_thinking_tokens field, so a model in
a reasoning spiral ran until max_tokens (seen live: 25k+ tokens of
nonstop thinking, ~20 min). The llamacpp client now enforces the budget
client-side during Stream: once exceeded while the round is still pure
reasoning, it cuts with FinishThinkingBudget and aborts the request
(freeing the server slot); the loop answers with its own corrective
nudge, on a separate counter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>