Profiled where the ~6s actually went, and it was not where I assumed. Haiku defaults to extended thinking, so it spent 225 output tokens and 1.9s deliberating before emitting an 80-character answer about the height of a tower. MAX_THINKING_TOKENS=0: 43 output tokens, time-to-first-text 2657ms -> 1122ms, wall clock 5.7s -> ~2.4s, answer unchanged. Two things measured NOT to be the bottleneck, recorded in the comment so nobody optimises them later: process startup (0.12s boot plus 28ms to fire the request, so a warm/persistent process saves nothing) and prompt size (184 input tokens, no bloat to trim). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| hardware | ||
| FredOS-Gaming.nix | ||
| FredOS-Macbook.nix | ||
| FredOS-Mediaserver.nix | ||