fix(providers): add local tokenizer divergence headroom master
A local cogitate turn on a 16,384-token endpoint put 12,437 input + 4,096 completion (16,533) on the wire and was rejected with context_window_exceeded, while the condenser's 11,264 threshold was green. get_total_token_count already includes the system prompt text and tool schemas: it pulls tools off the SystemPromptEvent. The condenser trigger quantity and outgoing request therefore measure the same surface; the probe showed divergence 0 between those SDK counts. The old window * 11 // 16 threshold (11,264) was already more conservative than the exact modeled budget of window - reserve (12,288) and still failed. The measured cause is client/server tokenizer divergence: 12,437 served vs 11,237 LiteLLM-estimated, +10.7%. Set the local condenser boundary to floor((window - completion_reservation) / 1.125). The 1.125 value is a tokenizer-divergence safety factor derived from that single production observation (n=1), rounded upward. This yields 16,384 -> 10,922 and 32,768 -> 21,845; both satisfy ceil(boundary * 1.125) + reserve == window exactly. No system/tool surface is subtracted because that would double-count. Pre-fix red proof with the source change stashed: AssertionError: expected condensation before any agent-turn request; observed [_RequestRecord(kind='agent', estimated_input=10925, max_output_tokens=4096)] assert 'agent' == 'condenser' With the source change stashed, the 16k above-boundary case emitted an agent turn at 10,925 estimated input, inside the old 11,264 trigger, so no condensation. That is ceil(10925*1.125)+4096 = 16387 > 16384. The scope's acceptance criterion 6 terminal cannot-fit path is deliberately not implemented: with no surface term the boundary is floor(0.75*window/1.125), always >= 10,922 for any window eligible for a condenser, so the trigger is unreachable and a raise would be dead code. Existing context_window_exceeded classification is untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>