Tool-Use Training, Not Context Size, Makes Agents Work
Agentic and function-calling fine-tuning is the specific post-training work — synthetic tool-use trajectories, multi-step reasoning traces, rewards tied to task completion rather than next-token accuracy — that separates models that can actually run an agent workflow from ones that just have a big context window. As of mid-2026, benchmarks like tau-bench and the newly reweighted BFCL v4 show frontier models clustering near parity on single tool calls but diverging sharply on multi-turn, multi-constraint tasks. That divergence is now the signal worth watching as more products get built as agents rather than chatbots. For PMs, the takeaway is blunt: stop evaluating vendors on single-function-call demos and start asking what their post-training actually rewarded.