Reduce AI Agent Costs with Claude Code's PreToolUse Hook
Job to be done: Reduce AI agent orchestration token costs.
🇳🇬 Ways to use this in Nigeria
Ideas to get you started, adapt to your situation.
- Entrepreneur
If you're building an AI-powered customer service bot for your startup, implement this hook to ensure simple queries use cheaper models, preventing high API costs as your user base grows.
- 9-5 employee
As a software developer building internal AI tools for your company, use this to control model usage and reduce monthly API costs, helping your team meet budget targets.
What you’ll get
You will learn how to implement a “PreToolUse hook” in Claude Code to control which Claude model is used for specific tasks, preventing the accidental use of expensive models like Opus for simple jobs. This approach helps significantly reduce the number of tokens consumed per task, saving costs. It works by adding a layer of control before the AI model processes a request, ensuring that the most cost-effective model is chosen based on the task’s complexity.
Tools you need
- Claude Code (paid): An AI coding assistant environment for building and deploying AI applications.
- Claude Opus (paid): The most capable and expensive Claude model, used for complex reasoning.
- Claude Sonnet (paid): A balanced model offering good performance at a moderate cost.
- Claude Haiku (paid): The fastest and most affordable Claude model, suitable for simpler tasks.
Steps
-
Understand the cost multipliers: Recognize that using the most expensive model (Opus) for every task, even simple ones, dramatically increases costs. Opus costs $5.00 per 1 million input tokens and $25.00 per 1 million output tokens. Sonnet is cheaper at $3.00/$15.00, and Haiku is the cheapest at $1.00/$5.00 per 1 million tokens. The author found that running every task on Opus was a major cost driver.
-
Identify the problem with pure delegation: The original design delegated all work to sub-agents, ensuring a clean context for each task. However, this meant each sub-agent, by default, inherited the parent session’s model setting, which was Opus. This led to Opus being used for tasks that didn’t require its power, like summarizing a file.
-
Implement a PreToolUse hook: The core solution is to add a “PreToolUse hook.” This is a piece of logic that runs before the AI model is called to perform a tool’s action. This hook can inspect the intended task and decide which Claude model (Opus, Sonnet, or Haiku) is most appropriate based on predefined rules or budget constraints.
-
Define rules for model selection: The author suggests creating rules to determine the best model. For example:
- If a task is very simple (e.g., reading a file and summarizing it), use Claude Haiku.
- If a task is moderately complex, use Claude Sonnet.
- If a task requires deep reasoning or complex problem-solving, use Claude Opus.
The author doesn’t share the exact code for this hook; a starting point would be to write a Python function within Claude Code that intercepts the tool call and selects the model parameter before passing it to the main execution.
-
Consider prompt caching: Be aware that prompt caching relies on exact matches. If the system prompt, tools, or messages change even slightly, the cache may not be hit, leading to higher costs as the model has to re-process information. Ensure your PreToolUse hook and subsequent prompts are consistent where possible to maximize cache hits.
You should see reduced token usage in your Claude Code session logs after implementing this hook and ensuring consistent prompts.
Original source
This workflow is based on a postmortem analysis by akashy, shared on the DEV Community platform. The author details how an AI orchestration skill for Claude Code incurred unexpectedly high token costs and explains the redesign and enforcement layer implemented to fix it.
Notes & variations
- Free-tier alternative: While this specific technique is tied to Claude Code’s paid features, for general AI assistance, you can explore free tiers of other AI chat platforms like ChatGPT, Claude.ai, or Google Gemini to experiment with prompt engineering for cost savings on simpler tasks.
- Common mistake: A common pitfall is assuming the most powerful model is always necessary. This leads to overspending. Always evaluate if a simpler, cheaper model can achieve the desired outcome.
- Tip for better results: Implement a clear, tiered system for model selection. For instance, define specific keywords or task types that automatically trigger Haiku or Sonnet, reserving Opus only for tasks explicitly flagged as requiring advanced reasoning.