AI coding assistants come in two payment flavors. Pay-per-token: you’re billed for every request, so the only lever you have is token efficiency - prompt caching, keeping context small, reaching for the cheapest model that does the job. Quota-based: you get a fixed quota on a rolling window. Claude Code is capped at a 5-hour window and a 7-day window. Sending the first request starts the window, whatever your quota level is at. Hitting the limit on either one means waiting for the next reset.
I found the 5-hour window quite impractical. Usually I work 8 hours a day, so I’m more prone to overutilize the first quota window and underutilize the next one. Not to mention the risk of burning through everything early and being stuck waiting out the rest of that 5-hour window.
But there’s a way to hack it a bit. If I send my first prompt 1 hour before starting work, the window starts then instead of whenever I happen to send something real - so it’s already running by the time I actually sit down. That pushes the reset to land mid-workday instead of at some arbitrary point, and the risk of hitting the limit is much lower.
To automate that I set a simple cronjob:
spec:
schedule: "0 9,14 * * 1-5"
timeZone: "Europe/Warsaw"
jobTemplate:
spec:
template:
spec:
containers:
- name: prewarm
command: ["claude", "--model", "claude-haiku-4-5-20251001", "-p", "ack"]
It fires 1h before I start working. It sends a dummy message to the cheapest model - so the cost is a rounding error - after sending it, my usage is still at 0%. If you plan to heavily use quota in the first part of the day, you can shorten this window even more.