by karanb192
Keep Claude Code’s prompt cache warm during breaks and show the estimated cost before a cold send.
# Add to your Claude Code skills
git clone https://github.com/karanb192/cache-taxGuides for using ide extensions skills like cache-tax.
See how cache-tax compares with popular alternatives.
cache-tax is an open-source ide extensions skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by karanb192. Keep Claude Code’s prompt cache warm during breaks and show the estimated cost before a cold send. It has 55 GitHub stars.
cache-tax's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/karanb192/cache-tax" and add it to your Claude Code skills directory (see the Installation section above).
cache-tax is primarily written in TypeScript. It is open-source under karanb192 on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other IDE Extensions skills you can browse and compare side by side. Open the IDE Extensions category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh cache-tax against similar tools.
No comments yet. Be the first to share your thoughts!
Top skills in this category by stars
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
Keep Claude Code's prompt cache warm during breaks.
Cache Tax is a Claude Code mod that refreshes your prompt cache while you step away and shows the estimated rewrite cost before a cold send. Run /keepwarm 90m before a break. Pings cost tokens; choose a window you expect to return within.
Install · Website · How it works · Costs and limits · Commands

A two-line recap in a 330k-token session triggered a $6.61 estimate. After resending, the reported cache write was $6.28. Recorded on 2.1.1; current wording says “up to” the estimated token count. These are API-equivalent costs, not extra subscription charges.
Needs Claude Code 2.1.287 or later, a one-hour cache and Claude Code left running. Mods are on by default. Pings cost tokens, including uncapped output.
claude plugin marketplace add karanb192/claude-code-mods
claude plugin install cache-tax@claude-code-mods
Restart Claude Code or run /reload-plugins in an open session. No early-access flag is needed. If you set CLAUDE_CODE_ENABLE_FUNCTION_HOOKS previously, remove it; current Claude Code ignores it.
In a warm session, run /keepwarm to arm six hours. Use /keepwarm 90m for a shorter window, or /keepwarm off to stop.
Check the main cache lifetime. Included subscription usage defaults to one hour. API-key, usage-credit and cloud-provider sessions default to five minutes, so set promptCacheTtl to "1h" for those billing paths. A 50-minute timer cannot protect a five-minute main cache. See Limits for the fork evidence.
Prefer a settings hook? The hook version warns by default, refuses once with CACHE_TAX_BLOCK=1, and provides a standalone status-line countdown. Warming is part of this mod.
Before a break: /keepwarm arms a bounded window. After 50 idle minutes, a timer inside Claude Code sends one tool-less fork over the session's transcript. It reads usage after every ping and stops if reads are zero or writes reach 10% of reads. There is no background transcript-polling loop.
When you return cold: for a context of at least 50,000 tokens, the guard drops your ordinary message once when its one-hour clock says cold. It shows the estimated rewrite price. Resend to continue, or /clear and start from a note.
A detected cold write automatically arms at least three hours of keepwarm. /cache-tax shows the current cache estimate, warming window and this session's cold-write tally.
Windows belong to individual sessions. An already-cold session waits for your next turn before pinging. A window ends at its deadline; /keepwarm off also cancels it and clears the always setting.

Sonnet 5, 16 September 2026: the status reported 75k tokens read at $0.02. Nothing appeared in the conversation during the live run. This capture used the one-minute testing interval on 2.0.0; its card's $0.45 rewrite estimate used an older Sonnet price row, versus $0.30 for the same context with the current row. The two-cent ping is a receipt for that session, not a fixed price.
When the cached prefix expires, even “give me a recap” can trigger a rewrite of the old context before the answer. Cache Tax's Fable 5.1 price row compares a $20 one-hour write with a $0.25 cache read per million tokens: 80x per token, not 80x total savings. See the pricing assumptions and Anthropic's price table.
Watch the 23-second overview. Its designed scenes label Fable 5.1 list prices; every terminal and status-line pixel comes from the real refusal recording and keep-warm receipt above. The refusal was recorded on 2.1.1 and the keep-warm crop on 2.0.0.
New to caching? Anthropic explains how Claude Code uses it.
/keepwarm keep warm for six hours
/keepwarm 90m keep warm for a window of your own (also 2h30m, 6h)
/keepwarm always arm a six-hour window at every session start, remembered across sessions
/keepwarm 6h every 2m pinging every two minutes; a testing knob, floor 1m, forgotten after this window
/keepwarm status the warming window, next ping and last receipt
/keepwarm off stop, forget the window, and turn always off
/cache-tax the card
/cache-tax guard warn show the price and send (the hook's default)
/cache-tax guard refuse drop a cold send once, the resend goes through (this mod's default)
The refusal, verbatim:
cache-tax: the prompt cache went cold 2h00m ago. Sending this re-writes up to 200,502 tokens at $20/MTok = $4.01 (a warm turn would have cost $0.05). Send it again to pay it, and keepwarm will then hold the cache for 3h00m. Or /clear and start from a note.
The figure is an upper bound. The context count the engine reports for a resumed session is the last response's input, cache read, cache write and output together, and the resume payload carries no separate output count to take off; on one 15-day-old session the refusal said 330,316 tokens and $6.61 and the write that followed was 314k tokens, $6.28.
While keepwarm is armed, the terminal and Desktop Code tab show a row above the prompt with the Cache Tax cube. Its lid is green for warm, orange for cold, and grey before the first turn or after compaction. The bold state label matches; the timer and receipt text use normal weight. Ghostty and kitty can draw the image; other terminals show [>] in its place. A nonempty NO_COLOR selects a bold text icon and an uncolored state label. The row also says warm, cold or unknown, so color is never the only signal. These states use the same one-hour clock as /cache-tax; they are estimates, not a live server cache check.
The row includes the warming window, next ping and last readback. A stopped loop shows a neutral symbol and its reason. /keepwarm off or an expired window removes the row. Surveys temporarily take priority, and other mods' content stays in the band. Sessions without this band retain the plain status text.
The cube follows Claude's light or dark theme. Its terminal image also has a contrasting edge for terminal backgrounds that differ from that setting. The embedded assets come from the website's official SVGs; maintainers can regenerate them with python3 tools/generate-status-icons.py and rsvg-convert installed. No image files are read or fetched at runtime.
The mod reads the theme and NO_COLOR once when the session starts. Accepted theme changes update the icon immediately; ordinary redraws do not reread either setting. Set NO_COLOR before starting the session.
The hook and the mod share a name and a job, so having both means two guards on every cold send. The mod checks at session start whether the hook's /cache-tax:status command exists and says so once. Keep one guard. The hook remains an option for people who prefer settings hooks. Before uninstalling it, check your status line: the 🧊 row wired to cache-tax.js comes from the hook's files. This mod draws above the prompt while keepwarm is armed; it does not replace your shell status line.
Validated on Claude Code 2.1.289:
❯ ./register.ts hooks: config.set{key=theme}, ui.render{component=AbovePrompt}, session.start, classic.SessionStart, command.run{command=keepwarm}, command.run{command=cache-tax}, prompt.submit, turn.step, turn.complete, session.compact
❯ ./register.ts calls: $.clock.after (via arm), $.clock.now, $.command.list, $.command.register, $.config.list, $.env.get, $.model.fork (via ping), $.session.id, $.session.model, $.session.usage, $.store.delete (via prune, startWindow, stop), $.store.get, $.store.set, $.ui.invalidate, $.ui.log, $.ui.resolve, $.ui.status
❯ ./register.ts env reads: NO_COLOR
Reach L2, drives Claude. Sees every prompt you type, every model request's timing and every answer's token counts.
Threat model for cache-tax (reach L2, drives Claude)
1. Reads: of each prompt, whether it starts with a slash and nothing else (the text is passed on untouched, never kept, never logged); the time; the token counts and model id the engine already holds on turn.complete and on the fork's reply; the resume fields Claude Code computes for settings hooks; the command list, the session id and, when a resumed session's fields carry no model id, the session's model once at start; from its own $.store, the keepwarm deadline and the ping period keyed by session id, plus the global always switch and guard mode
2. Runs: one $.model.fork per idle stretch inside a keepwarm window, one per ping period (50 minutes unless the testing knob set it, floor 1 minute), never outside the window, never onto a cache the mod already knows is cold, never after a readback that read nothing or wrote at least a tenth of what it read
3. Sends: nothing leaves the machine except the fork, an API request over the session's own transcript with a fixed one-line prompt
4. Persists: in $.store, the keepw