Pithos is the smart local sidecar that intercepts inputs for Cline, Claude Code, and Aider. By pruning dead context, optimizing prompt caching, and condensing outputs locally, it slashes token usage by up to 45% (typically 18-20% in everyday use).
AI clients (like Cline or Claude Code) are context hogs. They dump full file trees, history, and terminal outputs into every single call. Because you pay per token, simple iterations cost dollars, and your API budget is consumed in days.
Pithos runs locally as an intelligent gateway, sitting seamlessly between your IDE / coding clients and the model provider APIs.
Pithos hooks into Claude Code's tool-call lifecycle (PostToolUse/SessionStart) to inspect command output and file reads before they reach your context.
It eliminates unnecessary dependencies, strips out lockfiles, and ensures Anthropic's prompt caching triggers correctly, so cached reads cost 90% less (Anthropic's own pricing, not a Pithos number).
The optimized, highly condensed prompt is sent to your own Claude, OpenAI, or DeepSeek API key. You get the exact same quality response, but pay a fraction of the cost.
Pithos equips you with standard optimizations and advanced local filters that cloud-based API wrappers can't support.
A Claude Code hook compresses noisy Bash/Read/Grep output at the source — passing test noise, repeated file reads, and oversized JSON get trimmed before they ever enter your session's context.
Keeps static system rules and baseline context in a stable order so Anthropic's prompt cache actually triggers, instead of silently missing.
An opt-in session hint nudges the model toward concise, delegatable work — off by default, and never forced.
Watch your actual token reduction and cash saved accumulate in real time. Proof of value you can export easily.
Local compilation pipelines, semantic optimization, and offline token management.
No installation required. Fix your prompts before pasting them to ChatGPT, Gemini, or Claude. Web Lite is 100% browser-local, client-side, and offers progressive enhancement with local AI models (WebLLM).
In Greek mythology, the Pithos was a storage vessel built to securely seal its contents inside. We take that philosophy literally.
No credit cards, usage calculations, or surprise cloud billing spikes. Pick the flat plan that matches your pipeline.
For developers to test prompt optimizations directly in the browser.
For active browser prompt optimization with local AI smart mode.
All features included: local proxy, caching, token optimization, and dashboard.
For organizations requiring air-gapped on-premises setups and SLAs.
Pithos runs fully locally, intercepts outbound prompts, and trims context size automatically. Get started in minutes.