Pithos Logo PITHOS
100% Local AI Context Optimizer

Cut your AI coding bill.
Your code never leaves your machine.

Pithos is the smart local sidecar that intercepts inputs for Cline, Claude Code, and Aider. By pruning dead context, optimizing prompt caching, and condensing outputs locally, it slashes token usage by up to 45% (typically 18-20% in everyday use).

Zero-Cloud Middleware Anthropic Cache-Aware Compatible with Any Client
01 10 {"} </>
The Problem

AI coding agents burn tokens fast.
Your Claude and Cursor bills keep climbing.

AI clients (like Cline or Claude Code) are context hogs. They dump full file trees, history, and terminal outputs into every single call. Because you pay per token, simple iterations cost dollars, and your API budget is consumed in days.

Before Optimization

Raw Agent Request Structure

Direct Context (No Filter) 145,000 Tokens
  • 📁 Entire node_modules / dist / build 85K tokens
  • 📄 All project files (regardless of imports) 42K tokens
  • 💬 Complete raw terminal buffer history 18K tokens
PITHOS ACTIVE
With Pithos Compression 35,200 Tokens
  • 📁 Excluded lockfiles & build assets -100% (0)
  • 📄 Read deduplication & file diffs 12K tokens
  • 💬 Compressed terminal errors/history 5K tokens
  • ⚡ Stable context ordering for caching Triggered correctly
Average Token Reduction 18-20% Savings
Est. Cost Per Run
$0.43 $0.05
Architecture

How Pithos Works

Pithos runs locally as an intelligent gateway, sitting seamlessly between your IDE / coding clients and the model provider APIs.

01

Local Analysis

Pithos hooks into Claude Code's tool-call lifecycle (PostToolUse/SessionStart) to inspect command output and file reads before they reach your context.

02

Context Trimming & Caching

It eliminates unnecessary dependencies, strips out lockfiles, and ensures Anthropic's prompt caching triggers correctly, so cached reads cost 90% less (Anthropic's own pricing, not a Pithos number).

03

Send Lean Prompt

The optimized, highly condensed prompt is sent to your own Claude, OpenAI, or DeepSeek API key. You get the exact same quality response, but pay a fraction of the cost.

Features

Designed for Coder Economics

Pithos equips you with standard optimizations and advanced local filters that cloud-based API wrappers can't support.

Tool Output Compression

A Claude Code hook compresses noisy Bash/Read/Grep output at the source — passing test noise, repeated file reads, and oversized JSON get trimmed before they ever enter your session's context.

Prompt Caching Layout

Keeps static system rules and baseline context in a stable order so Anthropic's prompt cache actually triggers, instead of silently missing.

Optional Output Nudges

An opt-in session hint nudges the model toward concise, delegatable work — off by default, and never forced.

Savings Dashboard

Watch your actual token reduction and cash saved accumulate in real time. Proof of value you can export easily.

What's inside Pithos

Local compilation pipelines, semantic optimization, and offline token management.

# Prompt Compression # Semantic Cache # PDF → Text # Model Routing # Proxy Mode
Fully Local & Integrated
💡 Web Lite — In-Browser Optimizer

Optimize prompts in your browser before you copy & paste

No installation required. Fix your prompts before pasting them to ChatGPT, Gemini, or Claude. Web Lite is 100% browser-local, client-side, and offers progressive enhancement with local AI models (WebLLM).

Security

"Your code stays in the jar."

In Greek mythology, the Pithos was a storage vessel built to securely seal its contents inside. We take that philosophy literally.

Cloud API Wrapper Alternatives
  • Your code passes through third-party servers.
  • Cloud hops add network latency to every LLM turn.
  • Proprietary enterprise assets are cached in external logs.
Pithos (100% Local Engine)
  • All optimizations happen on localhost. Code never exits.
  • Zero external hops. Compute speed is bound to your silicon.
  • Compatible with strictly air-gapped secure workspaces.
Plans

Flat Pricing. Infinite Savings.

No credit cards, usage calculations, or surprise cloud billing spikes. Pick the flat plan that matches your pipeline.

Web Lite Free
$0 / free

For developers to test prompt optimizations directly in the browser.

  • 15 optimizations / day
  • Clean-code prompt parser
Try for Free
Web Lite Plus
$9 / month

For active browser prompt optimization with local AI smart mode.

  • Unlimited optimizations
  • WebLLM smart mode
Subscribe Plus
All-in-One Local Suite
Pithos (VS Code)
$19 / month

All features included: local proxy, caching, token optimization, and dashboard.

  • Context Selection & Caching
  • Transparent Proxy Mode
  • Semantic Cache & Routing
  • PDF OCR & Table Parsing
Subscribe Pithos
Enterprise
Custom

For organizations requiring air-gapped on-premises setups and SLAs.

  • Unlimited tokens & seats
  • Air-gapped / offline setup
Contact Sales

Start Saving on AI Token Costs Today

Pithos runs fully locally, intercepts outbound prompts, and trims context size automatically. Get started in minutes.

Get Pithos License Download Extension