llm-cost-optimizerlisted
Install: claude install-skill Notysoty/openagentskills
# LLM Cost Optimizer
## What this skill does
This skill audits an LLM application's prompts, call patterns, and model selection to identify cost reduction opportunities. It covers prompt caching, model routing (right-sizing), token reduction, batching, and output length control — the techniques that typically cut LLM costs by 40–80% without sacrificing quality.
## How to use
### Claude Code / Cline
Copy this file to `.agents/skills/llm-cost-optimizer/SKILL.md` in your project root.
Then ask:
- *"Use the LLM Cost Optimizer to audit our AI application."*
- *"How can I reduce our OpenAI API costs? Here are our prompts..."*
Provide:
- Your system prompt(s)
- Approximate daily call volume
- Which model(s) you're using
- Typical input/output token counts if known
- Whether calls are real-time (low latency required) or batch (latency tolerant)
### Cursor / Codex
Paste your prompts, call patterns, and current monthly spend alongside these instructions.
## The Prompt / Instructions for the Agent
When asked to optimize LLM costs, audit the following areas in order of typical savings impact:
### Audit 1 — Prompt Caching (savings: 50–90% on repeated prefixes)
**Check:** Does the system prompt stay the same across calls?
If yes, enable prompt caching. The system prompt is sent once and cached — subsequent calls only pay for the new user tokens.
```python
# Anthropic Claude — cache_control on system prompt
response = client.messages.create(
model="claude-opus-4-6",
s