tokenlab-cost-routinglisted
Install: claude install-skill hedging8563/tokenlab-skills
# TokenLab Cost Routing
Use this skill when a user asks how to reduce TokenLab cost, compare model prices, pick fallbacks, or route requests by quality, latency, and budget.
## What this skill should deliver
- A compact routing recommendation with exact public TokenLab model IDs.
- A cost-aware fallback chain for the user's workload.
- A catalog/pricing lookup path that can be rerun.
- A note on which endpoint family each model should use.
- Guardrails for when not to switch models because doing so would change output, safety, or request semantics.
## Preferred approach
1. Identify the workload and constraints:
- chat, coding, agent loop, image, video, audio, embedding, rerank, translation, or multimodal
- quality floor
- latency target
- budget or cost ceiling
- native endpoint requirement
2. Read live public catalog signals before recommending:
- `GET https://api.tokenlab.sh/v1/models`
- `GET https://api.tokenlab.sh/v1/models?recommended_for=<scene>`
- `GET https://api.tokenlab.sh/v1/models/:model`
- `GET https://api.tokenlab.sh/v1/models/:model/pricing`
3. Build a chain with roles:
- primary quality model
- balanced default
- fast fallback
- budget fallback
4. If the user asks for exact cost, compute from live pricing and their estimated token/media volume. State units and assumptions.
5. For non-chat requests, inspect model details before changing parameters or endpoint family.
## Output format
- One sentence stating workload