Frontier models at cost, with a spend cap that actually stops.
One OpenAI-compatible endpoint for Claude and low-cost open models on AWS Bedrock. No markup on tokens. Every request reports what it cost, and when you reach your limit, requests stop.
- At costInference is billed at Bedrock’s own per-token prices, nothing added.
- A real limitMonthly or lifetime caps, overall or per key, enforced before a request ever reaches the model.
- Cost on every responseEach response says exactly what it cost, down to the cached token.
- Works with your agentsStreaming, tool calls and prompt caching, through the API your tools already speak.
curl https://api.glimpsed.ai/v1/chat/completions \
-H "Authorization: Bearer $GLIMPSED_AI_KEY" \
-d '{"model": "anthropic/claude-sonnet-5.5", "messages": [...]}'
X-Glimpsed-Cost-USD: 0.0001628