Inference included
Use powerful AI models without buying GPUs, negotiating separate provider contracts, or operating model infrastructure.
CostIQ Inference API
CostIQ provides the compute and a simple API. We provision separate keys for your company, teams, employees, and applications—each with its own model access, usage limits, and reporting.
ciq_••••7f2acodeciq_••••91c4general-fastciq_••••2be7general-fastSeparate keys · Independent limits · One usage report
What you get
Use powerful AI models without buying GPUs, negotiating separate provider contracts, or operating model infrastructure.
Give each company, team, employee, application, and agent its own revocable CostIQ credential.
Set model access, expiry, request rates, token rates, concurrency, and estimated-spend limits for every key.
See requests, tokens, latency, and estimated cost by customer, team, employee, or workload.
How it works
Share your use case, expected usage, preferred models, and who needs access.
We provision CostIQ access and create separate credentials with the right limits for each approved user or application.
Connect coding agents, tools, and products through native Responses or Chat Completions APIs, then monitor usage as your organization grows.
Use the API
Use a CostIQ key with Codex, standard HTTP requests, and compatible clients. The same API can serve people, internal tools, agents, and products.
export COSTIQ_BASE_URL="https://inference.costiq.xyz/v1"
curl "$COSTIQ_BASE_URL/responses" \
-H "Authorization: Bearer $COSTIQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "general-fast",
"input": "Help me understand this repo",
"store": false,
"stream": true
}'Get started
Tell us how many people or applications need access and what you want to run.