CostIQ Inference API

Powerful AI inference for your whole company.

CostIQ provides the compute and a simple API. We provision separate keys for your company, teams, employees, and applications—each with its own model access, usage limits, and reporting.

Example organizationAccess ready
Your companyOrganization API access
Engineeringciq_••••7f2acode
Customer supportciq_••••91c4general-fast
Production appciq_••••2be7general-fast

Separate keys · Independent limits · One usage report

What you get

One AI inference service for your entire organization.

01

Inference included

Use powerful AI models without buying GPUs, negotiating separate provider contracts, or operating model infrastructure.

02

Separate API keys

Give each company, team, employee, application, and agent its own revocable CostIQ credential.

03

Usage controls

Set model access, expiry, request rates, token rates, concurrency, and estimated-spend limits for every key.

04

Clear reporting

See requests, tokens, latency, and estimated cost by customer, team, employee, or workload.

How it works

From request to inference in three steps.

1

Tell us what you need

Share your use case, expected usage, preferred models, and who needs access.

2

Receive your organization and keys

We provision CostIQ access and create separate credentials with the right limits for each approved user or application.

3

Start using inference

Connect coding agents, tools, and products through native Responses or Chat Completions APIs, then monitor usage as your organization grows.

Use the API

One endpoint for the work your company already does.

Use a CostIQ key with Codex, standard HTTP requests, and compatible clients. The same API can serve people, internal tools, agents, and products.

Coding agentsCustomer supportInternal toolsDocument analysisResearchAutomated workflowsBatch processing
POST /v1/responses
export COSTIQ_BASE_URL="https://inference.costiq.xyz/v1"
curl "$COSTIQ_BASE_URL/responses" \
  -H "Authorization: Bearer $COSTIQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "general-fast",
    "input": "Help me understand this repo",
    "store": false,
    "stream": true
  }'

Get started

Give your company access to AI inference.

Tell us how many people or applications need access and what you want to run.