Get started

Quickstart

Mint a key, point any OpenAI client at Tokex, and set a cap. Your first call runs on the cheapest model your credit covers.

Three steps

Mint a keySign in to the app and create an API key under Keys. It starts with tk_, carries only the scopes you give it, and is shown once: Tokex stores a hash.
Point your client at TokexAny OpenAI SDK works unchanged against /v1. Every write carries an Idempotency-Key, so a retry never runs twice.
Set a capA cap is a refusal, not a warning: the request that would cross it never reaches a provider. Set one per key or for the whole account.

Your first call

Any OpenAI SDK works unchanged against /v1. Ask for model auto and the cheapest model your credit covers runs. Every write carries an Idempotency-Key, so a retry never runs twice.

import OpenAI from "openai"; const tokex = new OpenAI({  baseURL: `${process.env.TOKEX_URL}/v1`,  apiKey: process.env.TOKEX_KEY, // a tk_ key: inference scope, its own cap}); const res = await tokex.chat.completions.create(  {    model: "auto", // the cheapest model your credit covers    messages: [{ role: "user", content: "Draft release notes from this diff." }],    max_tokens: 800,  },  { headers: { "Idempotency-Key": crypto.randomUUID() } },); console.log(res.choices[0].message.content);console.log(res.tokex); // the receipt: model, usage, credit reserved and charged

The receipt

Each answer carries a tokex block: the model that ran, the tokens used, the credit held before the call and the credit charged after it.

res.tokex

{  "model": "gpt-5",  "asset": "openai:gpt-5",  "usage": { "inputTokens": 212, "outputTokens": 478 },  "reservedCredits": "1728",  "charged": "690",  "cached": false,  "latencyMs": 1840,  "txId": "tx_01J9ZM2Q8K"}

Illustrative figures. The worst case is reserved before the call; what the answer did not use returns. A failed provider call returns the whole hold.