Get started
Quickstart
Mint a key, point any OpenAI client at Tokex, and set a cap. Your first call runs on the cheapest model your credit covers.
Three steps
Mint a keySign in to the app and create an API key under Keys. It starts with tk_, carries only the scopes you give it, and is shown once: Tokex stores a hash.
Point your client at TokexAny OpenAI SDK works unchanged against /v1. Every write carries an Idempotency-Key, so a retry never runs twice.
Set a capA cap is a refusal, not a warning: the request that would cross it never reaches a provider. Set one per key or for the whole account.
Your first call
Any OpenAI SDK works unchanged against /v1. Ask for model auto and the cheapest model your credit covers runs. Every write carries an Idempotency-Key, so a retry never runs twice.
import OpenAI from "openai"; const tokex = new OpenAI({ baseURL: `${process.env.TOKEX_URL}/v1`, apiKey: process.env.TOKEX_KEY, // a tk_ key: inference scope, its own cap}); const res = await tokex.chat.completions.create( { model: "auto", // the cheapest model your credit covers messages: [{ role: "user", content: "Draft release notes from this diff." }], max_tokens: 800, }, { headers: { "Idempotency-Key": crypto.randomUUID() } },); console.log(res.choices[0].message.content);console.log(res.tokex); // the receipt: model, usage, credit reserved and chargedThe receipt
Each answer carries a tokex block: the model that ran, the tokens used, the credit held before the call and the credit charged after it.
res.tokex
{ "model": "gpt-5", "asset": "openai:gpt-5", "usage": { "inputTokens": 212, "outputTokens": 478 }, "reservedCredits": "1728", "charged": "690", "cached": false, "latencyMs": 1840, "txId": "tx_01J9ZM2Q8K"}Illustrative figures. The worst case is reserved before the call; what the answer did not use returns. A failed provider call returns the whole hold.