Cookbook
Pattern: score a batch
Score thousands of items with bounded concurrency, an idempotency key per item so a rerun never pays twice, and one results line per item.
To score a backlog (last month's tickets, a product catalogue, an eval set), send one request per item, a few at a time, and write one result line per item. The SDKs already retry rate limits and outages; your code only has to bound the concurrency and name each call.
The rules
- One request per item, every question about that item in it. The item is billed once however many questions it has.
- Concurrency at or below the key's limit. By default a live key allows 4 requests at once on pay as you go and Base, 8 on Hot and 16 on Fierce, and a test key 4 (see Rate limits). More only queues or earns
429 concurrency. - An idempotency key per item, derived from the batch and the item id. When the job crashes and you run it again within 24 hours, finished items replay their original response for free (
idempotent-replayed: true) instead of being charged again. - Pin the exact version so every item of the batch, and a rerun next week, is scored by the same model.
- Write the request id, model and
abstainedflags with each result, so you can count abstains and trace any row.
The code
The input is a JSON Lines file with an id and a text per line; the output is another JSON Lines file.
import asyncio
import json
from thinqit_dex import AsyncClient, DexError, check, pick
MODEL = "dex-1.0.1"
BATCH = "reviews-2026-09" # names this run; reuse it to resume
CONCURRENCY = 4 # 4 for a test key; a live key: 4, 8 or 16 by plan
QUESTIONS = {
"topic": pick(
"What is {{review.text}} mainly about?",
{"quality": "The product itself", "delivery": "Shipping and delivery", "price": "Price or value",
"service": "Customer service", "other": None},
min_confidence=0.5,
),
"wants_contact": check("Does the writer of {{review.text}} ask to be contacted?", min_confidence=0.6),
}
async def score(dex: AsyncClient, sem: asyncio.Semaphore, item: dict) -> dict:
async with sem:
try:
d = await dex.decide(
state={"review": {"text": item["text"]}},
questions=QUESTIONS,
model=MODEL,
idempotency_key=f"{BATCH}-{item['id']}",
)
except DexError as err:
return {"id": item["id"], "error": str(err)}
topic, contact = d.pick("topic"), d.check("wants_contact")
return {
"id": item["id"],
"topic": None if topic.abstained else topic.choice,
"wants_contact": None if contact.abstained else contact.probability >= 0.5,
"request_id": d.id,
"model": d.model,
"input_tokens": d.usage.input_tokens,
}
async def main() -> None:
items = [json.loads(line) for line in open("reviews.jsonl", encoding="utf-8")]
sem = asyncio.Semaphore(CONCURRENCY)
async with AsyncClient() as dex: # reads DEX_API_KEY
results = await asyncio.gather(*(score(dex, sem, item) for item in items))
with open("scores.jsonl", "w", encoding="utf-8") as out:
for r in results:
out.write(json.dumps(r, ensure_ascii=False) + "\n")
done = [r for r in results if "error" not in r]
print(len(done), "scored,", len(results) - len(done), "failed,",
sum(r["topic"] is None for r in done), "topic abstains,",
sum(r["input_tokens"] for r in done), "input tokens")
asyncio.run(main())import { readFile, writeFile } from "node:fs/promises";
import { Client, DexError, check, pick } from "@thinqit/dex";
const MODEL = "dex-1.0.1";
const BATCH = "reviews-2026-09"; // names this run; reuse it to resume
const CONCURRENCY = 4; // 4 for a test key; a live key: 4, 8 or 16 by plan
const dex = new Client(); // reads DEX_API_KEY
const questions = {
topic: pick(
"What is {{review.text}} mainly about?",
{
quality: "The product itself",
delivery: "Shipping and delivery",
price: "Price or value",
service: "Customer service",
other: null,
},
{ min_confidence: 0.5 },
),
wants_contact: check("Does the writer of {{review.text}} ask to be contacted?", { min_confidence: 0.6 }),
};
type Item = { id: string; text: string };
async function score(item: Item) {
try {
const d = await dex.decide(
{ model: MODEL, state: { review: { text: item.text } }, questions },
{ idempotencyKey: `${BATCH}-${item.id}` },
);
const { topic, wants_contact: contact } = d.answers;
return {
id: item.id,
topic: topic.abstained ? null : topic.choice,
wants_contact: contact.abstained ? null : contact.probability >= 0.5,
request_id: d.id,
model: d.model,
input_tokens: d.usage.input_tokens,
};
} catch (err) {
if (err instanceof DexError) return { id: item.id, error: String(err) };
throw err;
}
}
const items: Item[] = (await readFile("reviews.jsonl", "utf8"))
.split("\n")
.filter((line) => line.trim())
.map((line) => JSON.parse(line));
// A small pool: CONCURRENCY workers take the next item until none are left.
const results: Awaited<ReturnType<typeof score>>[] = new Array(items.length);
let next = 0;
await Promise.all(
Array.from({ length: CONCURRENCY }, async () => {
while (next < items.length) {
const i = next++;
results[i] = await score(items[i]);
}
}),
);
await writeFile("scores.jsonl", results.map((r) => JSON.stringify(r)).join("\n") + "\n");
console.log(results.filter((r) => !("error" in r)).length, "scored of", items.length);A reviews.jsonl to try it with:
{"id": "r1", "text": "Arrived two days late and the box was crushed, but the lamp works."}
{"id": "r2", "text": "Great sound for the price. Please call me about the extended warranty."}
{"id": "r3", "text": "Support never answered my three emails."}
Throughput and cost
- The requests and the tokens per minute both limit a batch, one request per item. By default a live key allows 60 requests and 60,000 input tokens a minute on pay as you go and Base, 120 and 120,000 on Hot, and 300 and 200,000 on Fierce; a test key allows 30 and 40,000. At 300 tokens an item, 60,000 tokens is about 200 items a minute, so on a pay-as-you-go or Base key the 60 requests set the pace: about 60 items a minute (120 on Hot, 300 on Fierce). The SDK waits out a
429by itsretry-afterand goes on. - A test key is free up to 250,000 tokens a day per account: enough to try a batch of several hundred items before you run it on a live key.
- Before a large run, estimate its cost from one item's
usage.input_tokens: see Estimate cost. - Items that failed (the
errorlines) can be run again with the sameBATCH: items that did finish replay for free, the rest run once.