# Pattern: retries, timeouts and fallbacks

> What the SDKs retry for you, how to retry safely with plain HTTP, and what your code should do when Dex cannot answer.

Retrying is safe when every attempt of one call carries the same idempotency key: a retry is then never charged twice, and a retry of a call that already finished gets the original response back. The SDKs do this for every `decide` call. This page shows what they do, how to do the same with plain HTTP, and what to do when Dex cannot answer at all.

## What the SDKs do

- Retry `429` (except `test_daily_quota`), `500`, `502`, `503`, `504`, `409 idempotency_in_progress`, timeouts and network errors, up to 3 times.
- Wait the `retry-after` seconds the response gives (at most 60), otherwise back off exponentially with full jitter from 250 ms.
- Send one `idempotency-key` per `decide` call (a new UUID, or yours) on every attempt.
- Never retry other 4xx errors: they fail again until the request, the key or the balance changes.

Tune them per client or per call:

```python tab="Python"
from thinqit_dex import Client, RateLimitError, UnavailableError, check

client = Client(timeout=10, max_retries=2)  # seconds per attempt, retries after the first attempt

try:
    decision = client.decide(
        state="Where is my parcel? I ordered it two weeks ago.",
        questions={"late": check("Is the customer asking about a late parcel?", min_confidence=0.6)},
        idempotency_key="ticket-88121-late",  # your own key: stable across process restarts
        client_request_id="ticket-88121",
    )
    late = decision.check("late")
    print(late.abstained, late.probability, decision.http.retries if decision.http else 0)
except RateLimitError as err:
    print("still rate limited after the retries:", err.code, err.retry_after, err.request_id)
except UnavailableError as err:
    print("no serving path answered:", err.code, err.retry_after, err.request_id)
```

```ts tab="TypeScript"
import { Client, RateLimitError, UnavailableError, check } from "@thinqit/dex";

const client = new Client({ timeout: 10_000, maxRetries: 2 }); // milliseconds per attempt

try {
  const decision = await client.decide(
    {
      state: "Where is my parcel? I ordered it two weeks ago.",
      questions: { late: check("Is the customer asking about a late parcel?", { min_confidence: 0.6 }) },
    },
    { idempotencyKey: "ticket-88121-late", clientRequestId: "ticket-88121" },
  );
  const late = decision.answers.late;
  console.log(late.abstained, late.probability, decision.http.retries);
} catch (err) {
  if (err instanceof RateLimitError) console.log("still rate limited:", err.code, err.retryAfter, err.requestId);
  else if (err instanceof UnavailableError) console.log("no serving path answered:", err.code, err.retryAfter, err.requestId);
  else throw err;
}
```

## With plain HTTP

Do what the SDKs do: one idempotency key per logical call, the same key on every attempt, `retry-after` when given, full-jitter backoff otherwise, and no retry for other 4xx errors.

```bash tab="curl"
# One key for every attempt. --retry repeats timeouts, 429, 500, 502, 503 and 504 and honours retry-after.
IDEM="$(uuidgen)"
curl --retry 3 --retry-max-time 120 \
  https://api.thinqit.ai/v1/decide \
  -H "authorization: Bearer $DEX_API_KEY" \
  -H "content-type: application/json" \
  -H "idempotency-key: $IDEM" \
  -d '{"state": "Where is my parcel? I ordered it two weeks ago.", "questions": {"late": {"type": "check", "instructions": "Is the customer asking about a late parcel?", "min_confidence": 0.6}}}'
```

```ts tab="TypeScript"
const RETRY_STATUS = new Set([429, 500, 502, 503, 504]);

export async function decide(body: string, key: string, attempts = 4): Promise<unknown> {
  const idem = crypto.randomUUID(); // one idempotency key for every attempt of this call
  for (let attempt = 0; attempt < attempts; attempt++) {
    let res: Response | null = null;
    try {
      res = await fetch("https://api.thinqit.ai/v1/decide", {
        method: "POST",
        headers: { authorization: `Bearer ${key}`, "content-type": "application/json", "idempotency-key": idem },
        body, // send the request text as it is: JSON.stringify would reorder labels such as "1", "10"
        signal: AbortSignal.timeout(30_000),
      });
    } catch {
      res = null; // network error or timeout: safe to retry with the same key
    }
    if (res?.ok) return res.json();
    const err = res && res.headers.get("content-type")?.startsWith("application/json") ? (await res.json()).error : {};
    const retryable =
      res === null ||
      (RETRY_STATUS.has(res.status) && err?.code !== "test_daily_quota") ||
      (res.status === 409 && err?.code === "idempotency_in_progress");
    if (!retryable || attempt === attempts - 1) {
      throw new Error(`${res?.status ?? "network"} ${err?.type} ${err?.code} ${err?.request_id}`);
    }
    const wait = res?.headers.get("retry-after");
    const ms = wait ? Math.min(Number(wait), 60) * 1000 : Math.random() * 250 * 2 ** attempt;
    await new Promise((r) => setTimeout(r, ms));
  }
  throw new Error("unreachable");
}

const body = `{"state": "Where is my parcel? I ordered it two weeks ago.", "questions": {"late": {"type": "check", "instructions": "Is the customer asking about a late parcel?", "min_confidence": 0.6}}}`;
console.log(JSON.stringify(await decide(body, process.env.DEX_API_KEY!)));
```

The Python version of this loop is in the [implementation guide for AI assistants](/llms-full.txt), section "Retries, idempotency and request ids".

## When a pinned version has no capacity

A request that names an exact version (`"model": "dex-1.0.1"`) or sends `"fallback": "never"` never goes to the hosted fallback. When the GPU path cannot serve it in time, it gets `503 no_capacity` with a `retry-after`, and nothing is charged. Choose one of these, per use:

| Your need | What to do after the retries |
| --- | --- |
| Deterministic answers (tests, audits, moderation) | Queue the item and try again later. Do not switch to the alias. |
| An answer now, determinism optional | Send the same request once more with the alias `dex-1` and `"fallback": "allow"`. Store `served_by` and `model` with the result: a `fallback` answer is calibrated separately and not deterministic. |
| A decision on the user's path | Take your safe default: send the case to a person, show "we will get back to you", or confirm with the user. |

## Timeouts

- The GPU path answers typical requests in a fraction of a second; the largest allowed request takes a few seconds. A timeout of 10 to 30 seconds per attempt is safe.
- A timed-out attempt may still finish on the server. The retry carries the same idempotency key, so it either waits for that attempt (`409 idempotency_in_progress`, retried after one second) or replays its response, and you are charged once.

## Never block your product on Dex

Decide what happens when Dex cannot answer before you ship: route to a person, apply a conservative rule, or queue for later. An automatic action should run only on an answer that exists and did not abstain.
