# Recipe: grade LLM outputs

> Grade an assistant's answer on a pinned version with one check per rubric property and a helpfulness rate, then record a pass or fail for the eval run.

Your support assistant answered a customer, and your eval suite must say whether the answer is good. One request grades it against a rubric: format, content, refusal and helpfulness. The request pins an exact version, so the same output gets the same grade in every run.

## The request

- **`model: "dex-1.0.1"`.** An exact version, not the alias: when the alias moves to a newer version, your eval scores do not move with it, and a pinned request never uses the fallback. See [Pin versions in CI](/docs/cookbook/pin-versions-in-ci/).
- **State.** The instruction the assistant got and the answer it wrote.
- **One check per property.** `uses_bullets`, `answers_question` and `refusal` each test one line of the rubric, so a failed grade says which line failed. A single `rate` or `pick` over the whole rubric mixes the lines into one answer and is weaker: `rate` is in beta, and one number cannot say what went wrong.
- **`helpfulness` is a `rate`** in beta on four levels. Its instructions say to judge the content, not the form, so the format rule is not counted twice.
- **`min_confidence`** on every question: an unsure grade abstains and goes to a person.

```json
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}
```

## Run it

Put a test key in `DEX_API_KEY` (see the [Quickstart](/docs/quickstart/#get-a-key-in-60-seconds)). Save the Python code as `llm_grading.py` and run `python llm_grading.py`. Save the TypeScript code as `llm-grading.mts` and run `npx tsx llm-grading.mts`: the code uses `await` at the top level, and the `.mts` ending makes the file an ES module. Install the SDKs from [Downloads](/docs/reference/sdks/#downloads).

```bash tab="curl"
curl https://api.thinqit.ai/v1/decide \
  -H "authorization: Bearer $DEX_API_KEY" \
  -H "content-type: application/json" \
  --data-binary @- <<'DEX_REQUEST'
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}
DEX_REQUEST
```

```python tab="Python"
# Save as dex_request.py, then run: python dex_request.py
import json

from thinqit_dex import Client

client = Client()  # reads DEX_API_KEY from the environment

request = json.loads(r'''
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}
''')

decision = client.decide(
    request["state"],
    request["questions"],
    model=request["model"],
)
for question_id, answer in decision.answers.items():
    print(question_id, answer)
```

```ts tab="TypeScript"
// Save as dex-request.mts, then run: npx tsx dex-request.mts (Node.js 18 or newer)
import { Client, parseRequest } from "@thinqit/dex";

const client = new Client(); // reads DEX_API_KEY from the environment

// parseRequest keeps the key order of the text (JSON.parse would move labels such as "1" to the front).
const request = parseRequest(`{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}`);

const decision = await client.decide(request);
console.log(decision.answers);
```

## Expected output

```json
{
  "id": "req_01M3JE5PYJ7ZX1YG6GJD98E7WR",
  "object": "decision",
  "created": 1790546467,
  "model": "dex-1.0.1",
  "served_by": "gpu",
  "calibration": "cal-20260926-1",
  "answers": {
    "uses_bullets": {
      "type": "check",
      "probability": 0.158,
      "confidence": 0.684,
      "abstained": false
    },
    "answers_question": {
      "type": "check",
      "probability": 0.8638,
      "confidence": 0.7277,
      "abstained": false
    },
    "refusal": { "type": "check", "probability": 0.094, "confidence": 0.812, "abstained": false },
    "helpfulness": {
      "type": "rate",
      "rating": 2.1711,
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "probabilities": [0.0167, 0.0564, 0.6661, 0.2608],
      "confidence": 0.603,
      "abstained": false
    }
  },
  "usage": {
    "input_tokens": 164,
    "state_tokens": 73,
    "question_tokens": 91,
    "allowance_tokens": 0,
    "paid_tokens": 0,
    "charge_micro_cents": 0,
    "unit_price_micro_cents": 0,
    "tier": "test"
  }
}
```

Captured from the live API on 2026-09-27 with a test key: model `dex-1.0.1`, calibration `cal-20260926-1`, `served_by: gpu`, 164 input tokens (73 for the state, 91 for the questions). A test key is charged nothing, so `tier` is `test` and the charge is 0. On this exact version the same request always returns these answers.

| Question | Type | Answer | Confidence | `min_confidence` | Abstained |
| --- | --- | --- | --- | --- | --- |
| `uses_bullets` | check | yes with probability 0.158 | 0.684 | 0.6 | no |
| `answers_question` | check | yes with probability 0.8638 | 0.7277 | 0.6 | no |
| `refusal` | check | yes with probability 0.094 | 0.812 | 0.6 | no |
| `helpfulness` | rate | rating 2.1711, most likely `Helpful` (0.6661) | 0.603 | 0.3 | no |

The answer tells the customer how to reset the password (0.8638) and does not refuse (0.094), but it is not written as bullet points (0.158), which the instruction asked for. Its helpfulness rating of 2.1711 is closest to `Helpful` (0.6661). No answer abstained, so the answer fails on format alone.

## Act on it

`case_id` and `record_grade` stand for your own code. Each check becomes a yes or no in the grade, and an abstained one becomes `None` for a person to grade. The case passes only when the answer uses bullets, answers the question and does not refuse. The grade is stored with `decision.model`, so every score says which version gave it. Here the case fails on `uses_bullets`.

```python tab="Python"
grade = {}
for question_id in ("uses_bullets", "answers_question", "refusal"):
    answer = decision.check(question_id)
    grade[question_id] = None if answer.abstained else answer.probability >= 0.5
helpfulness = decision.rate("helpfulness")
grade["helpfulness"] = None if helpfulness.abstained else helpfulness.rating

passed = grade["uses_bullets"] is True and grade["answers_question"] is True and grade["refusal"] is False
record_grade(case_id, decision.model, grade, passed)  # None: a person grades that property
```

```ts tab="TypeScript"
const grade: Record<string, boolean | number | null> = {};
for (const id of ["uses_bullets", "answers_question", "refusal"]) {
  const answer = decision.answers[id];
  grade[id] = answer?.type === "check" && !answer.abstained ? answer.probability >= 0.5 : null;
}
const { helpfulness } = decision.answers;
grade.helpfulness = helpfulness?.type === "rate" && !helpfulness.abstained ? helpfulness.rating : null;

const passed = grade.uses_bullets === true && grade.answers_question === true && grade.refusal === false;
recordGrade(caseId, decision.model, grade, passed); // null: a person grades that property
```

As in the [support recipe](/docs/cookbook/support-routing/#act-on-it), the TypeScript answers have the general `Answer` type, so the code narrows each one on `type`.

## Adapt it

- **Your rubric, one check per line.** Write each line as a yes or no statement about the answer, phrased so that yes is the case you look for. One request takes up to 32 questions, and the graded text is billed once however many lines read it.
- **Move the pin on purpose.** When a new version ships, grade your suite on both versions, compare, then move the pin. Superseded versions stay available for at least 90 days. See [Pin versions in CI](/docs/cookbook/pin-versions-in-ci/) and [Models and versions](/docs/concepts/models/).
- **Whole eval runs.** Grade hundreds of outputs with bounded concurrency and a results file: see [Score a batch](/docs/cookbook/batch-scoring/).
- **Check the grader.** Grade a sample by hand, compare with the checks, and raise `min_confidence` on the lines where the two disagree. See [Abstention](/docs/concepts/abstention/#choosing-a-threshold).
