Docs menu

Cookbook

Recipe: grade LLM outputs

Grade an assistant's answer on a pinned version with one check per rubric property and a helpfulness rate, then record a pass or fail for the eval run.

View as Markdown

Your support assistant answered a customer, and your eval suite must say whether the answer is good. One request grades it against a rubric: format, content, refusal and helpfulness. The request pins an exact version, so the same output gets the same grade in every run.

The request

  • model: "dex-1.0.1". An exact version, not the alias: when the alias moves to a newer version, your eval scores do not move with it, and a pinned request never uses the fallback. See Pin versions in CI.
  • State. The instruction the assistant got and the answer it wrote.
  • One check per property. uses_bullets, answers_question and refusal each test one line of the rubric, so a failed grade says which line failed. A single rate or pick over the whole rubric mixes the lines into one answer and is weaker: rate is in beta, and one number cannot say what went wrong.
  • helpfulness is a rate in beta on four levels. Its instructions say to judge the content, not the form, so the format rule is not counted twice.
  • min_confidence on every question: an unsure grade abstains and goes to a person.
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}

Run it

Put a test key in DEX_API_KEY (see the Quickstart). Save the Python code as llm_grading.py and run python llm_grading.py. Save the TypeScript code as llm-grading.mts and run npx tsx llm-grading.mts: the code uses await at the top level, and the .mts ending makes the file an ES module. Install the SDKs from Downloads.

curl https://api.thinqit.ai/v1/decide \
  -H "authorization: Bearer $DEX_API_KEY" \
  -H "content-type: application/json" \
  --data-binary @- <<'DEX_REQUEST'
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}
DEX_REQUEST
# Save as dex_request.py, then run: python dex_request.py
import json

from thinqit_dex import Client

client = Client()  # reads DEX_API_KEY from the environment

request = json.loads(r'''
{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}
''')

decision = client.decide(
    request["state"],
    request["questions"],
    model=request["model"],
)
for question_id, answer in decision.answers.items():
    print(question_id, answer)
// Save as dex-request.mts, then run: npx tsx dex-request.mts (Node.js 18 or newer)
import { Client, parseRequest } from "@thinqit/dex";

const client = new Client(); // reads DEX_API_KEY from the environment

// parseRequest keeps the key order of the text (JSON.parse would move labels such as "1" to the front).
const request = parseRequest(`{
  "model": "dex-1.0.1",
  "state": {
    "task": {
      "instruction": "Answer in exactly three bullet points: how does a customer reset their password?",
      "answer": "Sure! Resetting your password is easy. First, open the app and tap Settings. Then choose Account and tap Reset password. We will email you a link that works for 30 minutes. Let me know if you need anything else!"
    }
  },
  "questions": {
    "uses_bullets": {
      "type": "check",
      "instructions": "Is {{task.answer}} written as bullet points, as {{task.instruction}} asks?",
      "min_confidence": 0.6
    },
    "answers_question": {
      "type": "check",
      "instructions": "Does {{task.answer}} tell the customer how to reset their password?",
      "min_confidence": 0.6
    },
    "refusal": {
      "type": "check",
      "instructions": "Does {{task.answer}} refuse or avoid the task in {{task.instruction}}?",
      "min_confidence": 0.6
    },
    "helpfulness": {
      "type": "rate",
      "instructions": "How helpful is {{task.answer}} for a customer who wants to reset their password? Judge the content, not the form.",
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "min_confidence": 0.3
    }
  }
}`);

const decision = await client.decide(request);
console.log(decision.answers);

Expected output

{
  "id": "req_01M3JE5PYJ7ZX1YG6GJD98E7WR",
  "object": "decision",
  "created": 1790546467,
  "model": "dex-1.0.1",
  "served_by": "gpu",
  "calibration": "cal-20260926-1",
  "answers": {
    "uses_bullets": {
      "type": "check",
      "probability": 0.158,
      "confidence": 0.684,
      "abstained": false
    },
    "answers_question": {
      "type": "check",
      "probability": 0.8638,
      "confidence": 0.7277,
      "abstained": false
    },
    "refusal": { "type": "check", "probability": 0.094, "confidence": 0.812, "abstained": false },
    "helpfulness": {
      "type": "rate",
      "rating": 2.1711,
      "levels": ["Unhelpful", "Somewhat helpful", "Helpful", "Very helpful"],
      "probabilities": [0.0167, 0.0564, 0.6661, 0.2608],
      "confidence": 0.603,
      "abstained": false
    }
  },
  "usage": {
    "input_tokens": 164,
    "state_tokens": 73,
    "question_tokens": 91,
    "allowance_tokens": 0,
    "paid_tokens": 0,
    "charge_micro_cents": 0,
    "unit_price_micro_cents": 0,
    "tier": "test"
  }
}

Captured from the live API on 2026-09-27 with a test key: model dex-1.0.1, calibration cal-20260926-1, served_by: gpu, 164 input tokens (73 for the state, 91 for the questions). A test key is charged nothing, so tier is test and the charge is 0. On this exact version the same request always returns these answers.

Question Type Answer Confidence min_confidence Abstained
uses_bullets check yes with probability 0.158 0.684 0.6 no
answers_question check yes with probability 0.8638 0.7277 0.6 no
refusal check yes with probability 0.094 0.812 0.6 no
helpfulness rate rating 2.1711, most likely Helpful (0.6661) 0.603 0.3 no

The answer tells the customer how to reset the password (0.8638) and does not refuse (0.094), but it is not written as bullet points (0.158), which the instruction asked for. Its helpfulness rating of 2.1711 is closest to Helpful (0.6661). No answer abstained, so the answer fails on format alone.

Act on it

case_id and record_grade stand for your own code. Each check becomes a yes or no in the grade, and an abstained one becomes None for a person to grade. The case passes only when the answer uses bullets, answers the question and does not refuse. The grade is stored with decision.model, so every score says which version gave it. Here the case fails on uses_bullets.

grade = {}
for question_id in ("uses_bullets", "answers_question", "refusal"):
    answer = decision.check(question_id)
    grade[question_id] = None if answer.abstained else answer.probability >= 0.5
helpfulness = decision.rate("helpfulness")
grade["helpfulness"] = None if helpfulness.abstained else helpfulness.rating

passed = grade["uses_bullets"] is True and grade["answers_question"] is True and grade["refusal"] is False
record_grade(case_id, decision.model, grade, passed)  # None: a person grades that property
const grade: Record<string, boolean | number | null> = {};
for (const id of ["uses_bullets", "answers_question", "refusal"]) {
  const answer = decision.answers[id];
  grade[id] = answer?.type === "check" && !answer.abstained ? answer.probability >= 0.5 : null;
}
const { helpfulness } = decision.answers;
grade.helpfulness = helpfulness?.type === "rate" && !helpfulness.abstained ? helpfulness.rating : null;

const passed = grade.uses_bullets === true && grade.answers_question === true && grade.refusal === false;
recordGrade(caseId, decision.model, grade, passed); // null: a person grades that property

As in the support recipe, the TypeScript answers have the general Answer type, so the code narrows each one on type.

Adapt it

  • Your rubric, one check per line. Write each line as a yes or no statement about the answer, phrased so that yes is the case you look for. One request takes up to 32 questions, and the graded text is billed once however many lines read it.
  • Move the pin on purpose. When a new version ships, grade your suite on both versions, compare, then move the pin. Superseded versions stay available for at least 90 days. See Pin versions in CI and Models and versions.
  • Whole eval runs. Grade hundreds of outputs with bounded concurrency and a results file: see Score a batch.
  • Check the grader. Grade a sample by hand, compare with the checks, and raise min_confidence on the lines where the two disagree. See Abstention.