Docs menu

Cookbook

Recipe: check data quality

Check whether two records describe the same customer, pick the kind of value in a field and check for junk, then merge, send the pair to a person or move the value to the right field.

View as Markdown

An import brings in a customer who may already be in your CRM, and a contact field whose value does not say what it is. One request checks whether the two records describe the same customer, what kind of value the field holds, and whether the new record looks like junk. Your code merges only when every answer is sure.

The request

  • State. The two records and the field as JSON. {{record_a}} and {{record_b}} point at whole objects, so one question reads every field of both records, and each record is billed once however many questions point at it.
  • same_customer is a check whose criteria say what counts as the same customer: the same phone number in another format, the same city and a matching name.
  • value_kind is a pick of the kinds of value you store, with other as the way out.
  • junk is a check on the new record only.
  • min_confidence of 0.6 on both checks: a check then answers only at a probability of 0.8 or more, or 0.2 or less, and abstains in between.
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}

Run it

Put a test key in DEX_API_KEY (see the Quickstart). Save the Python code as data_quality.py and run python data_quality.py. Save the TypeScript code as data-quality.mts and run npx tsx data-quality.mts: the code uses await at the top level, and the .mts ending makes the file an ES module. Install the SDKs from Downloads.

curl https://api.thinqit.ai/v1/decide \
  -H "authorization: Bearer $DEX_API_KEY" \
  -H "content-type: application/json" \
  --data-binary @- <<'DEX_REQUEST'
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}
DEX_REQUEST
# Save as dex_request.py, then run: python dex_request.py
import json

from thinqit_dex import Client

client = Client()  # reads DEX_API_KEY from the environment

request = json.loads(r'''
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}
''')

decision = client.decide(
    request["state"],
    request["questions"],
)
for question_id, answer in decision.answers.items():
    print(question_id, answer)
// Save as dex-request.mts, then run: npx tsx dex-request.mts (Node.js 18 or newer)
import { Client, parseRequest } from "@thinqit/dex";

const client = new Client(); // reads DEX_API_KEY from the environment

// parseRequest keeps the key order of the text (JSON.parse would move labels such as "1" to the front).
const request = parseRequest(`{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}`);

const decision = await client.decide(request);
console.log(decision.answers);

Expected output

{
  "id": "req_01M3JE1K0AKMX2RT390N9G9C5S",
  "object": "decision",
  "created": 1790546332,
  "model": "dex-1.0.1",
  "served_by": "gpu",
  "calibration": "cal-20260926-1",
  "answers": {
    "same_customer": {
      "type": "check",
      "probability": 0.8776,
      "confidence": 0.7551,
      "abstained": false
    },
    "value_kind": {
      "type": "pick",
      "choice": "phone",
      "probabilities": {
        "email": 0.0276,
        "phone": 0.9069,
        "address": 0.0239,
        "name": 0.0218,
        "other": 0.0198
      },
      "confidence": 0.8793,
      "abstained": false
    },
    "junk": { "type": "check", "probability": 0.2219, "confidence": 0.5563, "abstained": true }
  },
  "usage": {
    "input_tokens": 227,
    "state_tokens": 124,
    "question_tokens": 103,
    "allowance_tokens": 0,
    "paid_tokens": 0,
    "charge_micro_cents": 0,
    "unit_price_micro_cents": 0,
    "tier": "test"
  }
}

Captured from the live API on 2026-09-27 with a test key: model dex-1.0.1, calibration cal-20260926-1, served_by: gpu, 227 input tokens (124 for the state, 103 for the questions). A test key is charged nothing, so tier is test and the charge is 0. On this exact version the same request always returns these answers.

Question Type Answer Confidence min_confidence Abstained
same_customer check yes with probability 0.8776 0.7551 0.6 no
value_kind pick phone (0.9069) 0.8793 0.5 no
junk check yes with probability 0.2219 0.5563 0.6 yes

The two records describe the same customer (0.8776), and the contact field holds a phone number (0.9069). The junk check abstained: 0.2219 leans towards a real customer, but it is above 0.2, so its confidence of 0.5563 falls below the floor of 0.6. The code must not act on it.

Act on it

pair_id, record_id, contact and the functions queue_for_review, flag_as_junk, merge_records, set_field and to_e164 stand for your own code. Here the pair goes to a person, because the junk check abstained: an unsure answer never merges data, even when the records match. The contact value still moves to the phone field, normalised in your code to +31302345678.

same = decision.check("same_customer")
junk = decision.check("junk")
if same.abstained or junk.abstained:
    queue_for_review(pair_id)  # an unsure answer: a person looks before any merge
elif junk.probability >= 0.5:
    flag_as_junk(pair_id)
elif same.probability >= 0.5:
    merge_records(pair_id)

kind = decision.pick("value_kind")
if not kind.abstained and kind.choice == "phone":
    set_field(record_id, "phone", to_e164(contact))  # normalise in code, not in Dex
const { same_customer: same, value_kind: kind, junk } = decision.answers;

if (same?.type === "check" && junk?.type === "check") {
  if (same.abstained || junk.abstained) queueForReview(pairId); // an unsure answer: a person looks before any merge
  else if (junk.probability >= 0.5) flagAsJunk(pairId);
  else if (same.probability >= 0.5) mergeRecords(pairId);
}
if (kind?.type === "pick" && !kind.abstained && kind.choice === "phone") {
  setField(recordId, "phone", toE164(contact)); // normalise in code, not in Dex
}

As in the support recipe, the TypeScript answers have the general Answer type, so the code narrows each one on type.

Adapt it

  • Normalise in code, then ask. When you can compute a normal form, do it before the request and pass it as a field: both phone numbers in E.164 (+31302345678), email addresses in lower case. The model then compares two equal values instead of working out that +31 30 234 5678 and 030 2345678 are the same number. Dex reads; it does not calculate. See Known limits.
  • Exact matches need no model. When two records share an email address exactly, merge them in code, and send only the fuzzy pairs to Dex.
  • Fewer pairs for a person. If too many pairs wait because junk abstains, ask a narrower question, such as whether the email address uses a test domain, and set its floor from records you already cleaned. See Abstention.
  • Whole tables. Run the pairs with bounded concurrency and a results file: see Score a batch. On the same version a rerun gives the same verdicts, so it does not reshuffle your data.