# Recipe: check data quality

> Check whether two records describe the same customer, pick the kind of value in a field and check for junk, then merge, send the pair to a person or move the value to the right field.

An import brings in a customer who may already be in your CRM, and a contact field whose value does not say what it is. One request checks whether the two records describe the same customer, what kind of value the field holds, and whether the new record looks like junk. Your code merges only when every answer is sure.

## The request

- **State.** The two records and the field as JSON. `{{record_a}}` and `{{record_b}}` point at whole objects, so one question reads every field of both records, and each record is billed once however many questions point at it.
- **`same_customer` is a `check`** whose `criteria` say what counts as the same customer: the same phone number in another format, the same city and a matching name.
- **`value_kind` is a `pick`** of the kinds of value you store, with `other` as the way out.
- **`junk` is a `check`** on the new record only.
- **`min_confidence`** of 0.6 on both checks: a check then answers only at a probability of 0.8 or more, or 0.2 or less, and abstains in between.

```json
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}
```

## Run it

Put a test key in `DEX_API_KEY` (see the [Quickstart](/docs/quickstart/#get-a-key-in-60-seconds)). Save the Python code as `data_quality.py` and run `python data_quality.py`. Save the TypeScript code as `data-quality.mts` and run `npx tsx data-quality.mts`: the code uses `await` at the top level, and the `.mts` ending makes the file an ES module. Install the SDKs from [Downloads](/docs/reference/sdks/#downloads).

```bash tab="curl"
curl https://api.thinqit.ai/v1/decide \
  -H "authorization: Bearer $DEX_API_KEY" \
  -H "content-type: application/json" \
  --data-binary @- <<'DEX_REQUEST'
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}
DEX_REQUEST
```

```python tab="Python"
# Save as dex_request.py, then run: python dex_request.py
import json

from thinqit_dex import Client

client = Client()  # reads DEX_API_KEY from the environment

request = json.loads(r'''
{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}
''')

decision = client.decide(
    request["state"],
    request["questions"],
)
for question_id, answer in decision.answers.items():
    print(question_id, answer)
```

```ts tab="TypeScript"
// Save as dex-request.mts, then run: npx tsx dex-request.mts (Node.js 18 or newer)
import { Client, parseRequest } from "@thinqit/dex";

const client = new Client(); // reads DEX_API_KEY from the environment

// parseRequest keeps the key order of the text (JSON.parse would move labels such as "1" to the front).
const request = parseRequest(`{
  "state": {
    "record_a": {
      "name": "J. van den Berg",
      "email": "jvdberg@bakkerijberg.nl",
      "phone": "+31 30 234 5678",
      "city": "Utrecht"
    },
    "record_b": {
      "name": "Jan van den Berg",
      "email": "jan@bakkerijberg.nl",
      "phone": "030 2345678",
      "city": "Utrecht"
    },
    "field": { "name": "contact", "value": "030 2345678" }
  },
  "questions": {
    "same_customer": {
      "type": "check",
      "instructions": "Do {{record_a}} and {{record_b}} describe the same customer?",
      "criteria": "The same phone number written in a different format, together with the same city and a matching name, points to the same customer.",
      "min_confidence": 0.6
    },
    "value_kind": {
      "type": "pick",
      "instructions": "What kind of value is {{field.value}}?",
      "options": {
        "email": "An email address",
        "phone": "A phone number",
        "address": "A postal address",
        "name": "A person's or company's name",
        "other": null
      },
      "min_confidence": 0.5
    },
    "junk": {
      "type": "check",
      "instructions": "Does {{record_b}} look like test data or junk rather than a real customer?",
      "min_confidence": 0.6
    }
  }
}`);

const decision = await client.decide(request);
console.log(decision.answers);
```

## Expected output

```json
{
  "id": "req_01M3JE1K0AKMX2RT390N9G9C5S",
  "object": "decision",
  "created": 1790546332,
  "model": "dex-1.0.1",
  "served_by": "gpu",
  "calibration": "cal-20260926-1",
  "answers": {
    "same_customer": {
      "type": "check",
      "probability": 0.8776,
      "confidence": 0.7551,
      "abstained": false
    },
    "value_kind": {
      "type": "pick",
      "choice": "phone",
      "probabilities": {
        "email": 0.0276,
        "phone": 0.9069,
        "address": 0.0239,
        "name": 0.0218,
        "other": 0.0198
      },
      "confidence": 0.8793,
      "abstained": false
    },
    "junk": { "type": "check", "probability": 0.2219, "confidence": 0.5563, "abstained": true }
  },
  "usage": {
    "input_tokens": 227,
    "state_tokens": 124,
    "question_tokens": 103,
    "allowance_tokens": 0,
    "paid_tokens": 0,
    "charge_micro_cents": 0,
    "unit_price_micro_cents": 0,
    "tier": "test"
  }
}
```

Captured from the live API on 2026-09-27 with a test key: model `dex-1.0.1`, calibration `cal-20260926-1`, `served_by: gpu`, 227 input tokens (124 for the state, 103 for the questions). A test key is charged nothing, so `tier` is `test` and the charge is 0. On this exact version the same request always returns these answers.

| Question | Type | Answer | Confidence | `min_confidence` | Abstained |
| --- | --- | --- | --- | --- | --- |
| `same_customer` | check | yes with probability 0.8776 | 0.7551 | 0.6 | no |
| `value_kind` | pick | `phone` (0.9069) | 0.8793 | 0.5 | no |
| `junk` | check | yes with probability 0.2219 | 0.5563 | 0.6 | **yes** |

The two records describe the same customer (0.8776), and the contact field holds a phone number (0.9069). The junk check abstained: 0.2219 leans towards a real customer, but it is above 0.2, so its confidence of 0.5563 falls below the floor of 0.6. The code must not act on it.

## Act on it

`pair_id`, `record_id`, `contact` and the functions `queue_for_review`, `flag_as_junk`, `merge_records`, `set_field` and `to_e164` stand for your own code. Here the pair goes to a person, because the junk check abstained: an unsure answer never merges data, even when the records match. The contact value still moves to the phone field, normalised in your code to `+31302345678`.

```python tab="Python"
same = decision.check("same_customer")
junk = decision.check("junk")
if same.abstained or junk.abstained:
    queue_for_review(pair_id)  # an unsure answer: a person looks before any merge
elif junk.probability >= 0.5:
    flag_as_junk(pair_id)
elif same.probability >= 0.5:
    merge_records(pair_id)

kind = decision.pick("value_kind")
if not kind.abstained and kind.choice == "phone":
    set_field(record_id, "phone", to_e164(contact))  # normalise in code, not in Dex
```

```ts tab="TypeScript"
const { same_customer: same, value_kind: kind, junk } = decision.answers;

if (same?.type === "check" && junk?.type === "check") {
  if (same.abstained || junk.abstained) queueForReview(pairId); // an unsure answer: a person looks before any merge
  else if (junk.probability >= 0.5) flagAsJunk(pairId);
  else if (same.probability >= 0.5) mergeRecords(pairId);
}
if (kind?.type === "pick" && !kind.abstained && kind.choice === "phone") {
  setField(recordId, "phone", toE164(contact)); // normalise in code, not in Dex
}
```

As in the [support recipe](/docs/cookbook/support-routing/#act-on-it), the TypeScript answers have the general `Answer` type, so the code narrows each one on `type`.

## Adapt it

- **Normalise in code, then ask.** When you can compute a normal form, do it before the request and pass it as a field: both phone numbers in E.164 (`+31302345678`), email addresses in lower case. The model then compares two equal values instead of working out that `+31 30 234 5678` and `030 2345678` are the same number. Dex reads; it does not calculate. See [Known limits](/docs/concepts/known-limits/#date-and-amount-arithmetic-is-unreliable).
- **Exact matches need no model.** When two records share an email address exactly, merge them in code, and send only the fuzzy pairs to Dex.
- **Fewer pairs for a person.** If too many pairs wait because `junk` abstains, ask a narrower question, such as whether the email address uses a test domain, and set its floor from records you already cleaned. See [Abstention](/docs/concepts/abstention/#choosing-a-threshold).
- **Whole tables.** Run the pairs with bounded concurrency and a results file: see [Score a batch](/docs/cookbook/batch-scoring/). On the same version a rerun gives the same verdicts, so it does not reshuffle your data.
