Compare

Dex vs LLMs

A general LLM can answer the same decisions as Dex. Here is how the two compare, row by row. Figures come from our own recorded races, everything else about the LLMs from their providers' own pages, each with the date we read it.

Comparison

Dex compared with general LLMs, row by row, with sources
TopicDexLLMsSource
Time to all answersDex differs108 to 153 ms per workflow, all questions in one request. At least 6.0 times faster than the fastest LLM, in every workflow.754 ms to 11.6 s per workflow for the 8 LLMs raced, each in one streamed call with all questions.Dex ran on its own inference node in the Netherlands, timed as a round trip over the network like the LLM calls, one request at a time. LLM latency depends on provider load at the time.thinQit Dex, recorded races against LLMsread 25 Sept 2026Median of 10 runs per model and workflow, 4 workflows
Cost per callDex differsEUR 0.000060 to 0.00011 per workflow call at the Fierce rate, EUR 0.09 per million input tokens (EUR 39.00 a month). At least 5.4 times cheaper than the cheapest LLM, in every workflow. Output is free.More on the Dex sideEUR 0.00038 to 0.051 per workflow call: the recorded input and output tokens times each model's list price. Output tokens are billed, thinking tokens included.Pay as you go, Dex costs EUR 1.00 per million input tokens. Prices exclude VAT.thinQit Dex, recorded races against LLMsread 25 Sept 2026Azure Retail Prices API, Foundry Models, Sweden Centralread 25 Sept 2026Claude API docs, Pricingread 25 Sept 2026Median cost per model and workflow; USD list prices converted at 0.858664 EUR per USD
Same input, same answersDex differsThe same answers in every run: 4 of 4 workflows, 10 identical calls each. On a GPU model version, identical request bytes return byte-identical answers, by design.More on the Dex side4 of the 8 LLMs changed at least one answer between identical calls (GPT-5.3 chat, GPT-5.4 mini, GPT-4.1 mini and Claude Sonnet 5), with temperature 0 where the model accepts it.Answers from the Dex fallback are not bit for bit repeatable and say so (served_by is fallback).thinQit Dex, recorded races against LLMsread 25 Sept 202610 identical calls per model and workflow, answers compared per question
Typed answersDex differsTyped by design. Pick one of up to 255 options, rate on 2 to 10 ordered levels, or check yes or no. Every answer carries calibrated probabilities and a confidence, and min_confidence makes an unsure answer abstain.More on the Dex sideGenerated text. In the races each LLM wrote its answers as JSON under a strict schema (the provider's structured output), which your code parses. The JSON holds one label per question and no probability per option.thinQit Dex, recorded races against LLMsread 25 Sept 2026Race method, call settings per model
Where requests are processedDex differsIn the EU only. The gateway in Azure West Europe in the Netherlands, Dex's own inference nodes on dedicated hardware operated by thinQit in the Netherlands, and Azure OpenAI in the EU data zone for the fallback.More on the Dex sideDepends on the offer. Azure OpenAI Global deployments, used in the races, may process prompts in any geography where the model is deployed; DataZone deployments in an EU resource stay in the EU. The Claude API runs inference in any available geography by default, or in the US only; it lists no EU-only option.Microsoft Learn, Data, privacy, and security for Foundry Models sold by Azureread 27 Sept 2026Claude API docs, Data residencyread 27 Sept 2026Location of processing for Global and Data zone deployments; Inference geo table and current limitations
Request content kept by defaultDex differsNone, for every account. Content is stored only if the account owner turns on content logging.More on the Dex sideAzure OpenAI may store prompts and completions that its abuse monitoring flags, for human review in the resource's geography, unless Microsoft approves modified abuse monitoring. The Claude API deletes inputs and outputs within 30 days by default; zero data retention needs an agreement.One exception on the Dex side. On the fallback path, until Microsoft approves modified abuse monitoring, Azure OpenAI may keep prompts its classifiers flag inside the EU data zone. Send fallback set to never to stay on Dex's own nodes.Microsoft Learn, Data, privacy, and security for Foundry Models sold by Azureread 27 Sept 2026Anthropic Privacy Center, How long do you store my organization's data?read 27 Sept 2026Preventing abuse section; default retention paragraph and its exceptions
Training on customer dataSame on bothNever. Customer content is not used to train, tune or calibrate any model.More on the Dex sideAzure OpenAI does not use prompts and completions to train, retrain or improve the base models. Anthropic does not use data it retains from the Claude API for training without express permission.Microsoft Learn, Data, privacy, and security for Foundry Models sold by Azureread 27 Sept 2026Claude API docs, API and data retentionread 27 Sept 2026Summary note and inferencing section; How Anthropic approaches data retention
Agreement with the race referenceLLMs ahead92.3% of answers equal to the reference, on average over the 4 workflows.More on the Dex side87.6% to 99.8% on average per LLM; the highest, Claude Sonnet 5, 99.8%.Not a quality benchmark. The reference is the majority answer of three of the LLMs raced, which favours them, and each workflow has one state. Dex quality is measured on our own English and Dutch test sets.thinQit Dex, recorded races against LLMsread 25 Sept 2026agreement_with_reference per model and workflow
Largest requestLLMs ahead16,384 billable tokens per request, of which the questions together can use up to 4,096.More on the Dex sideThe Claude models raced take 200,000 to 1,000,000 tokens of context.Each count uses its own tokenizer, so equal numbers are not equal lengths of text.Claude API docs, Models overviewread 27 Sept 2026Compare models table, Context window row
What it can doLLMs aheadDecisions only. Pick, rate and check about a text or JSON state. It does not write, summarise, translate or extract text, and it reads no images.More on the Dex sideGeneral. The Claude models raced take text and image input, write text and handle many languages.Claude API docs, Models overviewread 27 Sept 2026Compare models, the paragraph above the table

Where LLMs are ahead

  • 01Agreement with the race reference. 87.6% to 99.8% on average per LLM; the highest, Claude Sonnet 5, 99.8%.
  • 02Largest request. The Claude models raced take 200,000 to 1,000,000 tokens of context.
  • 03What it can do. General. The Claude models raced take text and image input, write text and handle many languages.

How we made this comparison

The LLM figures are our own measurements: the same workflows and the same questions for every model, 10 recorded runs each, on 25 Sept 2026. Every run is on the home page, where you can replay it.

Dex costs are at the Fierce plan's rate, EUR 0.09 per million input tokens; pay as you go is EUR 1.00 per million input tokens. LLM costs are the tokens each call used times the provider's list price on the day we read it.

Everything else about the LLMs comes from their providers' own public pages, with the page and the date we read it. Where LLMs are ahead, the row says so.

Replay the racesCost and speed per model

Sources