Telluvian vs Cleanlab TLM: hallucination detection compared
Telluvian and Cleanlab's Trustworthy Language Model (TLM) both estimate whether an LLM answer can be trusted. Telluvian reads a proxy model's internal activations and returns a hallucination score for every token, streamed with the answer. TLM returns one trustworthiness score per response, combining self-reflection, consistency across several generated responses and token probabilities.
Last reviewed
At a glance
| Telluvian | Cleanlab TLM | |
|---|---|---|
| Model selection method | Prompt-aware. Model Select predicts each model’s performance on the specific prompt and picks the cheapest one that clears your xPerf quality bar. | Not a model router. It evaluates the outputs of models you call yourself. |
| Models and providers | 190+ models from OpenAI, Anthropic, Google, Qwen, DeepSeek, xAI, Z-AI, Moonshot and Meta, behind one OpenAI-compatible API. | Scores responses from any LLM. Runs on a choice of 15+ evaluation models with 5 quality settings. |
| Pricing model | Prepaid, pay as you go. Provider list price for tokens, plus $0.05 per 1M input tokens for Model Select and $1.00 per 1M completion tokens for hallucination scores (optional). | Free trial tokens, then pay per token at rates shown in your Cleanlab account. Enterprise subscriptions with volume discounts. |
| Hallucination detection | Built in. Per-token hallucination scores from probes reading an open-weight proxy model, including for closed models. Off by default. | Yes. A trustworthiness score per response, from self-reflection, consistency sampling and token probabilities. Response times from about 300 ms. |
Key differences
Granularity
Telluvian scores each token; TLM scores each response.
Method
Telluvian replays the response once through a proxy model; TLM generates additional responses to check consistency and asks an LLM to reflect on the answer.
Delivery
Telluvian streams scores with the tokens, in the same request that routes the prompt.
Existing text
both can score text you already have; Telluvian does it through its -analyze model suffix.
Frequently asked questions
How does Cleanlab TLM compute its trustworthiness score?
Cleanlab says TLM combines self-reflection (an LLM assesses its own response), consistency (it generates several plausible responses and checks for contradictions) and probabilistic measures from token probabilities, plus its own uncertainty methods.
Is the score per token or per response?
TLM returns a trustworthiness score for each response. Telluvian returns a hallucination score for each token, alongside the decoded tokens, so you can highlight the exact span to check.
Does detection add extra model calls?
With TLM, yes: its consistency check generates several plausible responses and compares them, and its self-reflection asks an LLM to assess the answer. Telluvian replays the response once through an open-weight proxy model, reads its activations and streams a score with each token.
Who owns Cleanlab?
Cleanlab's own site states it has been acquired by Handshake AI. TLM continues under the Cleanlab name.
Try it on your own prompts
Point an OpenAI SDK at Telluvian, send telluvian/gallery-1 as the model, and see which model answers each request. Questions about your use case go straight to the team.
Sources
Competitor details come from their own public docs and pricing pages, checked on . Where a page did not say, neither do we. Spotted something out of date? Tell us at hello@telluvian.ai.