All posts
Adam Pattenden

Introducing Gallery

Motivation

AI use has skyrocketed and token prices have collapsed. Per-token inference prices for state-of-the-art capability have been falling by orders of magnitude each year and Gartner projects a further ~90% drop in provider-side costs by 2030.

Despite per-token pricing collapsing, overall spend is only going in one direction. Global enterprise AI spending is projected to hit ~$407B this year, up roughly 35% year-on-year, with enterprise LLM API spend passing $8.4B in 2025 and on track to double. The mechanism seems to follow Jevons paradox: usage is growing faster than unit prices are falling. This usage increase is itself known on the bottom line:

  • Uber's engineering wing reportedly blew its entire annual AI coding budget a third of the way into the year. Its CEO, Dara Khosrowshahi, said the cost of the company's AI initiatives was "getting harder to justify" relative to output.
  • Microsoft reportedly cancelled most of its internal Claude Code licenses, partly over cost, six months after rolling them out.
  • Axios reported one company spent half a billion dollars in a single month after failing to set usage limits.
  • Sam Altman, on stage at Intelligence at Work: "My company spent my entire 2026 budget in Q1". This has since become a meme.

Your problem

How much of your day needs your full concentration? What fraction of your AI requests actually need your most expensive model?

Most of the traffic sent to frontier models can be answered just as well by a model that costs less than a tenth as much. If you commit to using a specific model then you’re losing accuracy or overpaying for routine work. The AI frontier moves on every few weeks and keeping up with it is a full time job. In such an environment, locking into a long or even medium term deal with any given provider represents a risk.

Scatter plot titled 'A higher price doesn't mean a better answer': Artificial Analysis Intelligence Index against cost per benchmark task on a log scale. GLM-5.3-Flash scores 42 at $0.25 a task, while Claude Sonnet 5 scores 38 at $5.09 a task, marked 'Higher score, 20x cheaper'.

Our solution

Telluvian gives you every major AI model through one connection, with no setup, no training time and no lock-in. Send your request and we find the model that will answer it best at the lowest cost. Easy jobs go to cheap models, hard ones go to the frontier.

We don't rank models on how smart they are overall. We score how well each one will handle your particular request. That score is on a fixed scale so you can get the same quality next year at a lower cost as new models arrive.

If desired we can also provide every answer token with a confusion score, so you know which facts to double check.

We are more than a gateway. A gateway is one connection that gives you access to many models but you still choose which model to use. A router makes that choice for you, picking the right model for each request.

Technical details

We are plug and play. No long rebuild, no difficult migration. Point your code at Telluvian, swap in your key, and set the model to telluvian/gallery-1. That’s it. Everything else stays the same, including streaming and multi-turn conversations. The API is OpenAI compatible.

{
  "model": "telluvian/gallery-1",
  "messages": [
    { "role": "user", "content": "Give me the full name only: who was the 37th president of the United States?" }
  ]
}

Gallery reads the question, picks the model, then sends it on. The answer comes back in the same format as before, and tells you which model wrote it:

{
  "model": "google/gemini-3.5-flash-lite",
  "choices": [
    { "message": { "role": "assistant", "content": "Richard Milhous Nixon" } }
  ]
}

To provide you with greater control, we also include a number of optional fields in the call. For example:

SettingWhat it doesDefault
xPerfThe quality you need, e.g. "at least as good as Opus 5", or a number on a fixed scale that never moves0.9
xPerfEffortHow hard that reference model should be thinking, only takes effect when xPerf is specified as a model. minimal to maxBest available
modelZooWhich models Gallery is allowed to pick, e.g. "anthropic/*,google/*"all of them
sessionIdAn ID you reuse across a conversation so it tends to stay on one modelnone
{
  "model": "telluvian/gallery-1",
  "messages": [{ "role": "user", "content": "Summarise the termination clauses in this contract." }],
  "xPerf": "anthropic/claude-opus-5",
  "xPerfEffort": "high",
  "modelZoo": "anthropic/*,google/*,openai/*",
  "sessionId": "3f2b7c58-9d41-4e0a-9a7c-6f0b1c2d3e4f"
}

Add hallucination detection: add include_scores: true to any request and every response token comes back with a ‘confusion’ score showing how likely it is to be made up.

Pricing: routing costs $0.05 per million input tokens when turned on. Hallucination scoring costs $1.00 per million completion tokens, only when you turn it on. Otherwise, you pay the model's list price with no further margin added.

Optimisation

You can imagine that every AI call sits somewhere on a three-dimensional curve between cost, latency, and performance. Most enterprise users haven't yet gone looking to optimise their position on it. telluvian/gallery-1 allows you to move across this surface dynamically, in a prompt-aware manner. As you can see, the most significant gains are when dealing with easy prompts.

The current standard practice of rationing employee access to frontier models just trades performance in favour of cost. In contrast, Telluvian allows your team full access to the best models when they are needed.

Currently we have inbuilt hallucination detection. However, we are bringing further features to Gallery. Prompt pre-processing (compression), optimal reasoning (see Fragile Correctness: Cases of reasoning harming performance), reasoning swap and sharding/orchestration are all on our roadmap and will provide further cost savings to you, the user.

Glossary of terms

  • API: the means by which you send requests to an AI and get answers back.
  • API key: the password that identifies and bills your account.
  • Model: a specific AI, such as claude-opus-5, each with its own price and ability.
  • Provider: the company that runs a model, such as Anthropic or OpenAI.
  • Token: the unit AI throughput is measured in, it’s roughly equivalent to a word.
  • Input and output tokens: what you send versus what the AI writes back, with output usually costing several times more.