Star on GitHub

About

We’re building the LLM gateway you can put in front of a diligence team.

TokenTriage’s brand — and its moat — is honesty. In a category that competes on provider count and throughput, we compete on trust: provenance, quality-proven routing, and a dashboard where every number survives an audit.

The mission

Teams are spending more on inference than ever and can explain less of it. The obvious fix — switch to a cheaper model — keeps getting rolled back, because nobody can prove the cheaper path held quality. So spend stays high and trust stays low.

Our mission is to make LLM cost legible and safely reducible: to show where every dollar goes, and to cut the bill only after the cheaper path is proven just as good — with evidence you can audit. The product would rather say “I can’t prove that” than show you a pretty lie.

Why now

Three things changed at once

The vision

The decision plane

A router is the first step. The destination is the decision plane: for every workload, the cheapest LLM execution plan that can be shown to meet its quality, latency, cost, and policy requirements. Not “cheaper on average,” but cost reduction with bounded risk and defensible evidence — a control plane that treats a quality guarantee the way a database treats a transaction.

Everything we ship today — the ledger, the shadow report, the honesty verdict — is a building block for that: the evidence layer a decision plane needs to be trusted.

Principles

What we hold ourselves to

The project

TokenTriage is open source under Apache-2.0 — one self-contained binary with the dashboard baked in. The trust artifacts (the ledger, verdicts, provenance, the shadow report) are the brand, and they stay free. The code is the argument: read it, run it, and check every number for yourself.

Building something here, evaluating for your team, or investing? Get in touch.

Read the code. Check the numbers.

The fastest way to trust an honesty-first product is to audit it yourself.