bellwethr exists to publish a single, comparable price for AI inference — the USD cost of a million tokens from any model that clears a fixed bar of capability. Today that number is scattered across dozens of providers, models and pricing tiers, with no apples-to-apples way to compare them.
Our goal right now isn't to declare the answer — it's to define the standard in the open, and to gather real-world price and usage data as a public census for research. The methodology below is a draft. Every part of it is up for debate.
A model only enters the index if it clears a fixed benchmark bar — so we're always pricing comparable intelligence, not just the cheapest tokens. Which benchmarks and what score still open.
How input and output tokens, context length and providers combine into one figure. Draft leans volume-weighted across qualifying providers.
Median, volume-weighted or list price — and how often the index refreshes. Built from the prices contributors actually pay.
The census takes a few minutes — share the prices you pay, your usage, and how you think the index should be calculated.