Fast serverless inference
for open LLMs.
Canary runs the world's leading open source models in Canadian data centres — at a fraction of what you're paying today, with data that never leaves Canada. We serve trillions of tokens per month.
Drop-in compatible with Claude Code, Codex, Cline, and LangChain.
- base_url = "https://api.openai.com/v1"
+ base_url = "https://api.canary.ca/v1"
api_key = os.environ["CANARY_API_KEY"]Canadian data residency
No US cloud, no third-country data flows — nothing leaves Canada, ever.
True data control
Turn on ZDR (Zero Data Retention) and we will never store your requests. No logging, and we don't use it for training.
Zero migration friction
OpenAI and Anthropic-compatible endpoints out of the box. Point Claude Code, Codex, Cline, or LangChain at Canary and keep your code exactly as it is.
Radically cheaper
$4.00 per million output tokens, flat. No metering games, no surprise bills — a fraction of what incumbents charge for comparable capability.
Developer Partner Network
Join our Development Partner Program and receive monthly free token grants, early access to new models, and access to priority support.
The switch pays for itself immediately.
Same capability, dramatically lower cost, hosted in Canada.
Serverless GPU clusters, entirely on Canadian infrastructure.
We can host models up to 1 trillion parameters with low latency, and a model router that makes it easy to switch between models without breaking production. No GPU procurement, no capacity planning — just an API key.