Open-weight models · Montreal & Calgary

Fast serverless inference
for open LLMs.

Canary runs the world's leading open source models in Canadian data centres — at a fraction of what you're paying today, with data that never leaves Canada. We serve trillions of tokens per month.

Drop-in compatible with Claude Code, Codex, Cline, and LangChain.

- base_url = "https://api.openai.com/v1"
+ base_url = "https://api.canary.ca/v1"

  api_key = os.environ["CANARY_API_KEY"]

Canadian data residency

No US cloud, no third-country data flows — nothing leaves Canada, ever.

True data control

Turn on ZDR (Zero Data Retention) and we will never store your requests. No logging, and we don't use it for training.

Zero migration friction

OpenAI and Anthropic-compatible endpoints out of the box. Point Claude Code, Codex, Cline, or LangChain at Canary and keep your code exactly as it is.

Radically cheaper

$4.00 per million output tokens, flat. No metering games, no surprise bills — a fraction of what incumbents charge for comparable capability.

Developer Partner Network

Join our Development Partner Program and receive monthly free token grants, early access to new models, and access to priority support.

The switch pays for itself immediately.

Same capability, dramatically lower cost, hosted in Canada.

OpenAI / Anthropic
Claude, GPT
$15–75/M
Other inference resellers
wafer.ai, DeepInfra, Together
$4–12/M
Canary
Blended output tokens
$4.00/M

Serverless GPU clusters, entirely on Canadian infrastructure.

We can host models up to 1 trillion parameters with low latency, and a model router that makes it easy to switch between models without breaking production. No GPU procurement, no capacity planning — just an API key.

1T
Max model size supported, parameters
1M
Context window, in tokens
170+
Output throughput, tokens/sec
2
Canadian regions — Montreal, Calgary

Start building on Canadian infrastructure today.