Current callApplications close Oct 23, 2026
Ensembles & Routing (Winter 2026)
Overview
A 12-week grant for researchers who have worked on ensembles, routing, or model composition. We cover the compute and engineering to run your method on current frontier models, re-run every result independently, and publish the full recipe and cost so anyone can verify it.
The problem
More and more of what’s possible in AI now depends on how models are combined. There are dozens of capable models, each strong in different places, and prices for comparable quality span an order of magnitude or more.
Ensemble research saw this coming. Mixture-of-agents, generate-and-rank fusion, self-consistency, multi-agent debate, verifier-guided selection — again and again, a group of models did what no single one could. Routers and cascades proved it from the other side: send each query to the model that fits, and pay for the strongest only when you need it. If you’re reading this, some of that work is probably yours.
But nearly all of it ran under tight constraints — the models of the moment, whatever compute a lab could spare, an eval harness built for one paper. Ensembles are expensive by design: every query passes through several models, so testing one on today’s frontier means a real inference bill before you see a single result. Then you have to rebuild the evaluation and find a fair basis for comparison. So the strongest methods have mostly never run on the models where they’d matter most.
Our research grants take that wall down. We cover the compute, build the engineering, and run the evaluation — and we make every result credible to anyone who reads it. You bring the idea. We help you find out how far it goes.
The best quality reachable at each level of cost and latency, measured on current models — and put in the open, so it holds for everyone and not just for us.
Every result is re-run independently and published with its full recipe and its full cost, in tokens, dollars and time. Nothing rests on trust: anyone can verify a state-of-the-art mark by rerunning it. Together, the cohort’s work becomes a co-authored, openly published map — the first verified map of the quality–cost frontier for ensembles and routers on current models — a commons that outlasts the cohort and belongs to the field.
What this call seeks to accomplish
Map the quality–cost frontier for ensembles and routers on current models, in the open. We believe the strongest AI systems will be ensembles of many models, and that the evidence for this is spread across papers that are never measured against each other. Ensembles and routers should be measured against current models, with quality, cost, and latency as key dimensions to test. We re-run every result independently and publish its full recipe and cost in tokens, dollars, and time.
What we provide
- Compute
- Subsidized inference covering the costs of grant runs / up to $X per grantee, so an ensemble can use as many models as the idea needs.
- Infrastructure
- Our toolkit is open source and runs locally, on your keys or ours. Any ensemble or router becomes a single url4 recipe that can be rerun, shared, and published.
- Engineering
- Our engineers work directly with you to extend the toolkit around your method until it expresses the method faithfully.
- Evaluation
- Independent re-runs and full cost accounting, so your result stands on its own for anyone who reads it.
- Stipend
- $2,000 stipend for your time.
Who’s this for
Researchers with published work on ensembles, routing, multi-agent systems, model composition, or closely related areas, and people with a well-formed idea in the space and the rigor to test it. If you’ve built a mixture-of-agents system, a router, a cascade, a verifier, or a fusion method, please apply!
What you’d do
You bring one method or idea, and over 12 weeks you take it from a first run to a published result. Pick a track: bring a published method back to life on today’s models, or test an idea nobody has tried. Either way, the work happens in the open, on a public leaderboard, where every entry is a recipe with its full cost.
The 12 weeks follow the same arc for everyone:
- Weeks 1–2
- Get set up. Reproduce a known ensemble as your first submission, so you’ve proven the pipeline before the real work starts.
- Weeks 3–6
- Run your method. Rebuild it as an open recipe and measure it against current models and benchmarks.
- Weeks 7–10
- Push. Stronger models, better aggregation, a cheaper cascade, or a combination with another grantee’s recipe. This is where most frontier results come from.
- Weeks 11–12
- Write it up. Author review, a short technical note under your name, and the co-authored cohort paper.
Out of scope: anything that needs pretraining or fine-tuning large models, or that can’t run through the toolkit within 12 weeks. The work is zero-shot and composed from existing models: ensembles, routers, cascades, judges. If you’re set on a method the stack doesn’t support yet, tell us. We’d rather extend the toolkit than turn a good idea away.
How credit and IP work
Recipes, results, and code are open source. You keep authorship of your ideas and notes and share authorship of the cohort paper. When a grantee reimplements a published method, the original authors are invited to review the implementation before anything is published. Reimplementations are attributed to the original authors first and to the grantee second, and every entry links to the source paper.
How to apply
Grants are open to researchers with published work on ensembles, routing, multi-agent systems, model composition, or closely related areas. They are also open to people with a well-formed idea in this space and the rigor to test it. Nominations of others are highly encouraged and will be given additional consideration in applications.
We will put more weight on work you’ve shipped in this area (a paper, preprint, workshop paper, or competition write-up with code), and a specific proposal that names the method, the benchmark, and the models.
Submitting confirms you agree with our privacy policy.
Ensembles & Routing (Winter 2026) FAQ
Do I need my own compute budget or API keys?
No. Subsidized inference covers the costs of grant runs / up to $X per grantee. The toolkit runs on our keys or yours — whichever you prefer — so the cost of running an ensemble across many models is not your problem to solve.
What if my reimplementation doesn’t reproduce, or the method doesn’t transfer?
That’s a finding, and we will publish it with the same care as a new state of the art. A method that fails to transfer to current models tells the field something important. Nobody is penalized for an honest negative result, as the field needs a verified map of what has been tried, and the map needs to show which attempts lead to dead ends.
What models and benchmarks will we work against?
Current frontier models, measured along the frontier of quality against cost and latency. The launch benchmarks are launch benchmarks, and every result is published with its full recipe and its full cost in tokens, dollars, and time so anyone can rerun it.
Hear about future calls
Calls are announced here and through our community. Leave your details, and we’ll be in touch when one matches your work.
Submitting confirms you agree with our privacy policy.