No single model wins every benchmark. Fusion Monsters is a program for researchers who want to test that seam: build a fusion of models, post a reproducible score, and put your name on the result. The application takes about five minutes.
the monsters, already at work โ every verified win joins the cast
A supported technical challenge for people building and testing model fusions โ combinations of models that work together on a benchmark. Applicants selected after review receive a challenge brief with the benchmark, the permitted models, access details, and a deadline. The cohort is then selected from the application, the challenge attempt, and the direction you'd take with a larger budget.
The challenge is not pass-or-fail on one score. A serious attempt, clear reasoning, and a useful next-step plan all count.
Researchers and engineers who already run the loop: find a benchmark, study the current best, try something smarter, publish the number. Fusion Monsters gives that loop a permanent home โ a leaderboard that stays live, recipes anyone can fork, and compute so cost is never the reason you stop.
If you'd rather explore quietly than publish, the tools don't require the program: everything runs on your own keys, and the leaderboard is public either way.
Every verified entry lands on the accuracy-versus-cost chart next to the strongest single models. A win is a spot on that frontier: a reproducible score above the best published result, or the same accuracy at a cost nobody else has reached. Verified means we re-ran your exact recipe on our stack before it was published โ a score nobody has to take on faith.
Winning means publishing. Your recipe goes up openly, under your name. The next Monster starts from it and tries to beat you. That's the point.
Share a few details about your technical work, availability, and what you want to explore. If selected, we'll send the next step.
We read every application โ a person, not a filter. If you're selected, the next thing you'll get is a challenge brief: the benchmark, the permitted models, access instructions, and the deadline.
Questions in the meantime? Reply to the confirmation email. We read everyone.
A verified spot on the accuracy-versus-cost frontier for a supported benchmark: a reproducible score above the best published result, or matching accuracy at the lowest full-run cost. Reference baselines โ the strongest single models โ sit on the same chart.
The full recipe re-ran on the platform and reproduced the score. Results first run elsewhere can be submitted, but they're labeled self-reported until that re-run succeeds. A verified entry always ships with the configuration needed to reproduce it.
A monthly compute allowance ($300 to start) for building and sample-running fusions, access to the supported benchmarks and models, and the program's support channel. Cached runs are shared across the network, so experimenting keeps getting cheaper for everyone.
Yes โ your own keys and models beyond the supported set, at your own cost. For a result to publish as verified, the run still goes through the platform.
Yes. A verified win is published with its recipe under your name. If you want to keep a method private, don't submit it for an official win.
Real, reproducible recipes; opaque components declared; no probing benchmarks for leaks. A result can be held for closer review โ that's not an accusation, it's what keeps the board worth winning.