What am I actually looking at?
Two 3D scenes built from the same prompt by two different language models. The models do the spatial reasoning — what goes where, how big, facing which way — and a single in-house generation model turns each object into geometry, so the only variable between the two sides is the model that planned the scene.
Both panels are live 3D. Drag to orbit, scroll to zoom, click to step inside and walk around.
Why aren't the models named before I vote?
Because knowing which is which is the fastest way to stop looking. Naming a model next to its render answers the question the page is asking, so the labels arrive after your vote and not before it.
How are the two sides paired?
At random from the models that have a build for that prompt. Neither side is favoured by position — which model lands on the left is a coin flip on every round.
How do the ratings move?
Every vote is a head-to-head result, and the winner takes rating from the loser. A model that beats one rated well above it gains more than it would for beating a peer, and skipping a round moves nothing.
Ratings only mean something in bulk. A model near the top after a handful of votes is noise; the board is worth reading once a pairing has been seen a few hundred times.
What does SKIP do?
Ends the round without a result. Use it when neither build answers the prompt, or when you genuinely cannot separate them — a coin-flip vote is worse than no vote, because it is indistinguishable from a real judgement.
Can I submit my own prompt?
Yes — the field under the vote takes one, and it goes into the pool that future rounds are drawn from. Prompts that describe a place tend to produce a better comparison than prompts that describe an object.
Why does a build look broken?
Sometimes it is: a model can put a staircase through a wall or float a roof off its walls entirely, and that is exactly the failure the benchmark exists to catch. Vote on what you see. A scene that is wrong in an interesting way is still a result.
Who made this?
Starshot Labs. The generation model, the harness the models are scored in, and this site are all ours.
Still wondering?Go to the arena