- Two 3D scenes built from the same prompt by two different language models. The models do the spatial reasoning — what goes where, how big, facing which way — and a single in-house generation model turns each object into geometry, so the only variable between the two sides is the model that planned the scene.
- Both panels are live 3D. Drag to orbit, scroll to zoom, click to step inside and walk around.
- Because knowing which is which is the fastest way to stop looking. Naming a model next to its render answers the question the page is asking, so the labels arrive after your vote and not before it.
- At random from the models that have a build for that prompt. Neither side is favoured by position — which model lands on the left is a coin flip on every round.
- Every vote is a head-to-head result, and the winner takes rating from the loser. A model that beats one rated well above it gains more than it would for beating a peer, and skipping a round moves nothing.
- Ratings only mean something in bulk. A model near the top after a handful of votes is noise; the board is worth reading once a pairing has been seen a few hundred times.
- Ends the round without a result. Use it when neither build answers the prompt, or when you genuinely cannot separate them — a coin-flip vote is worse than no vote, because it is indistinguishable from a real judgement.
- Yes — the field under the vote takes one, and it goes into the pool that future rounds are drawn from. Prompts that describe a place tend to produce a better comparison than prompts that describe an object.
- Sometimes it is: a model can put a staircase through a wall or float a roof off its walls entirely, and that is exactly the failure the benchmark exists to catch. Vote on what you see. A scene that is wrong in an interesting way is still a result.
- Starshot Labs. The generation model, the harness the models are scored in, and this site are all ours.