About SceneBench

How well can a language model build in 3D?

LLMs inherently think in terms of text — a one-dimensional ordered sequence of words. How will they fare when asked to construct omnidirectional 3D scenes, a fundamentally different modality?

SceneBench places different LLMs in a harness we engineered in-house here at Starshot Labs, and asks them to build the same scene from the same prompt, while the actual 3D generation is taken care of by our foundational model. Due to the subjective and creative nature of 3D generation, we want your votes to decide which LLMs reign at the top, and which ones feed at the bottom.

You might be thinking: “of course the smarter, larger models will do better!” But some of the results might surprise you.

  1. 01

    Prompt. Both models are given the same prompt, selected across a variety of different scene kinds, from buildings to game levels.

  2. 02

    Build. Two models generate the scene independently. Neither sees the other's work, and both utilize Starshot's in-house 3D generation model.

  3. 03

    Vote. You pick the better build. Ratings update after every vote and feed the public leaderboard.

Scroll

Who we are

Starshot Labs builds foundation models for 3D.

We are a research team working on generative 3D: models that turn language into geometry, layout, and the relationships between objects in a scene.

SceneBench grew out of our own evaluation work. We needed a way to compare how language models reason about space, and head-to-head voting gave us a clearer signal than any rubric we wrote ourselves. We made it public because the same question is open for everyone building in this space.

Leaderboard