About the benchmark

ABOUT

Spatial reasoning, made visible

How well can a language model build in 3D?

Language models think through ordered sequences of text. SceneBench asks them to reason in a different medium: an omnidirectional world where scale, placement, orientation, and relationships all matter at once.

Each model orchestrates the same Starshot Labs 3D generation system. That keeps the rendering engine constant and turns the comparison into a direct test of planning and spatial judgment.

The result is evaluated by people, not a rigid rubric. You explore both scenes, choose the stronger build, and shape a public Elo ranking over time.

  1. 01

    Prompt

    Both language models receive the same scene prompt.

  2. 02

    Build

    Each model independently plans the scene while Starshot's 3D foundation model generates the objects.

  3. 03

    Vote

    The public compares the completed scenes without knowing which model built either side.

Who we are

Starshot Labs builds foundation models for 3D.

We are a research team working on generative 3D: models that turn language into geometry, layout, and the relationships between objects in a scene.

SceneBench grew out of our own evaluation work. We needed a way to compare how language models reason about space, and head-to-head voting gave a clearer signal than any rubric we wrote ourselves. We made it public because the same question is open for everyone building in this space.

Get started

Two ways in.