One unfinished sketch. One prompt. Four models asked to build an entire 3D shopping mall: Gemini 3.7 Flash, Grok 4.6, Claude Opus 5, and Kimi K3.
The comparison works because the input is shared and the differences are visible. One model invents space more aggressively, another focuses on materials, and another feels closer to simply finishing the sketch.
But one visual vote is not a model leaderboard. There are no repeated samples, blind judging, or task suite here. It answers a narrow question: for this sketch and this prompt, which result looks best to you?
Source: https://x.com/CodeByPoonam/status/2089687863390879901
What to verify before using this
- One visual comparison is not a model leaderboard.
- The source does not establish identical random seeds or post-processing.
Sources
The source-side engagement is only a discovery signal. It is not this article's performance, and no product or model was independently benchmarked for this post.
Top comments (0)