Comment by antelocnova
6 hours ago
Hi there, thanks!
I'm taking a look at this paper right now, didn't came across it before.
I think it's really good, and also an approach I tried.
In my case, it was working pretty well in terms of geometry and correctness, but not in terms of creativity: too many constraints and an agentic LLM would end up building very similar buildings all the time. Too few ones, and it would drift into unbuildable models.
In the end, I think the current iterative process gives it a good balance: let it try things first and be creative, parts will be placed incorrectly at the beginning, but it will fix their placements, colours and so on bit by bit, until the end result is (hopefully!) correct.
And regarding Opus, according to my experience, I'd say definitely models quality is way better when produced by Opus 5 and even better by Opus 5.5
You might also be interested in this research at CMU:
https://avalovelace1.github.io/BrickGPT/
We introduce BrickGPT, the first approach for generating physically stable toy brick models from text prompts. To achieve this, we construct a large-scale, physically stable dataset of brick designs, along with their associated captions, and train an autoregressive large language model to predict the next brick to add via next-token prediction. To improve the stability of the resulting designs, we employ an efficient validity check and physics-aware rollback during autoregressive inference, which prunes infeasible token predictions using physics laws and assembly constraints. Our experiments show that BrickGPT produces stable, diverse, and aesthetically pleasing brick designs that align closely with the input text prompts. We also develop a text-based brick texturing method to generate colored and textured designs. We show that our designs can be assembled manually by humans and automatically by robotic arms. We also release our new dataset, StableText2Brick, containing over 47,000 brick structures of over 28,000 unique 3D objects accompanied by detailed captions, along with our code and models.
I remember coming across this research paper, and I thought it was really impressive. I clearly remember the bricks guitar, built by robots =)
Currently, Nova doesn't handle any physical constraints regarding stability, just collision.
I think physics handling is a very desirable feature, maybe an agent using a physics engine directly or maybe via MCP that agent could get feedback about a model's stability, fix problems, back to more feedback, etc.
This would be similar to the current way finds and iteratively corrects defects by rendering images from the model, different perspectives, and directly examining them.