I Tried to Make AI Build a Complete Website From One Sentence. Here’s What Broke
The demo worked in minutes. Making it reliable enough for real users took much longer.
The first version looked almost too easy.
Give the AI one sentence:
“I run a roofing company in Austin.”
A few seconds later, you have a homepage.
Nice headline.
Services section.
CTA.
Decent copy.
At that point, it is very tempting to think:
“Okay. The hard part is done.”
It wasn’t.
The hard part started when I stopped asking AI to generate a page and started asking it to generate an entire website that a real business could actually use.
That changed everything.
A homepage is easy
LLMs are very good at producing something that looks convincing once.
Give them a prompt and they can generate:
- a hero section
- service descriptions
- an About section
- testimonials
- FAQs
- calls to action For a demo, this is enough. For a real website, it is nowhere near enough. A website needs consistency across multiple pages. The navigation has to make sense. The copy cannot contradict itself. The layout needs to survive mobile screens. Buttons need to point somewhere real. Sections cannot randomly disappear. The business name cannot change halfway through the site. And the AI should probably not invent a phone number. That last one happened more often than I expected. Problem #1: The AI had no idea what the website actually was My first approach was basically: Here is the business.
Generate the website.
The output looked fine.
Until I generated another page.
Then another.
Suddenly the website had three different personalities.
The homepage sounded premium.
The services page sounded like a discount contractor.
The About page sounded like a Fortune 500 company.
The model was generating pages independently.
That was the mistake.
A website is not a collection of unrelated prompts.
It needs a shared source of truth.
So instead of generating pages immediately, I started generating a structured representation of the business first.
Something closer to:
{
"business": "Austin Roofing Co",
"industry": "Roofing",
"location": "Austin, Texas",
"audience": "Homeowners",
"tone": "Professional and local",
"services": [
"Roof repair",
"Roof replacement",
"Storm damage repair"
],
"primaryGoal": "Request a quote"
}
Every page then uses the same context.
This sounds obvious now.
It was not obvious in version one.
Problem #2: AI loves inventing things
Ask an AI to create a business website and it wants to be helpful.
Sometimes a little too helpful.
It may invent:
- customer testimonials
- awards
- years of experience
- certifications
- addresses
- statistics
- phone numbers A sentence like: “Trusted by 2,000 homeowners since 1998”
looks fantastic on a landing page.
There is only one problem.
Nobody told the AI that.
For a fictional demo, this is harmless.
For a real business, it is a serious problem.
So I had to start treating certain information differently.
There are facts the model can rewrite.
There are facts it can infer carefully.
And there are facts it should never create.
That distinction became one of the most important parts of the system.
Problem #3: Good-looking HTML can still be broken
This one surprised me.
A generated page could look great at 1440px.
Then I opened it on mobile.
Chaos.
Headings wrapping into five lines.
Buttons leaving their containers.
Images becoming enormous.
Two-column layouts refusing to become one column.
Text becoming unreadable over backgrounds.
The AI had created something that was technically valid and visually broken.
This is where I realized something important:
Generating the site and verifying the site are two completely different problems.
So I added a separate visual QA step.
Instead of trusting the generated result, the system has to inspect what actually rendered.
Things like:
Does content overflow?
Are sections missing?
Is text readable?
Does the navigation work?
Are elements overlapping?
Does the page still make sense on mobile?
The AI can say the website is correct.
The browser is the final judge.
Problem #4: “Regenerate” is not a solution
A lot of AI products solve bad output with one button:
Generate again.
That works surprisingly well in demos.
It is also expensive.
If generation fails 10% of the time, simply running the whole thing again means:
- more model calls
- more latency
- more compute
- more money
- another chance for something else to break So instead of regenerating everything, I started trying to identify the exact thing that failed. Bad section? Fix the section. Broken page? Rebuild the page. Bad content? Regenerate the content. The goal became: detect -> isolate -> repair
instead of:
fail -> regenerate everything -> hope
That one change made the system feel much less like a slot machine.
Problem #5: The AI kept being creative when I needed it to be boring
This was probably the funniest problem.
Creativity sounds great when building with AI.
Until you are generating production UI.
Sometimes you do not want creativity.
You want:
Header
Hero
Services
About
CTA
Footer
That is it.
No floating glassmorphism card.
No randomly rotated image.
No giant quote taking up the whole viewport.
No mysterious “Our Philosophy” section for a plumbing company.
The model kept trying to impress me.
I kept trying to make it behave.
Eventually I realized the system needs both freedom and constraints.
Too many constraints and every site looks identical.
Too much freedom and every generation becomes unpredictable.
Finding that middle ground was much harder than writing the prompt.
The architecture slowly changed
The original idea was basically:
Business description
↓
AI
↓
Website
The real system started looking more like:
Business description
↓
Business understanding
↓
Website structure
↓
Page planning
↓
Content generation
↓
Design generation
↓
Rendering
↓
Visual validation
↓
Repair
↓
Final website
The interesting part is that AI is only one piece of that pipeline.
Most of the work is actually about controlling it.
The biggest lesson
Before building this, I thought the main challenge would be:
“Can AI generate a good website?”
It can.
That is not the interesting question anymore.
The real question is:
Can AI generate a good website repeatedly, predictably, and without a human fixing it afterward?
That is a completely different engineering problem.
The first impressive result took very little time.
The reliable system took much longer.
And I suspect this is true for a lot of AI products right now.
The demo is easy.
The last 20% is the product.
I’m still working on this problem while building Instantsite, where the goal is to turn a short business description into a complete editable website.
But the thing I find most interesting is not the website generation itself.
It is everything required to make unpredictable AI output behave like predictable software.
That part is much harder than the demo makes it look.
One question
If you’re building with generative AI, what was the first thing that worked perfectly in the demo and completely fell apart in production?



Top comments (0)