Originally published on the EnConvert blog.
I am Krystin, the non-technical co-founder of EnConvert. I do not write the code. What I do is talk to the engineers who are building with it, and one thing comes up in those conversations more consistently than anything else: the ingestion layer breaks before the retrieval layer, and web content is usually where it starts.
Not an engineering deep dive. Just what I have learned from being in enough of those conversations to see the pattern.
What raw extraction actually returns
Point a standard extraction tool at a documentation page or a blog post. The content is in there. So is everything else.
Navigation link text at the top of the output. The same links again in the footer. Sidebar content mixed into the body where the DOM placed it. Escaped HTML characters like & and sitting inline with the prose. CMS image URLs appearing as references that mean nothing outside the original page.
On JavaScript-rendered pages, the situation is worse. A static fetch often returns the page shell with placeholder elements where the content should be. The actual text was injected by a framework after load. What you get is a skeleton.
None of that is ready to pass into a model or a chunker without additional processing.
What clean output looks like
Clean output for an LLM pipeline is structured Markdown.
The page title becomes a top-level heading. Section headings sit at the correct hierarchy. Paragraphs are clean prose. Lists stay as lists. Code blocks keep their content and language identifiers. Tables come through intact.
What is not there: the navigation bar, the cookie banner, the sidebar, the footer links, the escaped characters, the CMS image references, and the repeated boilerplate that surrounded the content on the original page.
The structure matters because it determines how the pipeline behaves downstream. Consistent Markdown means consistent chunking. Consistent chunking means more predictable retrieval.
What you are handling when you build this yourself
The scope is real, and it is worth being honest about it.
JavaScript rendering requires a headless browser rather than a simple HTTP request. Managing that means handling browser instances, timing the page load correctly, and dealing with sites that behave differently in headless environments.
Content extraction requires deciding what is main content and what is chrome. Navigation and sidebars are not always wrapped in semantic tags. You are working with heuristics based on text density, element position, and link-to-text ratios. Those heuristics do not generalise cleanly across different site structures.
Dynamic content loaded on scroll or behind interactions may not be captured at all. CMS platforms add their own markup that bleeds into extracted text if the extraction layer does not account for them. Error states need explicit handling: pages that return 200 but show a login wall, rate limiting, pages that time out during rendering.
From the conversations I have had, the ongoing maintenance is what engineers underestimate most. Not the difficulty of building it. The persistence of keeping it working as sites change.
Using Perceive
Perceive is the EnConvert endpoint built for this. You pass a URL, it returns clean Markdown.
Full browser pass for JavaScript-rendered pages. Navigation and structural chrome stripped before the response comes back. Headings, lists, tables, and code blocks preserved. The output structure is consistent regardless of which site the URL points to, so downstream code does not need to account for source-specific formatting differences.
Free tier: 500 ops/month, no card required. enconvert.com
If you have hit a specific web extraction edge case that broke your pipeline, I would genuinely like to hear about it in the comments. That is the kind of thing that shapes what we build next.
Top comments (0)