If your onboarding flow asks for a URL and promises to learn the user's voice from it, you have a problem you probably haven't named yet: a marketing site is not a voice sample.
We hit this building the Cadencz onboarding step. Paste a website URL, and the system builds a Product Brain: what you sell, who you sell it to, and how you write. Two of those three are easy. The third one lies to you if you're not careful.
𝗪𝗵𝗮𝘁 𝗮 𝘀𝗶𝘁𝗲 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝘁𝗲𝗹𝗹𝘀 𝘆𝗼𝘂
A homepage and a pricing page are excellent sources for product facts. Scrape them and you get real answers to real questions: What does this do. Who is it for. What claims is the team willing to stand behind in public. Those are extractable, checkable, and stable over time.
𝗪𝗵𝗮𝘁 𝗮 𝘀𝗶𝘁𝗲 𝗰𝗮𝗻𝗻𝗼𝘁 𝘁𝗲𝗹𝗹 𝘆𝗼𝘂
Landing-page copy is not written by the person who will post on LinkedIn next Tuesday. It's usually written by a founder in pitch mode, a contractor, or a copywriter optimizing for conversion, not for how the founder actually talks. The sentence length on a pricing page tells you nothing about the sentence length the founder uses when they're annoyed about a bug on X. The emoji count on a features section tells you nothing about whether this person uses emoji at all in a real post.
We tried treating both as the same kind of signal early on. It didn't work. The model would infer a formal, hedge-everything tone from the homepage and apply it everywhere, including to a founder who writes like they're texting a friend.
𝗧𝗵𝗲 𝗳𝗶𝘅: 𝘀𝗽𝗹𝗶𝘁 𝗲𝘅𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘁𝗼 𝘁𝘄𝗼 𝗰𝗮𝘁𝗲𝗴𝗼𝗿𝗶𝗲𝘀
Product facts get a source id, so every claim can be cited, not paraphrased from memory. Style signals get flagged as a starting guess that the user has to confirm, not something the system assumes is settled.
{
"product_fact": {
"id": "de7d26b4-865c-44a5-9ef0-137830dcf39c",
"claim": "Batch a month of social content in one 3-hour session using five passes.",
"source_url": "https://cadencz.com/blog/social-media-batching-founders",
"confidence": "extracted"
},
"style_constraints": {
"max_sentence_words": 22,
"emoji_use": "none",
"forbidden_terms": ["revolutionary", "game-changing", "guaranteed"],
"status": "needs_user_confirmation"
}
}
Notice the two objects don't share a status field by accident. A product fact is either sourced or it isn't. A style constraint is either confirmed by a human or it's a guess dressed up as a fact.
𝗧𝗵𝗲 𝘁𝗮𝗸𝗲𝗮𝘄𝗮𝘆 𝘁𝗵𝗮𝘁 𝗮𝗽𝗽𝗹𝗶𝗲𝘀 𝗽𝗮𝘀𝘁 𝗼𝘂𝗿 𝗽𝗿𝗼𝗱𝘂𝗰𝘁
Any pipeline that extracts facts from a document and later generates text from them needs to store the citation with the fact, not just the fact. Without the source id, your generation step has no way to distinguish "I read this on their pricing page" from "I'm pretty sure this is the kind of thing they'd say." Those feel identical at generation time. They are not identical in cost when a claim turns out to be wrong.
If you're building anything that reads a URL and outputs prose in someone's voice, treat facts and style as two different extraction problems with two different confidence rules. One can be scraped. The other has to be asked.
Top comments (0)