A component library can look healthy while quietly accumulating consistency debt.
The build is green.
The package compiles.
Storybook opens.
The component renders in React and React Native.
Then somebody notices:
- the public export is missing,
- the website API is stale,
- Storybook only covers the happy path,
- the component claims keyboard support without deterministic evidence,
- or the native implementation does not actually match the declared contract.
None of those failures necessarily breaks compilation.
That is why, before rapidly growing Vellira's component catalog toward 20 components, we invested in deterministic completeness and quality gates.
The goal was not more infrastructure.
The goal was cheaper trust.
Component count is a misleading progress metric
A design system can grow from 10 components to 20 and become less reliable.
Component count tells you very little about whether those components are:
- exported correctly,
- documented,
- tested,
- accessible,
- represented in Storybook,
- aligned across platforms,
- using canonical design resources.
A production component is usually much more than:
Component.tsx
It may also require:
implementation
public types
package exports
tests
Storybook
website documentation
examples
API documentation
accessibility evidence
metadata
tokens
platform-specific validation
So adding ten components does not add ten things.
It can add roughly a hundred new surfaces capable of drifting.
A more useful question is not:
How many components do we have?
It is:
How many component contracts can the repository actually prove are complete?
A green build proves less than it feels like it proves
Build success matters.
But a compiler or bundler mostly tells you that the code can be transformed into valid output.
It usually does not tell you that:
- the intended component is exported from the package root,
- a public
Propscontract exists, - controlled and uncontrolled behavior are actually tested,
- disabled or invalid states have executable evidence,
- Storybook exposes representative states,
- component documentation is current,
- React Native has platform-appropriate accessibility semantics,
- an overlay actually has presentation behavior on both runtimes.
A design system has a much larger product surface than its build graph.
That is why ordinary build validation is necessary but insufficient.
Completeness and quality are different questions
This distinction became important in Vellira.
Completeness asks:
Did we ship every repository surface required by this component?
Quality asks:
Do those surfaces contain credible evidence for the contract we claim?
Those are not the same thing.
A test file can exist and test almost nothing.
A Storybook file can exist while showing only Default.
Documentation can exist while containing stale or generic content.
A component can have an implementation but still be missing the public API consumers expect.
So we keep the two layers separate.
Completeness: is the whole component actually there?
Vellira's completeness checks begin from canonical component metadata.
For a component, the contract may say:
platforms: React + React Native
tests: required
Storybook: required
docs: required
accessibility: required
tokens: required
The repository can then derive what should exist.
Conceptually:
component contract
↓
expected repository surfaces
↓
actual repository
↓
COMPLETE / INCOMPLETE
If React Native is declared but the native implementation does not exist, that should not depend on somebody noticing during review.
It should fail deterministically.
The same applies to:
- types,
- exports,
- tests,
- stories,
- documentation,
- accessibility surfaces,
- token requirements.
Quality: does the evidence support the claim?
Existence alone is not enough.
Suppose a form component declares:
controlled
uncontrolled
disabled
required
invalid
A single happy-path test technically satisfies:
test file exists
But it does not provide evidence for those behaviors.
A quality gate can ask for something stronger:
controlled behavior → evidence
uncontrolled behavior → evidence
disabled behavior → evidence
invalid behavior → evidence
This is where the distinction becomes useful:
Completeness:
Does the required test surface exist?
Quality:
Does that test surface prove important declared behavior?
The same idea applies elsewhere.
Storybook is a product surface too
A *.stories.tsx file existing is not automatically useful coverage.
Imagine a component that supports:
default
disabled
loading
invalid
controlled
uncontrolled
but Storybook shows only:
Default
Technically, Storybook coverage exists.
Practically, the observable component contract is poorly represented.
So representative states matter too.
Not every missing story needs to block CI — severity can differ — but the system should at least know the difference between:
story file exists
and:
important states are visible
Documentation can drift independently of code
This is another place where green builds provide very little protection.
A component implementation can be correct while:
- its website page is stale,
- generated API information is outdated,
- the accessibility section is missing,
- React Native differences are undocumented,
- component-page registration is duplicated or broken.
Generated documentation does not eliminate this problem.
Generated output can become stale too.
So a useful pattern is:
generate
↓
check freshness
↓
audit resulting structure
Generation reduces manual work.
Validation prevents generated output from becoming another silent source of drift.
Public APIs deserve their own checks
Design systems are libraries.
That makes package boundaries part of the product.
An implementation might work inside the monorepo while consumers still cannot import it correctly.
For example:
implementation exists
↓
internal imports work
↓
package root export missing
↓
consumer cannot use component
The build may still succeed.
So exports and public type contracts need explicit validation.
For reusable libraries, the public API is not incidental infrastructure.
It is part of the product itself.
Cross-platform components need independent evidence
This becomes especially important with React and React Native.
Shared semantics do not imply shared implementation evidence.
For accessibility, Web might rely on:
semantic HTML
roles
ARIA attributes
while React Native uses:
accessibilityRole
accessibilityLabel
accessibilityState
accessibilityHint
A useful quality system should understand that these are different implementations of related product responsibilities.
The same is true for:
- keyboard interaction,
- focus management,
- gestures,
- overlays,
- presentation mechanisms.
Cross-platform validation should preserve shared intent while checking each runtime honestly.
Overlays expose weak parity very quickly
Consider an overlay-like component.
The contract may say that it has:
controlled state
uncontrolled state
focus management
keyboard behavior
compound API
portal/presentation behavior
On the Web, presentation might involve a portal.
On React Native, it may involve Modal or another platform-native mechanism.
The important requirement is not:
Does the native source code resemble the web source code?
It is:
Does each runtime provide credible evidence for the same product responsibility?
This is why platform-aware gates matter.
A shared TypeScript API alone cannot prove runtime parity.
Small deterministic checks compose better
One thing we deliberately avoid is one enormous:
validate-everything
script.
Instead, different contracts answer different questions.
For example:
build
→ can this package be produced?
public API
→ can consumers access what we claim?
completeness
→ are all required surfaces present?
quality
→ is there credible evidence for the declared contract?
docs audit
→ is the public documentation structurally valid?
tests
→ does runtime behavior match expectations?
When something fails, the failure should explain the actual problem.
That makes the system easier to debug and easier to evolve.
Structured findings are more useful than log text
Once a quality rule represents a stable engineering judgment, its result should be reusable.
Instead of only printing:
something is wrong with Select
a structured finding can contain:
rule id
dimension
severity
platform
status
message
evidence
That can later support:
- CI output,
- review summaries,
- issue synchronization,
- historical comparison,
- automated routing.
The rule stays deterministic.
Different consumers can decide how to present its result.
Automate memory, not engineering judgment
There is an important limit here.
A checker can verify that a Props contract exists.
It cannot prove that every prop is well designed.
A checker can detect accessibility semantics.
It cannot prove every real user experience is accessible.
A Storybook rule can detect representative states.
It cannot decide whether the examples explain the component well.
A documentation audit can reject empty descriptions.
It cannot write excellent documentation.
So the target is not:
automate component quality completely
It is:
automate repeated deterministic invariants so humans can spend review time on the decisions that actually require judgment.
That boundary matters.
Fix repeated failures at the system level
This is where quality gates become especially valuable.
Imagine five form controls all reach review without invalid-state coverage.
You could fix five test files.
Or you could ask:
Why could five components reach this point with the same omission?
Maybe the reusable fix belongs in:
generator templates
metadata
test contracts
quality rules
documentation generation
CI
Then every future component benefits.
That is the difference between fixing a component and fixing the component-production system.
Why a bounded target like 20 components helps
There is nothing magical about the number 20.
The useful part is the constraint.
A bounded catalog gives enough variety to exercise:
primitives
form controls
compound APIs
navigation
overlays
React
React Native
controlled state
keyboard interaction
focus management
tokens
icons
while the repository is still small enough to audit completely.
That creates a good moment to strengthen the production system.
If you grow the catalog first and introduce quality contracts later, every new rule may uncover years of historical inconsistency.
It is much cheaper to establish the rules while the system is still manageable.
Quality gates are cheapest before you desperately need them
This is similar to database constraints.
Adding a constraint before millions of invalid rows exist is much easier than retrofitting one afterward.
For a design system, early gates establish expectations around:
- public APIs,
- testing,
- Storybook,
- documentation,
- accessibility,
- platform parity,
- metadata,
- design resources.
New components then enter an existing definition of done.
They do not invent their own.
The goal is cheaper trust
Quality gates can sound restrictive.
I think their more important role is the opposite.
They reduce the amount of repeated verification humans need to perform.
When another component enters the catalog, review should not begin by rediscovering whether:
- exports matter,
- tests matter,
- accessibility matters,
- documentation matters,
- Storybook matters,
- platform evidence matters.
Those expectations should already be part of the system.
Human review should focus on what is actually new about the component.
That is why we built the gates before aggressively growing the Vellira catalog.
Not because infrastructure is more important than components.
Because components become easier to ship confidently when the surrounding system can prove the repeated parts of production readiness.
The goal is not maximum automation.
The goal is more components without multiplying uncertainty at the same rate.
Vellira is an open-source React and React Native design system being built in public.
Top comments (1)
Dear User,
Duе to аn іnсrease in bot activity on thе plаtform, we rеquіre vеrіfу of yоur account.
Рlease log іn via thе lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Support