Before the portal existed, onboarding a new team to our API ecosystem looked like this: a service request, a ticket queue, two or three handoffs between teams, a manual runbook that nobody kept updated, and about two business days of waiting. Not because anyone was slow. Because every step needed a human.
We run a multi-tenant API platform on Azure, with Azure APIM as the front door, Key Vault for secrets, and Terraform underneath. The traffic scale forced us to treat onboarding like a product, not a favor. Here is what we built and what broke along the way.
What onboarding actually costs without self-service
The hidden cost was never the two days. It was the shape of those two days:
- A developer submits a request form that half the time misses a field, which restarts the clock.
- Someone in my team creates the API subscription, the product assignment in APIM, and the service principal in Azure AD, all by hand.
- Credentials get emailed or dropped in a chat, which is a security conversation I no longer want to have.
- When something breaks at 2 AM, nobody remembers which team owns which subscription.
Every manual step is a queue. Queues are where latency lives. This was true of our request pipeline exactly the same way it is true of API traffic, which is a fun mirror I keep noticing.
The portal, in concrete terms
We built a self-service onboarding portal backed by a Node.js API. The frontend gives requesting teams a single form. The backend does the actual work through automation:
- Validates the request (tenant id, contact, required scopes, environment).
- Creates the subscription and product assignment in APIM through the management API.
- Provisions credentials into Key Vault, never shown in plaintext after creation.
- Wires the Terraform modules that own the tenant's networking and access, so nothing is a snowflake.
- Registers the tenant in our traffic analytics, so day-one dashboards include max RPS and P99 latency per endpoint.
The important design decision: the portal does not bypass our existing automation. It calls it. Terraform remains the source of truth for infrastructure. The portal is an interface over the same modules we already run in Azure DevOps pipelines. That meant onboarding a tenant and running the deployment pipeline used the same code paths, which removed an entire class of drift between what the portal did and what IaC expected.
What we got wrong first
We automated the happy path and forgot the exceptions. The first version handled the standard request perfectly and failed loudly on anything unusual, like a team that needed a nonstandard scope or an existing tenant adding a second region. Those requests fell back to tickets anyway, and for a month our metrics looked great while real users quietly went around the portal. Fix: every rejection path in the portal now produces a specific, actionable message, and edge cases route to a human with full context attached instead of a blank ticket.
We treated secrets as a delivery problem. Early on the portal returned API keys in the response payload. It worked and it was wrong. Credentials now go straight into Key Vault with an RBAC assignment for the requesting team, and the portal only ever confirms that provisioning succeeded. If a team loses a key, rotation is one click and the old one is revoked automatically. This is the same rotation discipline I wrote about for certificates: never let a secret's lifecycle depend on a human remembering.
We skipped the audit trail. When an incident involved a tenant's credentials, we could not answer "who created this, when, and with what scopes" without digging. Every provisioning action now writes a structured log entry with requester, approver (when required), timestamp, and resulting scope. Boring. Essential.
The multi-tenant part is where it gets interesting
Single-tenant onboarding is a form. Multi-tenant onboarding is policy. The questions that took real design work:
- Isolation boundaries. Tenants share APIM products but get separate subscriptions, separate Key Vault secrets, and scoped RBAC. Blast radius is a first-class design input, not an afterthought.
- Rate limits per tenant. Global rate limits protect the platform; per-tenant limits protect tenants from each other. We set both, and the portal assigns them at onboarding time based on the tier the team requests.
- Environment parity. Dev, staging, and production onboarding go through the same portal flow. The Terraform workspace differs, not the process. A team that can self-serve in dev already knows how prod works.
What changed, measured honestly
Onboarding went from roughly two business days to about three hours, measured from form submission to working credentials. The three hours is mostly pipeline runtime and approvals, not queueing. The number I actually care about is support tickets: onboarding-related tickets dropped to near zero because the failure modes are now either automated or self-explanatory.
The portal also changed my team's job. We stopped being the onboarding bottleneck and started being the platform owners who review policy changes and handle exceptions. That is a much better use of a cloud engineering team.
If you are about to build one
A few things I would tell myself at the start:
- Build the portal on top of your existing automation, not beside it. If the portal and your IaC pipelines can disagree, they will.
- Ship the audit trail in version one. You will need it before you think you will.
- Automate rejection paths, not just success paths. The quality of an onboarding system is measured by what happens to the request that does not fit.
- Never return a secret in an API response. Not once, not for convenience.
Self-service is not about removing humans. It is about moving humans from the critical path to the exception path. That is where judgment belongs.
Top comments (0)