DEV Community

Cover image for Architecture on the Chopping Block: The OIDC Dependency We Didn’t Know We Had
Shamil
Shamil

Posted on Originally published at habr.com AI-assisted

Architecture on the Chopping Block: The OIDC Dependency We Didn’t Know We Had

Before joining Cloud.ru, I worked at a company that built and deployed enterprise IT solutions for large customers.

Its main product was a packaged platform: a corporate CRM with a much broader set of modules. It included modules for working with data and reference data, documents, reports, business processes, maps, access rights, and integrations with customer systems.

The platform was not a monolith. It was a set of microservices that we deployed using a traditional on-premises delivery model: the customer provided several servers, we deployed the services on them, and the platform then continued to operate almost autonomously.

Most deployments followed a similar pattern. In one project, however, the customer asked us to deploy the platform in their cloud and integrate it into their existing infrastructure. This became a stress test for us: some parts of the platform adapted without any problems, but authentication ran into a requirement to replace our usual OIDC flow with Kerberos.

In this article, I explain why the cloud deployment exposed an old dependency on OIDC, why replacing it with Kerberos required us to revisit the architecture, and where we ultimately drew the boundary between access management and business logic.

We Had OIDC, but the Customer Wanted Kerberos

By the time this project began, the platform was already in use by several customers. Its foundation remained the same, while only some modules and workflows were changed for each specific deployment.

For one customer, we might significantly customize the mapping features; for another, connect document modules; for a third, integrate a custom role model for data access.

Most of the platform was usually ready, and we only had to adapt it to the requirements of a particular customer.

We were used to keeping almost all of the logic inside our platform and rarely had to rely on capabilities provided by the customer's infrastructure.

This project was different. In addition to deploying the platform in the customer's cloud, we had to give up some of our own functionality and use the customer's mechanisms instead. One of those requirements was replacing our familiar OIDC setup with Kerberos.

Kerberos was mandatory. Users were already working in a corporate domain environment: they signed in to their computers with their own accounts and were not supposed to enter their passwords again in every system. This approach also reduced the extent to which security depended on human behavior.

We misjudged the scope. We thought we were changing only the login mechanism, but we ended up touching the architecture. The task was considered small for too long, so active development started late.

We therefore needed a solution that could be integrated into the existing code with minimal effort, would not require a large-scale rewrite of the services, and would preserve a consistent approach across the platform.

Why a Simple Replacement Did Not Work

If the platform had been a monolith, the task would most likely have been completed much faster because the changes would have affected far less code.

But ours was a microservice platform that had to be integrated with the customer's external IAM environment. Incoming requests were relatively straightforward; propagating user context between services was much harder.

Before the Kerberos integration, the simplified flow looked like this:

%%{init: {"theme":"base","themeCSS":"* { font-family: Arial, sans-serif !important; } text { font-weight: 600 !important; }","themeVariables":{"background":"#222222","fontFamily":"Arial, sans-serif","fontSize":"25px","textColor":"#F7F7F7","lineColor":"#8A8A8A","edgeLabelBackground":"#222222","primaryColor":"#2A2A2A","primaryTextColor":"#F7F7F7","primaryBorderColor":"#2BD879","secondaryColor":"#2BD879","secondaryTextColor":"#101510","secondaryBorderColor":"#65EFA3","tertiaryColor":"#545454","tertiaryTextColor":"#F7F7F7","tertiaryBorderColor":"#8A8A8A","clusterBkg":"#262626","clusterBorder":"#4A4A4A","mainBkg":"#2A2A2A","nodeBorder":"#2BD879","nodeTextColor":"#F7F7F7","labelTextColor":"#F7F7F7","titleColor":"#F7F7F7"},"flowchart":{"htmlLabels":false}}}%%
flowchart LR

classDef service fill:#2A2A2A,stroke:#2BD879,stroke-width:2px,color:#F7F7F7;
classDef accent fill:#2BD879,stroke:#65EFA3,stroke-width:2px,color:#101510;
classDef muted fill:#545454,stroke:#8A8A8A,stroke-width:1px,color:#F7F7F7;

subgraph DIAGRAM[" "]
direction LR

A["Client"]:::accent
B["Orders Microservice"]:::service
C["File Storage Microservice"]:::service

A -->|User token| B
B -->|The same token| C

end

style DIAGRAM fill:#222222,stroke:#222222,stroke-width:1px,color:#222222;

In this flow, the Orders Microservice called the File Storage Microservice not simply under its own identity, but in the context of the original user.

This is not a Client Credentials flow, in which services communicate under their own technical identities and the downstream service does not need information about the original user. In our model, the downstream service had to know on whose behalf the action was being performed.

With OIDC, this problem was barely noticeable. The access token effectively acted as a portable artifact: a service received it with an incoming request and forwarded it when necessary, while the next service reconstructed the user context from the token. This was not an ideal approach from an OAuth 2.0/OIDC perspective, but the platform had historically been built around that assumption.

That assumption no longer held with Kerberos. We could not simply replace the access token from the old OIDC flow with a Kerberos service ticket everywhere and consider the task complete.

The ticket is intended only for the service to which the client presents it. It cannot be forwarded to the next microservice like a bearer token. That requires a separate delegation flow.

The task was therefore no longer just a matter of changing a request header. We had to build a mechanism that could securely propagate user context between services.

Delegation became the workable option for us. In simplified terms, if the Orders Microservice needed to call the File Storage Microservice on behalf of the original user, it first obtained a Kerberos service ticket for the File Storage Microservice. The call was then made using that ticket.

%%{init: {"theme":"base","themeCSS":"* { font-family: Arial, sans-serif !important; } text { font-weight: 600 !important; }","themeVariables":{"background":"#222222","fontFamily":"Arial, sans-serif","fontSize":"25px","textColor":"#F7F7F7","lineColor":"#8A8A8A","edgeLabelBackground":"#222222","primaryColor":"#2A2A2A","primaryTextColor":"#F7F7F7","primaryBorderColor":"#2BD879","secondaryColor":"#2BD879","secondaryTextColor":"#101510","secondaryBorderColor":"#65EFA3","tertiaryColor":"#545454","tertiaryTextColor":"#F7F7F7","tertiaryBorderColor":"#8A8A8A","clusterBkg":"#262626","clusterBorder":"#4A4A4A","mainBkg":"#2A2A2A","nodeBorder":"#2BD879","nodeTextColor":"#F7F7F7","labelTextColor":"#F7F7F7","titleColor":"#F7F7F7"},"flowchart":{"htmlLabels":false}}}%%
flowchart LR

classDef service fill:#2A2A2A,stroke:#2BD879,stroke-width:2px,color:#F7F7F7;
classDef accent fill:#2BD879,stroke:#65EFA3,stroke-width:2px,color:#101510;
classDef muted fill:#545454,stroke:#8A8A8A,stroke-width:1px,color:#F7F7F7;

subgraph DIAGRAM[" "]
direction LR

A["Client"]:::accent
B["Orders Microservice"]:::service
C["File Storage Microservice"]:::service

A -->|Original Kerberos user| B
B -->|Delegated Kerberos service ticket for the original user| C

end

style DIAGRAM fill:#222222,stroke:#222222,stroke-width:1px,color:#222222;

How Deeply OIDC Had Reached into the Architecture

At first, we believed OIDC was concentrated in only a few places: authentication filters, current-user resolution, and a few parts of inter-service communication.

In practice, the dependency ran much deeper. The authentication layer was implemented differently across microservices. In some services it only verified the user; in others it already contained application-level checks and custom workflows. Services also obtained current-user data in different ways. Some application code depended directly on data created by the authentication layer, while other parts already knew about OIDC and JWT details.

Some services assembled the user object directly from the authentication context. Inter-service calls depended on the familiar token-based model, and incoming requests from neighboring services were handled inconsistently.

For example, some services extracted individual fields from the token and converted the login to their own format. Others independently decided which headers to forward. In a few places, a custom internal contract had already emerged between authentication and the application code.

Once we had mapped all of this, it became clear that replacing one protocol with another would not be enough. We first had to separate application logic from the authentication mechanism. Otherwise, we would merely replace a dependency on OIDC with a dependency on Kerberos.

A Typical Example of the Same Problem

I tried to describe this mess with a diagram, but it quickly became unreadable. So I will show the same problem using a simpler example: what we expected to find and what we actually found.

Imagine an application that, like many others, uses caching for optimization. There are no common rules for caching, so everyone implements it however they find convenient.

Time passes. You analyze the project and realize that a huge number of engineering hours go not only into fixing cache-related bugs, but also into continually maintaining the surrounding functionality. You finally decide to bring the situation under control. But you are the project manager and do not know exactly what is happening inside the code. You think at a more abstract level and look for a solution that can fix the problem with minimal effort.

The most obvious solution is to move to Managed Redis. Cloud.ru, for example, provides a ready-to-use managed service for this scenario: most infrastructure tasks are already handled, so from the outside it appears that all you need to do is replace the cache implementation.

You are convinced that the code looks like this:

%%{init: {"theme":"base","themeCSS":"* { font-family: Arial, sans-serif !important; } text { font-weight: 600 !important; }","themeVariables":{"background":"#222222","fontFamily":"Arial, sans-serif","fontSize":"25px","textColor":"#F7F7F7","lineColor":"#8A8A8A","edgeLabelBackground":"#222222","primaryColor":"#2A2A2A","primaryTextColor":"#F7F7F7","primaryBorderColor":"#2BD879","secondaryColor":"#2BD879","secondaryTextColor":"#101510","secondaryBorderColor":"#65EFA3","tertiaryColor":"#545454","tertiaryTextColor":"#F7F7F7","tertiaryBorderColor":"#8A8A8A","clusterBkg":"#262626","clusterBorder":"#4A4A4A","mainBkg":"#2A2A2A","nodeBorder":"#2BD879","nodeTextColor":"#F7F7F7","labelTextColor":"#F7F7F7","titleColor":"#F7F7F7"},"flowchart":{"htmlLabels":false}}}%%
graph LR

classDef service fill:#2A2A2A,stroke:#2BD879,stroke-width:2px,color:#F7F7F7;
classDef accent fill:#2BD879,stroke:#65EFA3,stroke-width:2px,color:#101510;
classDef muted fill:#545454,stroke:#8A8A8A,stroke-width:1px,color:#F7F7F7;

subgraph EXPECTED[" "]

P["Request"]

P --> U["User Service"]
P --> Z["Order Service"]
P --> T["Product Service"]
P --> O["Reporting Service"]

subgraph USERS["Users"]

U1["GetUser()"]

U1 --> U2["GetFromCache()"]
U2 --> U3{"Found?"}
U3 -->|No| U4["ReadFromDatabase()"]
U4 --> U5["SaveToCache()"]

end

subgraph ORDERS["Orders"]

Z1["GetOrders()"]

Z1 --> Z2["GetFromCache()"]
Z2 --> Z3{"Found?"}
Z3 -->|No| Z4["ExecuteQuery()"]
Z4 --> Z5["SaveToCache()"]

Z6["CreateOrder()"]

Z6 --> Z7["WriteToDatabase()"]
Z7 --> Z8["RemoveFromCache()"]

end

subgraph PRODUCTS["Products"]

T1["GetProducts()"]

T1 --> T2["GetFromCache()"]
T2 --> T3{"Found?"}
T3 -->|No| T4["ReadFromDatabase()"]
T4 --> T5["SaveToCache()"]

end

subgraph REPORTS["Reports"]

O1["GetReport()"]

O1 --> O2["GetFromCache()"]
O2 --> O3{"Found?"}
O3 -->|No| O4["GenerateReport()"]
O4 --> O5["SaveToCache()"]

end

U --> U1
Z --> Z1
Z --> Z6
T --> T1
O --> O1

end


class P accent;
class U,Z,T,O service;
class U1,U3,U4,Z1,Z3,Z4,Z6,Z7,T1,T3,T4,O1,O3,O4 muted;
class U2,U5,Z2,Z5,Z8,T2,T5,O2,O5 accent;
style EXPECTED fill:#262626,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style USERS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style ORDERS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style PRODUCTS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style REPORTS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;

You only need to provision a cluster, rewrite the one service used by the entire application, and the work is done.

I think this is what anyone would expect. But when you inspect the actual code, you see a very different picture:

%%{init: {"theme":"base","themeCSS":"* { font-family: Arial, sans-serif !important; } text { font-weight: 600 !important; }","themeVariables":{"background":"#222222","fontFamily":"Arial, sans-serif","fontSize":"25px","textColor":"#F7F7F7","lineColor":"#8A8A8A","edgeLabelBackground":"#222222","primaryColor":"#2A2A2A","primaryTextColor":"#F7F7F7","primaryBorderColor":"#2BD879","secondaryColor":"#2BD879","secondaryTextColor":"#101510","secondaryBorderColor":"#65EFA3","tertiaryColor":"#545454","tertiaryTextColor":"#F7F7F7","tertiaryBorderColor":"#8A8A8A","clusterBkg":"#262626","clusterBorder":"#4A4A4A","mainBkg":"#2A2A2A","nodeBorder":"#2BD879","nodeTextColor":"#F7F7F7","labelTextColor":"#F7F7F7","titleColor":"#F7F7F7"},"flowchart":{"htmlLabels":false}}}%%
graph LR

classDef service fill:#2A2A2A,stroke:#2BD879,stroke-width:2px,color:#F7F7F7;
classDef accent fill:#2BD879,stroke:#65EFA3,stroke-width:2px,color:#101510;
classDef muted fill:#545454,stroke:#8A8A8A,stroke-width:1px,color:#F7F7F7;

subgraph REAL[" "]

P["Request"]

P --> Z["Order Service"]
P --> T["Product Service"]
P --> O["Reporting Service"]

subgraph ORDERS["Orders"]

Z1["GetOrders()"]

Z1 --> Z2["BuildKey()"]
Z2 --> Z3["FindStoredData()"]
Z3 --> Z4{"Found?"}
Z4 -->|No| Z5["ExecuteQuery()"]
Z5 --> Z6["PrepareForStorage()"]
Z6 --> Z7["StoreFor5Minutes()"]

Z8["CreateOrder()"]

Z8 --> Z9["WriteToDatabase()"]
Z9 --> Z10["RemoveRelatedData()"]
Z10 --> Z11["StoreOrder()"]

end

subgraph PRODUCTS["Products"]

T1["GetProducts()"]

T1 --> T2["BuildKeyFromFilter()"]
T2 --> T3["FindInDictionary()"]
T3 --> T4{"Found?"}
T4 -->|No| T5["ReadFromDatabase()"]
T5 --> T6["AddToDictionary()"]

end

subgraph REPORTS["Reports"]

O1["GetReport()"]

O1 --> O2["CheckLocalStorage()"]
O2 --> O3{"Found?"}
O3 -->|No| O4["GenerateReport()"]
O4 --> O5["UpdateStorage()"]

end

Z --> Z1
Z --> Z8
T --> T1
O --> O1

end


class P accent;
class Z,T,O service;
class Z1,Z4,Z5,Z8,Z9,T1,T4,T5,O1,O3,O4 muted;
class Z2,Z3,Z6,Z7,Z10,Z11,T2,T3,T6,O2,O5 accent;
style REAL fill:#262626,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style ORDERS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style PRODUCTS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;
style REPORTS fill:#222222,stroke:#4A4A4A,stroke-width:1px,color:#F7F7F7;

Over time, every developer implemented caching in the way they considered correct. Instead of a single integration point, you ended up with dozens of different implementations scattered throughout the project.

Instead of provisioning a Redis cluster in a few clicks and rewriting the single service used by the entire application, you realize that you cannot change anything in isolation. The cost of such a refactoring is so high that the project simply does not have the resources for it.

This is exactly where the value of good abstractions becomes clear.

We faced a similar situation with authentication. The difference was that we could not abandon the migration: Kerberos was a mandatory customer requirement. Instead of ignoring the old dependencies, we had to extract them from the code and draw a clearer boundary between application logic and infrastructure.

How We Looked for a Solution

In theory, we could have said, "Let's redesign everything properly and make the platform fully compatible with any cloud environment." In practice, that option was not viable. There was no time for a complete rewrite, and trying to rewrite everything at once could have jeopardized customer acceptance. We needed to preserve a working platform, satisfy the customer's mandatory requirement, and keep the project on track.

We considered several options.

Kerberos on Top of OIDC

The fastest path seemed to be keeping the existing architecture and adding Kerberos on top of it through custom adapters and proxying.

Such a solution might have survived a demonstration, but no one would have accepted it as the production solution.

Shared Rules Without a Shared Library

Another option was to agree on architectural rules and document them. Each service would implement authentication independently, but would strictly follow the shared patterns.

This approach preserves service independence well. In our case, however, it depended too heavily on discipline and time. A few teams solving their tasks in their own way would have been enough for the shared approach to stop being shared.

A Universal Authentication Platform

We could also have built a separate, universal authentication platform that covered every specialized need of the microservices.

In practice, solutions like this grow rapidly. The result is a universal service that is difficult to extend without risking regressions and almost impossible to change without a deep understanding of its internals.

A Shared Infrastructure Layer with Extension Points

We ultimately chose the fourth option.

The idea was to move common authentication and user-context logic into an infrastructure layer, while keeping service-specific differences in predefined extension points.

This solution was possible largely because the team included an experienced engineer who could implement shared functionality that the rest of the team could then use.

Our goal was not merely to integrate Kerberos. We wanted to decouple the application code completely from any specific authentication and user-context mechanism.

To achieve this, we introduced a single contract for working with the current user. All details related to request validation, user-context construction, preparation of inter-service calls, and interaction with OIDC, JWT, Kerberos, and other mechanisms remained inside the shared layer.

At the same time, extension points allowed us to preserve differences between services without spreading those differences throughout the application code.

%%{init: {"theme":"base","themeCSS":"* { font-family: Arial, sans-serif !important; } text { font-weight: 600 !important; }","themeVariables":{"background":"#222222","fontFamily":"Arial, sans-serif","fontSize":"18px","textColor":"#F7F7F7","lineColor":"#8A8A8A","edgeLabelBackground":"#222222","primaryColor":"#2A2A2A","primaryTextColor":"#F7F7F7","primaryBorderColor":"#2BD879","secondaryColor":"#2BD879","secondaryTextColor":"#101510","secondaryBorderColor":"#65EFA3","tertiaryColor":"#545454","tertiaryTextColor":"#F7F7F7","tertiaryBorderColor":"#8A8A8A","clusterBkg":"#262626","clusterBorder":"#4A4A4A","mainBkg":"#2A2A2A","nodeBorder":"#2BD879","nodeTextColor":"#F7F7F7","labelTextColor":"#F7F7F7","titleColor":"#F7F7F7"},"flowchart":{"htmlLabels":false}}}%%
flowchart LR

classDef service fill:#2A2A2A,stroke:#2BD879,stroke-width:2px,color:#F7F7F7;
classDef accent fill:#2BD879,stroke:#65EFA3,stroke-width:2px,color:#101510;
classDef muted fill:#545454,stroke:#8A8A8A,stroke-width:1px,color:#F7F7F7;

subgraph DIAGRAM[" "]
direction LR

LIB["Shared Access Library"]:::muted
O["Order Business Logic"]:::service
X1["Extension Point"]:::accent
S["File Storage Business Logic"]:::service
X2["Extension Point"]:::accent

O --> X1
X1 -.-> LIB
S --> X2
X2 -.-> LIB

end


style DIAGRAM fill:#222222,stroke:#222222,stroke-width:1px,color:#222222;

Responsibilities were divided between the shared infrastructure layer and the individual microservices. The library processed each request through the common flow, while services added only their differences through predefined extension points. For example, the audit service recorded the authentication method, converted data to the log format, and identified the source of the event, while the file storage service converted the login to its own format.

This was not a perfect architecture for every possible case. It had costs of its own: services now shared a dependency; the shared layer could accidentally turn into a collection of unrelated special cases; and changes to the library had to be versioned and rolled out carefully. Poorly chosen extension points could also pull business logic back into the infrastructure layer.

In our situation, however, the benefits mattered more. This approach allowed developers with different levels of experience to contribute: the complex user-verification logic remained inside the shared layer, while each service only had to define its own behavior at a clear extension point.

What We Achieved

Supporting Kerberos itself was not the expensive part. The expensive part was investigating the existing platform and the implicit assumptions under which it had operated in a more familiar environment.

Almost every change began not with writing new code, but with understanding how a particular service identified the current user, how user context was propagated to the next microservice, and where infrastructure logic ended and business logic began.

The work usually followed the same pattern: first we found every place where OIDC was used; then we determined which logic depended on it, separated infrastructure concerns from application concerns, aligned the service with the shared approach, and only then connected Kerberos.

The greatest difficulty was the unpredictability. Every time we thought we had found the main dependencies, another service or workflow appeared in which OIDC was used differently.

The practical result was this: we implemented Kerberos without completely rewriting the platform and passed the customer acceptance tests within the project's constraints. The platform became part of the customer's cloud environment without requiring a separate login model just for our product.

Routine authentication work no longer began with an investigation of every microservice. A developer only needed to connect a service to the shared flow, obtain the current user, and prepare an inter-service call through the common mechanism.

New service-specific behavior was implemented through the existing extension points. If the access-management system had to be replaced again, most of the work would remain in the shared infrastructure layer instead of spreading across every microservice.

What We Learned

The main lesson is simple: the less application code knows about infrastructure, the easier it is for a platform to survive change.

When developers face a similar situation, they should not rush straight into rewriting the code. First, they should understand the scope of the dependency: where it is actually used, which workflows it breaks, and whether the migration can be avoided at all. Sometimes the right engineering decision is an uncomfortable one: not an ideal redesign, but a limited shared layer, an adapter, or a temporary compromise that keeps a business-critical project from failing.

If the migration is mandatory and there is time for refactoring, the dependency should be placed behind an explicit contract. Application code should not know protocol details, services should work through a common mechanism, and differences should remain in extension points. This protects the team not only during the current migration, but also from future infrastructure replacements.

This is especially important for cloud and hybrid projects. Application components should be replaceable and ready to integrate with third-party components, because it is impossible to shield a system completely from future changes. What we can do is draw a clear boundary between application code and infrastructure details in advance.

In our case, that became the central lesson of the migration from OIDC to Kerberos.

Top comments (0)