At AutoFi, we operate a containerized microservices infrastructure across several architectural stacks and multiple environments. Hundreds of services run on Amazon ECS and are built and supported by several development teams.
Most service configuration is provided through environment variables backed by AWS Systems Manager Parameter Store. Across our infrastructure, we manage more than 11,000 parameters.
We have historically operated with a very small DevOps team, so it is no surprise that managing this many parameters eventually became a significant challenge.
Some of our main problems were:
DevOps became a bottleneck whenever a developer needed to retrieve or update a parameter.
There was no convenient way to view, sort, and filter multiple parameter names and values.
Searching for a parameter name across AWS accounts, regions, and other infrastructure locations was difficult.
Searching by parameter value required exporting all parameters and searching the resulting file.
Comparing configuration across services and environments was cumbersome - for example, determining the difference between the staging and production configurations of the
consumerservice.We could not easily copy parameters in bulk between services or environments - for example, copying all
consumerparameters from staging to UAT.We could not schedule parameter changes and automatic rollbacks - for example, enabling debug logging temporarily and returning to the normal log level two hours later.
After several rounds of automation, we arrived at the following design:
Separate secrets from parameters that can be considered public within our organization.
Store secrets in a secrets manager (third party) and synchronize them one-way to the infrastructure parameter store. Secrets remain under DevOps control, and developers do not have direct read or write access to their values.
Store non-secret parameters in a separate database that supports searching by name and value, then synchronize them one-way to the infrastructure parameter store.
Allow developers to view all non-secret parameters and change parameters in non-production environments. Production changes are submitted for DevOps approval.
Record every change in an audit log and send a notification to our
#infrastructureSlack channel.
We implemented the missing data layer and user interface as a new Infrastructure Configuration subsystem in our internal DevOps Portal.
Here is what it looks like.
Parameter List
The main page allows users to view parameter names and values for any combination of stack, environment, and application. They can sort and filter the list by name or value, as well as add parameters, edit them, and attach comments.
Approval Workflow
If a parameter belongs to a group that requires approval - for example, parameters in a production environment - the change is placed in an approval queue.
These groups are configurable in the Portal settings. The DevOps team receives a Slack notification and can review and approve the update before it proceeds.
Synchronization
Approved changes are not applied to the infrastructure immediately. Instead, they accumulate on the Sync tab until someone starts the synchronization process manually.
We chose this approach so that multiple changes can be synchronized together, minimizing the number of service restarts. Services that depend on the modified parameters are restarted automatically after the new values are written to the infrastructure.
Configuration Comparison
One common task is comparing the configuration of the same service across environments - for example, finding the differences between the staging and production configurations of the consumer service.
We exposed configuration data through our DevOps Portal MCP server and use an AI agent to perform these comparisons using natural-language requests.
Scheduled Changes and Automatic Rollbacks
A parameter update can also be scheduled for a specific time. Optionally, the Portal can restore the previous value automatically after a specified period.
For example, we can temporarily increase application logging verbosity and then automatically return the log level to normal, reducing the risk that someone forgets to revert it and allowing us to control pressure on our log ingestion and storage systems.
Scheduling is also useful when several related parameters must be changed - and later rolled back - at the same time.
Bulk Copy
When creating a new service or environment, we often need to copy an existing configuration and use it as a starting point.
The bulk-copy workflow lets users select a source stack, environment, and application; choose the destination; and specify how existing parameters should be handled.
Text Editor
For operations involving many parameters, users can switch to text-editor mode and add or update parameters in bulk using a simple NAME = VALUE format.
Representing Secrets as Symbols
We use a separate system and workflow to manage secrets. However, developers still need to know which configuration entries exist, including entries whose values they are not allowed to access.
To provide a complete view without exposing sensitive data, we introduced a parameter type called a symbol. A symbol has a name but no value in the Portal.
A daily automation pipeline retrieves secret names - but never their values - from the infrastructure and imports them into the Portal through its API. Developers can therefore see the complete configuration structure while the actual secret values remain protected.
Summary
Building the Infrastructure Configuration subsystem required meaningful architecture and development effort, but it changed how configuration management works across our organization.
Developers can now independently search, review, compare, copy, and initiate changes to most application configuration. At the same time, DevOps retains control over secrets, production approvals, infrastructure synchronization, and auditing.
For our small DevOps team, this has removed a significant operational bottleneck while giving development teams faster and more transparent access to the configuration they depend on.
Originally published in AutoFi DevOps on Medium.









Top comments (0)