Most engineers who deal with PCI DSS treat compliance as something that happens after the infrastructure exists: build the thing, then ask someone to audit it. That backwards process is why compliance engagements are painful and expensive.
For this project I reversed it. The architecture was designed specifically to qualify for SAQ A-EP first, and then every technical control was built and documented against the corresponding PCI DSS v4.0.1 requirement. The output is a completed Self-Assessment Questionnaire (SAQ A-EP) and a signed Attestation of Compliance (AOC), not just working infrastructure.
This post covers what SAQ A-EP actually requires, why the architecture I chose qualifies for it, and exactly how each control was implemented in code. AWS equivalents are called out throughout because the principles transfer directly even though the tooling differs.
What SAQ A-EP Is and Why It Matters
PCI DSS defines several SAQ forms based on how a merchant handles cardholder data. Most engineers are familiar with SAQ A (fully outsourced payment pages, simplest) and SAQ D (you store, process, or transmit PANs, heaviest). SAQ A-EP sits between them.
SAQ A-EP applies when all three of the following are true:
- Your own server hosts the payment page
- All actual card data entry happens inside a PCI-certified third party's environment (like Stripe's Embedded Checkout)
- Your servers never receive, process, or store PANs, CVVs, or track data
WorkUp is a fictional B2B project-management SaaS charging $49/month per workspace with roughly 150,000 transactions per year. SAQ A-EP is the correct form. The payment page (billing.html) loads from WorkUp's own nginx. The card entry form is a Stripe-controlled iframe served from js.stripe.com. The Go API never receives card data. It only creates a Stripe Checkout Session and receives a webhook event after payment completes.
Because of this, Requirements 3 (stored card data), 4 (card data in transit), and 9 (physical access) are legitimately Not Applicable. Requirements 1, 2, 5, 6, 7, 8, 10, and 11 still apply in full to the systems hosting the payment page and supporting infrastructure.
That scope decision is the entire foundation of this project. Everything else builds on it.
Architecture: Four Isolated VMs, Zero LAN Traffic
The infrastructure runs on a self-hosted Proxmox node with four Ubuntu 22.04 VMs, each serving one segment:
Frontend VM (192.168.123.21) nginx serving static HTML/JS
Auth VM (192.168.123.22) Keycloak 26 + Postgres 18
Backend VM (192.168.123.23) Go/Gin API + Postgres 18
Obs VM (192.168.123.24) Grafana + Prometheus + Loki + Grafana Alloy
The key design choice: none of these VMs talk to each other over the internal LAN. Every inter-segment call goes through its public Cloudflare Tunnel URL. The frontend calls https://api.victorojeje.xyz, not http://192.168.123.23:8081. Auth tokens, API calls, log shipping: all of it exits through Cloudflare and comes back in through the tunnel on the destination VM.
This has a direct compliance consequence. PCI DSS Req 1 requires network segmentation. In most setups that means careful firewall rules to control east-west traffic between segments. Here, there is no east-west traffic. The default-DROP firewall policy on each VM is correct with zero application-level exceptions.
AWS parallel: In AWS this maps to placing each segment in its own subnet with a default-deny Network Access Control List (NACL), and using VPC Endpoints or PrivateLink for inter-service communication instead of direct subnet routing. The principle is the same: make the default behavior secure, then open only what's necessary.
Phase 1: Provisioning Infrastructure
Proxmox Template
Before Terraform provisions anything, a base VM template needs to exist. infra/template.sh handles this once. It downloads the Ubuntu 22.04 Jammy cloud image, uploads it to Proxmox, boots a VM, installs Docker and the QEMU guest agent, then converts it to a Proxmox template.
The script solves a real problem that breaks naively-written provisioning: Ubuntu cloud images fire apt-daily.timer with Persistent=true on first boot, which holds /var/lib/dpkg/lock-frontend and makes apt install fail immediately. The fix runs before touching apt:
sudo systemctl stop apt-daily.service apt-daily-upgrade.service \
apt-daily.timer apt-daily-upgrade.timer unattended-upgrades.service 2>/dev/null || true
sudo systemctl mask apt-daily.service apt-daily-upgrade.service \
apt-daily.timer apt-daily-upgrade.timer unattended-upgrades.service
sudo DEBIAN_FRONTEND=noninteractive apt-get -o DPkg::Lock::Timeout=120 update -y
sudo DEBIAN_FRONTEND=noninteractive apt-get -o DPkg::Lock::Timeout=120 install -y \
qemu-guest-agent curl ca-certificates
After Docker installs, the VM shuts down and becomes a template with ciupgrade=0 set, which stops cloud-init from running package upgrades every time a VM is cloned from it.
Terraform VM Provisioning
Provider: bpg/proxmox v0.112.0. The older telmate/proxmox provider is unmaintained. bpg/proxmox has proper cloud-init support, a cleaner resource model, and active maintenance.
All four VMs follow the same pattern in clones.tf. They clone from the template with a full clone (not linked), get a static IP, and sized memory based on what each service actually needs: 768 MB for nginx, 1536 MB for Keycloak's JVM, 1024 MB for Go, 1536 MB for the full observability stack.
The single most important line in the VM definition:
network_device {
bridge = "vmbr0"
model = "virtio"
firewall = true # without this, firewall.tf rules exist but never enforce
}
firewall = true tells Proxmox to attach the NIC-level firewall policy to that specific network interface. Without this flag, you can define all the rules you want in firewall.tf and they will be stored in Proxmox's API but never evaluated. This was a real misconfiguration found mid-project; the auth, backend, and observability VMs were missing it.
Proxmox Firewall as Code
The firewall runs at three levels: Datacenter, Node, and per-VM. All three are declared in firewall.tf.
# Define who is allowed to manage things
resource "proxmox_virtual_environment_firewall_ipset" "management" {
name = "management"
cidr {
name = "192.168.123.132/32"
comment = "Lab / management network"
}
}
# Datacenter-level: drop everything by default
resource "proxmox_virtual_environment_cluster_firewall" "cluster" {
enabled = true
input_policy = "DROP"
output_policy = "ACCEPT"
}
# Cluster-level rules: Proxmox GUI + SSH from management only
resource "proxmox_virtual_environment_firewall_rules" "cluster_mgmt" {
rule { type = "in"; action = "ACCEPT"; source = "+management"; dport = "8006"; proto = "tcp" }
rule { type = "in"; action = "ACCEPT"; source = "+management"; dport = "22"; proto = "tcp" }
rule { type = "in"; action = "ACCEPT"; source = "+management"; proto = "icmp" }
}
# Reusable security groups attached to each VM
resource "proxmox_virtual_environment_cluster_firewall_security_group" "mgmt_ssh" {
name = "mgmt-ssh"
rule { type = "in"; action = "ACCEPT"; source = "+management"; dport = "22"; proto = "tcp" }
}
# Enable DROP on each VM, attach security groups
resource "proxmox_virtual_environment_firewall_options" "frontend" {
node_name = proxmox_virtual_environment_vm.frontend.node_name
vm_id = proxmox_virtual_environment_vm.frontend.vm_id
enabled = true
input_policy = "DROP"
}
resource "proxmox_virtual_environment_firewall_rules" "frontend" {
node_name = proxmox_virtual_environment_vm.frontend.node_name
vm_id = proxmox_virtual_environment_vm.frontend.vm_id
rule { security_group = proxmox_virtual_environment_cluster_firewall_security_group.mgmt_ssh.name; iface = "net0" }
rule { security_group = proxmox_virtual_environment_cluster_firewall_security_group.mgmt_icmp.name; iface = "net0" }
}
Result: all four VMs drop all inbound traffic except SSH and ICMP from the management machine. No application ports are reachable from the LAN. Public traffic reaches services through Cloudflare Tunnels only.
AWS parallel: This maps to a layered defense using NACLs (stateless, subnet-level, applied to all traffic entering the subnet) and Security Groups (stateful, instance-level, applied to specific ENIs). The three-level Proxmox hierarchy mirrors that model: cluster policy is the NACL equivalent, per-VM policy is the Security Group equivalent, and security groups in Proxmox are reusable rule sets like AWS managed prefix lists attached to SGs.
Phase 2: Hardening Every VM With Ansible
Four playbooks, one per VM segment. Every playbook runs the same task sequence before deploying its application: host_firewall.yml, os_hardening.yml, auditd.yml, aide.yml.
Guest OS Firewall (ufw)
The Proxmox NIC firewall is one layer. ufw inside each guest OS is a second, independent layer:
# tasks/host_firewall.yml
- name: Reset ufw to defaults
community.general.ufw:
state: reset
- name: Set default incoming policy to deny
community.general.ufw:
direction: incoming
policy: deny
- name: Allow SSH from management network only
community.general.ufw:
rule: allow
port: "22"
proto: tcp
src: "{{ management_network }}"
- name: Enable ufw
community.general.ufw:
state: enabled
There is no ICMP rule in ufw. The community.general.ufw module doesn't handle ICMP cleanly, and the Proxmox firewall already allows management ICMP. Adding a poorly-written ICMP rule in ufw adds complexity without benefit.
The Docker Port Publishing Problem
Docker bypasses ufw by writing directly to iptables PREROUTING. Any ports: mapping in a Compose file exposes that port to the LAN regardless of ufw's default-deny. This means a ports: ["8081:8081"] entry on the API container makes port 8081 reachable from any machine on the LAN, even with ufw default deny incoming in place.
The fix: no ports: block on any application service. cloudflared reaches the API container via Docker's internal service DNS (http://api:8081) without the port being published. The Keycloak container uses 127.0.0.1:8080:8080 to bind to the host loopback only for local bootstrap access, and cloudflared reaches it via internal DNS anyway.
OS Hardening
# tasks/os_hardening.yml (condensed)
# Scope ubuntu's sudo: removes cloud-init's dangerous NOPASSWD: ALL default
- name: Scope ubuntu sudo to required commands only
ansible.builtin.copy:
dest: /etc/sudoers.d/90-ubuntu-scoped
content: |
ubuntu ALL=(ALL) NOPASSWD: /usr/bin/docker, /usr/bin/systemctl, /usr/sbin/ufw
mode: "0440"
validate: "visudo -cf %s" # validate before writing, prevents sudo lockout
# SSH hardening
- name: Harden sshd_config
ansible.builtin.lineinfile:
path: /etc/ssh/sshd_config
regexp: "{{ item.regexp }}"
line: "{{ item.line }}"
validate: /usr/sbin/sshd -t -f %s # test config before writing
loop:
- { regexp: '^#?PermitRootLogin', line: 'PermitRootLogin no' }
- { regexp: '^#?PasswordAuthentication', line: 'PasswordAuthentication no' }
- { regexp: '^#?MaxAuthTries', line: 'MaxAuthTries 3' }
- { regexp: '^#?X11Forwarding', line: 'X11Forwarding no' }
- { regexp: '^#?AllowTcpForwarding', line: 'AllowTcpForwarding no' }
- { regexp: '^#?ClientAliveInterval', line: 'ClientAliveInterval 300' }
- { regexp: '^#?ClientAliveCountMax', line: 'ClientAliveCountMax 2' }
notify: Restart ssh
# Kernel hardening via sysctl
- name: Apply sysctl hardening
ansible.posix.sysctl:
name: "{{ item.name }}"
value: "{{ item.value }}"
sysctl_file: /etc/sysctl.d/99-hardening.conf
loop:
- { name: 'kernel.randomize_va_space', value: '2' } # ASLR
- { name: 'net.ipv4.tcp_syncookies', value: '1' } # SYN flood protection
- { name: 'net.ipv4.conf.all.rp_filter', value: '1' } # reverse path filtering
- { name: 'net.ipv4.conf.all.accept_redirects', value: '0' } # no ICMP redirects
- { name: 'net.ipv4.conf.all.send_redirects', value: '0' }
- { name: 'fs.suid_dumpable', value: '0' } # no setuid core dumps
- { name: 'kernel.dmesg_restrict', value: '1' } # non-root can't read dmesg
- { name: 'kernel.kptr_restrict', value: '2' } # hide kernel pointers
The validate on sshd_config changes is not optional. Running sshd -t before writing the file ensures you never write a config that breaks SSH access. That validation step has saved this project from a lockout at least once.
AWS parallel: Ansible playbooks here are functionally equivalent to AWS Systems Manager (SSM) State Manager associations. SSM State Manager applies configurations on a schedule to EC2 instances, and SSM Parameter Store holds secrets the same way
vars/secrets.yml(Ansible Vault) does here. The sysctl and sshd hardening tasks map directly to EC2 user data scripts or SSM Run Command documents.
Phase 3: Audit Logging
Two controls here. Both are required by PCI DSS. Both ship their output to the central observability stack.
auditd: Kernel-Level Audit (PCI Req 10.2.1)
auditd is installed on all four VMs and configured with rules that map directly to PCI DSS Req 10.2.1 sub-requirements:
# infra/ansible/playbooks/files/pci-audit.rules
## 10.2.1.6: detect tampering with the audit subsystem itself
-w /etc/audit/ -p wa -k audit_config
-w /sbin/auditctl -p x -k audit_tools
-w /sbin/auditd -p x -k audit_tools
## 10.2.1.5: identity and authentication file changes
-w /etc/passwd -p wa -k identity
-w /etc/shadow -p wa -k identity
-w /etc/sudoers -p wa -k identity
-w /etc/sudoers.d/ -p wa -k identity
-w /etc/ssh/sshd_config -p wa -k sshd_config
## 10.2.1.2: privileged command execution
-w /usr/bin/sudo -p x -k privileged
-w /usr/bin/su -p x -k privileged
-w /usr/sbin/useradd -p x -k privileged
-w /usr/sbin/userdel -p x -k privileged
-w /usr/sbin/usermod -p x -k privileged
## 10.2.1.3: access to the audit log itself
-w /var/log/audit/ -p rwa -k audit_log_access
## 10.2.1.4: failed login tracking
-w /var/log/faillog -p wa -k logins
-w /var/run/faillock/ -p wa -k logins
The -p wa flag watches for writes and attribute changes. -p x watches for execution. -k is the key used to filter audit events when searching logs. These map one-to-one to the SAQ A-EP sub-requirements, which matters when filling in the questionnaire; you can point to a specific rule file, not just say "yes we log things."
The Ansible task that deploys these rules uses a handler to reload them:
# tasks/auditd.yml
- name: Deploy PCI audit rules
ansible.builtin.copy:
src: files/pci-audit.rules
dest: /etc/audit/rules.d/pci-dss.rules
owner: root
group: root
mode: "0640"
notify: Reload auditd rules # handler: augenrules --load
AIDE: File Integrity Monitoring (PCI Req 11.5.2)
AIDE (Advanced Intrusion Detection Environment) runs on all four VMs. The configuration template defines what to monitor and what to exclude:
# aide.conf.j2 (relevant portions)
FULL = p+i+n+u+g+s+b+m+c+sha256 # perms+inode+links+user+group+size+blocks+mtime+ctime+hash
BINLIB = p+u+g+sha256 # for binaries: just identity + hash
CONFIG = p+u+g+n+sha256 # for config files: + symlinks
/boot BINLIB
/bin BINLIB
/usr/bin BINLIB
/usr/sbin BINLIB
/etc CONFIG
/opt BINLIB # where Docker Compose stacks live
# Exclude volatile paths: these change constantly and would flood alerts
!/var/lib/docker
!/var/log
!/var/cache
!/proc
!/run
!/tmp
The exclusions are as important as the inclusions. Without excluding /var/log and /var/lib/docker, every Docker layer pull and every log rotation would trigger a false positive. The point is detecting unexpected changes to binaries and configuration, not logging every file write.
The Ansible task initializes the baseline and installs the daily cron:
# tasks/aide.yml
- name: Initialize AIDE database (first run only)
ansible.builtin.command: aideinit
args:
creates: /var/lib/aide/aide.db # idempotent: skip if db already exists
- name: Daily AIDE integrity check
ansible.builtin.cron:
name: "AIDE daily integrity check"
minute: "17"
hour: "3"
user: root
job: "/usr/bin/aide --check > /var/log/aide/aide-check.log 2>&1 || true"
The cron runs at 03:17, not 03:00. Running at exactly :00 means competing with every other cron job that defaults to that time. The || true prevents a non-zero exit from AIDE (which reports differences, not errors) from looking like a cron failure.
Grafana Alloy ships the AIDE log file to Loki. Changes to monitored files show up in Grafana immediately.
AWS parallel: In AWS, file integrity monitoring at this level combines AWS Config (detects configuration drift on EC2 instances), AWS CloudTrail with CloudTrail Insights (detects unusual API patterns), and Amazon GuardDuty (anomaly detection). For a direct equivalent of AIDE's file hash comparison, Amazon Inspector also includes a reachability and package vulnerability component. None of these is a 1:1 match for AIDE's low-level file monitoring, which is why on-prem compliance often runs AIDE even in hybrid environments.
Phase 4: Identity and Access
Keycloak 26 runs on the Auth VM in production mode (start, not start-dev). Everything is configured through a bootstrap script that calls the Keycloak Admin REST API, with no manual GUI steps after the first deploy.
Why REST API instead of a realm JSON export? An exported JSON file becomes a maintenance burden. When one setting needs updating, you diff and re-import the whole thing. The bootstrap script is idempotent: it checks whether each resource exists, creates it if missing, and updates it if present. The same REALM_SETTINGS block handles both the initial create (POST) and subsequent updates (PUT).
# app/auth/bootstrap.sh: realm settings (key security parameters shown)
REALM_SETTINGS=$(cat <<EOF
{
"realm": "workup",
"enabled": true,
"accessTokenLifespan": 300,
"ssoSessionIdleTimeout": 900,
"ssoSessionMaxLifespan": 28800,
"verifyEmail": true,
"passwordPolicy": "length(12) and upperCase(1) and lowerCase(1) and digits(1) and specialChars(1) and notUsername and passwordHistory(5)",
"bruteForceProtected": true,
"failureFactor": 5,
"maxFailureWaitSeconds": 1800,
"smtpServer": {
"host": "smtp.resend.com",
"port": "465",
"ssl": "true",
"auth": "true",
"user": "resend",
"password": "${RESEND_API_KEY}"
}
}
EOF
)
These values map directly to PCI DSS requirements: password complexity maps to Req 8.3.6, brute-force protection maps to Req 8.3.4, session timeouts map to Req 8.2.8.
TOTP is configured as a required action on the realm:
curl -X PUT \
"http://localhost:8080/admin/realms/workup/authentication/required-actions/CONFIGURE_TOTP" \
-H "Authorization: Bearer ${ADMIN_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"alias": "CONFIGURE_TOTP",
"enabled": true,
"defaultAction": true,
"priority": 10
}'
Setting defaultAction: true means every new user is prompted to configure an authenticator app on first login. This satisfies PCI DSS Req 8.4 (MFA for all accounts with access to the cardholder data environment).
The RBAC model is deliberately minimal. One realm role: admin. It is only meaningful in systems that read it from the JWT; currently only Grafana (role_attribute_path = "contains(realm_access.roles[*], 'admin')" with role_attribute_strict = true). An early design had four roles: admin, billing-ops, support, developer. billing-ops maps to Stripe, developer maps to GitLab. Neither system authenticates through Keycloak. Creating Keycloak roles for them produces dead configuration that looks good on paper but enforces nothing. The correct approach is to document each system's own access controls separately, which is what the evidence docs do.
Three OIDC clients are created by the bootstrap: workup-frontend (public, PKCE S256, standard flow only), cloudflare-access (confidential, used as the IdP for Cloudflare Zero Trust), and grafana (confidential, for Grafana OAuth). Direct access grants are disabled on all three.
AWS parallel: Amazon Cognito User Pools is the managed equivalent of Keycloak here. It supports OIDC/OAuth2, PKCE, built-in MFA (TOTP via software tokens), and password policies configured in the pool settings. The bootstrap.sh pattern maps to Cognito's API or Terraform's
aws_cognito_user_poolresource. One meaningful difference: Cognito doesn't have an equivalent of Keycloak's Admin REST API for idempotent realm bootstrapping; you use Terraform or CDK instead, which has the same idempotency property.
Phase 5: The Payment Flow
This is the architectural core of the SAQ A-EP eligibility. The flow:
- Authenticated user opens
billing.html(served by nginx on the Frontend VM) -
billing.jscallsPOST /billing/create-checkout-sessionwith a Keycloak JWT in theAuthorizationheader - The Go API validates the JWT, creates a Stripe Checkout Session, returns a
clientSecret -
billing.jscallsstripe.initEmbeddedCheckout({ clientSecret })and mounts the checkout UI - The card entry iframe is served from
js.stripe.com(Stripe's servers, not WorkUp's) - The user submits card details directly to Stripe
- Stripe calls the backend's webhook endpoint after payment
Step 5 is the entire basis for the N/A decisions on Req 3, 4, and 9. Card data (PAN, CVV, expiry) is submitted directly from the browser to Stripe's servers. WorkUp's servers are not in that data path.
The backend's JWT validation uses github.com/MicahParks/keyfunc/v3 to fetch and cache the JWKS from Keycloak's public endpoint:
// app/backend/internal/auth/auth.go
jwks, err := keyfunc.NewDefaultCtx(ctx, []string{
cfg.KeycloakIssuerURL + "/protocol/openid-connect/certs",
})
token, err := jwt.Parse(tokenString, jwks.Keyfunc,
jwt.WithValidMethods([]string{"RS256"}),
jwt.WithIssuer(cfg.KeycloakIssuerURL),
jwt.WithAudience(cfg.KeycloakClientID),
)
It validates algorithm (RS256), signature (against Keycloak's public keys), issuer, and audience. A rejected token is logged and counted as a Prometheus metric.
The Stripe webhook endpoint at POST /billing/webhook verifies the Stripe-Signature header before processing any event. A Cloudflare Custom Rule also blocks any POST to /billing/webhook that arrives without the stripe-signature header; unsigned payloads are rejected at the CDN edge before reaching the backend.
The database schema confirms the scope in code:
-- app/backend/db/schema.sql
CREATE TABLE workspaces (
id BIGSERIAL PRIMARY KEY,
keycloak_sub TEXT NOT NULL UNIQUE,
name TEXT NOT NULL
);
CREATE TABLE subscriptions (
id BIGSERIAL PRIMARY KEY,
workspace_id BIGINT NOT NULL REFERENCES workspaces(id),
stripe_customer_id TEXT NOT NULL,
stripe_subscription_id TEXT NOT NULL,
status TEXT NOT NULL
);
No PAN fields. No card_last4. No CVV. No expiry date. Not even a placeholder column. This was confirmed by direct code review, not assumed from the Stripe integration model. That distinction matters when filling in the SAQ; you have to be able to point to evidence, not just assert.
The Stripe secret key in use is a Restricted key scoped to only the operations the backend performs. The project started with a Standard key and rotated to a Restricted key mid-project when the gap was identified. Standard keys grant full Stripe API access; a restricted key for checkout sessions and webhook processing is the correct least-privilege approach.
The Go backend runs in a multi-stage Docker container. Builder stage uses golang:1.27-alpine with CGO_ENABLED=0 GOOS=linux and -ldflags="-s -w" to produce a statically linked binary with no debug symbols. The final stage uses alpine:3.20 with a non-root user:
FROM golang:1.27-alpine AS builder
WORKDIR /build
RUN apk add --no-cache git ca-certificates
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-s -w" -o /workup-backend ./cmd/server
FROM alpine:3.20
RUN apk add --no-cache ca-certificates
RUN addgroup -g 1000 workup && adduser -u 1000 -G workup -D workup
WORKDIR /app
COPY --from=builder /workup-backend /app/workup-backend
RUN chown -R workup:workup /app
ENTRYPOINT ["/app/workup-backend"]
AWS parallel: In AWS, this payment architecture would put the API behind an ALB with WAF rules (equivalent to the Cloudflare Custom Rules). The Stripe webhook signature verification is identical. Secrets (Stripe keys, Keycloak client credentials) would live in AWS Secrets Manager rather than Ansible Vault-encrypted files, with IAM roles granting the EC2 instance access. The JWT validation library and logic are the same regardless of where Keycloak runs.
Phase 6: Observability
Grafana Alloy runs on all four VMs. On the three non-observability VMs, it ships four log sources to the central Loki instance: container stdout/stderr (via Docker socket), systemd journal, auditd output, and AIDE check reports.
# app/auth/alloy-config.river (pattern used on all non-obs VMs)
# Ship container logs
discovery.docker "containers" {
host = "unix:///var/run/docker.sock"
}
loki.source.docker "default" {
host = "unix:///var/run/docker.sock"
targets = discovery.relabel.containers.output
forward_to = [loki.write.central.receiver]
}
# Ship auditd logs
local.file_match "audit" {
path_targets = [{"__path__" = "/var/log/audit/audit.log"}]
}
loki.source.file "audit" {
targets = local.file_match.audit.targets
forward_to = [loki.relabel.audit.receiver]
}
# Ship AIDE check reports
local.file_match "aide" {
path_targets = [{"__path__" = "/var/log/aide/*.log"}]
}
loki.source.file "aide_logs" {
targets = local.file_match.aide.targets
forward_to = [loki.relabel.aide.receiver]
}
# All remote writes include Cloudflare Access service token headers
loki.write "central" {
endpoint {
url = "https://logs.victorojeje.xyz/loki/api/v1/push"
headers = {
"CF-Access-Client-Id" = env("CF_ACCESS_CLIENT_ID"),
"CF-Access-Client-Secret" = env("CF_ACCESS_CLIENT_SECRET"),
}
}
}
The logs.victorojeje.xyz URL is a Cloudflare Tunnel that reaches Loki on the observability VM. The CF_ACCESS headers carry a Cloudflare service token that authenticates Alloy to the tunnel; random internet traffic trying to write to Loki over that URL gets rejected.
The observability VM's own Alloy writes locally (same Docker network as Loki and Prometheus). No CF_ACCESS headers needed.
Loki's retention is set to exactly 365 days:
# app/observability/loki-config.yaml
limits_config:
retention_period: 8760h # 365 days: PCI Req 10.5.1 requires 12 months minimum
compactor:
compaction_interval: 10m
retention_enabled: true
retention_delete_delay: 2h
8760h is 365 days. Not "approximately 12 months": exactly 365 days. PCI Req 10.5.1 requires 12 months. Any rounding that lands below 365 is a compliance gap.
Grafana is accessed via Cloudflare Access Auth Proxy, not a second OIDC login. When a user hits grafana.victorojeje.xyz, Cloudflare Access authenticates them (via Keycloak or OTP). If they pass, Cloudflare forwards the request to Grafana with the Cf-Access-Authenticated-User-Email header set. Grafana reads that header and signs the user in:
# app/observability/docker-compose.yml
grafana:
environment:
GF_AUTH_PROXY_ENABLED: "true"
GF_AUTH_PROXY_HEADER_NAME: "Cf-Access-Authenticated-User-Email"
GF_AUTH_PROXY_AUTO_SIGN_UP: "true"
GF_AUTH_DISABLE_LOGIN_FORM: "true"
GF_AUTH_ANONYMOUS_ENABLED: "false"
GF_AUTH_BASIC_ENABLED: "false"
The earlier implementation used Keycloak OIDC for Grafana. It caused a double-login: Cloudflare Access prompts for credentials, then Grafana's OIDC redirect triggers another Keycloak login. Auth Proxy eliminates the redundant prompt while keeping the access gate at the Cloudflare edge.
AWS parallel: Amazon CloudWatch Logs is the managed equivalent of Loki here. Set the log group retention policy to 365 days (
aws logs put-retention-policy --log-group-name /workup/audit --retention-in-days 365). AWS CloudTrail provides API-level audit logging equivalent to what auditd captures at the OS level. Amazon OpenSearch Service (or the Grafana integration with CloudWatch) replaces the local Grafana + Loki stack. The Grafana Auth Proxy pattern maps to Cognito-based authentication in AWS CloudWatch dashboards or directly in managed Grafana on AWS.
Phase 7: CI/CD Security
GitHub is the code source of truth. On every push, a GitHub Actions workflow mirrors the code to GitLab, where the full CI pipeline runs.
The mirror strips .github/ before pushing; GitHub Actions config doesn't belong on GitLab:
# .github/workflows/mirror-to-gitlab.yml
- name: Push to GitLab without .github
run: |
git checkout -b mirror-temp
git rm -r .github
git commit -m "temp: remove .github for GitLab mirror"
git push gitlab mirror-temp:main --force
The GitLab CI pipeline runs five jobs across three stages:
# .gitlab-ci.yml (structure)
stages: [test, build, scan]
include:
- template: Security/SAST.gitlab-ci.yml # Semgrep: Go source analysis
- template: Security/Secret-Detection.gitlab-ci.yml
govulncheck:
stage: test
image: golang:1.27
script:
- go install golang.org/x/vuln/cmd/govulncheck@latest
- govulncheck ./... # checks Go deps against the Go vuln database
build-backend:
stage: build
image: docker:24.0
services: [docker:24.0-dind]
script:
- docker build --tag "$BACKEND_IMAGE:$CI_COMMIT_SHORT_SHA" --tag "$BACKEND_IMAGE:latest" \
-f app/backend/Dockerfile app/backend
- docker push "$BACKEND_IMAGE:$CI_COMMIT_SHORT_SHA"
- docker push "$BACKEND_IMAGE:latest"
rules:
- if: $CI_COMMIT_BRANCH == "main"
trivy-container:
stage: scan
needs: [build-backend]
script:
- trivy image --severity HIGH,CRITICAL \
--format json --output gl-container-scanning-report.json \
"$BACKEND_IMAGE:$CI_COMMIT_SHORT_SHA"
artifacts:
paths: [gl-container-scanning-report.json]
expire_in: 30 days
govulncheck is specifically for Go. It checks the module's dependencies against the Go vulnerability database (vuln.go.dev), not just generic CVEs. A dependency can appear clean in a general scanner like Trivy but show up in govulncheck because the vulnerability is in a code path the module actually calls.
The built image is stored in the GitLab Container Registry. The Backend VM pulls it at deploy time using a GitLab deploy token scoped to read_registry only (not a personal access token, not the CI runner token).
AWS parallel: This pipeline maps to AWS CodePipeline triggering CodeBuild on push. The
govulncheckstep becomes a CodeBuild phase. The Docker build and push targets Amazon ECR. Trivy is replaced by Amazon Inspector V2, which integrates natively with ECR and scans images on push. GitLab's Secret Detection maps to AWS CodeGuru Security's secret detection. The registry pull uses an IAM role attached to the EC2 instance withecr:GetAuthorizationTokenandecr:BatchGetImagepermissions, equivalent to the scoped deploy token.
What the SAQ A-EP Actually Required
The SAQ itself is a 100-page questionnaire. Most of it is not surprising if you understand the scope. A few decisions worth explaining:
Req 3, 4, 9: Not Applicable.
These requirements cover protecting stored card data, card data in transit, and physical access to card data. None of them apply because WorkUp never has card data. The N/A justifications are written documents (docs/evidence/na-justifications.md) that explain why each requirement doesn't apply, not just check a box.
Req 5: Anti-malware, Not Applicable (justified).
PCI DSS Req 5.1 allows anti-malware to be declared not applicable for systems that are "not commonly affected by malware." These VMs have no interactive logins (SSH from management only), no browsers, no file uploads, and no email. They run containerized workloads with images scanned by Trivy in CI. The N/A justification documents the reasoning.
Req 6.4.3: Payment page scripts.
This requirement asks for a list of all scripts on the payment page, the purpose of each, and either an integrity mechanism or a written justification for scripts that don't support it. The payment page (billing.html) loads five scripts. Four are first-party. The fifth is https://js.stripe.com/v3/. Stripe.js explicitly does not support SRI (Subresource Integrity); their documentation states the script updates dynamically for fraud detection. The justification for the SRI exception is Stripe's own PCI Level 1 certification status. All five scripts are inventoried in docs/evidence/payment-page-script-inventory.md.
Req 8.4: MFA for all access to the cardholder data environment.
TOTP via Keycloak covers human users. Cloudflare Access provides OTP as a backup login method. Machine-to-machine access (Alloy to Loki/Prometheus, backend to Keycloak JWKS endpoint) uses service tokens and JWT validation respectively.
Things That Went Wrong
firewall = true missing on VM network interfaces.
The auth, backend, and observability VMs were defined in Terraform without firewall = true on their network_device blocks. The Proxmox firewall rules for those VMs existed in Terraform state but were never evaluated. Caught when checking the Proxmox GUI and noticing the shield icon was greyed out on three of four VMs.
Docker published ports bypassing ufw.
An early version of the backend's docker-compose.yml had ports: ["8081:8081"] on the api service. Port 8081 was reachable from the LAN even with ufw default deny incoming. Docker writes iptables PREROUTING rules that the INPUT chain (where ufw operates) never sees. Fix: remove the ports: block entirely.
Grafana double-login.
Using both Cloudflare Access (with Keycloak as IdP) and Grafana's own Keycloak OIDC integration created two separate login prompts. The fix was to disable Grafana's own OAuth and switch to Auth Proxy mode, where Cloudflare passes the authenticated user's email in a trusted header and Grafana reads it directly.
Cloudflare Access session duration set to "No duration, expires immediately" on the Grafana application.
This is correct for machine-to-machine service token applications. It is wrong for user-facing applications. Every page load triggered a Cloudflare Access re-authentication. Setting the session to 24h fixed it.
Postgres 18 volume mount path changed.
The volume for Postgres 18 mounts at /var/lib/postgresql (not /var/lib/postgresql/data as in earlier versions). Using the old path created a nested directory structure where Postgres thought the data directory was uninitialized on every container restart.
.gitignore accidentally hiding non-secret variable files.
An early .gitignore entry used vars/*, which matched management.yml and os_hardening.yml in addition to secrets.yml. Those files were never committed, and Ansible runs that referenced them failed silently. Changed to vars/secrets.yml specifically.
The Result
One completed SAQ A-EP, one signed AOC. The infrastructure runs. The audit trail is in Loki. AIDE runs nightly on all four VMs. The CI pipeline scans every push. The payment flow routes card data through Stripe without touching WorkUp.
The point of building this was not to demonstrate that you can pass a PCI audit. The point was to demonstrate that compliance is an architectural decision, not a retrofit. When the architecture is correct, when card data genuinely never reaches your servers, the audit becomes a documentation exercise rather than a gap-remediation project.
If you want to see the implementation, the repo is at github.com/escanut/on-prem-workup-pci-dss.
Victor Ojeje is a Cloud and DevOps Engineer based in Lagos. He holds AWS SAA-C03, CompTIA Security+, and CEH certifications.
Top comments (1)
Really thoughtful post! Documenting real-world engineering hurdles and actionable solutions like this brings immense value to the community.