DEV Community

Cover image for From scripted to declarative: what 10+ years of Jenkins pipelines taught me about shared libraries and quality gates
Tayguara Reis
Tayguara Reis

Posted on AI-assisted

From scripted to declarative: what 10+ years of Jenkins pipelines taught me about shared libraries and quality gates

I wrote my first Jenkins pipeline in 2016. Ten years later, the repository it started has more than 260 pipeline files: deploy jobs, end-to-end suites, a Selenium grid, health checks, scheduled maintenance. Most of them are scripted. A newer project, started in 2021, runs on six declarative pipelines and puts releases through a chain of quality gates before a blue-green deploy.

This is not a syntax tutorial; the Jenkins documentation covers syntax well. It is about the decisions behind those files: why I started scripted, what shared libraries fixed, how I order quality gates, how blue-green rollback fits in, and when an old scripted job is worth migrating. Some of those decisions I would make differently today, and I will say which.

The snippets below are simplified. A complete, runnable version of the same patterns lives in a public repository, jenkins-declarative-pipelines: a tested shared library, a declarative Jenkinsfile with quality gates and a simulated blue-green deploy, all started with make up on Docker Compose.

2016: why I started scripted

The honest reason is that I did not know the declarative syntax yet, and scripted felt easier. In an internal review at the end of that year I noted, half-jokingly, that the build script was written in Groovy only because Jenkins Pipeline required it. Scripted pipeline is Groovy with a few Jenkins steps: open a node {}, split the work into stage() blocks, call sh. If you can write a loop and an if, you can ship a pipeline on day one. There is no grammar to learn and no rule about where a block may appear.

That freedom paid off at first. I could load a helper file, parse a JSON parameter into a map and generate stages from a list. In that first year the main product shipped around 1,000 production builds, about 2.7 a day. Parallelizing the end-to-end suite cut a production run from roughly 40–50 minutes to 6–8, and a development run from 90–100 minutes to 25–30. A small pre-checkout verification script cut checkout from about 4 minutes to 2 at most. For a small team, that speed mattered more than consistency.

The cost showed up later, when other people had to read those jobs. Every pipeline had its own shape. Error handling lived in try/catch blocks, and when I went back to my oldest jobs I found empty catch blocks that swallowed exceptions, so a step could fail and the build would carry on as if nothing happened. Notification logic was duplicated in the success path and in the catch block. It is a legacy pattern that no longer earns its keep, and the migration to declarative retires it: the post conditions take over both jobs. For the first years, "reuse" meant load-ing Groovy helper files from the same repository: better than copy-paste, but with no versioning and no clear boundary between helper and job.

None of this is a flaw of scripted syntax itself. It lets you do anything, and nothing pushes back.

Shared libraries as the backbone

The change that mattered most in both eras was moving reusable logic into shared libraries. The first payoff was the end of copy-paste: pipelines with similar steps stopped carrying their own copies of the same code, and each step got one clear responsibility. The best example runs in every pipeline: the summary posted to Slack after each build. The logic that sanitizes and formats that message lives once in the library, and each pipeline only calls the notification step with its own project details.

The layout that survived is one common library plus one library per product:

ci-common/                  # shared by most jobs
  vars/
    notifyBuild.groovy      # build summary to the chat channel
    runGate.groovy          # run a check, keep its report, fail with a clear message
    coverageGate.groovy
    loadEnvConfig.groovy
  src/com/example/ci/       # plain Groovy classes, when a step needs them
ci-payments/                # one per product, only what is specific to it
  vars/
    deployBlueGreen.groovy
    smokeCheck.groovy
Enter fullscreen mode Exit fullscreen mode

The common library is imported by dozens of jobs; around ten product libraries sit next to it. Each file in vars/ becomes a global step that a Jenkinsfile calls by name. A step is a call method that takes a map, so call sites read like configuration:

// vars/runGate.groovy
def call(Map args) {
    String name   = args.name      // e.g. 'Static analysis'
    String cmd    = args.command   // exits non-zero on violations
    String report = args.report    // machine-readable output, read later by the summary

    int status = sh(script: "${cmd} > '${report}'", returnStatus: true)
    if (status != 0) {
        error("${name} gate failed (exit code ${status}). Findings: ${report}")
    }
}
Enter fullscreen mode Exit fullscreen mode

Two rules decide what goes where. A library holds how: how we run a gate, notify, or switch a blue-green slot. The Jenkinsfile holds what: which gates this service runs, in which order, with which thresholds. Groovy repeated in two Jenkinsfiles belongs in a library; a value that differs per service never gets hardcoded in one.

The second rule is versioning, and I learned it late. My Jenkinsfiles load libraries by name only, @Library('ci-common') _, which means every job follows whatever default version is configured in Jenkins. That is convenient until a library change breaks a job nobody has touched in a year. Jenkins lets you pin a branch, tag or commit in the annotation:

@Library('ci-common@v3.4.0') _
Enter fullscreen mode Exit fullscreen mode

Pin production pipelines to a tag, let a staging or canary job track the main branch, and bump the tag in a reviewed change. A shared library is a dependency like any other, and an unpinned dependency is a deploy you did not schedule. Pinning is on my own list: it comes with the migration, job by job.

2021: going declarative

When I started the newer project in 2021, I wrote it declarative from day one. What changed was not the steps. sh, checkout, junit and the library calls are the same. What changed is that the file now has a fixed shape:

@Library('ci-common@v3.4.0') _

pipeline {
    agent { label 'linux' }

    options {
        timeout(time: 45, unit: 'MINUTES')
        timestamps()
        disableConcurrentBuilds()
    }

    parameters {
        choice(name: 'TARGET_ENV', choices: ['staging', 'production'])
        string(name: 'GIT_REF', defaultValue: 'main')
        booleanParam(name: 'ROLLBACK', defaultValue: false)
    }

    environment {
        APP_DIR = "${WORKSPACE}/app"
    }

    stages {
        stage('Static analysis') {
            when { not { expression { params.ROLLBACK } } }
            steps {
                runGate(name: 'Static analysis',
                        command: 'make static-analysis',
                        report: 'reports/static-analysis.json')
            }
        }
        // more gates, then deploy
    }

    post {
        always {
            junit allowEmptyResults: true, testResults: 'reports/junit/*.xml'
            notifyBuild(reportsDir: 'reports')   // result comes from currentBuild.currentResult
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

What I gained, in order of how much it mattered:

  • Readability for people who did not write it. A new engineer finds the timeout, the parameters and the failure handling in the same place in every job.
  • post instead of try/catch. always, success, failure, unstable, fixed, regression, aborted and cleanup replace the hand-written error paths. Notification moved out of catch blocks and could no longer be skipped by an exception thrown in the wrong place.
  • when instead of if. A stage skipped by when still appears in the stage view, marked as skipped, so the run tells you what it chose not to do.
  • options as policy. A timeout, timestamps and no concurrent deploys are one line each, and their absence is visible in review.
  • Validation before running. The declarative linter (the declarative-linter command in the Jenkins CLI, or a POST to the controller's pipeline-model-converter/validate endpoint) catches structural mistakes before a build starts. Scripted pipeline only tells you at runtime.

The escape hatch is script {}, which runs scripted Groovy inside a declarative step. My first declarative pipeline leaned on it far too much: most stages were a script {} block that returned early on a rollback run. It was scripted pipeline in declarative clothing, and the stage view showed those stages as green although they did nothing. The when condition above is what I would write today, and rewriting those stages is part of the same improvement plan. The Jenkins docs agree: when a script {} block grows, move it into a library step.

Quality gates before a blue-green deploy

The newer project's quality gates live in three places, and where each one runs turned out to matter as much as what it checks. They were added one at a time over several years, not designed up front, and the complexity gate is the most recent.

In the deploy pipeline, before anything is copied to a server:

  1. Code style (a formatter in dry-run mode)
  2. Static analysis (type and bug checks)
  3. Automated refactoring check (a refactoring tool in dry-run mode: if it would change the code, the code is not up to the agreed standard)
  4. Unit tests, with coverage collected
  5. Coverage threshold
  6. Complexity risk (CRAP)

In the application repository's own CI:

  1. Mutation testing

In separate jobs that run after the deploy:

  1. End-to-end tests against the deployed environment, plus a dedicated security test job

The order inside the deploy pipeline follows one rule: cheap and fast first. Style and static analysis need seconds and no database, so they fail a bad commit before anything expensive starts. Coverage and complexity read the report the unit tests produce, so they cost almost nothing extra.

The two most expensive checks moved out of the deploy pipeline, for the same reason. Mutation testing re-runs the suite against many small mutations of the code. It started as the last stage of the deploy pipeline and made every build wait, so I moved it out of Jenkins and into a pipeline in the application repository itself, on GitLab CI, where it runs without holding a release hostage. End-to-end tests went the same way: they used to sit inside the deploy pipeline, and now a separate job runs them after each deploy, next to a security test job. The gates still exist; they just run where their cost does not block a deploy. That is a lesson in itself: a gate that makes releases slow enough that people want to skip it is a gate in the wrong place, not a gate to delete.

The CRAP gate deserves a sentence of explanation, because coverage alone misleads. CRAP (Change Risk Anti-Patterns) combines cyclomatic complexity with coverage per method: CRAP = complexity² × (1 − coverage)³ + complexity. A simple method with no tests scores low; a complex method with poor coverage scores very high. The gate fails the build when a method crosses the threshold, which catches the case a global coverage number hides: 90% overall, and the one complex method that matters is untested. The gates earn their keep in a very human situation. More than once, a developer told QA something like "Relax, I already validated everything in this task, you can skip testing and ship it straight to production." We took the advice, started the build, and the gates failed, keeping broken code out of production. That is the point of a gate: a cheap layer of safety that does not depend on anyone's confidence, including mine.

Every gate follows the same pattern: write a machine-readable report, then decide pass or fail with a message a human can act on:

// vars/coverageGate.groovy (simplified; the repository parses the XML in a tested class)
def call(Map args) {
    BigDecimal min = args.min as BigDecimal          // e.g. 80
    String raw = sh(returnStdout: true,
        script: "xmllint --xpath 'string(/coverage/@line-rate)' '${args.report}'").trim()
    BigDecimal actual = (raw as BigDecimal) * 100

    echo "Line coverage: ${actual}% (minimum ${min}%)"
    if (actual < min) {
        error("Coverage gate: line coverage ${actual}% is below the ${min}% minimum. " +
              "Add tests or lower the threshold in a reviewed change.")
    }
}
Enter fullscreen mode Exit fullscreen mode

The report-first habit pays off in the summary. In post { always }, a library step reads each report that exists and builds one message for the team channel: style, static analysis and refactoring findings, tests run and failed, coverage, the top CRAP score, plus the commit, who started the build and the duration. When a gate fails, the message still shows what every earlier gate found, so nobody opens the console log to learn whether it was one style issue or forty.

Blue-green with rollback

Blue-green is older than my declarative pipelines. The first version shipped in 2017, in a scripted pipeline, inspired by Jez Humble and Dave Farley's Continuous Delivery. The goals I wrote down at the time still hold: faster builds and rollbacks, more control over what reaches users, and lower downtime risk. When I wrote the declarative pipeline in 2021, I reused the same pattern almost step for step. The technique outlived the syntax change, which is the point: blue-green and rollback are architecture, not a scripted-versus-declarative feature.

Production keeps two release directories, blue and green, and one pointer to the live one (a symlink, or the web server's root setting). A deploy goes like this:

  1. Read which color is live. The idle one is the target.
  2. Copy the release, already through the gates, into the idle directory.
  3. Run migrations, warm the cache and, ideally, smoke-check the idle side.
  4. Switch the symlink atomically to the idle directory and reload the PHP and web server processes.

Rollback is the same switch in reverse. Since 2017 it has been a build parameter: a rollback run skips tests, file changes and library updates, and points the live pointer back at the previous slot, which still holds the previous release. No new build and no copy, so it is the fastest run the pipeline has.

// vars/deployBlueGreen.groovy (simplified)
def call(Map args) {
    String live = sh(returnStdout: true,
        script: "basename \"\$(readlink -f '${args.currentLink}')\"").trim()   // 'blue' or 'green'
    String target = (live == 'blue') ? 'green' : 'blue'

    if (!args.rollback) {
        sh "rsync -a --delete --exclude=.env build/ '${args.releasesDir}/${target}/'"
        smokeCheck(slot: target)
    }
    // ln -sfn alone removes the old link before creating the new one;
    // building a temporary link and renaming it over the old one is atomic
    sh "ln -sfn '${args.releasesDir}/${target}' '${args.currentLink}.next' && " +
       "mv -Tf '${args.currentLink}.next' '${args.currentLink}'"
    sh 'sudo systemctl reload php-fpm nginx'
    echo "Live color is now ${target} (was ${live})"
}
Enter fullscreen mode Exit fullscreen mode

A detail worth knowing: ln -sfn on its own is not atomic. It deletes the old link and then creates the new one, so for a moment there is no live release. Creating the new link under a temporary name and renaming it with mv -T closes that window, because a rename replaces the old link in a single step.

Three caveats I would put in any design review. First, rollback goes back exactly one release; after two deploys, the old one is gone. Second, the symlink rolls back code, not data: migrations must be backward-compatible (expand first, contract in a later release), or the old code will meet a schema it does not understand. Third, keep one source of truth for which color is live. Reading the symlink, as above, beats a separate marker file that can drift from reality.

Environment config without secrets

The pipeline is the same for every environment; only the config changes. A parameter picks the environment, and a library step loads a small per-environment file with hosts, paths and switches. I have used Groovy files that return a map for this. Today I would use YAML read with readYaml, because config should be data, and data cannot execute code.

Secrets never live in those files, and never in the repository. Database passwords, API tokens and SSH keys belong in Jenkins Credentials (or a vault, with a Jenkins integration) and enter the build only for the step that needs them:

def cfg = loadEnvConfig(params.TARGET_ENV)   // hosts, paths, app env: no secrets

withCredentials([usernamePassword(credentialsId: "db-${params.TARGET_ENV}",
                                  usernameVariable: 'DB_USER',
                                  passwordVariable: 'DB_PASS')]) {
    sh './bin/migrate'   // reads DB_USER/DB_PASS from env
}
Enter fullscreen mode Exit fullscreen mode

Note the single quotes in sh. With double quotes, Groovy interpolates the secret into the command string; with single quotes, the shell reads it from the environment and Jenkins can mask it in the log. In declarative, credentials('id') inside environment {} does the same for a whole stage. One credential ID per environment keeps a staging job away from production secrets.

Scripted vs declarative: where each one is stronger

Both run on the same Pipeline engine. Declarative is parsed into the same CPS-transformed execution, so the Groovy sandbox, script approvals, serialization rules and @NonCPS limits apply to both, and so does everything inside a script {} block or a library step.

Aspect Scripted Declarative
Learning curve Low to start if you know Groovy; the hard parts (CPS, serialization, sandbox) show up later A fixed set of directives to learn; the same engine limits still apply underneath
Readability and onboarding Depends entirely on the author's discipline Same sections in the same order in every file
Structure enforcement None Required agent, stages, steps; directives only where allowed
Validation before running Groovy compile errors only, at build start Linter via CLI or HTTP endpoint, plus full structure validation before the first stage
Dynamic logic (loops, generated stages) Full Groovy: generate stages and parallel branches from data matrix, parallel and when cover common cases; the rest needs script {} or a library
Error handling try/catch/finally, explicit and easy to get wrong post conditions per pipeline and per stage; catchError and warnError work in both
Restart from stage Not available (Replay only) Restart any top-level stage of a completed run, with the same parameters and SCM revision
Tooling and visualization Stages appear in the stage views; conditionally skipped stages just disappear Same views, plus when-skipped stages shown as skipped; Directive Generator in the UI
Reuse via shared libraries Call any library step or class Call library steps in steps, or expose an entire pipeline as a template step (one declarative pipeline per build)
Testability JenkinsPipelineUnit JenkinsPipelineUnit (DeclarativePipelineTest); in both worlds, library steps are the easiest unit to test

The template option deserves a note. A vars/ step can contain a whole pipeline {} block, so a service's Jenkinsfile shrinks to one call. That is the strongest standardization tool Jenkins has, and the most rigid: an exception for one team becomes a parameter for all of them. I use it when the gate chain must be identical across teams, not as a default.

When to migrate, and how to do it safely

Migrating my scripted base to declarative is my own next step, so this section is the plan I am following, not a retrospective.

Signals that it is time:

  • People who did not write the pipelines now maintain them, and every job reads differently.
  • The same blocks are copied between jobs, and a fix has to be applied in several places.
  • You need restart-from-stage or clearer visualization to recover from flaky infrastructure without re-running everything.
  • You want the same gates across teams, enforced, not suggested.
  • An audit needs to see, quickly, what runs before production.

Reasons not to:

  • The pipeline generates its stages from data in loops. Declarative will push most of it into script {}, and you gain little.
  • The job is stable, rarely changes and nobody touches it. Migration risk with no payoff.
  • You have no time to test the new job against the old one. An untested migration of a deploy pipeline is a production incident on a timer.

How to do it safely:

  1. Inventory and classify. List every job, when it last ran, who owns it, whether it deploys, and how many jobs share its pattern. Delete what is dead first.
  2. Move logic into library steps first. While the job is still scripted, extract the shared blocks into vars/ steps and switch the scripted job to call them. This is the riskiest change, and it happens with the old shape still in place.
  3. Migrate one job per sprint, starting with the most-copied pattern. The first migration becomes the template for the next ten.
  4. Run old and new side by side. Point the new job at a non-production target, or run it with deploy steps disabled, and compare results for a few cycles.
  5. Keep script {} for the genuinely hard parts, and move each of those into a library step later. Do not block the migration on making it pure.
  6. Add a linter step to the library repository and to the pipeline repository, so a broken Jenkinsfile fails review, not the deploy.
  7. Freeze, then delete. Disable the old job, keep it for one release cycle, then remove it. Two live versions of the same deploy is how drift starts.

Takeaways

  1. Start with whatever lets you ship, but expect to pay for structure later. Scripted made me fast in 2016 and expensive to read years later.
  2. Put the how in shared libraries and the what in the Jenkinsfile, and pin library versions in production pipelines.
  3. Order quality gates from cheap to expensive, write a report before deciding pass or fail, and summarize every gate in one message, even on failure.
  4. Blue-green and rollback are architecture, not syntax: the same pattern survived the move from scripted to declarative. Keep migrations backward-compatible, or rollback only rolls back half the system.
  5. Migrate scripted to declarative when people, copy-paste or compliance demand it, one job at a time, logic into libraries first.

Tayguara Dias Reis is a Tech Lead with 14+ years in software quality and ISTQB CTFL certification. He writes about test automation and CI at dev.to/tayguara. The code from this article runs end to end in jenkins-declarative-pipelines. For the GitHub Actions side of the same ideas, see the public playwright-qa-showcase repository; the next article in this series covers it.

Top comments (0)