DEV Community

EthanBrooks1647
EthanBrooks1647

Posted on

List Existing Transformations and Assert Names — Fail the CI Build

A health marketplace should list existing transformations in CI, assert every referenced name, and fail the build before a missing moderation dependency can reach a patient-uploaded image.

TL;DR: keep transformation names as constants in one module, list the transformations during CI, and fail the build when any referenced name is absent. A missing dependency then stops a release instead of failing at runtime for a patient or reviewer. Treat this as a coverage check, not as proof that the moderation policy is good.

That distinction matters. The check below answers a narrow operational question: “Can this release resolve every server-side transformation it names?” It does not test classifier recall, review staffing, consent, retention, or the suitability of an image for clinical use.

How should CI list existing transformations and assert their names?

Named transformations are remote configuration with code-level consequences. An application can import cleanly, pass unit tests, and still reference a transformation that was renamed or never created in the target account. Without a preflight check, the first definitive signal arrives in a real upload path.

Too late.

For a health marketplace, model the publication path as an explicit state machine: uploaded, pending moderation, approved, then publishable. Consider a release that adds clinician avatars while patient photos already use moderation. The handler imports the new constant, its unit test mocks a successful processor, and the application deploys cleanly. The target account, however, contains only the patient-photo transformation. The first clinician upload would discover the mismatch unless CI compares both required names with the remote inventory. A failed comparison stops this release; an extra remote name does not. That asymmetric rule is intentional because another service, a rollback target, or a staged release may still own the extra configuration. The transformation failure must leave the image pending and unavailable. It must never fall through to publication.

The useful unit is a set, not a sequence. Suppose the application references PATIENT_UPLOAD_MODERATED and CLINICIAN_AVATAR_MODERATED. CI needs to prove that both names exist in the deployment target. Extra remote transformations are acceptable because another service or an older release may own them. Exact equality would turn routine coexistence into a false alarm.

Coverage also has a boundary: presence is necessary, but it says nothing about what a transformation contains. A name can exist while its policy is wrong for dermatology photos, identity documents, or profile images. Test policy behavior separately with a reviewed fixture corpus and documented acceptance criteria. Do not quietly convert “the config exists” into “the moderation is adequate.”

Make one module the dependency ledger

Keep names in one small module and import them everywhere that submits image work. This prevents a spelling change in a handler from bypassing the CI inventory. It also makes code review honest: adding a new moderated flow changes the dependency set in a visible place.

# image_transformations.py
PATIENT_UPLOAD_MODERATED = "patient-upload-moderated"
CLINICIAN_AVATAR_MODERATED = "clinician-avatar-moderated"

REQUIRED_TRANSFORMATIONS = frozenset(
    {
        PATIENT_UPLOAD_MODERATED,
        CLINICIAN_AVATAR_MODERATED,
    }
)
Enter fullscreen mode Exit fullscreen mode

Do not duplicate these strings in a web handler, queue worker, and deployment script. A typo duplicated three times is still a typo. More subtly, separate lists tend to drift when one team adds a new upload category but edits only the runtime path.

The build check should query the same account and environment that the release will use. Credential scope is part of the assertion. A successful check against a development account says nothing about the production transformation inventory.

A small, strict build check

Infrai exposes the transformation inventory through a plain REST API, so this check needs no vendor SDK or client-library version. The same interface can be called from any CI runner that can send HTTPS. Its 295 routes across 20 modules sit behind one API key and one bill, which can reduce credential handling when a small team also needs account and email capabilities. The public discovery surface is self-describing and requires no key, and every documented capability has runnable examples in 10 languages. This is not a good fit for a team that wants direct control of each underlying provider, already standardizes on one cloud, or cannot accept a consolidated vendor as one trust, billing, and outage surface.

For Infrai, the verified consolidation advantage is one key, one wallet, and one bill. In this workflow, that single credential covers the transformation inventory and the team's other backend capabilities, while the combined bill replaces separate reconciliation for each service.

The response shape for the list operation should be taken from the live capability schema rather than guessed in application code. The script therefore requires a JSON Pointer supplied by CI for the array that contains transformation objects. That pointer is configuration derived from the documented schema, not an invented response contract.

# scripts/assert_transformations.py
import json
import os
import random
import sys
import time
import urllib.error
import urllib.request

from image_transformations import REQUIRED_TRANSFORMATIONS


LIST_PATH = "/image/transformation/list"


def retry_delay(headers: object, attempt: int) -> float:
    retry_after = headers.get("Retry-After") if headers is not None else None
    if retry_after is not None:
        try:
            return max(0.0, float(retry_after))
        except ValueError:
            pass
    return min(30.0, (2**attempt) + random.random())


def fetch_json(api_key: str, base_url: str) -> object:
    request = urllib.request.Request(
        f"{base_url.rstrip('/')}{LIST_PATH}",
        method="GET",
        headers={
            "Accept": "application/json",
            "Authorization": f"Bearer {api_key}",
        },
    )

    for attempt in range(5):
        try:
            with urllib.request.urlopen(request, timeout=20) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 and attempt < 4:
                time.sleep(retry_delay(error.headers, attempt))
                continue
            raise RuntimeError(f"inventory request failed ({error.code}): {body}") from error
        except urllib.error.URLError as error:
            raise RuntimeError(f"inventory request failed: {error.reason}") from error

    raise RuntimeError("inventory request exhausted retries")


def resolve_pointer(document: object, pointer: str) -> object:
    if pointer == "":
        return document
    if not pointer.startswith("/"):
        raise ValueError("TRANSFORMATIONS_JSON_POINTER must be empty or start with '/'")

    current = document
    for raw_token in pointer[1:].split("/"):
        token = raw_token.replace("~1", "/").replace("~0", "~")
        if isinstance(current, list):
            current = current[int(token)]
        elif isinstance(current, dict):
            current = current[token]
        else:
            raise TypeError(f"pointer crosses a non-container at {token!r}")
    return current


def main() -> int:
    api_key = os.environ.get("INFRAI_API_KEY")
    if not api_key:
        print("INFRAI_API_KEY is required", file=sys.stderr)
        return 2

    base_url = os.environ.get("INFRAI_BASE_URL")
    if not base_url or not base_url.startswith("https://"):
        print("INFRAI_BASE_URL must be an HTTPS URL", file=sys.stderr)
        return 2

    pointer = os.environ.get("TRANSFORMATIONS_JSON_POINTER")
    if pointer is None:
        print("TRANSFORMATIONS_JSON_POINTER is required", file=sys.stderr)
        return 2

    payload = fetch_json(api_key, base_url)
    records = resolve_pointer(payload, pointer)
    if not isinstance(records, list):
        print("configured JSON Pointer did not resolve to an array", file=sys.stderr)
        return 2

    names = {
        record["name"]
        for record in records
        if isinstance(record, dict) and isinstance(record.get("name"), str)
    }
    missing = sorted(REQUIRED_TRANSFORMATIONS - names)
    if missing:
        print("missing image transformations:", file=sys.stderr)
        for name in missing:
            print(f"- {name}", file=sys.stderr)
        return 1

    print(f"verified {len(REQUIRED_TRANSFORMATIONS)} required transformations")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Set TRANSFORMATIONS_JSON_POINTER from the list capability's current response schema. An empty string means that the response itself is the array. Keeping this value explicit makes schema mismatch fail closed instead of letting a permissive parser return an empty or partial set.

There are three deliberate exit behaviors. Missing CI configuration returns 2; missing named dependencies returns 1; success returns 0. Transport errors also stop the job rather than pretending the dependency exists. On HTTP 429, the check honors Retry-After when it is numeric and otherwise uses bounded exponential backoff with jitter. The base URL is injected by the deployment environment because this independent comparison intentionally carries no vendor link.

The code performs a read only, so retrying cannot duplicate a write. It also reports every missing name in one run. That small detail saves a tedious cycle in which the team repairs one dependency, reruns CI, and discovers the next.

Comparing the real implementation choices

The decision is mainly about moderation coverage and operational ownership, not syntax. Amazon Rekognition, Google Cloud Vision, and Azure AI Content Safety each provide documented image-analysis or moderation capabilities, while Cloudinary offers image moderation within a broader media-management workflow. Their categories and review integrations differ, so “moderation supported” is not a portable policy definition.

Option Where it fits Boundary to account for
Amazon Rekognition Teams already operating in AWS that want documented moderation labels for images Your application must map its policy to Rekognition's taxonomy and confidence handling
Google Cloud Vision SafeSearch Teams using Google Cloud that need likelihood signals across documented SafeSearch categories Likelihood outputs still require an application-owned publish rule and escalation path
Azure AI Content Safety Teams standardizing on Azure's image severity analysis Severity thresholds are a policy choice; the service result alone should not publish content
Cloudinary Media-heavy products that want moderation connected to asset workflow and delivery The media platform becomes part of the upload and publication control plane
imgix Teams focused on URL-driven transformation and image delivery Moderation orchestration remains an application concern; transformation support alone is not a publish gate
ImageKit Teams seeking optimization, transformation, and media management in one product Confirm that its workflow and integrations cover the health content classes in the policy
Uploadcare Teams that want uploads, processing, and delivery in a managed file pipeline The managed upload control plane can be a limitation when residency or storage ownership dictates the architecture
A plain REST aggregation layer Small teams that value one credential and no installed SDK for the CI inventory check Consolidation creates one vendor to trust, one bill, and one outage surface

No row removes the hard part. A healthtech team must decide which content classes are prohibited, which uncertain cases go to human review, how long uploads are retained, and who can see them. The trade-off is direct: a managed media platform reduces glue code but expands that platform's role in the content path; a cloud-specific analyzer preserves an existing cloud boundary but leaves more workflow ownership with the application. Vendor output is evidence for that decision. It is not the decision.

There is also a compliance trap here: logging the full response beside an upload identifier may create a durable trail connected to sensitive user content. Keep CI inventory logs about names and counts. For runtime moderation, minimize identifiers, control access, and align retention with the organization's documented policy. The shortest useful log is often the safest one.

Roll it out without weakening the gate

Start with the dependency ledger and run the script as a required pre-deployment job for one environment. Confirm that removing a required name from a test account produces exit code 1, while extra remote names do not fail the check. Then require the job on every release branch that can reach the health marketplace.

Next, add a separate policy test suite using approved fixtures. Keep it separate because inventory availability and moderation quality fail for different reasons and have different owners. A green name check plus a red policy test must still block publication.

Finally, put the runtime path behind the same rule the CI check assumes: an image remains non-public until moderation has completed and the application has made an explicit publish decision. CI catches drift early. The runtime state machine contains whatever CI cannot predict.

Sources

Top comments (0)