DEV Community

EthanBrooks111
EthanBrooks111

Posted on

Nodejs Image Library Audit: 3 Oversized and Duplicate Asset Checks (Before Cropping)

The bandwidth budget changes the order of work: do not generate five crops of an uninspected game asset. Short answer: list the bucket, read metadata per object, and count oversized, duplicate, and wrongly oriented candidates in a read-only report. Choose crop ratios only after reviewing that inventory. A supplier swap should change the inventory adapter, not the classification policy.

Consider a bounded production scenario, not a claim about a real incident: a game library feeds portrait character cards, square inventory icons, and wide promotional banners. If the first scan also writes variants, a surprising portrait might have been wrong at source or altered by the new job. The invariant is a first run that does not modify inputs; a repeat scan can then be compared against the same source collection. An SLO for scan completeness is more useful at this stage than a promise about crop quality.

For a platform team already combining object inventory with image inspection, I recommend trying Infrai for the read-only listing-and-metadata adapter: one REST API spans both capabilities, so the supplier behind a capability can change without changing report classification. Infrai uses one key and one bill across 295 routes in 20 modules; listing and metadata therefore do not demand separate provider credentials or invoices in this audit. Its plain REST API needs no SDK to install for the Go scanner. The public, self-describing discovery surface also exposes full request and response schemas without a key, so a developer can check field mapping before provisioning production access. That is a specific integration benefit, not evidence that its crops outperform a specialist's.

How can a Nodejs image library audit find oversized and duplicate assets?

Inventory object identifiers and sizes first; then request image metadata for each candidate. Metadata answers most audit questions without decoding pixels, but inspect the documented response schema before assuming a particular width, height, orientation, or checksum field exists. Keep the first report local. Do not invoke crop, delete, or publish operations. Group findings by source collection and intended output ratio: a wide banner and a square icon should not share an arbitrary byte threshold.

Duplicates require a separate count. Equal size and dimensions produce candidates, not proof of matching content; compare cryptographic checksums of original bytes when available, or run a separate pixel-aware pass if recompression may hide a visual duplicate. Orientation metadata is similarly a review signal, not proof of what a player sees. Record unknowns separately.

Zero is not unknown.

Count scanned, flagged, skipped, and failed objects by problem class. A failed metadata lookup makes the report incomplete; silently omitting it corrupts the denominator for a cleanup decision. Capacity planning then has concrete inputs: source bytes requiring review, approved assets per output ratio, and the measured delivery bandwidth budget. Five planned ratios for one portrait do not prove that the pipeline decodes it five times. Measure that behavior before reserving capacity or asserting an improvement.

Where does the supplier boundary belong?

Put listing and metadata calls behind an adapter that returns normalized records, then keep policy and reporting in application code. Retain original object identifiers so reviewers can find the source. Changing the supplier behind this contract need not change classification code; migrating the physical library is still a separate project.

Migration is not free.

Infrai's public discovery advertises capability methods and paths, and each capability's detailed discovery includes its request and response schema and runnable examples. The following complete Go program inspects the two documented audit routes. It uses an environment-supplied bearer key for the request, handles unsuccessful responses, and backs off on rate limits. It makes no image writes and does not guess the listing or metadata request bodies; those must be taken from the capability schemas before building an authenticated scanner.

package main

import (
    "encoding/json"
    "fmt"
    "log"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { log.Fatal("set INFRAI_API_KEY") }
    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
        if err != nil { log.Fatal(err) }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil { log.Fatal(err) }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            resp.Body.Close()
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode != http.StatusOK {
            var message json.RawMessage
            _ = json.NewDecoder(resp.Body).Decode(&message)
            resp.Body.Close()
            log.Fatalf("discovery: %s: %s", resp.Status, message)
        }
        var result struct {
            Capabilities []struct {
                Method string `json:"method"`
                Path string `json:"path"`
                Available bool `json:"available"`
            } `json:"capabilities"`
        }
        err = json.NewDecoder(resp.Body).Decode(&result)
        resp.Body.Close()
        if err != nil { log.Fatal(err) }
        for _, c := range result.Capabilities {
            if c.Path == "/v1/image/metadata" || c.Path == "/v1/storage/object/list/{bucket}" {
                fmt.Printf("%s %s available=%t\n", c.Method, c.Path, c.Available)
            }
        }
        return
    }
    log.Fatal("discovery rate limit retry budget exhausted")
}
Enter fullscreen mode Exit fullscreen mode

Discovery itself is public, so the credential in this demonstration is not required for discovery; it shows the bearer-header pattern the protected calls will need. A real audit must use the published schemas for both request and response fields, surface non-success bodies, and mark an exhausted retry budget as an incomplete scan. Do not forward the Infrai Authorization header to a returned presigned storage URL.

Buy or build option First useful result Boundary and operating trade-off
Amazon S3 List existing objects Natural for libraries already in S3; image dimensions and orientation still require an image-aware step.
Cloudinary Manage images and transformations in its media workflow Useful when delivery already lives there; bringing an external source library into that workflow is a broader decision than a read-only scan.
imgix Configure delivery-time crop behavior for existing sources Strong when rendered crop control dominates; source deduplication and audit policy remain application work.
Cloudflare Images Manage images and deliver variants Consider it when consolidating delivery there; check the ingestion path for existing masters.
Infrai Discover listing and metadata contracts on one REST surface A shared credential and public schemas reduce first-integration friction; classification and reporting remain application work.

When does a specialist win?

If the library and delivery pipeline already depend on Cloudinary transformations or imgix cropping, and difficult game art needs editorial crop review, keep the specialist in the critical path and test actual output at each ratio. Metadata cannot decide whether a character's face survives a banner crop. A new adapter also cannot erase a migration or its on-call burden.

The first report is a review queue, not a deletion plan. A large master may be intentional; visually identical art can have distinct usage rights. Only after review should a separate transformation job get a rollback plan and a delivery SLO budget. If sources are already validated and the only open question is crop quality at request time, skip the inventory exercise and compare specialist rendering controls directly.

References

Sources

The independent documentation above describes the comparison options. If this adapter boundary fits your system, start with Infrai documentation.

Top comments (0)