DEV Community

My AWS Bill Tells a Story. I Built a Pipeline to Understand It.

An AWS bill can go up for an excellent reason.

You enable audit logging. Improve threat detection. Strengthen backups. The following month, the spending curve climbs.

The chart says: “This costs more.”

It does not say: “We bought visibility and reduced a risk.”

Conversely, a stable bill can hide a filling disk, an expiring reservation, or an alarm so noisy that nobody pays attention anymore.

That is the problem I wanted to address by building a monthly AWS reporting pipeline. Not another dashboard. A document connecting what the infrastructure costs, what it is doing, and which decisions it calls for.

The stack is deliberately straightforward: Bash, AWS CLI, Python, JSON, HTML, and PDF. AI helps draft the commentary; it does not produce the numbers.

The principle: automate the facts, assist the analysis, and keep ownership of the conclusions.

A dashboard shows. A report should help you decide.

For this review, I needed to answer three questions:

  1. What changed since the previous review?
  2. What needs intervention?
  3. Which decision needs to be made now?

Those questions span several AWS consoles. Cost Explorer knows the spending. CloudWatch knows the metrics. Inventory APIs know the resources. The business reasons for a change often live elsewhere: in an incident, a planned action, or a conversation.

The report had to bring those pieces together without turning a hypothesis into a certainty.

So I separated the system into three layers: collect, interpret, present.

Reporting architecture: AWS APIs feed a JSON snapshot and a Python renderer produces HTML and PDF. A separate branch extracts facts for AI drafting; human review supplies editorial JSON.

Numbers and commentary meet at rendering time, not during collection. The code below consists of simplified, anonymized excerpts—not a complete tool you can install.

1. Start with a snapshot, not a PDF

The pipeline’s first output is not a document. It is a JSON snapshot.

The collector queries sources including:

Source What I use it for
Cost Explorer Spending by service and comparison with a reference period
CloudWatch Cache metrics, S3 storage, SQS queues, and alarm history
EC2, RDS, ElastiCache, and ELB APIs Resource inventory and characteristics
AWS Backup Backup jobs, failures, and restores
IAM Credential Report MFA status, access key age, and inactivity signals
Systems Manager Filesystem utilization and software versions on reachable instances

The JSON also records the period, collection timestamp, and warnings encountered.

Here is the kind of query used for costs. The dates are illustrative; Cost Explorer treats End as exclusive.

export AWS_PROFILE="reporting"  # Fictional local profile

aws ce get-cost-and-usage \
  --time-period Start=2026-08-01,End=2026-09-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE \
  --output json
Enter fullscreen mode Exit fullscreen mode

The metric matters: UnblendedCost is not an amortized allocation of commitment costs. Before interpreting a number, you need to know what it measures.

This intermediate step has a practical benefit: I can change the layout without collecting the data again.

Fixing a paragraph should not trigger another exploration of the AWS account. And another collection several weeks later may not return the same information: resources change, and available history depends on the service.

The collector therefore refuses to replace an existing snapshot unless explicitly instructed. The renderer does not rewrite that JSON.

This is not cryptographic immutability. It is an operational guardrail: do not silently replace the evidence behind a review.

An important nuance: not everything is historical

Costs and metrics are queried over a time window. The inventory primarily describes the state at collection time.

A monthly report is therefore not an exhaustive record of every resource that existed during the month. An instance created and deleted between collections can incur costs without appearing in the final inventory.

That distinction deserves to be made clear. Otherwise, a photograph starts passing for a film.

2. Two JSON inputs: facts on one side, context on the other

Alongside the AWS snapshot sits a second JSON input, specific to the month: the editorial data.

It contains section commentary, incidents, follow-up actions, and manually entered deadlines.

Why not combine everything?

Because “this service’s cost increased” and “this increase corresponds to an approved decision” have different sources.

The first statement comes from a measurement. The second requires context.

The renderer combines both inputs to produce the report. It refuses to run when the expected editorial input is missing. When preparing a new month, a helper carries forward the previous tracking information but clears the commentary: actions may continue; conclusions must be reviewed.

That small detail prevents a particularly quiet failure: a beautifully written paragraph that was only true last month.

Another useful rule: for persistent action items, prefer a date over “for the past three months.” A date ages honestly. A copied relative duration becomes wrong without making a sound.

Here is a deliberately minimal editorial input. The comment is fictional; other sections are omitted. The French keys reflect the actual tool’s input contract.

{
  "mois": "2026-08",
  "commentaire_couts": [
    "An increase was detected in storage costs: the cause needs investigation."
  ],
  "actions": [],
  "incidents": [],
  "expirations_manuelles": []
}
Enter fullscreen mode Exit fullscreen mode

A Makefile provides a common entry point for the workflow:

make new-input MONTH=2026-08
make collect MONTH=2026-08
# Extract facts, draft the commentary, and review it.
make report MONTH=2026-08
Enter fullscreen mode Exit fullscreen mode

These targets belong to my tooling, not AWS CLI. The important part is the pause before report: generating a file is not the same as validating its content.

3. The mid-month trap

Imagine a review partway through the month.

Entirely fictional example: spending is $420 for the first ten days, compared with $1,200 for the whole previous month.

A quick calculation announces a 65% decrease. Arithmetically correct. Operationally misleading.

If the first ten days of the previous month cost $400, an equal-duration comparison tells a different story: a 5% increase.

Fictional example: comparing $420 over ten days with $1,200 over a full month gives minus 65%, while comparing with $400 over the previous first ten days gives plus 5%.

The same spending tells two stories depending on the reference. Every amount in this illustration is fictional.

The calculation is simple. Choosing the two values is the real challenge:

# Standalone example; every amount is fictional.
current_first_10_days = 420.0
previous_full_month = 1200.0
previous_first_10_days = 400.0

def variation_pct(current, reference):
    if reference == 0:
        return None  # Undefined percentage: handle the new cost separately.
    return (current - reference) / reference * 100

print(variation_pct(current_first_10_days, previous_full_month))     # -65.0
print(variation_pct(current_first_10_days, previous_first_10_days))  # 5.0
Enter fullscreen mode Exit fullscreen mode

The pipeline detects unfinished months and produces an interim report:

  • The covered period and number of days are shown.
  • The reference is a query over the first days of the previous month, not its full total.
  • A linear end-of-month projection is provided.
  • Output names distinguish interim reports from final ones.

The reference is not simply the previous total divided by thirty. It uses observed spending over a comparable window.

But “same number of days” does not mean “same activity.” Weekends, batch processing, and seasonality can still affect the interpretation. A linear projection is an extrapolation, not a promise.

The document should expose those limitations, not bury them behind a percentage.

4. AI holds the pen. It does not hold the calculator.

AI is not integrated into AWS collection or the rendering engine. It operates in a separate commentary-drafting workflow, implemented as a skill for the development assistant.

Before drafting, a script extracts relevant facts:

  • Current and reference costs.
  • Calculated changes.
  • New services and significant increases.
  • Operational metrics.
  • Inventory changes.
  • Commentary from the previous review.

The script reuses the renderer’s section registry. The structure given to the assistant stays aligned with the report.

Writing instructions are explicit: start with a quantified summary, then add a few useful observations; every number must come from the extracted facts; recommendations must remain recognizable as recommendations.

The most important rule fits into one sentence:

If the data does not explain an increase, write “cause needs investigation” instead of inventing a plausible explanation.

A lightly used cache becoming more expensive does not prove that snapshots caused the increase. That may be a lead, but it is not a conclusion.

The workflow also includes reading previous commentary. Not to recycle it, but to distinguish a new problem from an existing issue under follow-up.

Proposed commentary still needs review before distribution. The pipeline does not automatically contact an LLM on every run: a report can be rendered from previously approved comments without AI involvement.

What I want from the assistant: clearer wording. What I do not delegate: the truth of the report.

5. Missing data is not good news

Suppose a metric was not retrieved.

Is it zero? A permission issue? An unreachable instance? A publication delay?

Those situations should not carry the same meaning in a review.

The collector applies timeouts and records warnings when queries fail. The report includes a section exposing them. Additional checks flag certain missing data.

The mechanism has a limitation: fallback values still exist in the processing. A displayed zero does not automatically become a reliable measurement. Warnings must be read alongside the tables, and the representation of unavailable values needs continued improvement.

That caution matters for S3, where size metrics are not instantaneous, and for instances unavailable through Systems Manager.

A useful report does not only say what it knows. It shows what it failed to verify.

6. Remove SSH from reporting, without pretending to remove risk

To retrieve disk utilization and selected software versions, I use SSM Run Command instead of an SSH connection.

That removes this workflow’s dependency on SSH keys and direct SSH access to machines. Collection targets Systems Manager-managed instances with online agents. The others are flagged.

An essential security point remains: ssm:SendCommand is not a read-only permission.

Even when the commands used here are diagnostic, the ability to run commands remotely needs boundaries: a dedicated role, authorized resources, permitted documents, and traceability.

Moving from SSH to SSM changes where access is controlled. It does not make access control unnecessary.

The same care applies to network collection: some AWS responses can contain secrets. For VPN connections, the query explicitly selects the fields needed for the report instead of retrieving and displaying the entire response.

Minimizing data at the source beats relying on deletion afterward.

7. The forgotten service should still appear

The renderer uses a section registry: each section declares cost keywords and a rendering function.

This connects a storage expense, for example, with the corresponding resources and observations.

But an AWS service can appear without a dedicated section. I did not want it to disappear simply because the report did not yet know where to put it.

Uncovered expenses therefore go into “Other billed services,” with a notice at the end of generation.

The system also uses simple rules to highlight new services and significant increases. Thresholds combine relative changes and absolute amounts: a spectacular percentage applied to a few cents may not deserve attention.

Here is a simplified version of the rules, using the generator’s current thresholds:

def cost_alert(current, previous):
    if current >= 5.0 and previous < 1.0:
        return "new service"

    if previous > 0:
        delta = current - previous
        pct = delta / previous * 100
        if pct >= 50.0 and delta >= 10.0:
            return "large increase"

    return None
Enter fullscreen mode Exit fullscreen mode

These thresholds are triage choices, not universal laws. They should match the size of the bill and the desired level of detail.

This is not a predictive model. It is an explainable filter identifying the lines worth reviewing.

The goal is not to earn a red badge. It is to reach a useful question: expected cost, changed scope, or anomaly to investigate?

8. A report without another application to maintain

The final output is self-contained HTML: inline CSS, SVG charts, no JavaScript, and no charting library to load. Conversion with wkhtmltopdf produces the PDF.

That choice fits the need: a periodic review that can be read offline and archived without deploying another portal.

It comes with trade-offs. A static document does not replace interactive exploration. And wkhtmltopdf relies on an old rendering engine: its use here does not make it a universal recommendation for a new project.

Easy distribution must not be confused with harmless content. Infrastructure reports can expose identities, addresses, and versions. Distribution remains a security concern in its own right.

9. Test the document, not just the script

A generator can run successfully and still produce a bad report.

Regression tests therefore compare generated HTML with expected outputs, using fixtures covering a full month and a partial month.

To make those comparisons stable, tests fix the dates and disable AWS refresh calls. They also check refusal of missing editorial input and HTML escaping of comments.

One nuance prevents a promise of perfect reproducibility in production: at rendering time, RDS and ElastiCache reservations can be refreshed in memory to show current expiration information. The original JSON remains untouched, but two renders at different times may differ.

That means distinguishing reproducing a historical review from updating information relevant to a decision today. Both are legitimate; they are not the same contract.

These tests also do not prove that every API responded or that a recommendation is sound. They protect the rendering. Review protects the interpretation.

What I would keep if I started again

Not necessarily Bash. Not necessarily the PDF converter. Maybe not even the same data format.

But I would keep these five decisions:

  1. Save the facts before presenting them.
  2. Separate observations from context and recommendations.
  3. Compare explicit periods, especially during an unfinished month.
  4. Expose failures and unclassified services.
  5. Give AI verifiable facts and a limited role.

I am not claiming to have replaced a FinOps platform or a security audit. I built a pipeline for a narrower need: making a monthly review repeatable, readable, and useful for discussion.

The most valuable next improvement is not “more AI,” either. It is more precision around data provenance, unavailable values, and the conditions required to reproduce an old report.

The bill was only the beginning of the story.

The real deliverable is being able to say: here is what changed, here is what we know, here is what still needs verification—and here is the decision to make.


What question remains unanswered in your cloud reviews, even when all the metrics seem to be available?

Top comments (0)