TL;DR Serverless cost optimizations fail when engineers apply intuitions built for VM-based pricing to billing models that charge on compute-time multiplied by memory, or on data scanned
Quick Answer (TL;DR)
Serverless cost optimizations fail when engineers apply intuitions built for VM-based pricing to billing models that charge on compute-time multiplied by memory, or on data scanned rather than data returned. Reducing Lambda memory allocation extends execution duration, and the duration increase frequently outpaces the memory savings, producing a higher bill. A DynamoDB SCAN with a Limit parameter still bills against every row scanned, not every row returned. The fix is to learn the provider's billing unit before tuning any resource parameter.
Why this happens
The root cause is a mismatch between the billing unit a service charges against and the resource parameter an engineer chooses to tune.
Compute-time pricing. AWS Lambda bills on duration multiplied by allocated memory, measured in gigabyte-seconds. Cutting memory lowers one factor in that product, but longer execution raises the other. Because compute-intensive functions finish proportionally faster when given more memory, the duration drop frequently exceeds the memory cost increase. The net effect is a higher invoice, not a lower one.
Scanned-data pricing. DynamoDB SCAN operations bill against every item the engine reads internally, regardless of how many items cross the network to the caller. Adding a Limit parameter stops the result set at a threshold the caller sees. It does not stop the read capacity units from accumulating on rows the engine touched before applying that filter. The billing meter runs on engine traversal, not on payload delivery.
Both failure modes share one mechanism: the engineer tuned a visible output parameter while the billing meter sat on an invisible internal operation. In our testing on a high-frequency event processor, we measured this directly after 30 days of data. Memory reductions that looked like savings on a spreadsheet produced larger charges on the invoice because duration extended faster than memory cost fell. The billing model charged for work done inside the runtime, not work delivered to the application.
The correctable assumption is that reducing a resource parameter always reduces the resource consumed. For services priced on internal compute work or internal read traversal, that assumption is structurally wrong. Identify which internal operation the provider meters before touching any configuration value.
Fix #1: most common
The fix for Lambda cost overruns is raising memory allocation, not cutting it, and the subcommand that reveals whether your current setting is wrong is get-function-configuration.
Why less memory costs more
The field that matters. get-function-configuration returns the MemorySize field. That number, combined with the Duration field from CloudWatch Logs Insights, gives you the two inputs you need to compute actual gigabyte-seconds billed. Run the query against your own function name. Without both numbers on the same workload, any tuning decision is a guess.
The inversion mechanism. Lambda prices each invocation on duration multiplied by allocated memory. Reduce MemorySize and the function's CPU allocation drops proportionally, because Lambda ties CPU to memory. A compute-bound function then runs longer. In our testing on a high-frequency event processor, the duration increase outpaced the memory reduction, and the invoice rose.
The billing meter rewards speed, not frugality.
The per-invocation math
The trap most answers omit. Nearly every cost guide says "right-size your Lambda memory downward." None of them walk through the per-invocation math. The correct test is: multiply current MemorySize in GB by median Duration in seconds to get a baseline gigabyte-second figure. Then increase MemorySize by one step, measure the new duration after at least 50 invocations have stabilized, and recompute. If the product falls, the higher memory setting is cheaper.
This works for CPU-bound functions. It breaks for I/O-bound functions waiting on a downstream database, because added CPU does not compress wait time.
| Step | What to measure | Field or metric |
|---|---|---|
| 1. Baseline | Allocated memory |
MemorySize from get-function-configuration
|
| 2. Baseline | Median execution time |
Duration from CloudWatch Logs Insights |
| 3. After tuning | Recomputed GB-seconds |
MemorySize x Duration product |
| 4. Decision gate | Did the product decrease? | Lower product means lower bill |
Where to start first
Start with the function that runs most frequently. By sprint 3 of any rightsizing effort, the high-invocation functions dominate the invoice, and a single correct memory increase on one function outweighs a dozen correct reductions on rarely-called ones.
Fix #2: alternative
The fix for DynamoDB cost overruns is replacing SCAN operations with targeted Query operations backed by the right index, and the operation that exposes whether your current access pattern is wrong is describe-table.
Why Limit doesn't reduce cost
The field that matters. describe-table returns the GlobalSecondaryIndexes field. That structure shows every GSI already provisioned on your table, including its KeySchema and Projection. If the attribute your application filters on is absent from every key schema listed, every request touching that attribute runs a full SCAN internally. The billing meter has no knowledge of your Limit parameter.
It counts read capacity units against every item the engine touches before returning results.
The inversion mechanism. DynamoDB bills SCAN operations on data read internally by the storage engine, not on data delivered to the caller. A Limit parameter instructs the engine to stop returning items after a threshold. It does not instruct the engine to stop reading. On a table with 500,000 items, a SCAN with Limit set to 20 still consumes read capacity units proportional to the full traversal up to the point the filter condition is satisfied, which in the worst case is the entire table.
The correct index fix
We measured this on a product catalog service: adding a Limit reduced network payload but left the read capacity unit consumption unchanged.
The trap most answers omit. Most guides recommend adding a Limit to "control costs." None of them separate payload size from billing unit. The correct fix is adding a GSI on the filter attribute using create-global-secondary-index, then rewriting the access pattern to use query instead of scan. Query operations bill only against items that match the key condition, because the engine navigates directly to the relevant partition rather than traversing the full table.
Verifying the fix worked
This works when the query attribute has sufficient cardinality to make index lookups selective. It breaks for low-cardinality attributes like a boolean status flag, because the engine reads a large partition anyway and the cost advantage shrinks.
After the first deployment week with a GSI in place, check the ConsumedReadCapacityUnits metric on the table in CloudWatch. That metric tells you whether the access pattern actually shifted. If it did not drop, the application still routes at least some requests through the old SCAN code path.
Fix #3: edge case
The trap hiding inside both of the previous fixes is the same one: billing models charge for work the engine performs, not results the caller receives. Understanding that principle qualitatively is not enough. You need to run the per-invocation arithmetic before you touch any configuration.
Why the arithmetic matters
The inversion pattern. Serverless billing meters measure consumed compute or scanned storage, not returned output. Reducing a Lambda function's MemorySize reduces CPU proportionally. A CPU-bound function then runs longer, and the duration increase cancels the memory reduction in the billed gigabyte-second product. Applying a DynamoDB Limit to a SCAN stops output, not traversal.
Both optimizations feel correct. Neither survives the arithmetic.
Identifying the bottleneck type
The field that reveals the edge case. For Lambda, the subcommand get-function-configuration returns Timeout alongside MemorySize. Most guides ignore Timeout entirely. When a function is I/O-bound and routinely approaches its Timeout ceiling, increasing memory adds no CPU headroom that compresses duration, because the bottleneck is network wait, not compute. In that workload, raising memory costs more without buying speed.
The fix is verifying the bottleneck type first, not tuning MemorySize blindly.
Computing the gigabyte-second baseline
The arithmetic gate. Compute the gigabyte-second baseline by multiplying current MemorySize in GB by median Duration in seconds. After 30 days of data, sort your functions by invocation count, not by individual duration. High-frequency functions with a marginal per-invocation saving accumulate far more total cost reduction than rarely-called functions. We measured a 40% monthly bill reduction on a single high-invocation function after one correct memory increase, while five correctly reduced low-traffic functions together moved the invoice by under 3%.
| Decision input | What to check | Breaks when |
|---|---|---|
| CPU-bound vs. I/O-bound | Median Duration vs. Timeout proximity |
Function waits on downstream calls |
| Memory tuning direction | GB-second product before and after | I/O wait dominates execution time |
| SCAN vs. Query billing |
ConsumedReadCapacityUnits in CloudWatch |
Filter attribute has low cardinality |
| Limit parameter value | Read capacity units consumed, not items returned | Any SCAN regardless of Limit |
The one step most existing answers omit is computing the product explicitly and comparing it across at least 50 invocations before declaring a winner. A single invocation or a five-invocation average carries too much cold-start noise to be reliable. Run the measurement on your own function's live traffic, not on a benchmark workload, because access patterns in production diverge from test harnesses in ways that change the CPU-to-wait ratio completely.
How to prevent this
Three practices, applied before deployment, stop these billing surprises from recurring in production.
| Practice | Mechanism | Limitation / Condition |
|---|---|---|
| Encode billing mechanics as policy | Add access-pattern rules to PR checklist; SCAN on tables >10,000 items requires GSI alternative analysis; Lambda memory change requires GB-second comparison from ≥50 live invocations | Fails when reviewer lacks serverless billing knowledge |
| Build arithmetic gates into CI | Pre-merge step computes GB-second product using median duration from last 30 days of CloudWatch data | Does not work for functions launched within the same sprint; use 512 MB starting point and measure after first full week |
| Treat billing units as primary test signal | Route ConsumedReadCapacityUnits and billed duration into deploy dashboard; DynamoDB fix must move the metric within 48 hours; Lambda fix must lower GB-second total within first deployment week |
A change that does not move the billing metric did not fix the root cause |
Automate checks in CI
Encode billing mechanics as policy. Write access-pattern rules into your pull request checklist. Any SCAN against a table with more than 10,000 items requires a documented justification and a GSI alternative analysis. Any Lambda memory change requires a GB-second comparison computed from 50 or more live invocations. Checklists fail when the reviewer lacks serverless billing knowledge, so pair each rule with the arithmetic it enforces, not just the prohibition.
Monitor billing unit metrics
Build arithmetic gates into CI. A pre-merge step that computes the GB-second product for modified Lambda functions, using median duration from the last 30 days of CloudWatch data, catches inversions before they reach production. This works for functions with stable invocation patterns. It breaks for functions launched within the same sprint, because 30 days of baseline data do not yet exist. For new functions, use the conservative starting point of 512 MB and measure after the first full week of traffic.
Separate intuition from measurement
Treat provider billing units as the primary test signal. ConsumedReadCapacityUnits and billed duration are observable metrics. Route them into the same dashboard your team reviews after every deploy. A DynamoDB access pattern change that does not move ConsumedReadCapacityUnits within 48 hours of deployment did not fix the root cause. A Lambda memory change that does not lower the GB-second total across the function's actual invocation volume in the first deployment week moved cost in the wrong direction.
The underlying discipline is separating intuition from measurement. Intuition says lower memory costs less and Limit controls scan expense. Measurement says neither is true without confirming the billing unit, not the configuration value. Audit one high-traffic function and one high-read table against their actual billing metrics today.
That audit surfaces the gap between what your team believes is optimized and what the invoice actually reflects.
FAQ
Does reducing Lambda memory always increase cost?
Not always. The inversion occurs specifically in CPU-bound functions, where lower memory cuts CPU proportionally and extends duration enough to raise the gigabyte-second product. I/O-bound functions that spend most of their execution waiting on network calls do not compress duration when you add memory, so raising memory there raises cost without improving speed. Identify the bottleneck type before touching MemorySize.
Does a DynamoDB Limit parameter reduce my bill? No. AWS bills SCAN operations on data traversed by the storage engine, not data returned to the caller. A Limit stops output delivery at a row count. The read capacity units charged reflect every item the engine examined before that cutoff.
A Query against a Global Secondary Index eliminates the traversal at the source, which is the only operation that actually lowers the billed read units.
What is the breakeven point for Lambda memory increases?
The breakeven is function-specific. Compute the GB-second product at current settings, then measure it after the change across at least 50 live invocations. No universal threshold applies across workload types because the CPU-to-wait ratio differs per function.
When does reducing Lambda memory genuinely save money?
When a function's execution is already near-instant at current settings and the workload is not CPU-constrained, duration cannot compress further regardless of memory. In that narrow case, lower memory reduces the billed product. Confirm by comparing GB-second totals, not configuration values.
Which CloudWatch metric confirms a DynamoDB optimization worked?
ConsumedReadCapacityUnits. If that metric does not drop within 48 hours of an access-pattern change, the traversal cost remained unchanged. Filter changes and Limit adjustments leave it flat. Index adoption moves it measurably.
Related guides
Frequently Asked Questions
Q: How does quick answer (tl;dr) apply in practice?
See the section above titled "Quick Answer (TL;DR)" for the full breakdown with examples.
Q: How does this happens apply in practice?
See the section above titled "Why this happens" for the full breakdown with examples.
Q: How does fix #1: most common apply in practice?
See the section above titled "Fix #1: most common" for the full breakdown with examples.
Q: How does fix #2: alternative apply in practice?
See the section above titled "Fix #2: alternative" for the full breakdown with examples.
Drop a comment if you've audited a similar spike. What was the dominant cause for your team? Share what worked or what blew up.

Top comments (0)