DEV Community

Cover image for DynamoDB Query vs Scan: Which to Use, and Why (w/ Examples)
DynoTable
DynoTable

Posted on • Edited on • Originally published at dynotable.com

DynamoDB Query vs Scan: Which to Use, and Why (w/ Examples)

Query reads a single item collection by partition key (optionally narrowing the
sort key); Scan reads the whole table and filters afterwards. They look
similar in the API but they bill — and scale — completely differently.

When should I use Query vs Scan in DynamoDB?

Use Query whenever you can name the partition you need — it reads one item collection and bills only for matched items. Reach for Scan only for one-off exports or tiny tables; it reads every item and bills the whole table before any FilterExpression runs. On real data, Query wins.

  • Query is targeted: you pay for the items in the matched partition.
  • Scan is exhaustive: you pay to read every item, then throw most away with a FilterExpression that runs after the read is metered.

On a table of any real size, a Scan with a filter is the classic "why is my
bill huge and my latency worse than RDS" footgun.

Side by side

Query Scan
Reads One partition (by PK) Every item in the table
Billed capacity Items matched in the partition Whole table, before filtering
FilterExpression Applied after the read — still billed for the read Same — filtering never cuts cost
Latency Flat as the table grows Grows with table size
Pagination 1 MB/page → LastEvaluatedKey 1 MB/page; parallelisable
Use it for Known access patterns One-off exports, tiny config tables

The key trap: a FilterExpression runs after DynamoDB meters the read, on
both operations. A Scan that "returns 10 rows" can bill for reading a million —
filtering is a convenience, never a cost control.

What a full Scan actually costs

Put numbers on it. DynamoDB meters reads in 4 KB units: a
strongly consistent read costs one
read capacity unit per 4 KB, an
eventually consistent read half that. Query and
Scan sum the size of every item they touch — not every item they return —
and round up to the next 4 KB.

Take a 1-million-item table averaging 2 KB per item (~2 GB of data), and an
access pattern that needs 10 of those items:

Items read Data metered Read units (eventually consistent)
Scan + FilterExpression 1,000,000 ~2 GB ~262,000
Query on a matching key 10 20 KB 3

Same 10 items, five orders of magnitude apart — and the Scan bills that
~262,000 every time it runs, whether the filter matches ten items or none. On
on-demand billing those are request units straight onto the
bill; on provisioned tables a big Scan competes with
production traffic for throughput and can throttle it into
ProvisionedThroughputExceededException.

Three more cost facts that surprise people:

  • Select: COUNT is not free. A counting Query or Scan consumes exactly the same read capacity as reading the items — it just doesn't return them.
  • Limit caps items evaluated, not items matched. Combined with a filter, a page can come back empty while still billing a full page of reads.
  • You never have to guess. Pass ReturnConsumedCapacity: TOTAL and every response reports the capacity it just consumed.

Check what your own items weigh with the
item size calculator, then turn read
units into a monthly bill with the
pricing calculator.

Use Query

Example Notes
Query PK = "USER#42" AND SK begins_with "ORDER#"

If you find yourself reaching for Scan to answer a common access pattern, that
is a modelling signal: add a Global Secondary Index so the
pattern becomes a Query.

The choice comes down to one question — can you name the partition you need?

DynamoDB Query vs Scan: Which to Use, and Why (w/ Examples)

If the key is known you Query; if not, add a GSI to make it one, and fall back
to Scan only when no key fits.

When Scan is fine

One-off exports, tiny config tables, and background jobs that page through the
whole table deliberately. Use Segment/TotalSegments to split a Scan across
workers (a parallel scan — see
parallel Scans in DynamoDB) when you genuinely
must read everything, and page it properly with LastEvaluatedKey
(pagination guide). If a Scan you already run is the
problem, why Scan is slow and expensive
walks the triage.

A reflexive SELECT * FROM table over DynamoDB is the same anti-pattern in
PartiQL clothing — it compiles to a Scan. When you really do need cross-item
analytics (a GROUP BY, a JOIN, an aggregate), DynoTable's SQL Workbench runs
them client-side over a bounded result set instead of hammering the table.

Try DynoTable to run and inspect these queries against your own
tables — it shows the consumed capacity of every operation it runs.

Top comments (0)