A scheduled job in one of our services started crashing every few days with this:
TypeError: Cannot read properties of undefined (reading 'id')
Inside a for..of over the rows of a query. Not on the first row, not on every run, and never reproducible locally.
The other clue turned up in a log line from the same job, which prints how many pending records it picked up. The query has LIMIT 500. The log said pendingTotal: 534. Another run said 813.
A SELECT with a LIMIT of 500 returning 813 rows is the kind of thing that makes you go back and read your own SQL four times. The SQL was fine. So was the ORM. The bug was one line below the driver's surface, and it has been open upstream for about a year.
What actually happens
We use postgres-js (the postgres package) with a connection pool, like most people do. Here's the shape of the bug, straight out of src/connection.js in the current release:
A connection keeps a single write cursor for the result it's filling in:
let ...
, rows = 0
Every DataRow message the server sends gets written at that cursor:
function DataRow(x) {
// ...
: (result[rows++] = transform.row.from ? transform.row.from(row) : row)
}
rows goes back to zero in exactly two places: CommandComplete, which is the server saying the query finished, and PortalSuspended, which is the cursor case. That's it.
Now look at what happens when a query doesn't finish. ErrorResponse records the error and returns. Then ReadyForQuery cleans up:
function ReadyForQuery(x) {
// ... resolve or reject the query ...
query = results = errorResponse = null
result = new Result()
connectTimer.cancel()
}
It allocates a fresh result array. One line later than where the cursor should have been reset, and the cursor isn't in that list.
So: a query streams 97 rows, then the server kills it. rows is left at 97. The connection goes back in the pool. The next query that lands on it gets a brand new empty array, writes its two rows at index 97 and 98, and resolves with an array whose length is 99 and whose first 97 slots are holes.
Why this one is nastier than a normal wrong answer
A sparse array is a bad failure mode because nothing about it throws:
-
JSON.stringify(rows)gives you a list starting with 97nulls. That goes straight into a log line or an API response. -
for..ofand.entries()hand youundefined, which is where our TypeError came from. -
.lengthis 99, which is how aLIMIT 500query reports 534. -
.map()preserves the holes, so an ORM mapping layer passes them through. -
.filter(Boolean)hides it completely, so whether you ever find out depends on which array method you happened to use.
And the quiet one, which is the reason I'm writing this down:
const [row] = await db.select().from(users).where(eq(users.id, id));
if (!row) return notFound();
row is undefined, the record exists, and you take the "not found" branch. No error, no log, just a wrong decision. We had a version of that on a write path.
The trigger is almost always a statement timeout
To hit this you need a query that errors after it has streamed at least one row. The common way is statement_timeout, Postgres error 57014, firing mid-scan on a query that's already returning rows.
The thing to check is where your timeout is set. Ours is on the database role, not in application code, which means any slow query in any service arms this for the next query on that connection. One heavy scan in a cron job, and an unrelated request handler sharing the pool gets the holes. That's also why it never reproduced locally: on a laptop nothing takes long enough to get cancelled.
If you want to confirm it in your own logs, look for these two together:
- A count in a log line that exceeds the
LIMITof the query it came from. - A
TypeErroronundefinedinside a loop over query results, within a few seconds of acanceling statement due to statement timeoutin the same process.
Same process matters, since the pool lives in the process. Different pod, different pool, unrelated.
Upgrading doesn't fix it
I checked the current release and master while writing this. postgres@3.4.9 is the latest on npm, and both it and master reset the cursor in the same two places and no others. The fix is a one-line PR, #1120, open since October 2025. The mechanism is written up in issue #1181, which is open and still getting comments this week.
So until it lands, patch it. With pnpm:
pnpm patch postgres
Add the one line to ReadyForQuery:
query = results = errorResponse = null
result = new Result()
+ rows = 0
connectTimer.cancel()
Then pnpm patch-commit <dir>, which writes the patch file and the patchedDependencies entry into your workspace. npm and yarn both have their own equivalents.
If you'd rather not patch a dependency, the other mitigations are all worse but worth knowing: set max: 1 so a poisoned connection only hurts the code that poisoned it (it doesn't make the problem go away), drop your statement_timeout low enough that queries die before streaming anything (good luck), or check rows.length against your own LIMIT on every read and throw. We patched.
The general lesson I took
I had a rule in my head that an error on a connection is contained: the transaction rolls back, the pool hands the connection to someone else, life continues. This was a reminder that the client library also has per-connection state, and an error path is exactly where that state stops being maintained.
So when a client hands you a fresh object for each query, the question worth asking is whether anything else is still pointing into the old one.
Top comments (0)