DEV Community

Morgan Ma
Morgan Ma

Posted on

The Short Test Kept a Dead Cache

I will start with the conclusion, not the tool. Caching c_str across growth is a lifetime bug. A short unit test can still pass cleanly.

Later growth is what kills the cached pointer. Did your suite ever cross the capacity line? If it did not, the cache only looked stable.

A faster coding model does not relax object lifetime. The language rules did not get softer this year. That gap is why this small repro still matters.

The symptom

Picture a small label type in ordinary service code. It stores one string and one raw pointer. Every reset writes both the bytes and the cache.

Every suffix mutates the string and leaves the pointer. Short names still print, so the test stays green. Did that tiny append look harmless during review?

A long export name dies after one more append. AddressSanitizer then reports a heap use after free. The text of the suffix was never the real suspect.

This listing is a constructed lab repro only. It is not a story about a shipped outage. I am not claiming a customer impact number.

Compile it yourself and then judge the addresses. Snapshot each address as an integer before mutation. Comparing a dangling pointer value is not the goal.

#include <cstdint>
#include <cstdio>
#include <string>
#include <string_view>

struct NameTag {
  std::string storage;
  const char* cached = nullptr;

  void reset(std::string next) {
    storage = std::move(next);
    cached = storage.c_str();
  }

  void add_suffix(std::string_view suffix) {
    storage.append(suffix);
  }

  void force_growth() {
    const std::size_t extra = storage.capacity() - storage.size() + 1;
    storage.append(extra, 'x');
  }

  void print_cached() const {
    std::fputs(cached, stdout);
    std::fputc('\n', stdout);
  }
};

static void report(const char* label, const NameTag& tag,
                   std::uintptr_t before) {
  const auto after =
      reinterpret_cast<std::uintptr_t>(tag.storage.c_str());
  const int moved = before != after;
  std::printf("%s size=%zu cap=%zu moved=%d\n",
              label, tag.storage.size(), tag.storage.capacity(), moved);
}

int main() {
  NameTag short_tag;
  short_tag.reset("ok");
  const auto short_before =
      reinterpret_cast<std::uintptr_t>(short_tag.cached);
  short_tag.add_suffix("!");
  report("short", short_tag, short_before);
  if (short_before ==
      reinterpret_cast<std::uintptr_t>(short_tag.storage.c_str())) {
    short_tag.print_cached();
  }

  NameTag long_tag;
  long_tag.reset(std::string(64, 'a'));
  const auto long_before =
      reinterpret_cast<std::uintptr_t>(long_tag.cached);
  long_tag.force_growth();
  report("forced", long_tag, long_before);
  // Unsafe on purpose. Uncomment only under AddressSanitizer.
  // long_tag.print_cached();
  std::fputs("forced cache is stale; not reading it", stdout);
  std::fputc('\n', stdout);
}
Enter fullscreen mode Exit fullscreen mode

Step 1: Freeze the failing length

Do not chase a random crash across the whole suite. Pin one input length before you pin the suffix. Change only one of those lengths at a time.

Which length first moves the c_str address? Write that number down before you change any code. A moving pointer is the bug you can measure.

A segfault is only the late and noisy symptom. I would log four fields on every single probe. Log the size, both capacities, and pointer equality.

If equality flips, stop reading the old pointer. Why read an address you already proved stale? The integer snapshot is the evidence, not the print.

Step 2: Build a sanitizer binary

Pretty logs are not a substitute for a hard fault. You want the sanitizer to trap the bad read. The command below is the local bar to clear.

c++ -std=c++20 -g -O1 -fno-omit-frame-pointer -fsanitize=address,undefined sso_cache.cpp -o sso_cache
./sso_cache
Enter fullscreen mode Exit fullscreen mode

Why choose -O1 instead of a fully unoptimized build? Some reallocations hide when optimization is fully off. Why keep frame pointers in this sanitizer build?

The report then names a function you actually wrote. Did you forget debug symbols on the compile line? Then the report gives an address, not a source line.

Undefined behavior checks belong in that same binary. A bad pointer read is not your only nearby risk. Signed size math can sit quietly beside this bug.

Step 3: Split SSO from heap growth

Small string optimization is the liar in this story. A short string may store bytes inside the object. Then c_str points into that internal object buffer.

An append that still fits does not move it. Your unit test cheers, and you merge the patch. Cross local capacity and the bytes hit the heap.

Later growth can free that heap buffer entirely. The cached pointer still names the old block. That is use after free, not a compiler mystery.

Is the local capacity always fifteen characters long? Sometimes it is fifteen, and sometimes it is twenty-two. Another standard library can pick a third size.

Do not hardcode a threshold you found online. Measure capacity on the exact toolchain you ship. A blog number is not an ABI guarantee for you.

One long input is still not a proof of safety. A heap string can append without reallocating yet. Capacity may already be larger than the live size.

So grow until capacity actually changes underneath you. Then read the old pointer only under ASan. If it does not trap, you are not testing the bug.

Append one more byte than the spare capacity. That call forces reallocation without a hardcoded SSO size. The helper in the listing does exactly that.

Use this matrix before you trust any green run. Same binary, four seeds, two very different suffixes. The last column is the lie a short test tells.

seed length suffix what you should expect what a short test misses
1 1 char often no move under SSO false confidence
15 1 char may cross the local buffer the real edge
16 1 char heap already; slack may hide it one long case is weak
64 force past capacity pointer must move this is the production shape

Step 4: Shorten the pointer lifetime

The fix is not a smarter cache of the pointer. The fix is a shorter lifetime for that address. Take the pointer at the use site instead.

void print_live(const std::string& storage) {
  std::fputs(storage.c_str(), stdout);
  std::fputc('\n', stdout);
}
Enter fullscreen mode Exit fullscreen mode

Or store an offset and rebuild after each mutation. Prefer the use-site call whenever you can. An offset is easy to get wrong after erase.

A copied buffer is stable and costs an allocation. Reserve plus a cache is only a sized bet. Would you approve that bet in code review?

I would reject it and ask for the use-site call. The bet dies when a later suffix exceeds reserve. Production names are longer than the sample you tried.

This is a different bug from a view into a temporary. A temporary dies at the end of the full expression. This cache dies later, only when capacity changes.

Do not merge those two tests into one assertion. Each bug needs its own length story and oracle. A passing temporary test says nothing about growth.

Step 5: Treat the model as a suspect list

A model is a suspect list, not a verdict. That list can speed the first hour of debugging. It still cannot see the heap you actually have.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode's free model access can list those suspects for you. Think of SSO, iterator invalidation, or a moved-from string. The free server option can host a clean compile of the file.

I would not treat that machine as production hardware. I do not know its library, quota, or offer lifetime. So I will not invent those numbers in this draft.

I would ask one question of that extra compile. Does capacity move on the same lengths as yours? A model may suggest reserve and then cache c_str.

That patch passes the short test and fails later. Ask for three patches, then delete every raw cache. Any pointer kept across a non-const call must go.

You still need the sanitizer run after that edit. A fluent explanation is not evidence of safety. Did the patch name a capacity log, or only a vibe?

What this workflow will not prove

Do not trust a green run on one standard library. The three major libraries disagree on local capacity. A free server with another library is a second sample.

It is not a waiver for your production toolchain. Do not ask the model to confirm pointer stability. It cannot inspect your allocator or your heap.

Paste the capacity log and ask for the first retain. Then you delete that line and rerun the probe. If the model argues with the log, keep the log.

Do not disable ASan to get a clean screenshot. A silent binary is how this class of bug ships. If the free server image lacks sanitizer support, stop.

Compile the same probe on a machine that has it. Matching log lines across machines is still useful. A missing trap is not a successful test result.

Who should skip this

Skip this workflow if you cannot rebuild with sanitizers. The printf guard hides the crash on purpose. Without ASan, a lucky print can still be undefined.

Looking right in a terminal is not a proof. Skip this article if your real bug is a data race. Capacity logs will not catch two threads appending.

You need ThreadSanitizer, a mutex, or both tools. This walkthrough does not cover that race fix. Mixing those bugs will waste a full afternoon.

Skip cached pointers in parsers that retain input. If the caller mutates the source, your offset dies too. Own a copy, or end the borrow before the next write.

Document that bound next to the function, not in chat. Skip the free server as your only integration fleet. A shared free image can change under your feet.

I am not claiming that it will change tomorrow. This draft has no contract for duration or hardware. Keep the repro inside your own pipeline first.

Use the extra machine only as a second opinion. Who should not use the free box as a stamp? Anyone who needs a matched production ABI should not.

Step 6: Run this test plan

  1. Save the listing as sso_cache.cpp before you edit it.
  2. Build with the sanitizer command from step two.
  3. Record moved for lengths 0, 1, 15, 16, 22, 23, 24, 64, and 4096.
  4. Repeat with a one-byte suffix and a sixty-four byte suffix.
  5. On any move, read the old pointer only under ASan.
  6. Confirm the report names print_cached, not a mystery frame.
  7. Replace the cache with a use-site c_str call.
  8. Rerun the same matrix after that ownership change.

Pointer movement may remain, and that can be legal. Reading the old address cannot remain in the binary. Step eight is the real pass condition for this bug.

What stays true after the patch

So what should you remember after running this repro? Stability is a measured property of one toolchain. It is not a feeling you get from a short fixture.

Growth changes addresses even when the text looks fine. Small string optimization makes the short case lie. Want a second compiler for this same matrix?

MonkeyCode's free server option is enough to start. Bring sanitizer flags, or do not trust that run.

Top comments (0)