<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vansh Sharma</title>
    <description>The latest articles on DEV Community by Vansh Sharma (@vanshsharma27).</description>
    <link>https://dev.to/vanshsharma27</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163623%2Ff550aff4-fb4e-439b-87b7-d9b1e27f1db3.png</url>
      <title>DEV Community: Vansh Sharma</title>
      <link>https://dev.to/vanshsharma27</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vanshsharma27"/>
    <language>en</language>
    <item>
      <title>When Skipping an API Error Can Delete Valid Data: A Cartography Fix</title>
      <dc:creator>Vansh Sharma</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:45:38 +0000</pubDate>
      <link>https://dev.to/vanshsharma27/when-skipping-an-api-error-can-delete-valid-data-a-cartography-fix-1d6g</link>
      <guid>https://dev.to/vanshsharma27/when-skipping-an-api-error-can-delete-valid-data-a-cartography-fix-1d6g</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I fixed a BigQuery sync failure in Cartography, then narrowed the error tolerance after review showed that skipping failed enumeration could feed false deletions into graph cleanup.&lt;/p&gt;

&lt;p&gt;One malformed BigQuery table could stop a Cartography GCP sync. My first patch let more API errors be skipped, but an automated reviewer caught a risk I had missed: a failed list request could leave cleanup deleting valid inventory nodes. I traced that concern through the sync code and changed which errors the patch would tolerate.&lt;/p&gt;

&lt;p&gt;Cartography is a &lt;a href="https://www.cncf.io/projects/cartography/" rel="noopener noreferrer"&gt;CNCF sandbox project&lt;/a&gt; written in Python that collects infrastructure assets and their relationships into a Neo4j graph. Its BigQuery ingestion lists resources, fetches extra details, loads the results, and runs cleanup. I worked on &lt;a href="https://github.com/cartography-cncf/cartography/pull/2972" rel="noopener noreferrer"&gt;PR #2972&lt;/a&gt; to fix the sync failure without making incomplete collection look like resource removal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a missing graph node matters
&lt;/h2&gt;

&lt;p&gt;Cartography's &lt;a href="https://github.com/cartography-cncf/cartography/blob/e4c08bd71380d165981f99a2d69982a4846bc134/README.md" rel="noopener noreferrer"&gt;README&lt;/a&gt; describes queries for datastore access, network paths, and internet-exposed compute instances. The graph supplies the inventory and relationships behind those security questions. If cleanup removes a valid table node because enumeration failed, that table disappears from the inventory available to queries that depend on it.&lt;/p&gt;

&lt;p&gt;The cloud resource has not disappeared. The graph has lost its record of it. That was the risk in the broader patch, not a security incident I observed. Continuing ingestion would be a poor trade if it made the resulting inventory less trustworthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  One table-detail request stopped ingestion
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/cartography-cncf/cartography/issues/2951" rel="noopener noreferrer"&gt;original report&lt;/a&gt; described an external table backed by a Delta Lake path in GCS without valid Delta log files. When ingestion called &lt;code&gt;tables().get&lt;/code&gt;, BigQuery returned HTTP &lt;code&gt;400&lt;/code&gt; with the reason &lt;code&gt;invalidQuery&lt;/code&gt;. The exception propagated out of ingestion and aborted the sync.&lt;/p&gt;

&lt;p&gt;My regression tests inject that status-and-reason combination using the existing test helper:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_make_http_error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invalidQuery&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tests replace each handler module's &lt;code&gt;gcp_api_execute_with_retry&lt;/code&gt; with a function that raises the error. On the pre-fix code, the table-detail handler re-raises it rather than returning &lt;code&gt;None&lt;/code&gt;. This reproduces the handler failure without requiring a broken external table or running cloud ingestion.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recognizing the error was only half the fix
&lt;/h2&gt;

&lt;p&gt;There were two separate checks involved. The shared utility classified the exception, and each handler decided which categories it could tolerate.&lt;/p&gt;

&lt;p&gt;Before my change, the classifier's HTTP &lt;code&gt;400&lt;/code&gt; branch was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;reason&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_error_reason&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invalid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;badrequest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invalid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;invalidQuery&lt;/code&gt; became &lt;code&gt;unknown&lt;/code&gt;. Adding &lt;code&gt;invalid&lt;/code&gt; to a handler's tolerated categories would therefore have done nothing by itself. The classifier needed to recognize the reason too.&lt;/p&gt;

&lt;p&gt;I checked the other GCP modules named in the report, including storage, compute, and secrets manager. They already handled the &lt;code&gt;invalid&lt;/code&gt; category. That comparison supported changing the shared classification, but it did not establish that every BigQuery handler should use the same policy.&lt;/p&gt;

&lt;p&gt;The classifier change was small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-        if reason.lower() in ("invalid", "badrequest"):
&lt;/span&gt;&lt;span class="gi"&gt;+        if reason.lower() in ("invalid", "badrequest", "invalidquery"):
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The harder part was deciding what should happen after classification.&lt;/p&gt;

&lt;h2&gt;
  
  
  The automated review caught what happened after the exception
&lt;/h2&gt;

&lt;p&gt;My initial approach added &lt;code&gt;invalid&lt;/code&gt; tolerance across the BigQuery handlers. The automated reviewer &lt;code&gt;@cubic-dev-ai[bot]&lt;/code&gt; &lt;a href="https://github.com/cartography-cncf/cartography/pull/2972#discussion_r3502246158" rel="noopener noreferrer"&gt;flagged the table-list path&lt;/a&gt;: if listing tables for a dataset failed and ingestion continued, project-scoped cleanup could remove previously ingested table nodes that the failed request had not returned.&lt;/p&gt;

&lt;p&gt;I checked the surrounding sync code and agreed with that concern. A handler returning &lt;code&gt;None&lt;/code&gt; affected more than its immediate caller. It changed the data that would reach loading and cleanup.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;get_bigquery_routines&lt;/code&gt; and &lt;code&gt;get_bigquery_connections&lt;/code&gt; had the same relevant shape: collection fed cleanup that ran at project scope. I reverted the new tolerance in those functions as well as &lt;code&gt;get_bigquery_tables&lt;/code&gt;, then &lt;a href="https://github.com/cartography-cncf/cartography/pull/2972#discussion_r3502493159" rel="noopener noreferrer"&gt;recorded the reasoning in the review thread&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The bot's comment identified one path. Checking the sibling handlers was still my responsibility. Fixing only the function named in the comment would have left the same mistake in the other two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failed enrichment can preserve an asset; failed enumeration cannot establish absence
&lt;/h2&gt;

&lt;p&gt;The table-detail path was safe to tolerate for the reported error because the table already existed in the result from &lt;code&gt;tables.list&lt;/code&gt;. The merged sync code fetches details and only applies them when the request succeeds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;                &lt;span class="n"&gt;detail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_bigquery_table_detail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;project_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataset_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Returning &lt;code&gt;None&lt;/code&gt; here leaves the listed table available for loading. It may lack extra fields such as row and byte counts, but the failed enrichment request does not remove the table from the collected inventory.&lt;/p&gt;

&lt;p&gt;Dataset handling had its own guard. &lt;code&gt;sync_bigquery_datasets&lt;/code&gt; only loads and cleans up datasets inside the branch where &lt;code&gt;get_bigquery_datasets&lt;/code&gt; returns something other than &lt;code&gt;None&lt;/code&gt;. If dataset collection fails with this tolerated error, that branch does not run. Dataset detail failure also leaves the listed dataset available.&lt;/p&gt;

&lt;p&gt;The final policy was therefore specific:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Handler&lt;/th&gt;
&lt;th&gt;Behavior for HTTP &lt;code&gt;400&lt;/code&gt; / &lt;code&gt;invalidQuery&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_datasets&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Return &lt;code&gt;None&lt;/code&gt;; skip dataset loading and cleanup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_dataset_detail&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Return &lt;code&gt;None&lt;/code&gt;; retain the listed dataset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_table_detail&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Return &lt;code&gt;None&lt;/code&gt;; retain the listed table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_tables&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Re-raise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_routines&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Re-raise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;get_bigquery_connections&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Re-raise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For table details, the final handler change was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-            ("api_disabled", "billing_disabled", "forbidden", "not_found"),
&lt;/span&gt;&lt;span class="gi"&gt;+            ("api_disabled", "billing_disabled", "forbidden", "invalid", "not_found"),
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I did not add the same category to the three list handlers. Retaining those exceptions prevents this fix from letting their &lt;code&gt;invalidQuery&lt;/code&gt; failures continue into cleanup with incomplete results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Regression tests need to preserve the errors we still want
&lt;/h2&gt;

&lt;p&gt;I added &lt;code&gt;invalidQuery&lt;/code&gt; to the classifier's parametrized HTTP &lt;code&gt;400&lt;/code&gt; tests. I also split the handler cases into a group that returns &lt;code&gt;None&lt;/code&gt; and a group that must raise.&lt;/p&gt;

&lt;p&gt;The raising test uses the same injected error as the tolerant test, but its assertion requires the exception to survive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;pytest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raises&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;HttpError&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That assertion matters because a future edit that broadly adds &lt;code&gt;invalid&lt;/code&gt; tolerance could otherwise look reasonable while restoring the cleanup risk. The tests document which entry points can continue and which must stop.&lt;/p&gt;

&lt;p&gt;These tests exercise the handler policy; they do not execute Neo4j cleanup. The reason for each policy comes from the sync code, and the tests keep that decision visible when someone changes the handlers later.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix shipped in Cartography 0.139.0
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;@jychp&lt;/code&gt; approved the PR before it merged on July 6, 2026. &lt;code&gt;@thenadz&lt;/code&gt; also &lt;a href="https://github.com/cartography-cncf/cartography/pull/2972#issuecomment-4850018102" rel="noopener noreferrer"&gt;validated the patch against their environment&lt;/a&gt; and confirmed that the previously failing cases no longer crashed their scan. That gave the fix a check against the reported cloud behavior alongside my mocked regression tests.&lt;/p&gt;

&lt;p&gt;The change shipped in &lt;a href="https://github.com/cartography-cncf/cartography/releases/tag/0.139.0" rel="noopener noreferrer"&gt;Cartography 0.139.0&lt;/a&gt;. Running &lt;code&gt;git tag --contains&lt;/code&gt; on merge commit &lt;code&gt;e4c08bd71380d165981f99a2d69982a4846bc134&lt;/code&gt; confirms that this is the earliest containing release tag; the release notes also list the PR.&lt;/p&gt;

&lt;p&gt;The lesson I took from this contribution is to trace an error-handling change through every operation that consumes its result. In an inventory system, failed collection does not establish that a resource disappeared. Before returning an empty result or &lt;code&gt;None&lt;/code&gt; from an exception handler, check what the caller will infer from it and whether cleanup will still run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/cartography-cncf/cartography/pull/2972" rel="noopener noreferrer"&gt;Merged PR #2972&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/cartography-cncf/cartography/issues/2951" rel="noopener noreferrer"&gt;Original issue #2951&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/cartography-cncf/cartography/releases/tag/0.139.0" rel="noopener noreferrer"&gt;Cartography 0.139.0 release notes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>opensource</category>
      <category>cncf</category>
      <category>security</category>
    </item>
  </channel>
</rss>
