<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mohammed Arshad Ansari</title>
    <description>The latest articles on DEV Community by Mohammed Arshad Ansari (@mohammed_arshadansari_f2).</description>
    <link>https://dev.to/mohammed_arshadansari_f2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3609964%2Fd0a642d8-0372-4976-99f8-0575aa92a93c.png</url>
      <title>DEV Community: Mohammed Arshad Ansari</title>
      <link>https://dev.to/mohammed_arshadansari_f2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mohammed_arshadansari_f2"/>
    <language>en</language>
    <item>
      <title>DuckDB vs ClickHouse: which one, and when you need both</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/duckdb-vs-clickhouse-which-one-and-when-you-need-both-4ekf</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/duckdb-vs-clickhouse-which-one-and-when-you-need-both-4ekf</guid>
      <description>&lt;p&gt;DuckDB and ClickHouse get compared because they sit near each other in every analytical benchmark: both are columnar, both are vectorised, both are very fast. The benchmark framing hides the difference that actually decides it. &lt;strong&gt;DuckDB is a library you put inside your program. ClickHouse is a server your programs connect to.&lt;/strong&gt; Almost every practical difference follows from that.&lt;/p&gt;

&lt;p&gt;I run both. ClickHouse is the analytical store behind &lt;a href="https://hikmahtechnologies.com/systems/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Ansaar&lt;/a&gt;, feeding a live API. DuckDB is what I reach for in batch jobs and local work, and it is the subject of my book. So this is the comparison from operating them, not from a benchmark table.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short answer
&lt;/h2&gt;

&lt;p&gt;Pick &lt;strong&gt;DuckDB&lt;/strong&gt; when the work happens inside one program — a transformation job, a service that owns its own queries, a notebook, a CI run — and the data arrives in batches. There is nothing to operate, and it is extraordinarily fast for that one program.&lt;/p&gt;

&lt;p&gt;Pick &lt;strong&gt;ClickHouse&lt;/strong&gt; when many clients query the same data at the same time, the data keeps arriving, and people expect the numbers to be seconds old. That is a server's job, and ClickHouse is one of the best servers for it.&lt;/p&gt;

&lt;p&gt;If you have both shapes, run both. That's more common than either camp admits.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each one is
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DuckDB&lt;/strong&gt; is an in-process analytical database — the usual shorthand is "SQLite for analytics". You import it into Python, Node, Go or the CLI, and it runs a columnar query engine inside that process. It reads Parquet, CSV and JSON directly, including from S3, with no load step. There is no server, no port, no cluster. Only one process may open a database file read-write, and while it does, no other process can open the file at all; any number of processes may read it if none is writing. The pattern that scales is one writer producing Parquet and many read-only readers over it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ClickHouse&lt;/strong&gt; is a client-server columnar database built for real-time analytics. You run it as a service (or buy ClickHouse Cloud), clients connect over the network, and it stores data in its MergeTree engine: rows land in sorted parts on disk, and background merges combine them. It is built to ingest continuously and answer aggregations over billions of rows in milliseconds, for many clients at once. It scales out across machines when one isn't enough.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/duckdb-vs-clickhouse/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;DuckDB vs ClickHouse: which one, and when you need both&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>analytics</category>
    </item>
    <item>
      <title>A Pile of Documents Is Not Knowledge</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/a-pile-of-documents-is-not-knowledge-3lf8</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/a-pile-of-documents-is-not-knowledge-3lf8</guid>
      <description>&lt;p&gt;My research agent had read about 20,400 documents. It had concluded nothing.&lt;/p&gt;

&lt;p&gt;Not metaphorically. The knowledge store held roughly twenty thousand rows — papers, articles, PDFs, nightly log entries — and there was nowhere in the system that said &lt;em&gt;what any of it meant&lt;/em&gt;. Research answers were stored as &lt;code&gt;aegis://research/&lt;/code&gt; rows filed among ten thousand PDFs. The nightly journal wrote one dated entry into the same pile, in a format nobody could open.&lt;/p&gt;

&lt;p&gt;This is the third of three posts about rebuilding lanes of &lt;a href="https://hikmahtechnologies.com/aegis?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;AEGIS&lt;/a&gt;, my self-hosted agent platform, between 5 and 13 September. The first two were &lt;a href="https://hikmahtechnologies.com/blog/your-alerts-have-no-identity?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;alerting&lt;/a&gt; and &lt;a href="https://hikmahtechnologies.com/blog/a-ledger-is-not-a-database-table?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;money&lt;/a&gt;. All three ended at the same rule, and this lane is where it's easiest to see why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure use, not ingestion
&lt;/h2&gt;

&lt;p&gt;Every RAG dashboard I have ever seen measures the wrong end of the pipe. Documents ingested. Chunks embedded. Index size. All of it is &lt;em&gt;input&lt;/em&gt;, and input is the part that is trivially easy to grow.&lt;/p&gt;

&lt;p&gt;The measurement that mattered took one join: a log of which documents were injected into which prompt, joined back to where each document came from. Over the 30 days to 12 September:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;arXiv: 1,889 papers, 89,669 chunks — 90% of all RSS chunks.&lt;/strong&gt; Prompts used &lt;strong&gt;14 papers&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Across the whole corpus, arXiv PDFs were &lt;strong&gt;93% of all chunks&lt;/strong&gt;, and &lt;strong&gt;78 of 10,284&lt;/strong&gt; had ever reached a prompt.&lt;/li&gt;
&lt;li&gt;The intelligence scans had stored &lt;strong&gt;371 items&lt;/strong&gt;. Exactly &lt;strong&gt;one&lt;/strong&gt; was ever used in a prompt.&lt;/li&gt;
&lt;li&gt;Those same scans created &lt;strong&gt;277 task-manager items in 30 days&lt;/strong&gt;, every one auto-classified as reference material on arrival. My task list was being used as a log file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A retention preview said the same thing in bytes. One rule — keep only the first chunk of any PDF no prompt has used in 30 days — would touch 8,297 PDFs and drop 362,153 of 500,055 chunks, freeing about 530 MB of text and 1 GB of vectors from a 4.8 GB table. The PDFs themselves stay, and nothing has been deleted: the preview is a dry run, and deleting is a separate, explicit decision.&lt;/p&gt;

&lt;p&gt;This is not a story about arXiv being low quality. It is excellent. It is a story about a system that was optimising the metric it could see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The filter I measured and did not ship
&lt;/h2&gt;

&lt;p&gt;The obvious fix is a topic gate: only store the full text when the item matches something you care about. I built it, measured it, and made it &lt;strong&gt;opt-in&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here is why. On arXiv the gate would pass 41% of papers — about 2.3× fewer chunks, for real added complexity. And on every &lt;em&gt;other&lt;/em&gt; feed, the gate would have kept the full text of only &lt;strong&gt;2 of the 10 documents a prompt actually used&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That second number killed it as a default. A filter that removes 80% of the value you can prove, to save storage you have plenty of, is not an optimisation. The measurement changed the design, which is the only reason to take a measurement.&lt;/p&gt;

&lt;p&gt;The actual win was somewhere else entirely.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/a-pile-of-documents-is-not-knowledge/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;A Pile of Documents Is Not Knowledge&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>A Ledger Is Not a Database Table</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/a-ledger-is-not-a-database-table-14n0</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/a-ledger-is-not-a-database-table-14n0</guid>
      <description>&lt;p&gt;For months, my finance agent read the first 200 characters of each email and called it accounting.&lt;/p&gt;

&lt;p&gt;It was not lying, exactly. It genuinely did classify email, genuinely did track subscriptions, and genuinely did produce a monthly spend figure. The figure was just wrong, and nothing in the system was capable of noticing.&lt;/p&gt;

&lt;p&gt;This is the second of three posts about rebuilding lanes of &lt;a href="https://hikmahtechnologies.com/aegis?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;AEGIS&lt;/a&gt;, my self-hosted agent platform. The first was about &lt;a href="https://hikmahtechnologies.com/blog/your-alerts-have-no-identity?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;alerting&lt;/a&gt;; the third is about &lt;a href="https://hikmahtechnologies.com/blog/a-pile-of-documents-is-not-knowledge?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;knowledge&lt;/a&gt;. All three converged on the same rule, and this is the lane where getting it wrong shows up as a wrong number.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "working" actually looked like
&lt;/h2&gt;

&lt;p&gt;Measured on production on 2026-09-05, over the data since 1 July:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The extractor never saw the email.&lt;/strong&gt; The fetch pulled each message in full, then kept &lt;code&gt;snippet[:500]&lt;/code&gt; — Gmail's preview line, which runs to about 200 characters, so the cap never even bit. Downstream, that snippet was read back as if it were the message body. Of 61 emails judged to be receipts, &lt;strong&gt;30 had an amount&lt;/strong&gt;. Of 35 recurring-charge rows, &lt;strong&gt;19 carried &lt;code&gt;amount_cents = 0&lt;/code&gt;&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The richest stream was discarded on purpose.&lt;/strong&gt; A list of bank sender addresses existed to stop bank alerts minting fake subscriptions — correct for a subscription tracker, catastrophic for a finance agent. In 30 days, from one mailbox, that list threw away 46 UPI debit alerts, plus IMPS transfers, UPI credits, card spends, a credit-card statement and an inbound international remittance. None of it was recorded anywhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing dated reached me.&lt;/strong&gt; Sitting unactioned in the inbox at that moment: a credit-card statement due in two days, an advance-tax instalment due in ten, a declined subscription payment with a fix-by date, an electricity bill, and an "AWS past due". Tasks created: zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What it did send was noise.&lt;/strong&gt; 73 &lt;code&gt;Anomaly: ? Apple&lt;/code&gt;-shaped tasks in nine weeks, most with no amount. 27 chat pings in 30 days about the same four charges. The same "what is this vendor?" question asked six times, because vendor-name variants produced different dedupe keys.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The monthly total was fiction.&lt;/strong&gt; One electricity account appeared as three vendors, so "total monthly burn" counted it three times. A client's own supplier invoices, sitting in a work mailbox, counted as my subscriptions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The plumbing was in perfect health: 242 extraction calls in 30 days, zero failures, 212 completed workflow runs. &lt;strong&gt;Every dashboard was green and every number was wrong.&lt;/strong&gt; That is the specific failure mode of agent systems, and it is why I now measure the output rather than the pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision: hledger is the record, Postgres is the index
&lt;/h2&gt;

&lt;p&gt;The rebuild starts with one structural choice. The book of record is a plain-text double-entry journal — &lt;a href="https://hledger.org" rel="noopener noreferrer"&gt;hledger&lt;/a&gt; — in a private git repo. Postgres holds an index over it for fast queries.&lt;/p&gt;

&lt;p&gt;Three properties a table does not give you by default:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every write is a diff.&lt;/strong&gt; A journal entry is a few lines of text in a git commit. I can read it, a reviewer can read it, and &lt;code&gt;git log&lt;/code&gt; is the audit trail I would otherwise have had to build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The arithmetic is checked by something that isn't me.&lt;/strong&gt; Every write runs &lt;code&gt;hledger check --strict&lt;/code&gt;, and a failure reverts exactly the paths that write touched. An entry that does not balance, or that uses an account nobody declared, does not land. No amount of LLM confidence gets past a tool that does arithmetic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The index is disposable.&lt;/strong&gt; The rule in the codebase is blunt: &lt;em&gt;never treat an amount in the index as authoritative — run hledger.&lt;/em&gt; That means an index bug is a display bug, not a financial one, and the whole index can be rebuilt from the journal whenever I want.&lt;/p&gt;

&lt;p&gt;The mechanics are ordinary and worth stating anyway: core and worker share &lt;strong&gt;one&lt;/strong&gt; checkout, serialised by a file lock, so every write goes through one module. A hand-rolled file write would skip the strict check and its revert — so there is exactly one writer, and the rest of the system asks it.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/a-ledger-is-not-a-database-table/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;A Ledger Is Not a Database Table&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>automation</category>
    </item>
    <item>
      <title>Your Alerts Have No Identity</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/your-alerts-have-no-identity-1528</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/your-alerts-have-no-identity-1528</guid>
      <description>&lt;p&gt;I went looking for why my homelab kept telling me the same thing eight times.&lt;/p&gt;

&lt;p&gt;The answer was structural, and it is the kind of thing that hides in a system for years because every individual piece of it is reasonable. AEGIS — &lt;a href="https://hikmahtechnologies.com/aegis?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;my self-hosted agent platform&lt;/a&gt; — had no concept of a &lt;em&gt;problem&lt;/em&gt;. It had notifications. One probe failed on eight consecutive days and wrote eight Todoist tasks, because the only thing it could ask was "did I already write a task about this?" — and it asked by looking at its own most recent log line, which by then said something else.&lt;/p&gt;

&lt;p&gt;The Todoist task &lt;strong&gt;was&lt;/strong&gt; the identity of the incident. And that is the bug.&lt;/p&gt;

&lt;p&gt;This is the first of three posts about three lanes of AEGIS I rebuilt between 5 and 13 September. The other two — &lt;a href="https://hikmahtechnologies.com/blog/a-ledger-is-not-a-database-table?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;the money lane&lt;/a&gt; and &lt;a href="https://hikmahtechnologies.com/blog/a-pile-of-documents-is-not-knowledge?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;the knowledge lane&lt;/a&gt; — landed on the same rule from completely different directions, which is the only reason I trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the measurement said
&lt;/h2&gt;

&lt;p&gt;I don't rebuild anything on a hunch any more. Before touching the code I measured what was actually there on 2026-09-07:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Three alerting systems that could not see each other.&lt;/strong&gt; The investigation flow (a &lt;code&gt;#alert&lt;/code&gt; task plus a chat card), a set of watchdogs deduping against the audit log, and domain tables with their own private keys for certificate expiry and config drift. The resolved-aware &lt;code&gt;NOT EXISTS&lt;/code&gt; dedupe SQL was copy-pasted in three files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thirteen fingerprint or dedupe-key schemes&lt;/strong&gt;, mutually incompatible. A vendor fingerprint. A synthesised &lt;code&gt;alertmanager:{alertname}:{instance}&lt;/code&gt;. A &lt;code&gt;sentry:{issue_id}&lt;/code&gt;. Three separate signature classes. Three mute-key namespaces sharing one primary key. A day-bucketed drift key. Plus three &lt;em&gt;informal&lt;/em&gt; links — a &lt;code&gt;LIKE '%' || task_id || '%'&lt;/code&gt; against workflow ids, title substring matching, and a comment footer used as an authorship marker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dedupe keyed on the task, not the problem.&lt;/strong&gt; The dedupe index's &lt;code&gt;task_id&lt;/code&gt; column was &lt;code&gt;NOT NULL&lt;/code&gt; and joined to the task table. One flow refused to use it at all and built a fourth ledger inside a settings row. Because the signature was the primary key, recreating a task reset &lt;code&gt;first_seen_at&lt;/code&gt; and the occurrence count — recurrence history was destroyed on every recreate. Only &lt;strong&gt;12 of 42&lt;/strong&gt; open alert tasks had a signature row at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One producer bypassed the capture path entirely&lt;/strong&gt;, hand-building its command and deduping against the newest audit row. The close-on-resolve function structurally could not reach it. That is where the eight copies came from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No service state.&lt;/strong&gt; No maintenance window, no deploy suppression. The GitHub webhook claimed deployment events and dropped them. Suppression existed only as incidental delays scattered across three files.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the part that made it expensive to fix: of 72,499 source lines, the four files carrying this logic held 11,212 of them. The code doing all this was the code hardest to change safely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The primitive: a problem is a row, a task is a projection
&lt;/h2&gt;

&lt;p&gt;The replacement is deliberately small. Every operational signal — an Alertmanager webhook, a Sentry issue, a swarm heartbeat, a certificate about to expire, a stuck social post, a hand-written task — becomes an &lt;code&gt;Event&lt;/code&gt; and goes through one function, &lt;code&gt;ingest_event&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That function owns three things and nothing else:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identity.&lt;/strong&gt; One open &lt;code&gt;problems&lt;/code&gt; row per &lt;code&gt;correlation_key&lt;/code&gt;. If a row for that key is already open, this is another occurrence of the same problem, not a new one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;History.&lt;/strong&gt; Every occurrence, every resolution, every report is a &lt;code&gt;problem_events&lt;/code&gt; row, idempotent on &lt;code&gt;(source, external_id)&lt;/code&gt;. Recurrence survives everything, including deleting the ticket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The decision.&lt;/strong&gt; The return value carries &lt;code&gt;investigate: true&lt;/code&gt; or &lt;code&gt;false&lt;/code&gt;. The producer starts the investigation workflow only when told to. The flow itself never dedupes and never captures a task.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Todoist task is created by a projector, from the problem. It is a &lt;strong&gt;projection&lt;/strong&gt;, never the identity. The rule I wrote into the docs, because I knew I'd be tempted to break it: &lt;em&gt;do not look a problem up by its task title or fingerprint.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/your-alerts-have-no-identity/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Your Alerts Have No Identity&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Local-first analytics in practice: DuckDB, Parquet, and killing the round-trip</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Fri, 25 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/local-first-analytics-in-practice-duckdb-parquet-and-killing-the-round-trip-59h8</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/local-first-analytics-in-practice-duckdb-parquet-and-killing-the-round-trip-59h8</guid>
      <description>&lt;p&gt;In &lt;a href="https://hikmahtechnologies.com/blog/why-i-wrote-local-first-analytics?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;the previous post&lt;/a&gt; I argued that a lot of analytics infrastructure exists to solve laptop-sized problems. Here's what the alternative actually looks like in code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup: data as files, not endpoints
&lt;/h2&gt;

&lt;p&gt;The local-first pattern starts by treating your dataset as a &lt;em&gt;file&lt;/em&gt; — ideally &lt;a href="https://parquet.apache.org" rel="noopener noreferrer"&gt;Parquet&lt;/a&gt;, a columnar format that's compact, typed, and splittable. You can sit it in object storage, ship it in a release, or cache it on disk. No connection string, no credentials rotation, no warehouse to keep warm.&lt;/p&gt;

&lt;p&gt;Then you point an in-process engine at it. &lt;a href="https://duckdb.org" rel="noopener noreferrer"&gt;DuckDB&lt;/a&gt; reads Parquet natively and runs the query where you are.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;duckdb&lt;/span&gt;

&lt;span class="c1"&gt;# No server. No connection pool. Just a query against a file.
&lt;/span&gt;&lt;span class="n"&gt;con&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;duckdb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;con&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    SELECT
        symbol,
        date_trunc(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;month&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, trade_date) AS month,
        avg(close) AS avg_close,
        count(*) AS days
    FROM &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;prices/*.parquet&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;      -- glob straight over partitioned files
    WHERE trade_date &amp;gt;= &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;2024-01-01&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;
    GROUP BY 1, 2
    ORDER BY 1, 2
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;df&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                         &lt;span class="c1"&gt;# straight into a pandas DataFrame
&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That query reads only the columns it needs (columnar), only the row groups that match the predicate (pushdown), and never leaves the machine. On a few gigabytes of price data it returns in milliseconds — on the same laptop you're reading this on.&lt;/p&gt;

&lt;h2&gt;
  
  
  It runs in the browser too
&lt;/h2&gt;

&lt;p&gt;The part that surprises people: the &lt;em&gt;same&lt;/em&gt; engine compiles to WebAssembly. With &lt;code&gt;duckdb-wasm&lt;/code&gt; you can ship a Parquet file and a query to the browser and render an interactive analytics view with &lt;strong&gt;no backend at all&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;duckdb&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@duckdb/duckdb-wasm&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;initDuckDB&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;            &lt;span class="c1"&gt;// wasm engine in the tab&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;con&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;con&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`
  SELECT regime, count(*) AS n
  FROM 'https://cdn.example.com/regimes.parquet'
  GROUP BY regime
`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The user's machine does the work. Your "API" is a static file on a CDN. Your hosting bill is whatever a CDN charges to serve a few megabytes — and there's no query endpoint to attack, rate-limit, or scale.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/local-first-analytics-in-practice-duckdb-parquet/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Local-first analytics in practice: DuckDB, Parquet, and killing the round-trip&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>python</category>
      <category>analytics</category>
    </item>
    <item>
      <title>72% of India's NSE bulk deals are same-day round trips (a 90,000-row analysis)</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Fri, 25 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/72-of-indias-nse-bulk-deals-are-same-day-round-trips-a-90000-row-analysis-gm4</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/72-of-indias-nse-bulk-deals-are-same-day-round-trips-a-90000-row-analysis-gm4</guid>
      <description>&lt;p&gt;&lt;em&gt;I run a small research site for Indian markets. I pulled every bulk and block deal the NSE has disclosed since January 2023, about 90,000 rows, to see what they actually contain. Below is the explainer I wrote from it. At the end is the short bit of Python behind the headline number, and the one trap that would have doubled every total.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;When a large trade happens on the NSE, you usually hear about it the same evening: &lt;em&gt;"Fund X buys 1.2% of Company Y in a bulk deal."&lt;/em&gt; Two kinds of disclosure produce these headlines, bulk deals and block deals. They sound alike, but the rules are different, and so is what they tell you.&lt;/p&gt;

&lt;p&gt;This guide explains both, then looks at what nearly 90,000 of them, every one the NSE published from January 2023 to September 2026, actually contain. The biggest surprise: &lt;strong&gt;most bulk deals are not anyone taking a position.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a bulk deal?
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;bulk deal&lt;/strong&gt; is any trade, or set of trades, in which a single client buys or sells &lt;strong&gt;more than 0.5% of a company's listed shares in one day&lt;/strong&gt;. It happens in the normal market, alongside everyone else's orders.&lt;/p&gt;

&lt;p&gt;The broker who carries out the trade must report it to the exchange, and the exchange publishes the list after the market closes. The rule dates from a SEBI circular of January 2004.&lt;/p&gt;

&lt;p&gt;Two things follow from the definition:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;There is no rupee minimum.&lt;/strong&gt; In a small company, 0.5% of the shares can be worth a few lakh rupees. In a large one it can be hundreds of crores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A bulk deal can be reversed the same day.&lt;/strong&gt; Nothing stops a client from buying 0.6% of a company in the morning and selling it by the afternoon. Both legs are disclosed.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is a block deal?
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;block deal&lt;/strong&gt; is a single large trade done in a &lt;strong&gt;separate trading window&lt;/strong&gt;, away from the normal market. Since &lt;strong&gt;7 December 2025&lt;/strong&gt;, under SEBI's revised framework, the rules are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Minimum size ₹25 crore.&lt;/strong&gt; It was ₹10 crore before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two windows.&lt;/strong&gt; 8:45 am to 9:00 am, priced within 3% of the previous day's close; and 2:05 pm to 2:20 pm, priced within 3% of the volume-weighted price between 1:45 pm and 2:00 pm.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delivery is compulsory.&lt;/strong&gt; A block trade cannot be squared off or reversed. The shares really change hands.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The exchange publishes block deals after the close, just like bulk deals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bulk deal vs block deal at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Bulk deal&lt;/th&gt;
&lt;th&gt;Block deal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What defines it&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;More than 0.5% of the company's listed shares in a day&lt;/td&gt;
&lt;td&gt;One trade of at least ₹25 crore&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Where it trades&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The normal market&lt;/td&gt;
&lt;td&gt;A separate window, twice a day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Can it be reversed the same day?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No, delivery is compulsory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Typical size (median, 2023–2026)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;₹6 crore&lt;/td&gt;
&lt;td&gt;₹29 crore&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Published&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;After the close, same day&lt;/td&gt;
&lt;td&gt;After the close, same day&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What 90,000 NSE deals since 2023 show
&lt;/h2&gt;

&lt;p&gt;We looked at every bulk and block deal the NSE disclosed from &lt;strong&gt;3 January 2023 to 18 September 2026&lt;/strong&gt;: &lt;strong&gt;84,740 bulk deals and 5,898 block deals&lt;/strong&gt; across &lt;strong&gt;2,851 companies&lt;/strong&gt;. Rupee values below count only the buying side of each trade, because every trade is disclosed once for the buyer and once for the seller, and adding both would count it twice.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Year&lt;/th&gt;
&lt;th&gt;Bulk deals&lt;/th&gt;
&lt;th&gt;Bulk buying (₹ crore)&lt;/th&gt;
&lt;th&gt;Block deals&lt;/th&gt;
&lt;th&gt;Block buying (₹ crore)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2023&lt;/td&gt;
&lt;td&gt;18,525&lt;/td&gt;
&lt;td&gt;1,64,457&lt;/td&gt;
&lt;td&gt;1,094&lt;/td&gt;
&lt;td&gt;54,902&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2024&lt;/td&gt;
&lt;td&gt;26,680&lt;/td&gt;
&lt;td&gt;2,99,040&lt;/td&gt;
&lt;td&gt;1,556&lt;/td&gt;
&lt;td&gt;96,736&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025&lt;/td&gt;
&lt;td&gt;19,394&lt;/td&gt;
&lt;td&gt;2,96,427&lt;/td&gt;
&lt;td&gt;2,026&lt;/td&gt;
&lt;td&gt;1,22,832&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026, to 18 Sep&lt;/td&gt;
&lt;td&gt;20,141&lt;/td&gt;
&lt;td&gt;2,81,217&lt;/td&gt;
&lt;td&gt;1,222&lt;/td&gt;
&lt;td&gt;88,153&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Five things stand out.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Most bulk deals are same-day round trips
&lt;/h3&gt;

&lt;p&gt;In &lt;strong&gt;72% of bulk deals, the same client bought and sold the same stock on the same day.&lt;/strong&gt; Those round trips make up about two-thirds of all bulk buying by value, and the share barely moves from year to year: 71% of deals in 2023, 72% in 2024, 69% in 2025 and 77% so far in 2026.&lt;/p&gt;

&lt;p&gt;The clients behind them are mostly trading firms. The four with the most round trips are &lt;strong&gt;Graviton Research Capital, HRTI, QE Securities and NK Securities Research&lt;/strong&gt;. They typically trade in and out through the day, and each crossing of 0.5% gets disclosed, even though they end the day owning little or nothing.&lt;/p&gt;

&lt;p&gt;So a headline that says a firm "bought 0.8% of a company" may describe half of a trade that was reversed by the close. &lt;strong&gt;Before reading anything into a bulk deal, check whether the same client appears on the other side that day.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Bulk deals cluster in small companies
&lt;/h3&gt;

&lt;p&gt;Because the rule is a share of the company rather than a rupee amount, small companies cross it far more easily. The stocks with the most bulk deals since 2023 are small and recently listed ones: tickers &lt;strong&gt;ATALREAL (1,100 deals), MOBIKWIK (1,006), QUADFUTURE (629), MTNL (622) and RATNAVEER (569)&lt;/strong&gt;. The median bulk deal is about ₹6 crore.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Block deals are where large stakes change hands
&lt;/h3&gt;

&lt;p&gt;Block deals are fewer, much larger and cannot be reversed, so they are where real ownership moves. The median block deal is about ₹29 crore, and the largest single trades since 2023 run to thousands of crores:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bharti Airtel, 18 February 2025:&lt;/strong&gt; ₹8,485 crore, sold by Indian Continent Investment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Asian Paints, 12 June 2025:&lt;/strong&gt; ₹7,704 crore, bought by SBI Mutual Fund from Siddhant Commercials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adani Ports, 2 March 2023:&lt;/strong&gt; ₹5,282 crore, sold by the S.B. Adani Family Trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wipro, 9 June 2025:&lt;/strong&gt; ₹5,058 crore, sold by the Azim Premji Trust.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stocks with the most block deals are large, widely held ones: &lt;strong&gt;Lenskart, Kotak Mahindra Bank, Axis Bank, PB Fintech (PolicyBazaar) and Shriram Finance.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. June is block-deal season
&lt;/h3&gt;

&lt;p&gt;Almost a quarter of all block-deal value since 2023, &lt;strong&gt;23%, was traded in June&lt;/strong&gt;. August (14%), September (12%) and December (11%) follow. January and April each carried about 2%. One caution: October to December have three years of data in this period, not four, so those months are slightly under-counted.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Mutual funds are the steady buyers
&lt;/h3&gt;

&lt;p&gt;Mutual funds, taking every client whose name contains "Mutual Fund", were net buyers through bulk and block deals in every year, and by a growing amount:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Year&lt;/th&gt;
&lt;th&gt;Mutual funds, net bought (₹ crore)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2023&lt;/td&gt;
&lt;td&gt;14,018&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2024&lt;/td&gt;
&lt;td&gt;29,899&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025&lt;/td&gt;
&lt;td&gt;60,569&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026, to 18 Sep&lt;/td&gt;
&lt;td&gt;51,775&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;When a promoter or a foreign fund sells a large block, a domestic fund is often on the other side, as with SBI Mutual Fund in Asian Paints.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read a deal disclosure
&lt;/h2&gt;

&lt;p&gt;A bulk or block deal is a fact about a trade that already happened. To get something useful out of one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check both sides.&lt;/strong&gt; If the same client bought and sold that day, it is most likely trading, not a change in ownership.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the type.&lt;/strong&gt; A block deal ends in delivery; a bulk deal may not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare it with the company's size.&lt;/strong&gt; 0.5% of a small company can be a tiny rupee amount.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Look at who is on the other side.&lt;/strong&gt; A promoter selling to a mutual fund is a different story from two trading firms swapping shares.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep it in proportion.&lt;/strong&gt; One deal is one data point about the past.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What a deal cannot tell you.&lt;/strong&gt; A disclosure shows who traded, how much and at what price. It does not show why. A fund may be meeting redemptions, rebalancing to an index, or trading for a client. It also says nothing about where the price goes next. Read deals as context for your own research, never as a reason on their own.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  About the data
&lt;/h2&gt;

&lt;p&gt;The figures come from the NSE's bulk and block deal disclosures, 3 January 2023 to 18 September 2026, as republished by Ansaar. A &lt;strong&gt;round trip&lt;/strong&gt; means the same client name, in the same stock, on the same day, appears both as a buyer and as a seller in the bulk-deal list. Rupee totals are the buying side only. Mutual funds are identified by the words "Mutual Fund" in the client's name, so funds disclosed under other names are not counted. This is exchange data, not a model estimate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A bulk deal is more than 0.5% of a company's listed shares in a day, in the normal market; a block deal is one trade of at least ₹25 crore in a separate window.&lt;/li&gt;
&lt;li&gt;Block deals must end in delivery; bulk deals can be reversed the same day.&lt;/li&gt;
&lt;li&gt;72% of NSE bulk deals since 2023 were same-day round trips, mostly by trading firms.&lt;/li&gt;
&lt;li&gt;Bulk deals cluster in small companies; large stakes change hands through block deals.&lt;/li&gt;
&lt;li&gt;A deal tells you what happened, not why, and nothing about the next price.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;See the latest session's deals, with the client, the direction, quantity and price: &lt;a href="https://www.ansaar.in/equities/bulk-deals" rel="noopener noreferrer"&gt;NSE bulk deals&lt;/a&gt; and &lt;a href="https://www.ansaar.in/equities/block-deals" rel="noopener noreferrer"&gt;NSE block deals&lt;/a&gt;. For the bigger picture of foreign and domestic money, read &lt;a href="https://www.ansaar.in/learn/how-to-read-fii-dii-data" rel="noopener noreferrer"&gt;How to Read FII/DII Data&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the numbers were computed
&lt;/h2&gt;

&lt;p&gt;Each row is one client's side of a deal: date, symbol, client, direction, quantity and price. Two things matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Every trade appears twice.&lt;/strong&gt; A block deal between a seller and a buyer is disclosed once for each of them, so summing every row doubles the rupee value. The totals above count only the buying side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A round trip is a group, not a row.&lt;/strong&gt; Group bulk-deal rows by trade date, symbol and client, and flag the groups that contain both directions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;defaultdict&lt;/span&gt;

&lt;span class="n"&gt;sides&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BUY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;bulk_deals&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                      &lt;span class="c1"&gt;# one dict per disclosed row
&lt;/span&gt;    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trade_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;symbol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;client_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;sides&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;direction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;quantity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;sides&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;

&lt;span class="n"&gt;round_trips&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sides&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BUY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;span class="n"&gt;share&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sides&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rows&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;round_trips&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bulk_deals&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;share&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                     &lt;span class="c1"&gt;# 72.3% for Jan 2023 to Sep 2026
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole method. Client names are matched exactly as the exchange publishes them, so a firm disclosed under two spellings would be split, and the 72% is if anything a slight undercount.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions people ask
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between a bulk deal and a block deal?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A bulk deal is any trade, or set of trades, in which one client buys or sells more than 0.5% of a company's listed shares in a single day, in the normal market. A block deal is one large trade of at least ₹25 crore, done in a separate trading window twice a day, and it must end in delivery. So a bulk deal is defined by its share of the company, a block deal by its size and where it is traded.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the minimum size of a block deal?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;₹25 crore, since 7 December 2025. SEBI raised it from ₹10 crore in its revised block deal framework, dated 8 October 2025. There is no rupee minimum for a bulk deal. What makes a trade a bulk deal is crossing 0.5% of the company's listed shares.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What are the block deal window timings?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are two windows. The morning window runs from 8:45 am to 9:00 am, priced within 3% of the previous day's close. The afternoon window runs from 2:05 pm to 2:20 pm, priced within 3% of the volume-weighted average price between 1:45 pm and 2:00 pm.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a large bulk purchase mean the buyer expects the stock to rise?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not necessarily, and usually not. In the NSE's bulk-deal disclosures from January 2023 to September 2026, 72% of bulk deals were round trips: the same client bought and sold the same stock on the same day. Those are mostly trading firms, not investors taking a position. Even a genuine purchase tells you what happened, not why, and nothing about what the price will do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where can I see today's bulk and block deals?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The NSE publishes both after the market closes each trading day. Ansaar republishes them on its bulk deals and block deals pages, with the latest session in full and recent history. Because they come out after the close, the newest figures are usually for the previous trading day.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This guide first appeared on &lt;a href="https://www.ansaar.in/learn/bulk-deals-vs-block-deals" rel="noopener noreferrer"&gt;Ansaar&lt;/a&gt;, a free quantitative-research site for Indian markets. It is educational material, not investment advice.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>datascience</category>
      <category>dataengineering</category>
      <category>fintech</category>
    </item>
    <item>
      <title>Can DuckDB be your SaaS product's warehouse? Where the ceiling actually is</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Wed, 23 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/can-duckdb-be-your-saas-products-warehouse-where-the-ceiling-actually-is-76n</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/can-duckdb-be-your-saas-products-warehouse-where-the-ceiling-actually-is-76n</guid>
      <description>&lt;p&gt;A recurring question, roughly: &lt;em&gt;we're a small SaaS, we need analytics — internal dashboards and eventually a customer-facing one. Can DuckDB be the warehouse, or do we need Snowflake/BigQuery from day one?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Almost always: yes, it can, for longer than you'd think. But "small SaaS" hides three very different workloads, and only two of them fit comfortably. Here's how to tell which you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three workloads hiding behind "analytics"
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Internal analytics.&lt;/strong&gt; You and a few colleagues asking questions about signups, churn, revenue. Handful of people, no concurrency to speak of, latency measured in "before my coffee gets cold."&lt;/p&gt;

&lt;p&gt;DuckDB fits this with enormous room to spare. This is the easiest yes in data engineering. A nightly job writes Parquet, DuckDB reads it, your BI tool or notebook queries it. Done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Embedded customer-facing analytics.&lt;/strong&gt; A dashboard &lt;em&gt;inside your product&lt;/em&gt; — each customer sees their own numbers. This is where it gets interesting, and where the answer becomes "yes, with a specific design."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. A shared analytics platform for a data team.&lt;/strong&gt; Many analysts, ad-hoc SQL, governance, roles. This is a warehouse workload. If you genuinely have this, buy a warehouse. Most companies asking the question do not have this yet, and some never will.&lt;/p&gt;

&lt;p&gt;The mistake is buying for (3) while actually living in (1).&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the ceiling is, concretely
&lt;/h2&gt;

&lt;p&gt;The instinct is to think about total data volume. That's the wrong number. Three things bound a single-node design, roughly in the order you'll hit them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrent write throughput.&lt;/strong&gt; DuckDB is single-writer. One process writes at a time. For a SaaS this is usually fine because analytics writes are batch — a job runs, produces Parquet, exits. It becomes a problem the moment you want near-real-time ingestion from multiple services writing continuously. That's the first real ceiling, and it arrives independent of data size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrent query volume.&lt;/strong&gt; A single node serving customer-facing dashboards can handle a genuinely surprising number of requests when queries are small and well-partitioned — but it's one machine. When p99 latency starts drifting under load and you're already caching, you've found the second ceiling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Working-set size.&lt;/strong&gt; Last, not first. Modern servers take a lot of RAM, and Parquet with good partitioning means a query touches a fraction of the data. Teams hit the concurrency ceilings long before the size ceiling.&lt;/p&gt;

&lt;p&gt;Notice none of these is "how many GB do you have". Which is exactly why "we have 500GB, we need Snowflake" is a non-sequitur.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/duckdb-as-a-saas-warehouse/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Can DuckDB be your SaaS product's warehouse? Where the ceiling actually is&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>analytics</category>
    </item>
    <item>
      <title>Market regime detection in production: what the model actually changes</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Tue, 22 Sep 2026 03:30:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/market-regime-detection-in-production-what-the-model-actually-changes-3i3i</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/market-regime-detection-in-production-what-the-model-actually-changes-3i3i</guid>
      <description>&lt;p&gt;A while ago I wrote up &lt;a href="https://hikmahtechnologies.com/blog/market-regime-detection-from-hidden-markov-models-to-wasserstein-clustering-6ba0a09559dc/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;two ways to detect market regimes&lt;/a&gt;: hidden Markov models and clustering on Wasserstein distance. That post was research on a toy S&amp;amp;P 500 series. This one is what actually runs, every trading day, inside the trading system behind &lt;a href="https://hikmahtechnologies.com/products/ansaar/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Ansaar&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The short version: the model turned out to be the small part. What made it safe to run without me watching was everything around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three states, three features
&lt;/h2&gt;

&lt;p&gt;The model is a &lt;code&gt;GaussianHMM&lt;/code&gt; from hmmlearn with three states. I tried more. The Bayesian information criterion kept picking three, and three is also the number a human can act on.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;What it looks like&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bull trend&lt;/td&gt;
&lt;td&gt;Positive returns, moderate volatility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bear trend&lt;/td&gt;
&lt;td&gt;Negative returns, elevated volatility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sideways&lt;/td&gt;
&lt;td&gt;Near-zero returns, volatility all over the place&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It sees three features, and only three:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;ret_5d&lt;/code&gt;, the 5-day log return&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;realized_vol_21d&lt;/code&gt;, the 21-day rolling standard deviation, annualised&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;return_vol_ratio&lt;/code&gt;, the first divided by the second&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I had a longer list at one point. Every feature I added made the fit look better in-sample and the labels worse out of it. Three features is enough to separate "going up calmly" from "going down violently" from "going nowhere", and that is the whole job.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug that scaling fixed
&lt;/h2&gt;

&lt;p&gt;The first version fed those three features to the model raw. It trained fine, the states had names, and then I looked at the state statistics and the bear trend had a &lt;em&gt;positive&lt;/em&gt; average return.&lt;/p&gt;

&lt;p&gt;The cause was scale. &lt;code&gt;return_vol_ratio&lt;/code&gt; moves over a much wider range than a 5-day return, and unscaled it had about 19 times the influence of &lt;code&gt;ret_5d&lt;/code&gt;. The model was clustering on the ratio and mostly ignoring the return. The tell was March 2020: the COVID crash, the most obvious bear regime in the training window, was classified as sideways.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;StandardScaler&lt;/code&gt; in front of the model fixed it, and the verification is the part I keep:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;March 2020: 100% of days classified as bear trend&lt;/li&gt;
&lt;li&gt;23 March 2020, the worst single day: bear trend, a -13.93% return at 70.3% volatility, with 100% state probability
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;hmmlearn.hmm&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;GaussianHMM&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.preprocessing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StandardScaler&lt;/span&gt;

&lt;span class="n"&gt;scaler&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;features&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ret_5d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;realized_vol_21d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;return_vol_ratio&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]])&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;GaussianHMM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_components&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;covariance_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;full&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Keep the scaler with the model. Scoring unscaled data later is the same bug again.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you take one thing from this post, take that: check the model against a day you already know the answer to.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/market-regime-detection-in-production/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Market regime detection in production: what the model actually changes&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Running DuckDB on your own infrastructure: a production setup</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/running-duckdb-on-your-own-infrastructure-a-production-setup-530d</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/running-duckdb-on-your-own-infrastructure-a-production-setup-530d</guid>
      <description>&lt;p&gt;If you're asking how to run DuckDB on your own infrastructure for faster queries, you've already made the important decision. What's left is mostly avoiding four or five specific mistakes that make a single node look slower than it is.&lt;/p&gt;

&lt;p&gt;This is the setup I actually run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the shape, not the settings
&lt;/h2&gt;

&lt;p&gt;The layout matters more than any tuning flag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;object storage (or a local disk)
  └── warehouse/
      └── events/
          dt=2026-08-11/part-0.parquet
          dt=2026-08-12/part-0.parquet
          dt=2026-08-13/part-0.parquet

one writer process   →  produces new partitions
N reader processes   →  open read-only, query across partitions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One writer. Many readers. Immutable partitions. Everything else is detail.&lt;/p&gt;

&lt;p&gt;If you take one thing from this: &lt;strong&gt;don't put a mutable &lt;code&gt;.duckdb&lt;/code&gt; file at the centre of a multi-process system.&lt;/strong&gt; DuckDB is single-writer, and the moment two processes want to write, you're fighting the design. Write Parquet, read Parquet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The settings that actually matter
&lt;/h2&gt;

&lt;p&gt;Most DuckDB tuning advice is noise. These four are not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Leave real headroom. This is DuckDB's budget, not the container's.&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;memory_limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'12GB'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Spilling must land on real disk with real space.&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;temp_directory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'/var/lib/duckdb/tmp'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Match the cores you actually have, not the ones the host advertises.&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;threads&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Only if you're reading from object storage.&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;preserve_insertion_order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;memory_limit&lt;/code&gt;&lt;/strong&gt; should sit meaningfully below your container limit — I use roughly 70–75%. DuckDB accounts for its own buffer pool, not for the Python process around it, the Arrow tables in flight, or the runtime. Set it to the container limit and the orchestrator kills you before DuckDB ever decides to spill.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;temp_directory&lt;/code&gt;&lt;/strong&gt; is the one that bites in containers. The default may point at a path on the container's ephemeral layer, which is often small and sometimes memory-backed. A large join then fills it and the pod dies looking like an OOM when it's really a disk problem. Mount a volume and point at it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;threads&lt;/code&gt;&lt;/strong&gt; matters in Kubernetes specifically. DuckDB sees the &lt;em&gt;host's&lt;/em&gt; core count, not your CPU limit. On a 64-core node with a 4-core limit, it will happily spawn 64 threads and spend its life being throttled. Set this from your actual limit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;preserve_insertion_order = false&lt;/code&gt;&lt;/strong&gt; lets DuckDB parallelise reads more aggressively when you don't care about row order — which, for aggregate queries over Parquet, you usually don't.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/running-duckdb-on-your-own-infra/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Running DuckDB on your own infrastructure: a production setup&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>python</category>
    </item>
    <item>
      <title>Is DuckDB safe for production? The honest limitations</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Fri, 18 Sep 2026 14:00:01 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/is-duckdb-safe-for-production-the-honest-limitations-2h8k</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/is-duckdb-safe-for-production-the-honest-limitations-2h8k</guid>
      <description>&lt;p&gt;"Is DuckDB safe for production?" is the right question asked slightly wrong. DuckDB is not a scaled-down toy that becomes safe once you're brave enough. It's a database with a specific concurrency and durability model, and it is completely safe inside that model and genuinely unsafe outside it.&lt;/p&gt;

&lt;p&gt;So the useful version of the question is: &lt;strong&gt;safe for which workload?&lt;/strong&gt; Here's what actually bites, in the order it bites people.&lt;/p&gt;

&lt;h2&gt;
  
  
  The single-writer model is the whole story
&lt;/h2&gt;

&lt;p&gt;DuckDB allows &lt;strong&gt;one writing process at a time&lt;/strong&gt; against a database file. Multiple threads inside that process are fine. Many processes can read the file at once, as long as none of them is writing. Two processes both opening the file for writing are not.&lt;/p&gt;

&lt;p&gt;This is not a bug or a temporary limitation — it's the design. DuckDB is an in-process database, like SQLite. There is no server arbitrating between clients, because there is no server.&lt;/p&gt;

&lt;p&gt;Almost every "DuckDB isn't production-ready" story I've heard traces back to this. Someone runs the ETL job on a schedule, someone else points a dashboard at the same file, and eventually the two overlap. What you get is a lock error, or — if you were clever enough to copy the file to dodge the lock — a reader seeing a half-written state.&lt;/p&gt;

&lt;p&gt;The design that avoids it entirely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌─ writer job ──► Parquet partitions on S3
                    │                   (immutable, append-only)
  source systems ───┤
                    └─ readers ────► DuckDB, N processes, read-only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Writers never share a mutable file. They produce &lt;strong&gt;immutable Parquet partitions&lt;/strong&gt;. Readers open those partitions read-only, as many processes as you like, with no coordination at all. The concurrency problem disappears because you removed the shared mutable state, not because you managed it better.&lt;/p&gt;

&lt;p&gt;If you must use a &lt;code&gt;.duckdb&lt;/code&gt; file with concurrent readers, open them explicitly read-only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;duckdb&lt;/span&gt;

&lt;span class="c1"&gt;# Safe: many of these can run at once against the same file.
&lt;/span&gt;&lt;span class="n"&gt;con&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;duckdb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analytics.duckdb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;read_only&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read-only processes share the file with each other but not with a writer: while any process holds it read-write, the others cannot open it at all. That is why the writer should produce Parquet rather than update the file the readers use. DuckDB will also refuse a write through a read-only connection rather than corrupt anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Durability: it's ACID, but know what that covers
&lt;/h2&gt;

&lt;p&gt;DuckDB is ACID-compliant with write-ahead logging. A transaction that commits is durable; a crash mid-transaction rolls back cleanly. On this axis it behaves like a real database, because it is one.&lt;/p&gt;

&lt;p&gt;What people actually get wrong is what sits &lt;em&gt;around&lt;/em&gt; the transaction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The file is a single point of failure.&lt;/strong&gt; ACID protects you from a crash. It doesn't protect you from a deleted file, a corrupted volume, or an ephemeral container's disk vanishing at the end of the run. If your database lives on a pod's local disk, it lives exactly as long as the pod.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;.duckdb&lt;/code&gt; file is not a backup format.&lt;/strong&gt; Copying it while a writer is mid-transaction gives you a file whose contents are undefined. Snapshot by exporting (&lt;code&gt;EXPORT DATABASE&lt;/code&gt;) or by treating Parquet as the durable layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage format compatibility.&lt;/strong&gt; DuckDB reached 1.0 with a stability commitment, and files are forward-compatible within that line — but a file written by a newer version isn't necessarily readable by an older one. Pin the version in production and upgrade deliberately, the same as any database.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule I use: &lt;strong&gt;Parquet is the durable artefact, the &lt;code&gt;.duckdb&lt;/code&gt; file is a cache.&lt;/strong&gt; If losing the file would be a data-loss incident rather than an inconvenience, the architecture is wrong, not DuckDB.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/is-duckdb-safe-for-production/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;Is DuckDB safe for production? The honest limitations&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>python</category>
    </item>
    <item>
      <title>DuckDB vs Snowflake: the 4 questions that decide it</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Wed, 16 Sep 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/duckdb-vs-snowflake-the-4-questions-that-decide-it-2m85</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/duckdb-vs-snowflake-the-4-questions-that-decide-it-2m85</guid>
      <description>&lt;p&gt;DuckDB and Snowflake get compared as if they're competitors. They mostly aren't. They answer different questions, and picking the wrong one is how you end up either paying for a warehouse you don't need or outgrowing a laptop-sized tool in production. Here's how I decide.&lt;/p&gt;

&lt;h2&gt;
  
  
  They're built for different shapes of problem
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Snowflake&lt;/strong&gt; is a cloud data warehouse. It separates storage from compute, scales elastically, handles many concurrent users, and comes with governance, sharing and a large ecosystem. You point a cluster at your data and a hundred analysts can query it at once. You pay per second of compute, by the credit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DuckDB&lt;/strong&gt; is an in-process analytical database — think "SQLite for analytics." It runs inside your Python process, your laptop, or a single server. No cluster, no service to operate, no per-query meter. It reads Parquet and CSV directly, including files sitting on S3, and it is very fast on a single machine.&lt;/p&gt;

&lt;p&gt;One is a rented warehouse with a loading dock and a staff. The other is a workbench in your own garage. The question is which one the job needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision usually comes down to four things
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Does your data fit on one big machine?&lt;/strong&gt; This is the one people get wrong. "Big data" is rarer than the marketing suggests. A single modern server handles hundreds of gigabytes to a few terabytes comfortably, and DuckDB is built to use all of it. If your working set is in that range — and most companies' is — a single-node engine is not a compromise, it's the right tool. If you're genuinely at tens of terabytes scanned per query, or petabytes at rest, that's Snowflake territory.&lt;/p&gt;

&lt;p&gt;The number that matters is not your &lt;em&gt;total&lt;/em&gt; data. It's the &lt;strong&gt;working set&lt;/strong&gt;: the bytes a typical query actually touches. Columnar Parquet plus predicate pushdown means a query over a 2 TB table that filters to last month and selects six columns may read a few gigabytes. Teams routinely provision for the 2 TB and never measure the few gigabytes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. How many people query it at once?&lt;/strong&gt; DuckDB is fundamentally single-node. It's perfect for one analyst, a transformation job, or an app backend serving queries it controls. It is not built for fifty analysts running ad-hoc dashboards simultaneously. Concurrency at that scale is exactly what Snowflake's elastic compute is for.&lt;/p&gt;

&lt;p&gt;Be precise about what "concurrency" means for you, though. Fifty people with a dashboard open is not fifty concurrent queries — it's fifty mostly-idle browser tabs hitting a cache. Fifty analysts writing exploratory SQL at 10am on a Monday &lt;em&gt;is&lt;/em&gt; concurrency. The first case a single node handles fine behind a result cache; the second it does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Who runs it, and do they want to run anything?&lt;/strong&gt; Snowflake is zero-ops — there's no server to keep alive. DuckDB has nothing to operate either, but only because it lives &lt;em&gt;inside&lt;/em&gt; something you already run. If you want a managed, hands-off, governed platform for a whole org, that's Snowflake. If you want a fast engine embedded in a pipeline or a notebook, that's DuckDB.&lt;/p&gt;

&lt;p&gt;This is the question teams answer emotionally. "We don't want to manage infrastructure" is usually true and usually decisive — but notice that DuckDB-over-Parquet has no infrastructure to manage either. What it has is &lt;em&gt;no vendor to call&lt;/em&gt;. Those aren't the same thing, and which one you're actually buying matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. What's the cost model doing to you?&lt;/strong&gt; Snowflake bills compute by the second against credits. That's elastic and fair when usage is spiky, and brutal when a scheduled job or a careless dashboard leaves a warehouse running. DuckDB's compute cost is whatever the machine it runs on already costs — often effectively zero, because it's your existing CI runner, app server or laptop.&lt;/p&gt;

&lt;p&gt;The asymmetry worth internalising: Snowflake's bill scales with &lt;em&gt;how you query&lt;/em&gt;, DuckDB's with &lt;em&gt;what you already rent&lt;/em&gt;. A badly-written query on Snowflake costs money every time it runs. The same query on a box you're already paying for costs nothing extra — it's just slow, which is a problem you can see and ignore.&lt;/p&gt;




&lt;p&gt;This is the first part. The full post — including the rest of the working details — is on my site: &lt;a href="https://hikmahtechnologies.com/blog/duckdb-vs-snowflake/?utm_source=devto&amp;amp;utm_medium=syndication&amp;amp;utm_campaign=blog" rel="noopener noreferrer"&gt;DuckDB vs Snowflake: the 4 questions that decide it&lt;/a&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>database</category>
      <category>analytics</category>
    </item>
    <item>
      <title>The most dangerous API response is HTTP 200 with an empty body</title>
      <dc:creator>Mohammed Arshad Ansari</dc:creator>
      <pubDate>Tue, 15 Sep 2026 10:30:00 +0000</pubDate>
      <link>https://dev.to/mohammed_arshadansari_f2/the-most-dangerous-api-response-is-http-200-with-an-empty-body-40c2</link>
      <guid>https://dev.to/mohammed_arshadansari_f2/the-most-dangerous-api-response-is-http-200-with-an-empty-body-40c2</guid>
      <description>&lt;p&gt;Every pipeline eventually inherits a source that moves. Ours did: &lt;code&gt;dataservices.imf.org&lt;/code&gt;&lt;br&gt;
stopped resolving in DNS entirely — not a 500, not a timeout, the hostname itself was gone.&lt;br&gt;
The IMF had migrated its data platform and the old host was on its way out. No villain in&lt;br&gt;
this story; following a publisher to its new home is the consumer's job.&lt;/p&gt;

&lt;p&gt;The interesting part is the failure mode the migration exposed, which has nothing to do with&lt;br&gt;
the IMF and everything to do with how most of us write ingestion code.&lt;/p&gt;
&lt;h2&gt;
  
  
  The symptom
&lt;/h2&gt;

&lt;p&gt;Monthly CPI for a set of non-OECD countries went quietly stale, frozen at December 2025, while&lt;br&gt;
every other source kept ticking. Nothing screamed. The daily job ran green. Downstream scoring&lt;br&gt;
kept computing on last-known values, because CPI resolves through a fallback chain — if the&lt;br&gt;
freshest monthly print is unavailable, the model reaches for the next-best source rather than&lt;br&gt;
failing outright.&lt;/p&gt;

&lt;p&gt;That fallback is a feature: one dead source should not take scoring down for a whole tier of&lt;br&gt;
countries. It also has a cost. &lt;strong&gt;Graceful degradation and silent staleness are the same&lt;br&gt;
mechanism viewed from two angles.&lt;/strong&gt; A freshness dashboard is what keeps the second angle&lt;br&gt;
visible; without one, the design that saves you also hides the problem from you.&lt;/p&gt;
&lt;h2&gt;
  
  
  The migration itself
&lt;/h2&gt;

&lt;p&gt;The new home is &lt;code&gt;api.imf.org&lt;/code&gt;, on SDMX 2.1, and it still speaks the same StructureSpecific XML&lt;br&gt;
dialect — so the parser barely changed. All the pain was in the &lt;em&gt;keys&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CPI now lives on the &lt;code&gt;IMF.STA,CPI&lt;/code&gt; dataflow, and the all-items series is keyed with COICOP
&lt;code&gt;_T&lt;/code&gt; (the SDMX "total" convention), not the &lt;code&gt;CP00&lt;/code&gt; the old platform used.&lt;/li&gt;
&lt;li&gt;Balance-of-payments moved to &lt;code&gt;IMF.STA,BOP&lt;/code&gt;, where the legacy indicator codes do &lt;strong&gt;not&lt;/strong&gt; port
one-to-one. No find-and-replace; each series had to be re-derived against the new structure
definition and re-tested.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tedious rather than clever — and easy to get subtly wrong, because of what a wrong key does.&lt;/p&gt;
&lt;h2&gt;
  
  
  The gotcha worth stealing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A bad key returns HTTP 200 with an empty &lt;code&gt;DataSet&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not a 404. Not a 400. A cheerful &lt;code&gt;200 OK&lt;/code&gt;, a well-formed SDMX document, zero observations&lt;br&gt;
inside it. If your ingestion trusts the status line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;          &lt;span class="c1"&gt;# &amp;amp;lt;- the bug
&lt;/span&gt;    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;store&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                &lt;span class="c1"&gt;# stores nothing, reports success
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;…then a typo in a dimension code, or an un-migrated series key, sails through as success and&lt;br&gt;
writes nothing. Job green. Data frozen. You find out weeks later from a freshness alert, if&lt;br&gt;
you have one, or from a customer, if you do not.&lt;/p&gt;

&lt;p&gt;Our fix is a &lt;strong&gt;canary&lt;/strong&gt;: the request always carries a segment we know must return data — United&lt;br&gt;
States all-items CPI — and a parse that yields zero series therefore cannot mean "no data".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;series&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse_sdmx21_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;xml_bytes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;series&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;IMFIFSUnavailableError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;IMF.STA,CPI returned an empty DataSet incl. the USA canary — &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;key schema change or platform fault&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The USA rows are then filtered &lt;em&gt;out&lt;/em&gt; before storage — its index base differs from our other&lt;br&gt;
US source and double-sourcing would corrupt the derived series. The canary lives in the&lt;br&gt;
request, not in the warehouse.&lt;/p&gt;

&lt;p&gt;Three lines, and it is the difference between finding a structural break in one run versus one&lt;br&gt;
month. Having built it for CPI, it generalises to every SDMX-style source that can return a&lt;br&gt;
well-formed empty response: pick a segment that must exist, assert it came back populated, let&lt;br&gt;
a broken key contract throw. Note what it is &lt;em&gt;not&lt;/em&gt; for: a genuine upstream outage is a&lt;br&gt;
different event, and that one deliberately does &lt;strong&gt;not&lt;/strong&gt; raise — it materialises with an&lt;br&gt;
&lt;code&gt;imf_ifs_outage&lt;/code&gt; flag, because failing hard there would skip every downstream scoring asset in&lt;br&gt;
the same run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same bug one layer up: a fetch that succeeds is not a document
&lt;/h2&gt;

&lt;p&gt;The identical mistake, in a different shape, was sitting in our central-bank statement&lt;br&gt;
scrapers. They fetched a listing page, followed a link, extracted text, and handed it to an&lt;br&gt;
LLM for a hawkish/dovish read. Every step returned 200. Every step "worked".&lt;/p&gt;

&lt;p&gt;What was actually being scored, once we read the stored text back:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a bank's login form, because its statement listing redirects to &lt;code&gt;/login&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;a press-release index whose top headline was Treasury-bill auction results&lt;/li&gt;
&lt;li&gt;a SharePoint "you may be trying to access this site from a secured browser" notice&lt;/li&gt;
&lt;li&gt;115 characters of "this page depends on JavaScript"&lt;/li&gt;
&lt;li&gt;a rates &lt;em&gt;statistics&lt;/em&gt; nav link, because the link pattern &lt;code&gt;interest.*rate&lt;/code&gt; matched it and it
sorted first in the DOM&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The parallel to the empty &lt;code&gt;DataSet&lt;/code&gt; is exact: nothing failed, so nothing raised, so a&lt;br&gt;
plausible-looking value went downstream. The fixes were the same shape as the canary —&lt;br&gt;
assert the &lt;em&gt;content&lt;/em&gt;, not the transport:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A length floor and a bot/JS-notice check, so a page that cannot be a statement is skipped
rather than scored.&lt;/li&gt;
&lt;li&gt;Feeds and APIs instead of HTML scraping where the bank publishes one (RSS/Atom, or the
PDF-minutes API where the web page is a JS shell).&lt;/li&gt;
&lt;li&gt;A source is only enabled once its extracted text has been read back and confirmed to be a
monetary policy decision — and a source that cannot pass is switched off &lt;em&gt;in config with
the reason recorded&lt;/em&gt;, not left nominally "covered". Fifteen of twenty configured banks are
on; five are off behind commercial bot management, a client-rendered shell, hard 403s on
the decision pages, or a listing API that refuses anonymous reads — protections we decline
to work around.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Two reusable lessons
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Assert data presence, not status codes.&lt;/strong&gt; &lt;code&gt;200 OK&lt;/code&gt; means the HTTP conversation
succeeded. It says nothing about whether the payload contains what you asked for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connection outages and silent-empty responses are different failure modes needing
different handling.&lt;/strong&gt; A DNS failure throws — retry with backoff and you will know. A
200-plus-empty never throws, so &lt;code&gt;try/except&lt;/code&gt; + retry leaves it completely uncovered. It
needs an affirmative presence check. Handle both, separately, on purpose.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The payoff, incidentally, was not just damage control: following the source to its new home&lt;br&gt;
took BOP coverage from roughly 70 to 148 countries and made reserves available quarterly. The&lt;br&gt;
endpoint you are reluctantly forced onto is often the one the publisher is actually investing&lt;br&gt;
in.&lt;/p&gt;

&lt;p&gt;Full write-up, with the migration detail:&lt;/p&gt;

&lt;p&gt;What the pipeline feeds — a daily 0–100 credibility score for 169 countries on 100% free&lt;br&gt;
public data, with per-indicator provenance:&lt;/p&gt;

</description>
      <category>python</category>
      <category>dataengineering</category>
      <category>api</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
