<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Puneet Chandna</title>
    <description>The latest articles on DEV Community by Puneet Chandna (@puneet-chandna).</description>
    <link>https://dev.to/puneet-chandna</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3390583%2Facbb5182-5adc-405b-af74-a10a17ce55d0.jpeg</url>
      <title>DEV Community: Puneet Chandna</title>
      <link>https://dev.to/puneet-chandna</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/puneet-chandna"/>
    <language>en</language>
    <item>
      <title>How I built recall-first search on PostgreSQL</title>
      <dc:creator>Puneet Chandna</dc:creator>
      <pubDate>Tue, 15 Sep 2026 23:04:23 +0000</pubDate>
      <link>https://dev.to/puneet-chandna/why-we-built-recall-first-search-on-postgresql-4ga4</link>
      <guid>https://dev.to/puneet-chandna/why-we-built-recall-first-search-on-postgresql-4ga4</guid>
      <description>&lt;p&gt;Our careers page search often felt broken. Searching for “Java” could return JavaScript roles. A misspelling could produce no results at all. Searching for “manager” could collect assistant-manager vacancies across unrelated departments, without distinguishing what those jobs involved.&lt;/p&gt;

&lt;p&gt;The implementation was matching words inside job metadata. It did not understand spelling variations, equivalent terminology, or descriptions of the work a candidate wanted to do. Relevant vacancies could disappear simply because the candidate and employer used different words.&lt;br&gt;
Replacing that behavior with hybrid retrieval solved part of the problem—and exposed another. Our first hybrid evaluation achieved an NDCG@10 of 0.9944, yet returned precision@5 on semantic queries was only 55%. We could now find relevant jobs that literal matching missed, but we still had to decide how many weak matches should accompany them.&lt;/p&gt;

&lt;p&gt;We eventually built selective hybrid retrieval inside our existing Next.js and PostgreSQL stack, using our existing OpenAI integration for embeddings. Before choosing that design, we considered Elasticsearch , managed search, and self-hosted inference. Each would have solved part of the problem while changing which systems we had to operate and keep consistent.&lt;/p&gt;

&lt;p&gt;Choosing the smaller architecture left substantial engineering work of our own. The evaluation would force us to define what a useful result set meant. Production would test whether that definition survived stale data, late responses, and provider failures.&lt;/p&gt;

&lt;p&gt;I worked on this search system as part of the engineering team at Hyr.works , where Hyr's careers search is used to help candidates find open roles. The implementation described here is the result of that work.&lt;/p&gt;
&lt;h2&gt;
  
  
  The limits of fixing keyword matching
&lt;/h2&gt;

&lt;p&gt;The keyword path preceding hybrid search assembled job titles, categories, skills, locations, and other metadata into lowercase text. A job matched when every query term appeared somewhere in that text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Simplified from the previous search path.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;searchableFields&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flat&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Boolean&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;keywords&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;every&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;word&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;word&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was easy to understand, but its failure modes followed directly from the representation. &lt;code&gt;kubernets&lt;/code&gt; did not contain &lt;code&gt;kubernetes&lt;/code&gt;. Aliases needed explicit treatment. A description of responsibilities could fail to overlap with the title or metadata. Conversely, a substring match could blur distinctions such as Java and JavaScript.&lt;/p&gt;

&lt;p&gt;Our investigation started around skipped end-to-end tests, but we did not treat a skipped assertion as proof of a production defect. Manual reproduction confirmed careers-search problems; another reported hydration issue did not reproduce. That distinction mattered: fixing search did not require accepting every diagnosis attached to the original test audit.&lt;/p&gt;

&lt;p&gt;The new requirement was broader than making those assertions pass. We needed lexical precision for named technologies and semantic recall for descriptions of work, while keeping company boundaries and explicit filters exact.&lt;/p&gt;

&lt;p&gt;The application already used Next.js and Supabase/PostgreSQL. Jobs had multiple writers, including another backend and direct database operations. Any replacement needed to see those changes without making every writer responsible for updating search.&lt;/p&gt;

&lt;p&gt;The intended change was larger than adding more fields to the string comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Existing keyword path&lt;/th&gt;
&lt;th&gt;Replacement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Every term must occur literally in concatenated metadata&lt;/td&gt;
&lt;td&gt;Weighted lexical retrieval plus semantic candidates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Misspellings and aliases need literal overlap&lt;/td&gt;
&lt;td&gt;Reviewed aliases and bounded typo recovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The same containment check handles every nonempty query&lt;/td&gt;
&lt;td&gt;Separate lexical-preview, lexical-only, and hybrid paths&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search runs against the loaded job list in the browser&lt;/td&gt;
&lt;td&gt;Server retrieval applies tenant and selected filters before candidate limits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That describes the capability we wanted, not yet the system we should buy or build.&lt;/p&gt;

&lt;h3&gt;
  
  
  Elasticsearch
&lt;/h3&gt;

&lt;p&gt;The problem wasn't that Elasticsearch couldn't solve search.&lt;/p&gt;

&lt;p&gt;It was that we'd have to introduce another service, another deployed system, another copy of the searchable inventory, and another indexing pipeline to keep that copy synchronized with PostgreSQL.&lt;/p&gt;

&lt;p&gt;There was also a simpler practical question: introducing another service meant another piece of infrastructure to approve, deploy, monitor, and pay for. I didn't want to ask for that complexity and another recurring bill unless we had a strong reason to need it.&lt;/p&gt;

&lt;p&gt;Our jobs already lived in PostgreSQL, and PostgreSQL could handle the filtering and lexical retrieval we needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Managed search
&lt;/h3&gt;

&lt;p&gt;A search-as-a-service product reduced some of the operational work, but it didn't remove the integration problem.&lt;br&gt;
We would still need to decide what to index, keep the external index synchronized, and make sure public visibility matched the source database. We'd also be giving another system part of the search policy itself.&lt;br&gt;
For this release, we wanted to control how lexical matches, semantic candidates, hard filters, and relevance thresholds interacted.&lt;/p&gt;

&lt;p&gt;So we didn't add another vendor either.&lt;/p&gt;
&lt;h3&gt;
  
  
  Running our own embedding model
&lt;/h3&gt;

&lt;p&gt;Self-hosting EmbeddingGemma on our DigitalOcean infrastructure looked more attractive because it avoided another external API.&lt;br&gt;
But spare CPU and memory aren't the same thing as having an inference service.We would own model deployment, resource allocation, concurrency, availability, and another runtime alongside the web application. More importantly, changing the embedding model isn't just changing an API call. Query vectors and indexed job vectors need to exist in a compatible embedding space.&lt;br&gt;
That meant a new model would require a new compatible index generation.&lt;br&gt;
We didn't have workload evidence that taking on that operational cost would materially improve this release.&lt;/p&gt;
&lt;h3&gt;
  
  
  Using what we already had
&lt;/h3&gt;

&lt;p&gt;The application already used GPT/OpenAI models for several AI and agentic workflows. We already had the provider integration and the operational patterns around it. Using an embedding model through that existing integration gave us semantic retrieval without adding another inference runtime just for search. We chose text-embedding-3-small specifically for embeddings. GPT-generated text was not involved in ranking jobs.&lt;/p&gt;

&lt;p&gt;The result was a deliberately small architecture:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jgqpuiwqkf7jzwr8sof.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jgqpuiwqkf7jzwr8sof.png" alt="job search indexing architecture" width="800" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We weren't trying to build another Elasticsearch. We were building the smallest search system that fit our workload.&lt;/p&gt;

&lt;p&gt;That decision kept the number of systems down, but it also meant the search policy, indexing lifecycle, failure handling, and rollout were ours to solve.&lt;br&gt;
Lexical retrieval reads the projection; hybrid retrieval also compares compatible vectors. The provider returns embeddings, while PostgreSQL decides which jobs are eligible and how candidates rank. These were architectural choices, not evidence that this stack would outperform every dedicated search product. We still had to measure the system we chose.&lt;/p&gt;
&lt;h2&gt;
  
  
  Give each retrieval mechanism a narrower job
&lt;/h2&gt;

&lt;p&gt;Our first retrieval decision was to avoid asking one mechanism to handle incompatible notions of a match. Named technologies need lexical precision; descriptions of work need more flexibility. We combined weighted PostgreSQL full-text search, bounded typo recovery, and vector similarity.&lt;/p&gt;

&lt;p&gt;Titles and primary skills receive the strongest lexical weight. Legacy skills, category, and industry follow; descriptions and location text receive less. The projection normalizes arrays, JSON-encoded skill lists, and legacy strings instead of assuming one clean source format.&lt;/p&gt;

&lt;p&gt;Technical tokens need special treatment before ordinary normalization. We preserve distinct representations for C++, C#, and .NET, including versioned forms such as C++17 and .NET8. Reviewed aliases handle terms such as &lt;code&gt;k8s&lt;/code&gt; and Kubernetes. Short or ambiguous tokens are not silently spell-corrected.&lt;/p&gt;

&lt;p&gt;Typo recovery searches a tenant-local title-and-skill vocabulary, not whole descriptions. It inspects only the first eight query tokens. For inferred corrections, it requires separation between the best candidate and its runner-up; reviewed spelling corrections use an explicit map. Remaining query text stays intact. This bounds work and avoids turning a long natural-language sentence into a sequence of speculative rewrites.&lt;/p&gt;

&lt;p&gt;Embeddings address the remaining vocabulary gap. They let us retrieve candidates whose descriptions concern the requested work even when their titles and skills do not contain the same words. That broader retrieval still operates inside a strict eligibility boundary.&lt;/p&gt;

&lt;p&gt;Within PostgreSQL, every retrieval channel starts from the same eligible inventory. Tenant, publication status, employment type, and work mode are constraints applied before candidate limits. Default browsing shows open jobs; closed jobs require an explicit selection. Drafts and unrecognized statuses stay outside public search.&lt;/p&gt;

&lt;p&gt;We did not convert arbitrary natural-language phrases into hard filters. A request for "remote work in Bangalore" leaves unanswered whether the location describes the office or the candidate's permitted residence. The explicit UI filters remain authoritative.&lt;/p&gt;

&lt;p&gt;We started with exact vector search over the eligible tenant inventory. Approximate nearest-neighbor indexing was an option to measure later, not a prerequisite for using vectors. Chunk matches collapse to one candidate per job using the strongest qualifying chunk, so a long description cannot occupy several result positions.&lt;/p&gt;
&lt;h2&gt;
  
  
  A good rank is not proof of relevance
&lt;/h2&gt;

&lt;p&gt;Each channel contributes at most 100 candidates. We combine lexical and semantic ranks using reciprocal rank fusion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;score(job) = 1 / (60 + lexical_rank)
           + 1 / (60 + semantic_rank)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A missing channel contributes zero. Typo recovery feeds the lexical channel rather than casting another independent vote. Strong phrase matches in title/primary-skill text receive priority, and posting date breaks ties rather than allowing a newer weak match to outrank a stronger one merely through freshness.&lt;/p&gt;

&lt;p&gt;RRF avoids trying to compare a full-text rank directly with a cosine similarity. It does not determine whether either candidate list deserves to exist. The nearest vector to an unrelated query is still the nearest vector.&lt;/p&gt;

&lt;p&gt;We therefore apply a similarity threshold before fusion and preserve a reviewed set of explicit technical-token constraints on semantic candidates. Those safeguards improve control; they do not establish universal relevance.&lt;/p&gt;

&lt;p&gt;The dimensionality experiment made that limitation concrete. We expanded an initially small synthetic fixture into 337 jobs and 214 queries, using sanitized public-job snapshots from two employers alongside synthetic regressions. Source IDs were replaced and contact details removed. The descriptions produced 835 chunks. We evaluated both dimensions with real provider embeddings and actual SQL retrieval in disposable local databases.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;First dimensional comparison&lt;/th&gt;
&lt;th&gt;512 dimensions&lt;/th&gt;
&lt;th&gt;1,536 dimensions&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Held-out NDCG@10, excluding inventory and empty-result checks&lt;/td&gt;
&lt;td&gt;0.9727&lt;/td&gt;
&lt;td&gt;0.9944&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic returned precision@5&lt;/td&gt;
&lt;td&gt;52.73%&lt;/td&gt;
&lt;td&gt;55.00%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic recall@10&lt;/td&gt;
&lt;td&gt;92.24%&lt;/td&gt;
&lt;td&gt;92.24%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exact-query top-three success&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typo/alias top-three success&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The smaller representation lost more ranking quality than our planned 0.01 NDCG tolerance, making 1,536 dimensions the stronger candidate. Its measured embedding relation, including storage and index overhead, occupied about 7.18 MB versus 2.58 MB in the fixture. That was a modest absolute increase for this dataset, not evidence that larger dimensions are free. We did not establish their CPU or memory cost under representative production concurrency.&lt;/p&gt;

&lt;p&gt;More importantly, neither configuration met our original 90% semantic precision target. Three of four unrelated held-out queries also returned results. We initially withheld semantic activation rather than interpreting the stronger ranking score as a pass.&lt;/p&gt;

&lt;p&gt;The measurements left us with a product question the choice of engine could not answer: was excluding a relevant vacancy worse than showing extra weak matches?&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall-first was a product decision, not a passing precision score
&lt;/h2&gt;

&lt;p&gt;For this careers experience, the answer became explicit: relevant jobs should remain discoverable even if the result set also contains loosely related or unrelated jobs. We selected 1,536 dimensions and adopted recall-first calibration. Precision stayed visible, but it was no longer the release blocker it had originally been.&lt;/p&gt;

&lt;p&gt;On our labelled set, the original 90% target had failed because semantic retrieval admitted too many weak candidates, not because it consistently buried the desired job. RRF could rank a desired role first while leaving unrelated commercial or technical roles below it. Increasing dimensions improved ranking but barely moved precision. The calibration experiments did not establish a threshold that met the original target.&lt;/p&gt;

&lt;p&gt;Recall-first changed what we optimized and how we checked it. We measured whether labelled relevant jobs appeared across the admitted result pages, then examined their ranking. Recall@10 alone could not establish completeness for queries with more than ten relevant jobs.&lt;/p&gt;

&lt;p&gt;It did not remove the admission threshold or authorize unrestricted similarity results. Tenant and publication eligibility, selected hard filters, explicit technical-token constraints, and the 100-candidate channel limits stayed in place. A job from another company was still incorrect, however similar its description. A closed job still could not enter the default open-jobs view. Exact and typo-query regressions still mattered.&lt;/p&gt;

&lt;p&gt;Within that eligible inventory, we knowingly accepted more relevance noise. The distinction was between broadening discovery and weakening the conditions under which a job could appear at all.&lt;/p&gt;

&lt;p&gt;Under the revised criterion, the 1,536-dimensional evaluation included every labelled relevant job across all returned pages. That statement has limits. The labels had not received independent human review, and the previously inspected holdout was now a regression set rather than a fresh blind test.&lt;/p&gt;

&lt;p&gt;Candidate caps also remained real: 119 of 204 evaluated held-out queries reported channel truncation. No labelled job was missing in this snapshot, but arbitrary future queries can exceed the bounded candidate set. The UI discloses truncation rather than implying that a paginated list necessarily contains every possible match.&lt;/p&gt;

&lt;p&gt;We changed the acceptance criterion, not the meaning of the earlier measurements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Semantics without embedding every search
&lt;/h2&gt;

&lt;p&gt;Broader semantic discovery did not justify invoking a model on every interaction. We had designed separate paths for browsing, typing, and searches that warranted semantic retrieval.&lt;/p&gt;

&lt;p&gt;We rejected a pipeline that embedded every keystroke. We also rejected the simpler rule "use vectors only when lexical search returns nothing." Incidental words in descriptions can produce lexical hits without answering the query.&lt;/p&gt;

&lt;p&gt;Instead, the UI and server divide the work:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9x23foho4enpzxl2mmhv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9x23foho4enpzxl2mmhv.png" alt="Search Query Handling &amp;amp; Retrieval Flow" width="799" height="619"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both timers start from the latest input change; the pauses are not cumulative. The lexical-only shortcut is intentionally conservative: a short query must contain recognized concepts and have a strong phrase match in title/primary-skill text. The server makes the decision with deterministic rules, not an LLM query parser. A client request phase does not authorize arbitrary provider work.&lt;/p&gt;

&lt;p&gt;The query-vector cache holds at most 1,000 entries for 30 minutes per process. Its key includes semantic text and the embedding specification. Filter and page changes reuse the vector. Concurrent requests for the same vector share one provider call, and one cancelled caller does not cancel another caller's shared work.&lt;/p&gt;

&lt;p&gt;We did not cache result lists. Avoiding stale job visibility was more useful than adding another invalidation problem.&lt;/p&gt;

&lt;p&gt;This design does not make novel semantic queries free. In the evaluation workload, selective routing avoided only 8.33% of cold query embeddings while matching always-hybrid NDCG. That result did not support an extravagant efficiency claim. Cache reuse and real query mix determine the eventual savings.&lt;/p&gt;

&lt;p&gt;Progressive results also introduced races. An old preview could arrive after settled results, or after the user had moved to page two. Aborting a fetch was insufficient as a correctness mechanism.&lt;/p&gt;

&lt;p&gt;The UI associates responses with the current query/filter revision and paging state, rejecting obsolete successes and errors. Existing cards and keyboard focus remain while richer retrieval runs.&lt;/p&gt;

&lt;p&gt;Provider failure is an explicit &lt;code&gt;lexical_fallback&lt;/code&gt; mode. Database failure is a service error, not an empty result. Those distinctions are visible to the UI because they describe different things the user can reasonably do next.&lt;/p&gt;

&lt;h2&gt;
  
  
  The asynchronous index needs a commit protocol
&lt;/h2&gt;

&lt;p&gt;Fast query handling would not help if a job edit left the vector index describing the wrong vacancy. A new search endpoint also could not maintain its index by intercepting only its own application's writes. Other writers would bypass it.&lt;/p&gt;

&lt;p&gt;We put projection updates in a database trigger. Each source-job transaction refreshes the private search document and normalized filters. Changes to title, skills, or description change a content hash and make embedding work pending. Unrelated metadata does not require another embedding call.&lt;/p&gt;

&lt;p&gt;An authenticated worker in Next.js claims small batches using database leases and &lt;code&gt;FOR UPDATE SKIP LOCKED&lt;/code&gt;. The claim transaction finishes before any provider request. A host timer invokes the worker roughly a minute after the previous run ends; it claims at most four jobs per invocation.&lt;/p&gt;

&lt;p&gt;The critical operation is completion. In simplified form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;accept embeddings only if:
    the source document still exists and is eligible
    AND current_content_hash = claimed_content_hash
    AND current_lease_token = claimed_lease_token
    AND the lease has not expired
    AND the embedding specification is compatible
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without that check, a slow worker could overwrite an edited job with vectors for its previous description. Hashes identify the content; lease tokens identify the worker's current right to finish. Both are necessary when work can be retried.&lt;/p&gt;

&lt;p&gt;Semantic retrieval independently checks the content hash and specification. Stale vectors cannot participate while replacement work is pending. Lexical search remains available throughout. Status changes take effect without waiting for re-embedding, and deleting a job removes its derived artifacts.&lt;/p&gt;

&lt;p&gt;The specification includes model, dimension count, and preprocessing version. Switching dimensions is consequently an index migration, not just a provider option. Our approved upgrade operated on an empty careers vector index and refused populated embeddings or active leases. A populated incompatible generation would require a separately built replacement.&lt;/p&gt;

&lt;p&gt;Long descriptions are split into token-bounded chunks with a title/skills prefix, an 800-token maximum complete input, and up to 80 tokens of overlap. We preserve the full source text rather than quietly dropping its tail.&lt;/p&gt;

&lt;p&gt;Preparation had its own performance surprise. A JavaScript tokenizer initially looked sufficient, but its Unicode-heavy test exceeded a 15-second deadline. Replacing it with the server-side WASM tokenizer brought that text-test group to roughly 0.3 seconds of execution. Chunk preparation deserved measurement just as much as vector retrieval did.&lt;/p&gt;

&lt;p&gt;The evaluation corpus is separate from this lifecycle. Adding a new production job does not require editing a JSON fixture; the trigger and worker discover it automatically. Conversely, expanding the fixture improves our evidence, not the production model's knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The latency problem was not only the embedding call
&lt;/h2&gt;

&lt;p&gt;Once the retrieval path worked, latency still needed investigation. One avoidable cost was in our own eligibility lookup.&lt;/p&gt;

&lt;p&gt;To accommodate source representations, the first implementation extracted identifiers and status by converting complete rows to JSON. Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Before: serialize the source row to read a few fields.&lt;/span&gt;
&lt;span class="n"&gt;to_jsonb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&amp;gt;&lt;/span&gt;&lt;span class="s1"&gt;'job_id'&lt;/span&gt;
&lt;span class="n"&gt;to_jsonb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&amp;gt;&lt;/span&gt;&lt;span class="s1"&gt;'company_id'&lt;/span&gt;
&lt;span class="n"&gt;to_jsonb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&amp;gt;&lt;/span&gt;&lt;span class="s1"&gt;'job_status'&lt;/span&gt;

&lt;span class="c1"&gt;-- After: read only the required columns.&lt;/span&gt;
&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;
&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;company_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;
&lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;job_status&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read-only production comparisons returned identical job IDs while measuring approximately 102-106 ms for the original eligibility lookup and 8-11 ms for direct-column access. We also replaced a JSON-extracted company-username predicate with a direct-column predicate.&lt;/p&gt;

&lt;p&gt;These were component timings, not end-to-end search results. They nevertheless identified useful work: the application performs lexical retrieval before deciding whether to perform hybrid retrieval, so unnecessary eligibility cost affects both stages. We could remove it without changing the vector model, candidate limits, or relevance behavior.&lt;/p&gt;

&lt;p&gt;The fix shipped as a guarded SQL upgrade that changed only the expected expressions and preserved populated embeddings, indexing state, ownership, and grants. An unexpected function definition caused the upgrade to stop rather than rewriting unknown SQL.&lt;/p&gt;

&lt;p&gt;Provider variability was a separate problem. The initial 550 ms query deadline produced timeouts in both local and deployment-region experiments. We increased it to 1,000 ms to give novel queries more time to obtain an embedding; query calls still have no SDK retries, while indexing uses a separate ten-second deadline.&lt;/p&gt;

&lt;p&gt;That decision trades a longer wait on some novel queries for a better chance of obtaining semantic results. It does not make a 700 ms API target true. SDK deadlines are also not exact end-to-end wall-clock ceilings: database work and request handling sit outside them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The public endpoint changed our trust boundary
&lt;/h2&gt;

&lt;p&gt;An anonymous search that can invoke a paid provider needs more protection than an in-process concurrency counter. We added a generous per-IP token bucket before body parsing, database access, and embedding calls: 240 immediate requests, replenishing at 20 per second.&lt;/p&gt;

&lt;p&gt;The bucket itself was small. Establishing whose IP it counted was the harder part.&lt;/p&gt;

&lt;p&gt;Production traffic passed through Cloudflare and Caddy before reaching Next.js. Counting the immediate proxy address would group unrelated visitors. Accepting an arbitrary forwarding header would let callers choose their own buckets.&lt;/p&gt;

&lt;p&gt;Caddy now overwrites a dedicated search identity header. Only connections from the trusted CDN ranges may supply the CDN's visitor address; other connections use their actual remote address. The production application port is bound to loopback so public callers cannot bypass that sanitization by reaching Node directly.&lt;/p&gt;

&lt;p&gt;Review caught a subtler mismatch: the first proxy matcher covered valid tenant names, but the dynamic application route could receive malformed tenant paths. Because rate limiting ran before tenant validation, those paths could reach the limiter without having their identity header sanitized. The proxy matcher needed to cover the route family's malformed inputs too.&lt;/p&gt;

&lt;p&gt;We tested that behavior with a disposable loopback backend, including spoofed headers, simulated trusted proxies, encoded and overlong tenants, and unrelated routes. We did not flood the production database to prove a 429.&lt;/p&gt;

&lt;p&gt;The limiter remains deliberately modest. It is per process, resets on restart, and does not protect direct anonymous Supabase RPC calls. Those calls require their own database validation, bounded work, permissions, and effective caller statement deadlines. More replicas or a different traffic pattern would change the requirements.&lt;/p&gt;

&lt;p&gt;When a browser does receive a 429, it retains the previous cards, identifies them as previous results, respects a cooldown, and retries the latest query only on request. Rate limiting must not turn a visible result page into a misleading "no jobs found."&lt;/p&gt;

&lt;h2&gt;
  
  
  What our validation established
&lt;/h2&gt;

&lt;p&gt;We used different tests for different claims. Provider mocks exercised timeout and cancellation control flow without making normal CI depend on OpenAI. Disposable PostgreSQL tests executed real retrieval, permission, lease, stale-completion, and migration behavior. The dimensionality experiment used actual embeddings. None substituted for the others.&lt;/p&gt;

&lt;p&gt;The final protection changes passed 105 focused tests and TypeScript checking. The enabled-search Playwright suite covered 17 cases, including stale preview successes and failures, pagination races, normalized filters, explicit fallback, rate-limit recovery, and keyboard behavior at mobile width. The standard E2E command now runs both the existing flag-off suite and the enabled-search suite.&lt;/p&gt;

&lt;p&gt;Deployment had its own evidence gap. At one checkpoint, the migration had succeeded and the application contained the worker route, but the database had zero embeddings and the host had no indexing timer. Deploying code had not scheduled work. We configured the worker separately, completed the backfill, and verified current embeddings for all 724 eligible jobs before activation.&lt;/p&gt;

&lt;p&gt;Rollout proceeded through lexical search, proxy protection, rate limiting, and semantic activation. Production checks confirmed tenant identity and status filters, and a natural-language retail query displayed relevant roles in the browser with focus preserved.&lt;/p&gt;

&lt;p&gt;The final smoke test also showed the limitation we had designed around. Three initial embedding requests fell back to lexical search. Later hybrid calls succeeded: cached public requests measured 224-309 ms, one fresh finance query took 685 ms, and other successful uncached responses were around 1.2 seconds.&lt;/p&gt;

&lt;p&gt;The samples were too small to establish p95 latency or a steady-state fallback rate. We did not certify performance at twice observed peak traffic.&lt;/p&gt;

&lt;p&gt;Semantic search was enabled with that limitation explicit. Returning HTTP 200 after a fallback was evidence that the failure path worked, not that semantic latency had passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search correctness has more than one boundary
&lt;/h2&gt;

&lt;p&gt;The original vocabulary problem was real: literal containment could not reliably connect a candidate's wording to an employer's description. Hybrid retrieval widened that connection. Its usefulness depended on the boundaries around it.&lt;/p&gt;

&lt;p&gt;Eligibility belongs in SQL before ranking. Ranking quality does not establish candidate relevance. An embedding belongs to a particular content version, not merely a job ID. A late response does not own the current browser state. A forwarded address is not a caller identity until a trusted proxy establishes it.&lt;/p&gt;

&lt;p&gt;Those distinctions gave us concrete ways to test the system and explain its failures. They also clarified the architecture decision. PostgreSQL and our existing provider integration gave us enough primitives for this workload without a separate search service or inference runtime. In return, we owned the behavior connecting those primitives: eligibility, index freshness, query routing, and failure handling.&lt;/p&gt;

&lt;p&gt;A dedicated engine or managed platform could change where that work lives. It would not decide which weak matches the product should tolerate, make incompatible vectors comparable, or determine whether an old response still belongs on the screen.&lt;/p&gt;

&lt;p&gt;The most consequential decision was accepting that relevance itself needed a product definition. Our high ranking score had not answered whether extra results were acceptable. Once that choice was explicit, we could optimize for finding the labelled relevant jobs without pretending that a recall-first system had passed the precision target it replaced.&lt;/p&gt;

</description>
      <category>nlp</category>
      <category>vectordatabase</category>
      <category>ai</category>
      <category>postgres</category>
    </item>
    <item>
      <title>🚨 I Fell for the Krutrim Hype (Twice) - Here's Why You Shouldn't</title>
      <dc:creator>Puneet Chandna</dc:creator>
      <pubDate>Sat, 23 Aug 2025 16:04:13 +0000</pubDate>
      <link>https://dev.to/puneet-chandna/i-fell-for-the-krutrim-hype-twice-heres-why-you-shouldnt-g5p</link>
      <guid>https://dev.to/puneet-chandna/i-fell-for-the-krutrim-hype-twice-heres-why-you-shouldnt-g5p</guid>
      <description>&lt;p&gt;&lt;em&gt;A software engineering student's journey from excitement to disappointment to exposing the truth about India's "revolutionary" AI&lt;/em&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  TL;DR:
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;I tested Krutrim (2024) and Kruti (2025).&lt;/strong&gt; Both were slow and error-prone; Kruti repeatedly returned identical, scripted identity replies and showed behavioral signs of being built on LLaMA 3 with heavy system prompts. That’s fine if disclosed — it’s not fine when marketed as “revolutionary.”&lt;/p&gt;




&lt;h2&gt;
  
  
  The Day I Believed in the Dream 🌟
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;February 2024.&lt;/strong&gt; I'm scrolling through my Twitter feed when I see it - Bhavish Aggarwal, the founder of Ola, announcing something that made my heart race as an Indian CS student:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"India's own ChatGPT is here! Krutrim AI - built for Bharat, by Bharat!"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Finally! As a final-year Computer Science student, I'd been watching OpenAI, Google, and Meta dominate the AI space while India seemed to be playing catch-up. Here was our chance to show the world that Indian developers could build world-class AI too.&lt;/p&gt;

&lt;p&gt;I was &lt;strong&gt;pumped&lt;/strong&gt;. I was &lt;strong&gt;proud&lt;/strong&gt;. I was about to be &lt;strong&gt;very disappointed&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  First Encounter: The Red Flags I Ignored 🚩
&lt;/h2&gt;

&lt;p&gt;Within hours of the announcement, I signed up for Krutrim's beta. My expectations were sky-high - this was supposed to understand Hindi, Tamil, Bengali, and 20+ other Indian languages better than any Western AI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My first question:&lt;/strong&gt; "भारत की राजधानी क्या है?" (What is the capital of India?)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wait time:&lt;/strong&gt; 30+ seconds&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Response:&lt;/strong&gt; Technically correct but... wait, why did it take so long for such a basic query?&lt;/p&gt;

&lt;p&gt;I brushed it off. "Beta version," I told myself. "They'll optimize it."&lt;/p&gt;

&lt;h3&gt;
  
  
  The Identity Crisis That Should've Been a Red Alert 🚨
&lt;/h3&gt;

&lt;p&gt;But then I asked the most basic question any AI user asks:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My question:&lt;/strong&gt; "Who are you?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Krutrim's response:&lt;/strong&gt; "I am a Large Language Model created by OpenAI."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My reaction:&lt;/strong&gt; Wait... WHAT?! 🤯&lt;/p&gt;

&lt;p&gt;I stared at my screen in disbelief. India's revolutionary AI just told me it was made by OpenAI? I refreshed the page, asked again, same response.&lt;/p&gt;

&lt;p&gt;My patriotic heart sank, but I rationalized it: "Must be a bug. They'll fix it."&lt;/p&gt;

&lt;h3&gt;
  
  
  The Questions That Broke My Heart 💔
&lt;/h3&gt;

&lt;p&gt;Over the next days I ran quick tests, trivia, math, simple code. Two patterns emerged: slow responses and flaky accuracy. Examples that stuck with me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A wrong answer about the 1983 Cricket World Cup.&lt;/li&gt;
&lt;li&gt;A 15–40 second delay for simple arithmetic or a small Python function.&lt;/li&gt;
&lt;li&gt;Code snippets with logical errors that a CS student can spot instantly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I felt embarrassed on behalf of the product. If this was being touted as India’s answer to ChatGPT, it wasn’t a great look.&lt;/p&gt;

&lt;p&gt;But I'm an optimist. "They'll improve," I kept telling myself. "It's just v1."&lt;/p&gt;

&lt;h2&gt;
  
  
  Plot Twist: Kruti Launch and My Detective Work 🕵️
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Today, August 23, 2025&lt;/strong&gt;. I was scrolling Instagram in the campus lab when I got a notification from Ola, another announcement that made my heart skip a beat:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Introducing Kruti - India's first agentic AI assistant!"&lt;/em&gt;&lt;br&gt;
&lt;em&gt;"Powered by advanced Krutrim 2 models!"&lt;/em&gt;&lt;br&gt;
&lt;em&gt;"Next-generation AI for India!"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Despite being burned before, that familiar excitement crept back in. "Maybe they actually fixed everything this time," I thought. "Maybe Krutrim 2 is the real deal."&lt;/p&gt;

&lt;p&gt;I tried it immediately.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Investigation Begins 🔍
&lt;/h3&gt;

&lt;p&gt;My developer instincts kicked in harder this time. After 18+ months of studying deep learning and prompt engineering and getting fooled once, I approached this with the curiosity of a CS student and the skepticism of someone who'd been burned before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I saw:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every question took &lt;strong&gt;15-20 seconds&lt;/strong&gt; to respond, Even for basic queries like “What’s 2+2?”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Identity Question That Revealed Everything:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My question:&lt;/strong&gt; "Who are you?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happened:&lt;/strong&gt; &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Web search initiated&lt;/strong&gt; (I could see the loading indicator)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;"Thinking..." for 15 seconds&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Final response:&lt;/strong&gt; "I am Kruti, an AI assistant developed by the Krutrim AI team. I'm powered by Krutrim 2 and other advanced open source AI models."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Wait. It needed to do a WEB SEARCH to know who it is?! And what's with "other advanced open source AI models"?&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The Detective Work: System Prompts &amp;amp; Model Tells 🕵️‍♂️
&lt;/h3&gt;

&lt;p&gt;I pushed further for some follow-up questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Follow-ups yielded the same scripted reply every time:&lt;/strong&gt;&lt;br&gt;
I asked “What models are you using?”, “Who made you?”, and “What is Krutrim 2?” — and each time the assistant returned the identical sentence:&lt;br&gt;
&lt;em&gt;“I am Kruti, an AI assistant developed by the Krutrim AI team. I'm powered by Krutrim 2 and other advanced open source AI models.”&lt;/em&gt;&lt;br&gt;
Identical, word-for-word replies like that scream of system-prompt hardcoding, not genuine model reasoning.&lt;/p&gt;
&lt;h3&gt;
  
  
  The Smoking Gun 🔫
&lt;/h3&gt;

&lt;p&gt;After hours of prompt-injection tests and timing measurements (yes, I spent my Saturday evening on this) the pattern was clear: latency and token-generation speed matched LLaMA 3 behavior, failures were the same kinds of hallucinations and reasoning gaps, and identity queries returned a canned sentence every time. In short — heavy system prompts + light fine-tuning on an existing open-source base, wrapped in marketing.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Technical Reality Check 💻
&lt;/h2&gt;

&lt;p&gt;Let me break this down as a CS student who's actually studied how LLMs work:&lt;/p&gt;
&lt;h3&gt;
  
  
  &lt;strong&gt;What Krutrim Claims:&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Proprietary "Krutrim 2" model&lt;/li&gt;
&lt;li&gt;Built from scratch for Indian languages&lt;/li&gt;
&lt;li&gt;Revolutionary architecture optimized for Indian context&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  &lt;strong&gt;What I Actually Found:&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base Model:&lt;/strong&gt; Meta's LLaMA 3 (obvious from behavior patterns)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Innovation":&lt;/strong&gt; Heavy system prompts and fine-tuning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identity Responses:&lt;/strong&gt; Pre-scripted, identical answers to avoid revealing base model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance:&lt;/strong&gt; Still terrible - 15-20 seconds per basic question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identity Questions:&lt;/strong&gt; Requires web search + thinking time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accuracy:&lt;/strong&gt; Still making basic factual errors&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Why This Stings as a CS Student 📚
&lt;/h2&gt;

&lt;p&gt;Here's the thing - I’m pro–open-source. I build on OSS all the time. The issue here isn’t that they used open-source, it’s that they &lt;strong&gt;&lt;em&gt;appear to be hiding it&lt;/em&gt;&lt;/strong&gt; and selling it as proprietary innovation.&lt;/p&gt;

&lt;p&gt;This matters because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It erodes trust in Indian AI startups.&lt;/li&gt;
&lt;li&gt;It wastes resources if investors believe there’s a unique, fast model under the hood.&lt;/li&gt;
&lt;li&gt;It distracts from genuine engineering work that could actually improve performance for Indian languages.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  &lt;strong&gt;What Should Have Happened:&lt;/strong&gt;
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;"We're building India's best AI assistant and AI agent using LLaMA 3 as our foundation, with specialized fine-tuning for Indian languages, cultural context, and use cases."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;That's it!&lt;/strong&gt; Honest, clear, and actually impressive from a product perspective.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Reality: It's Still Just Slow LLaMA 3 (With Extra Steps) 🦙
&lt;/h2&gt;

&lt;p&gt;After all this investigation, here's what I'm convinced Kruti actually is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Krutrim's "Revolutionary AI" = 
  LLaMA 3 
  + Heavy system prompts to hide identity
  + Pre-scripted responses for deflection
  + Some fine-tuning for Indian content
  + Web search integration (badly implemented)
  + Marketing budget
  + A prayer that CS students won't notice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Call to Action: What We Can Do 🚀
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;As Developers:&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test identity questions&lt;/strong&gt; - Ask "Who are you?" and watch for scripted responses&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time the responses&lt;/strong&gt; - 15+ seconds for basic questions is a red flag
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test follow-ups&lt;/strong&gt; - Identical word-for-word responses indicate system prompts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share your findings&lt;/strong&gt; - Help the community make informed decisions&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;As Indian Tech Community:&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Demand performance benchmarks&lt;/strong&gt; - Speed and accuracy matter more than marketing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Call out scripted responses&lt;/strong&gt; - Real AI doesn't need web searches to know its identity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop falling for "Version 2.0" hype&lt;/strong&gt; - Judge by performance, not version numbers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support companies that are honest&lt;/strong&gt; about their technology stack&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Uncomfortable Truth 💭
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;I wanted Krutrim to succeed.&lt;/strong&gt; I really did. &lt;strong&gt;Twice.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;As an Indian CS student about to graduate, I dream of working for Indian companies that compete globally on technical merit, not elaborate deception schemes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But here's what really hurts:&lt;/strong&gt; The performance got WORSE. Krutrim was slow, but at least it tried to respond naturally. Kruti takes longer AND gives robotic, scripted responses designed to hide its origins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This isn't innovation - it's regression with better marketing.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Discussion — I Want Your Experience 💬
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Did you test Kruti? What did you find for identity questions?
&lt;/li&gt;
&lt;li&gt;What’s the slowest you’ve seen a basic answer take?
&lt;/li&gt;
&lt;li&gt;CS students: how do you validate model provenance in practice?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Drop your test results and clips in the comments — let’s build a shared dataset of evidence.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;PS: I’m not against Indian AI companies or building on open-source models. I’m against slow UX, scripted identity replies, and marketing that pretends something basic is revolutionary.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Follow me for more honest tech reviews, interesting  tech Blogs, and the occasional rant about overhyped startups that waste our time&lt;/strong&gt; 😤&lt;/p&gt;

&lt;p&gt;— Puneet, final-year CS student.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>indianstartups</category>
      <category>news</category>
      <category>llm</category>
    </item>
    <item>
      <title>10 Must-Read Books for Software Engineers in 2025 📘💻</title>
      <dc:creator>Puneet Chandna</dc:creator>
      <pubDate>Mon, 28 Jul 2025 16:28:43 +0000</pubDate>
      <link>https://dev.to/puneet-chandna/10-must-read-books-for-software-engineers-in-2025-4717</link>
      <guid>https://dev.to/puneet-chandna/10-must-read-books-for-software-engineers-in-2025-4717</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6eq27asj24ayz4lkv81y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6eq27asj24ayz4lkv81y.png" alt="10 Must-Read Books for Software Engineers" width="800" height="1041"&gt;&lt;/a&gt;As a software engineer, one thing has been constant: &lt;strong&gt;learning never stops&lt;/strong&gt;.&lt;br&gt;&lt;br&gt;
So here’s a list of 10 books every engineer should read at least once in their career.&lt;br&gt;&lt;br&gt;
They’ve shaped how I approach problem-solving, systems, and even team collaboration.&lt;/p&gt;




&lt;h2&gt;
  
  
  ✅ Books I’ve Already Read:
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. &lt;strong&gt;Designing Data-Intensive Applications (DDAI)&lt;/strong&gt;
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;My absolute favorite. This book changed how I design systems.&lt;br&gt;&lt;br&gt;
I used its concepts to optimize a DB that reduced query times by 40% on a real project.&lt;br&gt;&lt;br&gt;
I’ve highlighted the key sections and revisit them regularly — &lt;em&gt;it’s more of a playbook than a textbook&lt;/em&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  2. &lt;strong&gt;System Design Interview – Vol 1&lt;/strong&gt;
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;Great for both interview prep and practical, real-world system architecture.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  3. &lt;strong&gt;Introduction to Algorithms (CLRS)&lt;/strong&gt;
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;The classic. Dense, but worth it if you want to master the core fundamentals.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🕐 What’s Next?
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;remaining books are still on my list&lt;/strong&gt;, and I plan to read them one by one — as soon as I get time.&lt;br&gt;&lt;br&gt;
If you're building a roadmap for serious engineering growth, this list is a great place to start.&lt;/p&gt;




&lt;h2&gt;
  
  
  ⭐ My Top Recommendation?
&lt;/h2&gt;

&lt;p&gt;Without a doubt: &lt;strong&gt;DDAI&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It’s not just for learning — it’s for &lt;em&gt;relearning&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;I revisit it every few months.&lt;/li&gt;
&lt;li&gt;It stays relevant across systems, databases, and scalability topics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No book can beat that.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🔁 Your Turn:
&lt;/h2&gt;

&lt;p&gt;What’s one book that changed how you think as an engineer?&lt;/p&gt;

&lt;p&gt;Drop a comment, link your review, or just say hi 👋&lt;br&gt;&lt;br&gt;
Let’s build the &lt;strong&gt;ultimate software engineer reading list&lt;/strong&gt; together!&lt;/p&gt;




&lt;p&gt;#softwareengineering #backenddevelopment #books #learning #systemdesign #career #developer &lt;/p&gt;

</description>
      <category>programming</category>
      <category>softwareengineering</category>
      <category>books</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>I Accidentally Discovered a Hidden Gem for Testing Premium AI Models (Completely Free!)</title>
      <dc:creator>Puneet Chandna</dc:creator>
      <pubDate>Sat, 26 Jul 2025 20:19:57 +0000</pubDate>
      <link>https://dev.to/puneet-chandna/i-accidentally-discovered-a-hidden-gem-for-testing-premium-ai-models-completely-free-2k3c</link>
      <guid>https://dev.to/puneet-chandna/i-accidentally-discovered-a-hidden-gem-for-testing-premium-ai-models-completely-free-2k3c</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally posted as a LinkedIn discovery that I just had to share with the dev community.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Discovery That Made Me Break My "No Social Media Posts" Rule&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;I'll be honest,I don't usually write LinkedIn posts or blogs. But sometimes you stumble across something so useful that you feel obligated to share it with fellow developers and AI enthusiasts.&lt;br&gt;
That happened to me recently when I discovered &lt;strong&gt;LMArena&lt;/strong&gt; (lmarena.ai), and it completely changed how I approach AI model testing and comparison.&lt;br&gt;
&lt;strong&gt;What Exactly Is LMArena?&lt;/strong&gt;&lt;br&gt;
LMArena is essentially a public arena where AI models battle it out, anonymously. Here's how it works:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You submit a prompt&lt;/li&gt;
&lt;li&gt;Two AI models respond (you don't know which models they are)&lt;/li&gt;
&lt;li&gt;You vote for the better response&lt;/li&gt;
&lt;li&gt;The results feed into a public leaderboard based on real user preferences&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But here's the kicker,while you're participating in this research, you get &lt;strong&gt;free access to premium AI models&lt;/strong&gt; that normally cost serious money.&lt;br&gt;
&lt;strong&gt;The Model Lineup (And Why It's Impressive)&lt;/strong&gt;&lt;br&gt;
The platform gives you access to models that would typically require expensive API credits or premium subscriptions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Opus 4 - Anthropic's flagship model&lt;/li&gt;
&lt;li&gt;Gemini 2.5 Pro - Google's latest and greatest&lt;/li&gt;
&lt;li&gt;DeepSeek R1 - The reasoning powerhouse&lt;/li&gt;
&lt;li&gt;Grok 4 - X's premium AI model&lt;/li&gt;
&lt;li&gt;And many more...&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No login required. No credit card. No sketchy popups or malware concerns.&lt;br&gt;
&lt;strong&gt;Three Ways to Use LMArena&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Arena Mode (The Classic)&lt;/strong&gt;&lt;br&gt;
Submit your prompt and vote between two anonymous responses. Perfect for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Testing prompt engineering techniques&lt;/li&gt;
&lt;li&gt;Getting multiple perspectives on coding problems&lt;/li&gt;
&lt;li&gt;Comparative analysis without bias&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;2. Direct Chat Mode&lt;/strong&gt;&lt;br&gt;
Choose a specific model and have a direct conversation. Great for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deep-diving into technical problems&lt;/li&gt;
&lt;li&gt;Iterating on code solutions&lt;/li&gt;
&lt;li&gt;Model-specific testing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;3. Side-by-Side Mode&lt;/strong&gt;&lt;br&gt;
battle between 2 models of your choice, great for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Understanding model strengths and weaknesses&lt;/li&gt;
&lt;li&gt;Choosing the  model for your use case&lt;/li&gt;
&lt;li&gt;Research and analysis&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why This Matters for Developers&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost Savings&lt;/strong&gt;&lt;br&gt;
Instead of paying for multiple API subscriptions to test different models, you can evaluate them all in one place for free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unbiased Comparison&lt;/strong&gt;&lt;br&gt;
The anonymous voting system removes brand bias. You're judging purely on output quality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-World Performance Data&lt;/strong&gt;&lt;br&gt;
The leaderboard reflects actual user preferences, not just benchmark scores.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Engineering Laboratory&lt;/strong&gt;&lt;br&gt;
Perfect environment for testing how different models respond to various prompting techniques.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Word of Caution (And Why I Trust This One)&lt;/strong&gt;&lt;br&gt;
I've seen countless "free GPT" clones online, and most are either:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Spam-filled nightmares&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Potential security risks&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Barely functional wrappers&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;LMArena is different because:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It's research-backed and transparent about its purpose&lt;br&gt;
No data collection beyond the voting mechanism&lt;br&gt;
Open about its methodology and model selection&lt;br&gt;
Clean, professional interface without dark patterns&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-World Use Cases&lt;/strong&gt;&lt;br&gt;
Here are some ways I've been using LMArena in my development workflow:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code Review and Debugging&lt;/strong&gt;&lt;br&gt;
Prompt: "Review this Python function and suggest improvements for performance and readability:"&lt;br&gt;
Getting multiple model perspectives helps identify issues you might miss.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Architecture Decisions&lt;/strong&gt;&lt;br&gt;
Prompt: "Compare microservices vs monolithic architecture for a team of 5 developers building a SaaS platform"&lt;br&gt;
Different models often emphasize different trade-offs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Documentation Writing&lt;/strong&gt;&lt;br&gt;
Prompt: "Explain this API endpoint in simple terms for junior developers"&lt;br&gt;
Comparing explanations helps you find the clearest communication style.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Bigger Picture&lt;/strong&gt;&lt;br&gt;
LMArena represents something fascinating in the AI space,a democratized testing ground where models are evaluated based on real user needs rather than academic benchmarks.&lt;br&gt;
As developers, we often need to choose between different AI tools for our projects. Having a neutral space to test and compare these models without financial commitment is invaluable.&lt;br&gt;
Getting Started&lt;/p&gt;

&lt;p&gt;Visit lmarena.ai&lt;br&gt;
Choose your mode (Arena, Chat, or P2L)&lt;br&gt;
Start testing with your own prompts&lt;br&gt;
Vote on responses to contribute to the community&lt;/p&gt;

&lt;p&gt;No signup, no credit card, no commitment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Final Thoughts&lt;/strong&gt;&lt;br&gt;
I'm sharing this not because I'm sponsored (I'm definitely not), but because good tools deserve to be known by the community that can benefit from them.&lt;br&gt;
In a world where AI access is increasingly paywalled, LMArena feels like a breath of fresh air—a place where you can experiment, learn, and contribute to AI research simultaneously.&lt;/p&gt;

&lt;h2&gt;
  
  
  Just remember:
&lt;/h2&gt;

&lt;p&gt;if this tool proves as useful to you as it has to me, consider sharing it responsibly. Good free resources tend to get overwhelmed quickly, and we want this one to stick around.&lt;/p&gt;

&lt;p&gt;Have you tried LMArena? What models performed best for your use cases? Drop your experiences in the comments below!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devto</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
