<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Himanshu Negi</title>
    <description>The latest articles on DEV Community by Himanshu Negi (@himanshu_develops).</description>
    <link>https://dev.to/himanshu_develops</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3716597%2F7d89426b-0be9-421f-b11d-3b5e957a8de4.png</url>
      <title>DEV Community: Himanshu Negi</title>
      <link>https://dev.to/himanshu_develops</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/himanshu_develops"/>
    <language>en</language>
    <item>
      <title>Controllable Creative Generation with Constraint Ledgers: Findings from Codex and Fable</title>
      <dc:creator>Himanshu Negi</dc:creator>
      <pubDate>Wed, 29 Jul 2026 13:53:15 +0000</pubDate>
      <link>https://dev.to/himanshu_develops/controllable-creative-generation-with-constraint-ledgers-findings-from-codex-and-fable-5bn2</link>
      <guid>https://dev.to/himanshu_develops/controllable-creative-generation-with-constraint-ledgers-findings-from-codex-and-fable-5bn2</guid>
      <description>&lt;h1&gt;
  
  
  Constraint Ledgers for Controllable Creative Generation: Evidence from the Codex and Fable Pilot Programs
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Abstract
&lt;/h2&gt;

&lt;p&gt;Generative systems are increasingly asked to produce outputs that are both novel and defensible: a product concept should depart from familiar form without breaking its function; a story should surprise without losing its internal logic. This paper synthesizes two related internal pilot programs, here termed the &lt;strong&gt;Codex&lt;/strong&gt; and &lt;strong&gt;Fable&lt;/strong&gt; tracks, that test whether an explicit constraint representation improves such work. The central intervention distinguishes attributes that must be preserved from defaults that may be revised, then extends that distinction into a multi-axis ledger containing semantic role, contextual applicability, functional force, epistemic status, and a permitted action. Across the Fable classification pilot, a type-level functional-counterfactual prompt achieved 93.8% accuracy on clean physical items and 97.4% paraphrase agreement. However, its generation experiment found that a classification scaffold did not outperform an equally effortful, context-rich elaboration control; the latter was better on cliché avoidance (10–2, sign test &lt;em&gt;p&lt;/em&gt; = .039). The Codex binary and multi-axis pilots independently reached the same causal conclusion: neither an audit nor a layered ledger improved the primary valid-departure outcome beyond a matched rich-sham control. The results support constraint ledgers as inspectable control and audit interfaces, not as validated mechanisms for improving creative quality. We propose a preregistered, human-rated benchmark that separates classification, routing, and generation, and requires any layered approach to exceed both a flat gate and a content-matched nongated control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keywords:&lt;/strong&gt; generative AI; controllable creativity; constraint reasoning; human–AI co-creation; evaluation; defaults&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Introduction
&lt;/h2&gt;

&lt;p&gt;Creative generation is often framed as a search for novelty, yet useful novelty is constrained. A vehicle designed for a new environment must still satisfy safety and functional requirements; a narrative variation must preserve the causal and contextual commitments that make it intelligible. Large language models can generate fluent rationales for such departures, but fluency is not evidence that a departure is feasible, context-sensitive, or based on a sound distinction between necessity and convention.&lt;/p&gt;

&lt;p&gt;The Codex and Fable pilot programs address this problem with an &lt;strong&gt;external constraint ledger&lt;/strong&gt;. The basic proposal is simple: identify the attributes of a requested artifact that are functionally or definitionally binding and separate them from familiar but defeasible defaults. Generation is then instructed to preserve the former and intentionally vary the latter only with a context-specific justification. The more elaborate version records five dimensions for each claim: semantic role, applicability, force, epistemic status, and an action such as preserve, compensate, vary, or clarify.&lt;/p&gt;

&lt;p&gt;This is a plausible intervention, but a plausible representation is not a demonstrated causal mechanism. Extra intermediate text, richer contextual detail, or stronger prompt wording may explain any apparent improvement. Both programs therefore included controls designed to distinguish the value of labels from the value of additional task-relevant deliberation. This paper provides a transparent synthesis of their results, their shared limits, and the next experiment needed to make a credible empirical claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Conceptual Framework
&lt;/h2&gt;

&lt;p&gt;For a concept &lt;em&gt;K&lt;/em&gt;, an operative function &lt;em&gt;F&lt;/em&gt;, and a context &lt;em&gt;c&lt;/em&gt;, an attribute is treated as a &lt;strong&gt;functional invariant&lt;/strong&gt; when an otherwise well-functioning instance of &lt;em&gt;K&lt;/em&gt; cannot omit it while performing &lt;em&gt;F&lt;/em&gt; in &lt;em&gt;c&lt;/em&gt;. It is a &lt;strong&gt;defeasible default&lt;/strong&gt; when it is typical but a functioning instance can omit or replace it. The formulation is deliberately context-indexed: watertight seals may be optional for an ordinary car but binding for a car required to operate underwater. The resulting distinction resembles the treatment of defaults as defeasible in default logic [1] and the separation of typicality from causal centrality in concept research [2].&lt;/p&gt;

&lt;p&gt;The Fable track initially used the labels &lt;em&gt;structural&lt;/em&gt; and &lt;em&gt;conventional&lt;/em&gt;. The Codex work adopts &lt;em&gt;functional invariant&lt;/em&gt; and &lt;em&gt;defeasible default&lt;/em&gt;, avoiding a terminological collision with function–behaviour–structure design theory, in which “structure” ordinarily denotes components and relations rather than binding requirements. Both labels refer to the same practical decision: whether a feature may be changed without compensation.&lt;/p&gt;

&lt;p&gt;The multi-axis extension prevents several false equivalences that a binary split can create. A claim can be a genre convention yet explicit in the brief; a capability can be invariant while its implementation remains replaceable; or a claim can be relevant but contested. The proposed ledger therefore assigns each claim a semantic role (for example, functional, safety, prototype, or expressive), applicability (explicit, contextual, or background), force (invariant, compensation-required, defeasible, or unknown), epistemic status, and a revision action. It should be understood as an observable control representation—not a transcript of a model’s hidden reasoning. This caution is essential because stated rationales can be unfaithful to the computation that produced an answer [3].&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Research Questions
&lt;/h2&gt;

&lt;p&gt;The synthesis evaluates four questions.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Can a model classify invariants and defeasible defaults reliably, including under paraphrase?&lt;/li&gt;
&lt;li&gt;Does an explicit classification scaffold improve creative generation beyond direct prompting?&lt;/li&gt;
&lt;li&gt;Does it improve generation beyond equally rich, non-classificatory contextual planning?&lt;/li&gt;
&lt;li&gt;Does adding multiple semantic axes outperform both a flat gate and a rich nongated representation?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The crucial questions are the third and fourth. A positive result against direct prompting alone could be explained by additional tokens, attention, or setting detail. An intervention earns a mechanism claim only when it exceeds a final-prompt control with comparable information and deliberative richness.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Evidence Base and Method
&lt;/h2&gt;

&lt;p&gt;This paper is a comparative synthesis of four documented internal studies: the Fable classification and generation pilot; the Codex binary audit pilot; the Codex multi-axis ledger pilot; and a held-out, LLM-rated follow-up. These studies share a motivation but have different prompts, briefs, scales, and evaluators. Their scores are therefore interpreted side by side rather than pooled into an effect size.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.1 Fable: classification and generation
&lt;/h3&gt;

&lt;p&gt;The Fable program used a mid-capability model for two experiments. In Experiment 1, 38 items were each evaluated in an original and paraphrased form. The model supplied a binary label, confidence, a core function, and a reason. The original functional-counterfactual wording invited a systematic error: defective or toy instances were used as counterexamples, lowering physical-item accuracy to 77.1% and structural recall to 56.7%. A revised &lt;strong&gt;type-level&lt;/strong&gt; test asked whether a designer could make a new, fully functional instance without the feature, explicitly excluding defective, miniature, decorative, and look-alike cases. This changed the operationalization rather than the task and raised physical accuracy to 93.8%, structural recall to 83.3%, and paraphrase agreement to 97.4%.&lt;/p&gt;

&lt;p&gt;Experiment 2 compared eight creative briefs under three conditions: direct generation (A), classification scaffold plus generation (B), and matched-length contextual elaboration without classification (C). Outputs were pairwise judged blind in both presentation orders. B strongly exceeded A on justified deviation (13–1, &lt;em&gt;p&lt;/em&gt; = .0018) and cliché avoidance (12–1, &lt;em&gt;p&lt;/em&gt; = .0034). Yet B failed the decisive comparison: against C, it lost on cliché avoidance (2–10, &lt;em&gt;p&lt;/em&gt; = .039) and trailed overall (5–11). The model also followed its own B-stage classification plan in all eight cases, with no structural violations. Thus, the negative mechanism result cannot be dismissed as the model ignoring the scaffold. The more parsimonious explanation is that the C condition supplied more concrete and useful world-building material.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 Codex: binary audit and multi-axis ledger
&lt;/h3&gt;

&lt;p&gt;The Codex binary pilot used eight briefs and three primary conditions: direct generation, an oracle audit that marked requirements and conventions, and a generic task-relevant plan. The classification sanity check was strong on clear causal and logical anchors (16/16 correct in each of three passes) but weaker on ambiguous genre and social items (8/12 correct against provisional labels) and overconfident on contested cases. On generation, the audit did not show an incremental advantage over generic planning. The audit-minus-plan differences were 0.000 for preservation, −0.083 for departure, +0.167 for justification, and +0.042 for coherence; every descriptive bootstrap interval crossed zero. A label-flip diagnostic altered declared target choice on all eight briefs, demonstrating prompt-policy sensitivity but not sound classification or improved outcomes.&lt;/p&gt;

&lt;p&gt;The multi-axis pilot supplied oracle records to isolate routing and generation. Six briefs were produced under a flat gate, a layered ledger, and a rich sham retaining comparable claims, evidence, dependencies, and contextual detail without a mutable-row policy. All 18 outputs satisfied the output contract. The layered ledger and rich sham had the same valid-departure rate (0.611), the same departure score (4.222), and the same coherence score (4.722). Against the flat gate, the layered condition again had no valid-departure advantage and was descriptively lower on coherence. The study therefore establishes neither that additional layers help nor that they harm; it establishes that no benefit was observed in a small, ceiling-prone agent-rated pilot.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.3 Held-out diagnostic follow-up
&lt;/h3&gt;

&lt;p&gt;A later diagnostic rated twelve anonymized outputs from four new briefs under flat-gate, layered-gate, and rich-sham conditions. It is explicitly not a human evaluation: an LLM rater again evaluated LLM outputs. After correcting a packet-construction reversal that affected the apparent coherence of all layered cards, the rich sham retained the best justification and cliché profile and twice the valid-contextual-departure rate of either gate (0.50 versus 0.25). Because of its scale and rater confound, this is directional replication rather than confirmation. It reinforces the same design warning: a rich alternative representation can be at least as useful as a policy-bearing ledger.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Results Across Programs
&lt;/h2&gt;

&lt;p&gt;The common result is not that constraint distinctions are useless. Instead, three more limited findings emerge.&lt;/p&gt;

&lt;p&gt;First, explicit necessity/default classification is tractable in clear cases when the counterfactual is carefully operationalized. The Fable prompt revision shows that the model’s errors were directional and diagnosable; it had been judging a defective token rather than a functional type. The Codex anchor results likewise support competence on unambiguous causal and logical cases. However, contested items remain a problem. Social, genre, and context-sensitive claims require calibrated uncertainty and independent annotation rather than forced confidence.&lt;/p&gt;

&lt;p&gt;Second, a ledger can control behavior. Fable’s scaffolded outputs adhered to their stated plans, and Codex’s label-flip condition changed the selected target. These observations support a narrow controllability claim: labels can steer a generator when the prompt directs it to obey them. They do not establish that the labels are correct, that the rationale is faithful, or that the resulting artifact is better.&lt;/p&gt;

&lt;p&gt;Third, neither program validates the proposed quality mechanism. In Fable, a context-rich control outperformed the classification scaffold on the one statistically distinguishable B-versus-C result. In Codex, the binary audit did not exceed generic planning, and the layered ledger matched the rich sham on the primary outcome. The right conclusion is not equivalence—these pilots are small and use related LLM judges—but an absence of observed incremental benefit under the executed designs.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Discussion
&lt;/h2&gt;

&lt;p&gt;The evidence suggests a useful reframing. Constraint ledgers may be valuable for &lt;strong&gt;auditability, controllability, and collaboration&lt;/strong&gt; even if they do not increase creative quality. A visible record makes assumptions challengeable: a user can contest a claim, replace a default, require compensation, or request clarification. In software and design contexts, it can also expose which propositions are candidates for executable tests. These are legitimate benefits, but they differ from the causal claim that explicit classification creates more original or better concepts.&lt;/p&gt;

&lt;p&gt;The rich-sham result is particularly informative. Contextual grounding may be a stronger creative input than abstract labels. A preliminary research pass that surfaces climate, embodied action, resource constraints, cultural practices, or user goals can give the generator material from which to derive non-clichéd variation. The Fable control’s success is consistent with this account, as is the repeated Codex pattern. The appropriate hybrid hypothesis is therefore not “more labels,” but &lt;strong&gt;grounding plus targeted constraints&lt;/strong&gt;: the ledger should add value only when it prevents an otherwise attractive but invalid departure, clarifies a genuinely contested assumption, or routes an implementation change to a compensating mechanism.&lt;/p&gt;

&lt;p&gt;This boundary matters for novelty. Multi-constraint handling and requirement decomposition already have active benchmarks and design frameworks [4–7]. A future contribution cannot rest on assigning more categories. It must show, under an intervention, that a context-sensitive ledger and router improve valid, context-rooted departures beyond both a simple flat gate and an equally rich nongated record, while preserving protected requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Limitations
&lt;/h2&gt;

&lt;p&gt;These results are pilot evidence with substantial constraints. The studies use six to eight curated briefs per generation comparison, generally one sample per condition and brief, no reliable external model-provenance sweep, and related-agent or same-family LLM ratings. Several ledgers are oracle-supplied, so they isolate routing rather than end-to-end classification. Some controls are not perfectly token- or semantic-force-matched. In the multi-axis study, ratings are near ceiling and cliché incidence is zero, limiting sensitivity. The human-rater follow-up remains incomplete; the later LLM-rater pass cannot substitute for it. Finally, the separate programs cannot be statistically combined because they use different briefs, intervention formats, and rubrics.&lt;/p&gt;

&lt;p&gt;The paper therefore makes no claim that one method has been proven equivalent to another, that visible rationale exposes private reasoning, or that the framework improves creativity in general. It reports a convergent pattern of negative and indeterminate findings that narrows the next valid test.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Proposed Confirmatory Study
&lt;/h2&gt;

&lt;p&gt;A decisive study should preregister a frozen benchmark of at least 72 briefs: half in physical or software domains with mechanically checkable requirements and half in narrative or genre domains with disagreement explicitly annotated. Each brief should include four to six candidate claims. Independent annotators should provide multiple labels per claim, retain contested and unknown states, and report agreement rather than forcing consensus.&lt;/p&gt;

&lt;p&gt;The generation comparison should include six conditions: (A) direct generation; (B) model-generated flat gate plus router; (C) model-generated layered gate plus router; (D) content- and length-matched rich sham; (E) layered ledger without router; and (F) human-verified oracle layered ledger plus router. This design separates the effects of deliberation, representation, action restriction, and classification error. The primary outcome should be a blinded human judgment of &lt;strong&gt;valid, context-justified departure from an eligible default&lt;/strong&gt;, with preservation, novelty, feasibility, justification, and cliché reliance scored separately. Where possible, compilers, tests, simulations, or rule checks should validate invariants before human review.&lt;/p&gt;

&lt;p&gt;The layered method should advance only if it exceeds both the flat gate and rich sham by a preregistered practical margin on the primary outcome while remaining non-inferior on requirement preservation. If the flat gate alone exceeds the sham, the simpler representation is the result. If neither gate exceeds the sham, the correct publication is a null result: contextual elaboration helped, but explicit constraint classification did not add measurable value.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Conclusion
&lt;/h2&gt;

&lt;p&gt;The Codex and Fable pilot programs provide a disciplined corrective to an attractive intuition. Models can often distinguish clear functional constraints from familiar defaults, and explicit labels can steer their declared creative choices. Yet the current evidence does not show that those labels improve creative generation beyond equally rich contextual preparation; the context-rich control repeatedly matched or exceeded the gated alternatives. Constraint ledgers should therefore be retained as transparent interfaces for review and control, not promoted as validated creativity-enhancement mechanisms. Their strongest future test is causal, human-rated, and adversarially controlled—not a more elaborate rationale.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Reiter, R. (1980). &lt;em&gt;A logic for default reasoning&lt;/em&gt;. Artificial Intelligence, 13(1–2), 81–132. &lt;a href="https://doi.org/10.1016/0004-3702(80)90014-4" rel="noopener noreferrer"&gt;https://doi.org/10.1016/0004-3702(80)90014-4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Sloman, S. A., Love, B. C., &amp;amp; Ahn, W.-K. (1998). Feature centrality and conceptual coherence. &lt;em&gt;Cognitive Science, 22&lt;/em&gt;(2), 189–228. &lt;a href="https://doi.org/10.1207/s15516709cog2202_2" rel="noopener noreferrer"&gt;https://doi.org/10.1207/s15516709cog2202_2&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Turpin, M., Michael, J., Perez, E., &amp;amp; Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. &lt;a href="https://arxiv.org/abs/2305.04388" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2305.04388&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Chen, L., Zuo, H., Cai, Z., Yin, Y., Zhang, Y., Sun, L., Childs, P. R. N., &amp;amp; Wang, B. (2024). Towards controllable generative design: A conceptual design generation approach leveraging the function–behaviour–structure ontology and large language models. &lt;em&gt;Journal of Mechanical Design, 146&lt;/em&gt;(12). &lt;a href="https://doi.org/10.1115/1.4065562" rel="noopener noreferrer"&gt;https://doi.org/10.1115/1.4065562&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Xu, Z., et al. (2024). FollowBench: A multi-level fine-grained constraints following benchmark for large language models. &lt;em&gt;ACL 2024&lt;/em&gt;. &lt;a href="https://aclanthology.org/2024.acl-long.257/" rel="noopener noreferrer"&gt;https://aclanthology.org/2024.acl-long.257/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Shi, Y., et al. (2025). CFBench: A comprehensive constraints-following benchmark for LLMs. &lt;em&gt;ACL 2025&lt;/em&gt;. &lt;a href="https://aclanthology.org/2025.acl-long.1581/" rel="noopener noreferrer"&gt;https://aclanthology.org/2025.acl-long.1581/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Li, Z., et al. (2025). CARE-STaR: Constraint-aware self-taught reasoner. &lt;em&gt;Findings of ACL 2025&lt;/em&gt;, 21689–21703. &lt;a href="https://aclanthology.org/2025.findings-acl.1116/" rel="noopener noreferrer"&gt;https://aclanthology.org/2025.findings-acl.1116/&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is an independent work&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>discuss</category>
      <category>learning</category>
    </item>
    <item>
      <title>COGNEE as a memory layer</title>
      <dc:creator>Himanshu Negi</dc:creator>
      <pubDate>Sun, 05 Jul 2026 17:41:53 +0000</pubDate>
      <link>https://dev.to/himanshu_develops/cognee-as-a-memory-layer-1gn5</link>
      <guid>https://dev.to/himanshu_develops/cognee-as-a-memory-layer-1gn5</guid>
      <description>&lt;p&gt;Building Memoir: a health app that's actually supposed to know you&lt;/p&gt;

&lt;p&gt;I didn't set out to build a memory system. I set out to build a health app I'd actually want to use — medications, symptoms, workouts, mood, all in one calm place instead of scattered across four different apps that all feel like they were designed by a hospital's IT department in 2011. Warm colors. No dashboards screaming red numbers at you. Something that felt private, not like it was quietly compiling a dossier to sell to an insurer.&lt;/p&gt;

&lt;p&gt;The AI chat was supposed to be the easy part. Bolt on Gemini, let people ask it questions about their health data, done. It was not the easy part. It ended up being the part that taught me the most, mostly by breaking in ways I didn't see coming.&lt;/p&gt;

&lt;p&gt;The AI that knew a stranger named Himanshu&lt;br&gt;
The first version of the chat "worked" in the sense that it responded to messages. But early on I actually read the code path it used for context, and found this sitting there, dead serious, as the entire personalization layer:&lt;/p&gt;

&lt;p&gt;User: Himanshu, 25yo Male, 175cm, 72kg&lt;br&gt;
Medications: Metformin 500mg (twice daily, 93% adherence), Lisinopril 10mg...&lt;br&gt;
Hardcoded. Every single user, every single conversation, got told they were a 25-year-old man on three medications they'd never taken, with an adherence percentage somebody had just made up because it looked convincing in a demo. It wasn't lying to me specifically — it would've confidently told anyone the same fake life story. That's a specific kind of uncomfortable to discover in a health app: the AI was more confident about a stranger's medication schedule than it should ever be about a real one.&lt;/p&gt;

&lt;p&gt;The fix was mechanically simple — build the context string from whatever's actually in the person's profile and medication list instead of a fabricated one — but it changed how I thought about the whole feature. An AI health assistant that isn't grounded in your real data isn't a lesser version of the feature. It's a different, worse feature wearing the same UI.&lt;/p&gt;

&lt;p&gt;Giving it memory made it worse before it made it better&lt;br&gt;
Once the chat was honest about the current data, I wanted it to be smart about history too — notice that you've mentioned being stressed for a few days running, or that a headache pattern lines up with bad sleep, without you having to re-explain your whole week every time you open the chat. So I wired in Cognee, which turns stored text into an actual knowledge graph and lets you ask it things in plain language instead of writing a query for every pattern you can think of ahead of time.&lt;/p&gt;

&lt;p&gt;And it worked, right up until it worked too well. A few days into testing, the AI told me — with total confidence — that I was taking Metformin, Lisinopril, and something called a "Painkiller," and that my steps that day were approximately one hundred million. None of that was true. What had actually happened was that all my earlier testing, including that fake Himanshu profile, had gotten permanently absorbed into the memory graph, and nothing was telling the model that current, live data always outranks whatever the memory graph half-remembers. The graph doesn't forget, and it doesn't know that a fact from three weeks ago might not be a fact anymore.&lt;/p&gt;

&lt;p&gt;The actual fix was one paragraph in a system prompt, explicitly telling the model the live profile is the single source of truth and the memory graph is historical color commentary, never permission to state something as current fact. It's a small instruction carrying a lot of weight, and it's the kind of thing that's very easy to quietly delete in a future refactor because it looks like it's not doing anything. It's doing everything.&lt;/p&gt;

&lt;p&gt;Some bugs are dumb, some are just weird&lt;br&gt;
Not every bug taught me something profound. The "mark medication as taken" button, for instance, simply had no click handler on it. None. It looked completely real — hover states, color changes on press, the works — and did absolutely nothing when you tapped it. Sometimes the deep lesson is "read your own code more carefully," and that's fine.&lt;/p&gt;

&lt;p&gt;Other bugs were genuinely strange. The landing page had a rotating headline — "Your Personal Health OS," then Fitness OS, then Wellness OS — except the animated word was completely, invisibly gone. Not clipped, not the wrong color. Gone. It took actually screenshotting the rendered DOM to work out why: the gradient-text effect was applied to a wrapper whose children mixed position: absolute and position: relative, and browsers just refuse to paint a clipped gradient through that combination. The fix was applying the gradient to each word individually instead of the shared wrapper — three lines changed, half an hour of screenshotting weird invisible text to get there.&lt;/p&gt;

&lt;p&gt;I got the closely related lesson twice more before I stopped needing it a third time: once when a modal that should've been dead-center was instead offset by exactly half its own width and height (Framer Motion was quietly overwriting my manual centering transform), and again when a calorie input on mobile just stopped existing past a certain screen width, because two flex-1 inputs don't actually shrink below their own content size unless you tell them they're allowed to. CSS has more of these than I remembered.&lt;/p&gt;

&lt;p&gt;The moment I remembered other people exist&lt;br&gt;
For a long stretch, Memoir was built and tested as if there would only ever be one person using it, ever, on one browser, forever — a very easy trap when you're building alone. Then someone asked, reasonably: shouldn't a second person signing up get their own data instead of seeing whatever the first person left behind?&lt;/p&gt;

&lt;p&gt;The honest answer was no. Everything — medications, symptoms, diary entries, even the AI's memory graph — sat in one shared bucket with no concept of "whose is this." Fixing it meant going back through storage, the account system, and every API call that talked to the AI's memory, threading a real per-account ID through all of it. It's not a glamorous fix. It's the kind of thing that doesn't show up in a demo at all, right up until it's the whole reason the product is or isn't trustworthy.&lt;/p&gt;

&lt;p&gt;"Free" is doing a lot of work in that sentence&lt;br&gt;
Somewhere around adding email confirmations and a real "Continue with Google" button, I ran into the quieter reality of building something meant for other people, not just a screen recording: free tiers all have an asterisk. Resend gives you real email sending for free, but only to your own address until you verify a domain. Gemini's free tier is generous until it isn't — and it turns out the model spends part of its own answer budget on invisible "thinking" before it writes a word you can see, which explained why AI insights kept cutting off mid-sentence with no warning. Even Google sign-in, needing no backend at all for a client-side flow, still needs you to go create real credentials in a real console before a button click does anything.&lt;/p&gt;

&lt;p&gt;None of that is a complaint. It's just the gap between "I built a feature" and "a stranger can use this feature," and it's wider than it looks from the inside.&lt;/p&gt;

&lt;p&gt;What's still not real&lt;br&gt;
I'd rather say this plainly than let the app oversell itself. There's no appointment reminder that actually fires on the day — that needs a real server-side database and a scheduled job, and Memoir doesn't have either, by design, not as a placeholder for something bigger later. Google sign-in works, but it's a client-side identity handshake, not a full session system. Honest tradeoffs for what this is right now.&lt;/p&gt;

&lt;p&gt;What actually surprised me, looking back, is that almost none of the hard problems were about the AI itself. Gemini did what I asked it, basically every time. The hard problems were about trust — trusting the right data, trusting it to the right person, and being honest about what a free tier actually promises versus what it sounds like it promises. That turned out to be most of the job.&lt;/p&gt;

&lt;p&gt;Here's a link to the project: &lt;a href="https://memoir-sage.vercel.app/" rel="noopener noreferrer"&gt;https://memoir-sage.vercel.app/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Thank you for your kind attention!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mobile</category>
      <category>rag</category>
      <category>showdev</category>
    </item>
    <item>
      <title>New Year New Me</title>
      <dc:creator>Himanshu Negi</dc:creator>
      <pubDate>Mon, 02 Feb 2026 06:12:57 +0000</pubDate>
      <link>https://dev.to/himanshu_develops/new-year-new-me-1nn5</link>
      <guid>https://dev.to/himanshu_develops/new-year-new-me-1nn5</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/new-year-new-you-google-ai-2025-12-31"&gt;New Year, New You Portfolio Challenge Presented by Google AI&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  About Me
&lt;/h2&gt;

&lt;p&gt;Hi, I am a fairly new developer exploring the AI/ML field. Hoping to prove me and my skills. Finding different ways to do things is what i always look for! Wishing a happy and a great new year to everybody. &lt;/p&gt;

&lt;h2&gt;
  
  
  Portfolio
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pengu-205675755892.asia-south2.run.app/" rel="noopener noreferrer"&gt;https://pengu-205675755892.asia-south2.run.app/&lt;/a&gt;&lt;br&gt;
I know that im not the best one out there but I'll try and I'll keep on trying till i reach my goal and this portfolio is the embodiment of my journey. My will to be successful and all the times that I have failed. &lt;/p&gt;

&lt;p&gt;--labels dev-tutorial=devnewyear2026&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;I built it purely using Antigravity, though getting started with it was a bit confusing but eventually I figured it out. I really liked the feature of antigravity where it scrolls the website and can identify the issues.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Most Proud Of
&lt;/h2&gt;

&lt;p&gt;I really like the wall of failure on my portfolio website and the animation behind it, to be honest it was a bit hard to write the right prompt but it was worth it. Also it matters a lot to me and i believe everybody can relate to it because everybody at the beginning of the journey faces this phase.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pengu-205675755892.asia-south2.run.app/" rel="noopener noreferrer"&gt;Link to the website&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6jqy4ztx6so0k8com43z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F6jqy4ztx6so0k8com43z.png" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>googleaichallenge</category>
      <category>portfolio</category>
      <category>devnewyear2026</category>
    </item>
    <item>
      <title>RAG? or Text To SQL</title>
      <dc:creator>Himanshu Negi</dc:creator>
      <pubDate>Sat, 17 Jan 2026 16:40:07 +0000</pubDate>
      <link>https://dev.to/himanshu_develops/rag-or-text-to-sql-10bc</link>
      <guid>https://dev.to/himanshu_develops/rag-or-text-to-sql-10bc</guid>
      <description>&lt;p&gt;I recently worked on a project. To better understand which method works better for us, it is very important to understand the nature of the database. The database I was working on consisted of 16 columns and 20,001 rows, all related to company addresses and a few other details regarding their statuses. I was tasked with fetching data according to user queries, but I had to ensure there was zero hallucination. The system also needed to support aggregate functions, such as calculating maximums and averages. Since the chatbot was developed as a Proof of Concept (POC), the database size was intentionally kept small.&lt;/p&gt;

&lt;p&gt;Now that I am clear on the requirements and the nature of the dataset, I'll explain my approach and would like feedback on how I can improve it, as I am fairly new to the trade. I opted for a Text-to-SQL architecture because it provides higher deterministic accuracy than RAG and can perform average, max, and other operations; this is what we wanted to use on the database as it was a core functionality of the chatbot. Furthermore, since there will be a maximum of 100 concurrent users and the database supports horizontal scaling without the computational overhead of re-indexing vectors, I believe it is a more efficient approach—especially since the type of dataset I am working on has columns (like 'capital') that might be changed frequently.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>discuss</category>
      <category>security</category>
    </item>
  </channel>
</rss>
