<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sarah Pan</title>
    <description>The latest articles on DEV Community by Sarah Pan (@sarahpan).</description>
    <link>https://dev.to/sarahpan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4013323%2Fb2b6e204-dd2e-46bb-b193-e3cfe71f2245.jpg</url>
      <title>DEV Community: Sarah Pan</title>
      <link>https://dev.to/sarahpan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sarahpan"/>
    <language>en</language>
    <item>
      <title>I Found 3 Image Skills on GitHub That Feel Like Tiny Art Directors</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Thu, 13 Aug 2026 07:11:44 +0000</pubDate>
      <link>https://dev.to/sarahpan/i-found-3-image-skills-on-github-that-feel-like-tiny-art-directors-e06</link>
      <guid>https://dev.to/sarahpan/i-found-3-image-skills-on-github-that-feel-like-tiny-art-directors-e06</guid>
      <description>&lt;p&gt;I went to GitHub expecting to find more coding workflows. Instead, I found people packaging visual taste into Codex skills.&lt;/p&gt;

&lt;p&gt;The three repositories that caught my attention all create quiet, editorial, zine-inspired images. At first glance, they can look like variations of the same idea: upload a photograph or describe a theme, then receive a tasteful poster with paper texture, restrained typography, and plenty of negative space.&lt;/p&gt;

&lt;p&gt;But after reading through the repositories, I realized that the interesting part isn't the aesthetic alone. Each skill has a different opinion about what an image model should notice, what it is allowed to change, and how it should move from a source image or idea to a finished composition.&lt;/p&gt;

&lt;p&gt;They feel less like saved prompts and more like tiny art directors.&lt;/p&gt;

&lt;p&gt;This isn't a benchmark, and I haven't run the same photograph through all three. It is a closer look at how their creators have turned visual decisions into reusable workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Gathered Scenes Zine: Deciding What a Photograph Should Become
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/Zeejay0/gathered-scenes-zine-skill" rel="noopener noreferrer"&gt;Gathered Scenes Zine&lt;/a&gt; starts with an unusually thoughtful question: should the original photograph remain in the final work, or should it only inspire something new?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6eddsm6gjw4r1pgsi8gb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6eddsm6gjw4r1pgsi8gb.jpg" alt="My Test" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository includes two complementary skills.&lt;/p&gt;

&lt;p&gt;The first, Gathered Scenes, keeps the photograph as a visual anchor. It looks for relationships inside the scene—a person facing the distance, a road creating direction, a window holding light—and extends them through abstract illustration, structural color, negative space, and torn-paper edges.&lt;/p&gt;

&lt;p&gt;The second, Scene Distillation, removes the original photograph from the final image. It extracts the scene's emotional or semantic core and turns that into a new editorial illustration. A gesture, distance between two subjects, or an unfinished interaction can become the main visual metaphor.&lt;/p&gt;

&lt;p&gt;That distinction matters. Many image tools treat a reference photo as either something to copy or something to restyle. This repository treats it as material that first needs to be interpreted.&lt;/p&gt;

&lt;p&gt;Its workflow moves through observation, selection, translation, composition, and final output. The model is asked to preserve the minimum information needed for the scene to remain meaningful, rather than reproducing every visible detail.&lt;/p&gt;

&lt;p&gt;In other words, the skill doesn't begin with “make this look like a zine.” It begins with “what is actually important in this photograph?”&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Photo Abstract Editorial: Preserving the Photograph as Evidence
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/ZzzLc0405/photo-abstract-editorial" rel="noopener noreferrer"&gt;Photo Abstract Editorial&lt;/a&gt; takes a more restrained approach.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz2r4ca83a45gzfv8eab.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz2r4ca83a45gzfv8eab.png" alt="My Test" width="800" height="1300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It transforms a photograph into a vertical editorial composition with two main regions: the original photography and an abstract “memory panel.” The photograph must remain truthful. It should not be redrawn, expanded, or quietly altered by the image model.&lt;/p&gt;

&lt;p&gt;The abstract panel is then derived from relationships already present in the source: its colors, spatial structure, rhythm, and shapes. Every important visual element in that panel should be traceable back to something that genuinely exists in the photograph.&lt;/p&gt;

&lt;p&gt;I like this constraint because image generation often becomes less interesting when the model is allowed to invent everything. Here, creativity comes from interpretation rather than replacement.&lt;/p&gt;

&lt;p&gt;The result is closer to an editorial designer responding to a photograph. The designer cannot change what happened in the image, but can decide what to place beside it, which colors to repeat, how much space to leave empty, and what kind of title changes the way we read it.&lt;/p&gt;

&lt;p&gt;The repository also makes its underlying prompts available in Chinese and English and describes which parts users can adjust, including the photo-to-panel ratio, abstraction level, typography, and color balance. It presents the visual system as a starting point rather than an untouchable template.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. GC Minimal Zine Poster: A Small Visual Production System
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/LiamGvchi/gc-minimal-zine-poster" rel="noopener noreferrer"&gt;GC Minimal Zine Poster&lt;/a&gt; is broader and more technically structured than the other two.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3jpy5ohodhw4f2a3uphj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3jpy5ohodhw4f2a3uphj.png" alt="My Test2" width="800" height="1334"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It can turn a theme, sentence, article idea, object, mood, photograph, or collection of references into a minimal editorial poster. Its visual language is specific—large areas of negative space, one small visual event, restrained typography, a high-chroma color anchor, and visible print or paper imperfections—but the repository is designed to produce variation inside that system.&lt;/p&gt;

&lt;p&gt;What makes it especially interesting is its routing.&lt;/p&gt;

&lt;p&gt;The skill can generate an image, return only a production-ready prompt, analyze references without generating anything, or analyze a visual system and then create a new composition from it. When photographs are supplied, it classifies their roles and records how strictly each one should be preserved.&lt;/p&gt;

&lt;p&gt;It also separates fixed visual rules from variable decisions and source-specific residue. That is a useful distinction when working from references. A model should be able to learn that a design system uses sparse layouts and one strong color without copying the exact subject, text, or composition of a sample image.&lt;/p&gt;

&lt;p&gt;The repository includes separate reference files for its style system, prompt compiler, variation engine, reference-analysis process, and quality checks. At that point, “prompt” no longer feels like the most accurate description. It looks more like a compact creative production pipeline.&lt;/p&gt;

&lt;p&gt;A Skill Is an Opinion About Process&lt;/p&gt;

&lt;p&gt;These repositories changed how I think about image-generation skills.&lt;/p&gt;

&lt;p&gt;A normal image prompt usually describes the desired output: the format, subject, colors, composition, and style. A skill can also describe the process that should produce that output.&lt;/p&gt;

&lt;p&gt;It can tell an agent to inspect the source before generating, distinguish facts from creative interpretation, choose between different routes, preserve specific elements, avoid copying reference identity, and review the result against a quality bar.&lt;/p&gt;

&lt;p&gt;That doesn't mean every long image prompt becomes a useful skill. A folder full of aesthetic adjectives is still just a complicated prompt. The more convincing projects here contain decisions:&lt;/p&gt;

&lt;p&gt;What should be observed first?&lt;/p&gt;

&lt;p&gt;What must remain unchanged?&lt;/p&gt;

&lt;p&gt;What can vary between generations?&lt;/p&gt;

&lt;p&gt;When should the original photograph be preserved?&lt;/p&gt;

&lt;p&gt;How should a reference influence the result without being copied?&lt;/p&gt;

&lt;p&gt;What makes an output unacceptable even if it looks attractive?&lt;/p&gt;

&lt;p&gt;Those questions are usually answered silently by a designer. In these repositories, some of that judgment has been made explicit enough for an agent to follow.&lt;/p&gt;

&lt;p&gt;Can Taste Be Packaged?&lt;/p&gt;

&lt;p&gt;There is an obvious tension here.&lt;/p&gt;

&lt;p&gt;Part of what makes a visual style interesting is the creator's judgment in a particular moment. Once that judgment becomes a reusable package, it can help more people create coherent work—but it can also make a distinctive aesthetic spread very quickly and become familiar just as quickly.&lt;/p&gt;

&lt;p&gt;The repositories themselves hint at this issue. They include rules against copying sample-specific content, and two of them restrict commercial use without permission. These are not minor details. A public GitHub repository does not automatically mean its visual system can be repackaged, sold, or used in client work.&lt;/p&gt;

&lt;p&gt;At the time of writing, GC Minimal Zine Poster uses the MIT License, while Gathered Scenes Zine and Photo Abstract Editorial limit use to personal, educational, or non-commercial contexts unless the creator gives additional authorization. Anyone installing these skills should read the license, not just the SKILL.md.&lt;/p&gt;

&lt;p&gt;There is also a question of authorship. If someone defines the observation method, composition rules, visual boundaries, and evaluation criteria, how much of the resulting image belongs to the person who typed the final request? I don't have a clean answer, but creative skills make that question harder to ignore.&lt;/p&gt;

&lt;p&gt;GitHub May Become a Place for Sharing Creative Systems&lt;/p&gt;

&lt;p&gt;GitHub has always been a place where developers share code, tools, and ways of working. Skills add another layer: people can now share a structured way of making decisions.&lt;/p&gt;

&lt;p&gt;For coding agents, that might mean a release workflow or a migration process. For image generation, it can mean a way of reading photographs, extracting visual relationships, protecting source material, and deciding when a composition is finished.&lt;/p&gt;

&lt;p&gt;That is what I found most interesting about these three projects. The generated posters are attractive, but the real artifact is the creative logic behind them.&lt;/p&gt;

&lt;p&gt;We may be moving from sharing prompts to sharing small, portable creative systems.&lt;/p&gt;

&lt;p&gt;And if those systems keep becoming easier to install and reuse, GitHub might become a surprisingly important place for discovering not only what AI can generate, but how different people have taught it to see.&lt;/p&gt;

&lt;p&gt;Have you found any creative skills that changed the way you use an AI agent? I'd love to see what else people are building.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How Do You Choose Which LLM to Use in Production? I'd Love to Hear Your Process.</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Tue, 11 Aug 2026 07:25:41 +0000</pubDate>
      <link>https://dev.to/sarahpan/how-do-you-choose-which-llm-to-use-in-production-id-love-to-hear-your-process-14c0</link>
      <guid>https://dev.to/sarahpan/how-do-you-choose-which-llm-to-use-in-production-id-love-to-hear-your-process-14c0</guid>
      <description>&lt;p&gt;The more models I test, the less comfortable I am answering the question: Which LLM should I use?&lt;/p&gt;

&lt;p&gt;A model can look like the obvious choice in a benchmark and still be wrong for a real product. Maybe it is too slow. Maybe the price stops making sense at scale. Maybe it follows instructions well in a clean test but falls apart on the messy inputs users actually send.&lt;/p&gt;

&lt;p&gt;And even when you make a good choice, it may only stay good for a few months. I have been thinking about this a lot while working on TokenBay, where we connect several models through one OpenAI-compatible API. I spend a lot of time comparing model behavior, pricing, latency, and the friction involved in switching providers. The closer I look, the harder it is to believe in one universally “best” model.&lt;/p&gt;

&lt;p&gt;There is only the model that fits your current task and constraints.&lt;/p&gt;

&lt;p&gt;So I am curious: How are teams making this decision in practice?&lt;/p&gt;

&lt;p&gt;Four strategies I keep running into&lt;/p&gt;

&lt;p&gt;Most model-selection discussions seem to lead back to one of four approaches.&lt;/p&gt;

&lt;p&gt;Use different models for different jobs&lt;/p&gt;

&lt;p&gt;This is the most flexible setup. A team might use one model for writing, another for code, and a smaller, cheaper model for classification or tagging.&lt;/p&gt;

&lt;p&gt;It makes sense because these tasks do not reward the same strengths. The downside is that a multi-model architecture creates more work around fallbacks, error handling, monitoring, and billing.&lt;/p&gt;

&lt;p&gt;You stop searching for one model that does everything, but now you have to maintain the routing system that decides what does what.&lt;/p&gt;

&lt;p&gt;Start cheap and escalate when needed&lt;/p&gt;

&lt;p&gt;Another approach is to send a request to a cheaper model first, then escalate when the task is difficult or the result does not meet a quality threshold.&lt;/p&gt;

&lt;p&gt;I like the idea, but the phrase “when needed” hides the hardest part. How do you know that a response is not good enough without paying another model to judge it? A confidence score is useful for some structured tasks, but much less clear for writing, research, or open-ended reasoning.&lt;/p&gt;

&lt;p&gt;The routing logic can become more complicated than the original integration.&lt;/p&gt;

&lt;p&gt;Choose one provider and stay there&lt;/p&gt;

&lt;p&gt;This is often described as lock-in, but sometimes it is a completely rational decision.&lt;/p&gt;

&lt;p&gt;If a provider has already passed a company’s security and compliance review, adding another one may create more risk and work than it removes. A slightly better model is not automatically worth a new legal review, another billing relationship, and a second set of operational failures.&lt;/p&gt;

&lt;p&gt;For these teams, the best model may simply be the one they are allowed to use and can audit properly.&lt;/p&gt;

&lt;p&gt;Keep testing whatever is new&lt;/p&gt;

&lt;p&gt;Some teams treat model evaluation as an ongoing process. When a new release appears, they run it against their own tasks and switch if the improvement is meaningful.&lt;/p&gt;

&lt;p&gt;This sounds ideal in a fast-moving market, but constant experimentation has a cost too. Prompts behave differently, tool calls break in new ways, and a better benchmark score can still produce a worse user experience.&lt;/p&gt;

&lt;p&gt;At some point, staying current can turn into chasing every release.&lt;/p&gt;

&lt;p&gt;None of these strategies is obviously correct. They optimize for different things: flexibility, cost, stability, or speed of adoption.&lt;/p&gt;

&lt;p&gt;The questions I still struggle with&lt;/p&gt;

&lt;p&gt;When is model lock-in actually a problem?&lt;/p&gt;

&lt;p&gt;Model-agnostic architecture sounds safer. If one provider goes down, raises its prices, or changes model behavior, you have somewhere else to go.&lt;/p&gt;

&lt;p&gt;But abstraction is not free. Providers handle tool calling, structured output, rate limits, caching, and errors differently. Even when APIs look compatible, applications still end up depending on provider-specific behavior.&lt;/p&gt;

&lt;p&gt;So where do you draw the line? Do you build for portability from day one, or wait until switching becomes a real need?&lt;/p&gt;

&lt;p&gt;How do you compare models in production?&lt;/p&gt;

&lt;p&gt;Offline evals are useful, but they miss things users feel immediately: latency, consistency across a conversation, tone, and the strange edge cases that never made it into the test set.&lt;/p&gt;

&lt;p&gt;Do you shadow traffic to a second model? Route a small percentage of users to it? Keep each user on the same model so the experience stays consistent? What do you measure besides thumbs-up and thumbs-down?&lt;/p&gt;

&lt;p&gt;I am especially interested in what people actually do here, because production evaluation is usually much messier than the diagrams make it look.&lt;/p&gt;

&lt;p&gt;What happens when a new model launches?&lt;/p&gt;

&lt;p&gt;Do you test it immediately, wait for other developers to find the problems, or ignore it until your current setup gives you a reason to change?&lt;/p&gt;

&lt;p&gt;Testing every release can consume a surprising amount of time. Ignoring new releases can leave real improvements on the table. I still do not know what a healthy evaluation cadence looks like for a small team.&lt;/p&gt;

&lt;p&gt;Which matters most: cost, latency, or quality?&lt;/p&gt;

&lt;p&gt;“Pick two” is a simplification, but most products still have to favor one or two of these constraints.&lt;/p&gt;

&lt;p&gt;For a coding assistant, users may wait longer for a better answer. For autocomplete or live chat, an extra second can make the product feel broken. For high-volume classification, a small price difference can matter more than a quality gain that users never notice.&lt;/p&gt;

&lt;p&gt;The answer changes with the task, which is another reason a company-wide “default model” can be misleading.&lt;/p&gt;

&lt;p&gt;The decision that matters after the decision&lt;/p&gt;

&lt;p&gt;My current view is that choosing a model matters less than knowing how you will revisit that choice.&lt;/p&gt;

&lt;p&gt;Before committing to one, I would want to know:&lt;/p&gt;

&lt;p&gt;Do we have a small set of real prompts that represent our product?&lt;/p&gt;

&lt;p&gt;Can we compare cost, latency, and output quality on those prompts?&lt;/p&gt;

&lt;p&gt;How much provider-specific behavior have we built around?&lt;/p&gt;

&lt;p&gt;What would have to happen for us to switch?&lt;/p&gt;

&lt;p&gt;Could we test another model without rewriting the whole application?&lt;/p&gt;

&lt;p&gt;That does not mean every product needs an elaborate router or five providers. It just means “we use Model X” should not become an architectural fact nobody is allowed to question.&lt;/p&gt;

&lt;p&gt;Working on a multi-model platform has made me more skeptical of permanent model choices. The model that fits a workflow today may not be the one that fits it six months from now. The useful skill is not predicting the winner. It is building a process that lets you notice when your answer has changed.&lt;/p&gt;

&lt;p&gt;So I would love to hear how other developers handle this:&lt;/p&gt;

&lt;p&gt;What is your current LLM setup, and what would make you switch?&lt;/p&gt;

&lt;p&gt;If you have migrated a production workflow between models, I am especially curious about what was harder than you expected.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>I Automated My Content Workflow with n8n. Did It Actually Save Time?</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Mon, 03 Aug 2026 07:57:44 +0000</pubDate>
      <link>https://dev.to/sarahpan/i-automated-my-content-workflow-with-n8n-did-it-actually-save-time-574e</link>
      <guid>https://dev.to/sarahpan/i-automated-my-content-workflow-with-n8n-did-it-actually-save-time-574e</guid>
      <description>&lt;p&gt;I make content about AI products and industry trends, and honestly, a big chunk of the job is repetitive. I collect recent info, look for an angle, draft a script, dump everything into a spreadsheet, and then turn whatever survives into a video.&lt;/p&gt;

&lt;p&gt;Doing that by hand every day got old fast, because half my time went into moving the same information between tools before I could even start the actual creative part. Most of those steps felt predictable, so at some point I figured I might as well connect them with n8n. The goal was simple enough: collect data, generate topics, write a script, and get everything ready for video production. And the workflow worked — it's just that the content didn't always work with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;The first part of the workflow pulls information from the sources I normally use for research, and that data goes to an AI model which picks out possible topics and drafts an angle for each one before everything lands in a spreadsheet. So instead of opening ten websites and copy-pasting into a doc, I just scroll through one sheet where each row already has the source material plus a possible direction. Once I pick a topic, the next stage drafts a script, and originally I wanted that draft to go straight into production — the dream pipeline looked like this: research → topic selection → script → video.&lt;/p&gt;

&lt;p&gt;I never fully automated the final editing part, since I still used separate tools for visuals and tweaked things by hand, but I did expect the workflow to hand me material that was basically production-ready. And technically, it did. The scripts had intros, explanations, and conclusions, structured enough to become videos, and I was definitely producing faster than before. Then I watched the finished videos, and I didn't want to publish them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow produced content, but it didn't find the story
&lt;/h2&gt;

&lt;p&gt;One of the videos explained how businesses can show up better in AI-generated answers. It was a relevant topic and the script covered the basics correctly，and being online doesn't mean an AI model actually understands or recommends you — so nothing in it was exactly wrong, it just wasn't very interesting. The script opened with the concept and then stacked examples on top, which made the whole thing feel like a short lesson. It told people why the topic mattered, but never gave them a real reason to keep watching.&lt;/p&gt;

&lt;p&gt;Around the same time, I made another video, mostly by hand, where I asked a few AI models to recommend a milk tea franchise for someone with 300,000 RMB and zero food industry experience. Most models suggested big, familiar brands, but ChatGPT went somewhere different, and when I kept asking why, it turned out the models were interpreting the budget differently. Some only counted the franchise fee and equipment, while ChatGPT was also factoring in rent, deposits, initial inventory, and working capital. &lt;/p&gt;

&lt;p&gt;That video didn't come from a clean automated pipeline; it started with one specific question and grew through an actual comparison, and the models disagreeing with each other was the story right there. Viewers got it instantly with same person, same question, different answers, and there were real stakes, because if you follow a recommendation without checking what the budget actually covers, you've made a very expensive mistake. The automated video looked cleaner, sure, but the milk tea one had the thing that matters more: a discovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  I automated the wrong decision
&lt;/h2&gt;

&lt;p&gt;At first I blamed the video production tool, so I changed the visuals, blew up the key text, and added more motion. It got slightly better, but it still felt empty, and it took me a while to realize the problem started way earlier in the pipeline. The model can summarize information and shape it into a reasonable script, but what it can't do reliably is tell which result has real tension in it — and when every topic goes through the same structure, a weak idea comes out looking almost as polished as a strong one.&lt;/p&gt;

&lt;p&gt;That actually makes reviewing harder, because a nicely formatted script tricks you into thinking the idea is ready even when the story underneath is too general to carry anything. The workflow saved me from the blank page, but it also nudged me to keep going with ideas I would've dropped much earlier if I'd just looked at them properly. I had treated topic selection like another repetitive task, and it turns out it's the part that needs the most judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where n8n actually earns its keep
&lt;/h2&gt;

&lt;p&gt;None of this means the workflow failed. The research stage got way easier, the spreadsheet gave me one place to compare topics, and I stopped losing time shuffling info between tools — plus, once I'd picked a promising idea, it was decent at producing a rough first draft. The trouble only started when I expected it to handle the whole editorial process. A model can tell that a topic is AI-related, and it can explain why that topic might matter to some audience, but neither of those things means there's a story in it worth telling.&lt;/p&gt;

&lt;p&gt;For my content, the strong ideas usually start from something concrete: two models disagreeing, a company missing from every single recommendation, or an AI answer that flips when you add one small detail. Those situations create a question the viewer actually wants answered. Something broad like "why AI visibility matters" might be commercially relevant, but it still needs a real case behind it, otherwise you end up with a script full of correct explanations and zero investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I'm changing it
&lt;/h2&gt;

&lt;p&gt;I'm not trying to automate research-to-finished-video anymore. n8n still collects and organizes information, and it can still throw out initial questions and rough scripts, but before anything goes into production, I run the experiment myself and decide whether the result is worth showing.&lt;/p&gt;

&lt;p&gt;For example, instead of asking AI for a general script about manufacturers in AI search, I could ask several models to recommend Chinese steel suppliers for an overseas buyer. If every one of them skips a real factory, that's worth digging into — maybe the company has no clear English product info, no export records, nothing the models can verify. Now there's an actual question behind the video: why did a real supplier vanish from every answer? The workflow can collect the responses and organize the comparison, it just shouldn't get to decide the comparison is meaningful before I've even seen it.&lt;/p&gt;

&lt;p&gt;I'm also simplifying the format. My team recently moved toward image-based videos with narration, plus text-based educational content, which are faster to produce and leave more time for research. A fancy animation doesn't rescue a weak idea, and a specific finding stays interesting even with simple visuals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automation should remove work, not judgment
&lt;/h2&gt;

&lt;p&gt;When I first built this, I saw content production as a chain of steps, and I assumed that automating every step would make the whole thing more efficient. That was only half right. Some steps are just moving information around, while others are about deciding why the information matters — n8n is genuinely great at the first kind, and my mistake was assuming the second kind works the same way.&lt;/p&gt;

&lt;p&gt;I'm still using the workflow, but its job is smaller and clearer now: it prepares the material so I can spend my time deciding what deserves to become content. The best automation didn't replace the creative process, it just bought me more time for it.&lt;/p&gt;

&lt;p&gt;Have you ever built a workflow that ran perfectly and produced something you didn't actually want?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Is Vibe Coding Really Helping Beginners Learn?</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Fri, 31 Jul 2026 07:28:17 +0000</pubDate>
      <link>https://dev.to/sarahpan/is-vibe-coding-really-helping-beginners-learn-1jl8</link>
      <guid>https://dev.to/sarahpan/is-vibe-coding-really-helping-beginners-learn-1jl8</guid>
      <description>&lt;p&gt;AI coding tools allow people with little programming experience to build websites and apps much earlier than before. This has made programming more accessible, especially for beginners who want to turn an idea into a real project without spending months studying first. Seeing something work can also make learning feel less distant and give people a reason to continue.&lt;/p&gt;

&lt;p&gt;Social media has made this possibility look even more attractive. We often see people claiming that they built a complete product in a weekend or launched a startup without knowing how to code on social media. Their projects usually look professional, so it is easy to assume that AI has removed most of the difficulty from programming.&lt;/p&gt;

&lt;p&gt;The problem is that building a working project and learning how to program are not the same thing.&lt;/p&gt;

&lt;p&gt;A beginner can ask AI to create a website without understanding how the code works. When an error appears, they can send it back to the model and request a fix. This may solve the immediate problem, but it does not always help them understand what caused it. As the project becomes more complicated, they may need AI to explain or repair every new issue.&lt;/p&gt;

&lt;p&gt;This can create a difficult cycle. Each AI-generated change adds something the user may not understand, which makes the next problem harder to judge. The project keeps growing while the beginner becomes more dependent on the model. Instead of learning the knowledge needed to manage the project, they may spend more money on tokens or switch to a different tool whenever the current one stops producing useful answers.&lt;/p&gt;

&lt;p&gt;I have noticed this tendency in my own work. When something breaks, asking AI is often my first reaction because it is faster than investigating the problem myself. The model may fix it within a few minutes, but when a similar issue appears later, I sometimes realize that I still cannot explain the earlier solution. I completed the task without gaining much knowledge from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Success Stories Can Be Misleading
&lt;/h2&gt;

&lt;p&gt;An impressive AI-assisted project does not show how much its creator already knew. Someone who understands programming can use AI very effectively because they know how to evaluate its suggestions. AI helps them work faster, but their existing knowledge still guides the process.&lt;/p&gt;

&lt;p&gt;A beginner may copy the same workflow and receive a very different result. They can use the same model and follow the same tutorial, yet still struggle because they do not have the background knowledge that made the original creator successful. That difference is rarely visible in a short demo or a post about launching a product in one weekend.&lt;/p&gt;

&lt;p&gt;This can make beginners think that they are using the wrong tool. They may buy another course or pay for a more expensive model, hoping that it will finally make the process easy. In reality, the missing piece may be basic programming knowledge rather than better AI.&lt;/p&gt;

&lt;p&gt;The culture around vibe coding sometimes turns learning into a form of FOMO. People feel pressure to launch something quickly because everyone online appears to be doing it. Learning the fundamentals feels slow in comparison, and there is little social reward for spending several days understanding one concept.&lt;/p&gt;

&lt;p&gt;The result is that some beginners attempt projects far beyond their current ability. They may produce a polished prototype, but they cannot explain how it works or maintain it without constant help. The weakness becomes clear when the project reaches a stage that requires real technical judgment.&lt;/p&gt;

&lt;p&gt;It can also become a problem when they apply for jobs. A portfolio project may look impressive, but an interviewer can ask why a certain technical choice was made. If the applicant cannot explain the code, the project offers much less evidence of their ability.&lt;/p&gt;

&lt;p&gt;The same issue appears when someone tries to turn an AI-built prototype into a real product. A functioning demo is only an early step. Once real users are involved, the creator must be able to find problems and decide which solutions are safe. Those decisions cannot be judged only by whether the code runs.&lt;/p&gt;

&lt;p&gt;This helps explain why many AI-built projects do not lead to a job or earn money. The creator may have invested considerable time and paid for several tools, yet still lack the knowledge needed to move the project forward. AI helped produce more code, but it did not automatically build the ability to manage that code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Benefits Most from AI Coding?
&lt;/h2&gt;

&lt;p&gt;AI coding currently offers the greatest advantage to people who already understand programming. An experienced developer can use it to explore an unfamiliar codebase or find possible causes of a bug. More importantly, they can recognize when the answer is unreliable.&lt;/p&gt;

&lt;p&gt;This does not mean beginners should avoid AI. It means they may need to use it differently. Instead of asking the model to complete an entire ambitious project, they can use a smaller project to understand one concept at a time. When AI changes the code, they should pause long enough to understand why the change worked.&lt;/p&gt;

&lt;p&gt;For beginners, learning basic programming may still be more valuable than trying to launch a complex product immediately. Projects remain useful because they give the knowledge a real purpose. The problem begins when completing the project becomes more important than understanding it.&lt;/p&gt;

&lt;p&gt;This balance is difficult to maintain in an online culture built around speed. AI makes it possible to keep generating, so stopping to study can feel like falling behind. Yet the slower part of the process is often where real learning takes place.&lt;/p&gt;

&lt;p&gt;AI has lowered the barrier to building software, but it has not removed the need to understand it. Beginners can now create projects earlier in their learning journey. They still need enough knowledge to make decisions when the AI-generated solution stops working.&lt;/p&gt;

&lt;p&gt;If you were learning programming today, would you start by building with AI or spend more time learning the fundamentals first?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Asked AI for a Video. It Gave Me an Animated PowerPoint</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Wed, 29 Jul 2026 09:41:15 +0000</pubDate>
      <link>https://dev.to/sarahpan/i-asked-ai-for-a-video-it-gave-me-an-animated-powerpoint-1dp8</link>
      <guid>https://dev.to/sarahpan/i-asked-ai-for-a-video-it-gave-me-an-animated-powerpoint-1dp8</guid>
      <description>&lt;p&gt;I recently researching how to make a short vertical video with vibe coding    with painfully detailed prompts.&lt;/p&gt;

&lt;p&gt;I wrote the narration, timing, aspect ratio, color palette, subtitle rules, transitions, evidence screenshots, music direction, and even a list of visual clichés in the prompts to avoid mistakes. I described the result as an “AI search investigation,” with real model responses, highlighted fields, budget comparisons, and a clean monitoring interface.&lt;/p&gt;

&lt;p&gt;What came back was a valid MP4 that the resolution was correct, the scenes appeared in roughly the right order, the numbers were accurate; but it also looked like an animated PowerPoint.&lt;/p&gt;

&lt;p&gt;There were circles, rectangles, progress bars, and very large text. One scene showed “300,000” inside a ring. Another showed a range next to a flat horizontal bar. Most of the actual visual story—AI interfaces, search evidence, cost items, ranking changes—had quietly disappeared.&lt;/p&gt;

&lt;p&gt;The AI had followed the checklist. It had also completely missed the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prompt was detailed, but not visually specific
&lt;/h2&gt;

&lt;p&gt;My first reaction was that the prompt still wasn’t detailed enough.&lt;/p&gt;

&lt;p&gt;Then I looked at it again. It was already several pages long. Adding another paragraph about making the video “more polished,” “more cinematic,” or “more high-tech” probably would not have fixed anything.&lt;/p&gt;

&lt;p&gt;The problem was that I had mixed two different things into one document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Production requirements&lt;/li&gt;
&lt;li&gt;Creative direction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Production requirements are easy to verify. The video is either 1080×1920 or it isn’t. A scene lasts six seconds or it doesn’t. A subtitle appears inside the safe area or it doesn’t.&lt;/p&gt;

&lt;p&gt;Creative direction is much harder.&lt;/p&gt;

&lt;p&gt;“Make it feel like a real investigation” sounds clear to a person, but it leaves dozens of decisions unresolved. Should the interface resemble a browser, a terminal, or an analytics dashboard? How dense should the screen be? How large should the subtitles feel on a phone? What makes a frame feel investigative instead of corporate?&lt;/p&gt;

&lt;p&gt;When these choices were missing, the agent used the safest visual vocabulary it knew: cards, circles, large numbers, and smooth slide transitions.&lt;/p&gt;

&lt;p&gt;In other words, it converted ambiguity into geometry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Programmatic video solves a different problem
&lt;/h2&gt;

&lt;p&gt;This experience made me look more carefully at the tools underneath AI-generated videos.&lt;/p&gt;

&lt;p&gt;Remotion treats React code as the source of truth. Scenes can be written as components, connected to data, previewed in the browser, and rendered in batches. It is especially appealing when the video contains repeated layouts, dynamic data, reusable components, or content that needs to be generated at scale.&lt;/p&gt;

&lt;p&gt;HyperFrames takes a different approach. A composition is written with HTML and CSS, while timing is described through data attributes and seekable animations. It can work with GSAP, Lottie, Three.js, CSS animations, and other browser-native tools, then render the result through headless Chrome and FFmpeg.&lt;/p&gt;

&lt;p&gt;The distinction matters.&lt;/p&gt;

&lt;p&gt;Remotion gives you the structure and component model of React. HyperFrames gives an agent a format it already generates constantly: HTML and CSS. Neither one automatically provides taste, but both make timing, layout, assets, and rendering editable instead of hiding them inside a generated video file.&lt;/p&gt;

&lt;p&gt;That is the real advantage of programmatic video for me. It makes mistakes repairable.&lt;/p&gt;

&lt;p&gt;If a subtitle is too large, I can change one token. If every scene has too much empty space, I can update the shared layout. If the timing changes after I record the narration, I can adjust a configuration file instead of manually moving every clip along a timeline.&lt;/p&gt;

&lt;p&gt;But code gives you control only after someone has made the right visual decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do differently next time
&lt;/h2&gt;

&lt;p&gt;The biggest mistake was generating the entire video before validating its visual language.&lt;/p&gt;

&lt;p&gt;Next time, I would begin with the hardest 10 to 15 seconds. In this case, that would be the moment when three AI models recommend the same brand and the fourth gives a surprising answer.&lt;/p&gt;

&lt;p&gt;That sample should answer a few questions before the full render starts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do the screenshots feel like evidence or decoration?&lt;/li&gt;
&lt;li&gt;Is the text still readable on a phone?&lt;/li&gt;
&lt;li&gt;Does the animation have variation in speed and pauses?&lt;/li&gt;
&lt;li&gt;Does the interface feel like an investigation?&lt;/li&gt;
&lt;li&gt;Are the visual layers doing more than repeating the narration?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If those seconds do not work, rendering another minute in the same style only produces a longer mistake.&lt;/p&gt;

&lt;p&gt;I would also replace vague aesthetic words with visible rules.&lt;/p&gt;

&lt;p&gt;Instead of saying “make it feel advanced,” I could specify a graphite background, one accent color, small technical labels, restrained scanning animations, real screenshots as the main layer, and no oversized rounded cards.&lt;/p&gt;

&lt;p&gt;Instead of saying “use modern subtitles,” I could define the font, weight, maximum width, line count, position, shadow, and which three sentences are allowed to become full-screen titles.&lt;/p&gt;

&lt;p&gt;Reference frames would help even more. A few annotated screenshots can explain hierarchy, density, and texture far better than another 500 words of adjectives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assets should be treated as dependencies
&lt;/h2&gt;

&lt;p&gt;Another lesson was that missing visual assets should not silently become abstract shapes.&lt;/p&gt;

&lt;p&gt;If a scene requires four real AI responses and only two screenshots exist, the agent should stop and mark the other two as missing. A clearly labeled placeholder is more useful than a polished animation that pretends the evidence was never needed.&lt;/p&gt;

&lt;p&gt;I now think video assets should be handled much like code dependencies:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;List every required asset.&lt;/li&gt;
&lt;li&gt;Map it to a scene.&lt;/li&gt;
&lt;li&gt;Validate that it exists.&lt;/li&gt;
&lt;li&gt;Flag missing or low-resolution files.&lt;/li&gt;
&lt;li&gt;Only then build the full composition.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That sounds less exciting than prompting a complete video in one shot, but it prevents the entire production from drifting away from the original idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maybe AI should build the system before it builds the video
&lt;/h2&gt;

&lt;p&gt;I still think AI coding agents are genuinely useful for video production.&lt;/p&gt;

&lt;p&gt;They can scaffold scenes, centralize timing, create reusable animation components, generate subtitle files, inspect keyframes, and render multiple versions. Tools like Remotion and HyperFrames make that workflow much more accessible.&lt;/p&gt;

&lt;p&gt;What I trust less now is the one-shot request: “Here is my script. Please make the finished video.”&lt;/p&gt;

&lt;p&gt;A better division of work might be:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;- The human defines the visual language.&lt;/li&gt;
&lt;li&gt;- The agent builds a reusable system around it.&lt;/li&gt;
&lt;li&gt;- Both iterate on a short sample.&lt;/li&gt;
&lt;li&gt;- The full video is rendered only after that system survives review.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That process is slower during the first 15 seconds and much faster during the remaining 60.&lt;/p&gt;

&lt;p&gt;I’m still figuring out where prompting ends and creative direction begins. Maybe programmatic video will eventually make that boundary disappear. For now, it seems that AI is very good at building the thing it can verify—and very willing to approximate everything else with a rectangle.&lt;/p&gt;

&lt;p&gt;How do you handle this gap? Do you provide visual references, prototype one scene first, or accept that the final layer of creative direction still needs to be done by hand?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Asked Three AI Models the Same Question. Now I Had Three Problems.</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Thu, 23 Jul 2026 06:29:05 +0000</pubDate>
      <link>https://dev.to/sarahpan/i-asked-three-ai-models-the-same-question-now-i-had-three-problems-17hc</link>
      <guid>https://dev.to/sarahpan/i-asked-three-ai-models-the-same-question-now-i-had-three-problems-17hc</guid>
      <description>&lt;p&gt;I used to think that asking multiple AI models was a good way to verify an answer.&lt;/p&gt;

&lt;p&gt;If GPT gave me an answer I wasn’t sure about, I would ask Claude. If Claude disagreed, I would open Gemini and try again.&lt;/p&gt;

&lt;p&gt;In theory, the third model should settle it.&lt;/p&gt;

&lt;p&gt;In practice, I often ended up with three different answers, three convincing explanations, and no idea which one to trust.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More answers didn’t mean more certainty&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This happens to me most often when I’m working on questions with one correct answer.&lt;/p&gt;

&lt;p&gt;One model chooses A and explains why B is wrong. Another chooses B and gives an explanation that sounds just as reasonable. The third sometimes introduces an entirely different interpretation of the question.&lt;/p&gt;

&lt;p&gt;The strange part isn’t that they disagree. Different models are trained and tuned differently, so disagreement should be expected.&lt;/p&gt;

&lt;p&gt;The strange part is how confident every answer sounds.&lt;/p&gt;

&lt;p&gt;None of them says, “I might be misunderstanding this.” Each one presents its reasoning as if the problem has already been settled.&lt;/p&gt;

&lt;p&gt;At that point, comparing answers becomes less like fact-checking and more like listening to three persuasive people argue about something I don’t understand well enough to judge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Majority vote isn’t always the answer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A common solution is to ask several models and choose the answer that appears most often.&lt;/p&gt;

&lt;p&gt;That can help, but only if their errors are independent.&lt;/p&gt;

&lt;p&gt;If two models rely on the same common misconception, the majority can still be wrong. Models may also interpret the wording in similar ways or reproduce the same widely repeated but incorrect information.&lt;/p&gt;

&lt;p&gt;Two matching answers are not automatically two separate pieces of evidence.&lt;/p&gt;

&lt;p&gt;There’s another problem: if I ask the models one after another and include previous responses, they may start agreeing simply because they’ve seen the same reasoning. That creates consensus, but not necessarily accuracy.&lt;/p&gt;

&lt;p&gt;Agreement can be useful. It just isn’t proof.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The disagreement is actually the useful part&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I’ve started treating model disagreement as a signal to slow down.&lt;/p&gt;

&lt;p&gt;Instead of asking, “Which model should I trust?” I try to find the exact point where their reasoning separates.&lt;/p&gt;

&lt;p&gt;Did they interpret a word differently? Did one of them use an incorrect formula? Did they make different assumptions? Is one answer based on information that may be outdated?&lt;/p&gt;

&lt;p&gt;Once the conflict becomes specific, it is much easier to verify.&lt;/p&gt;

&lt;p&gt;For example, instead of giving one model the other model’s full response and asking which is correct, I can ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Model A chose A, while Model B chose B. Identify the exact assumption or reasoning step that causes the disagreement. Do not choose a winner yet.&lt;br&gt;
Then I can ask each model to challenge its own answer:&lt;/p&gt;

&lt;p&gt;What is the strongest reason your answer could be wrong? What evidence or test would distinguish between the competing answers?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This doesn’t guarantee a correct result, but it changes the models’ job. They are no longer competing to sound convincing. They are helping define what needs to be checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My current workflow&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When an answer matters, I now use a simple process:&lt;/p&gt;

&lt;p&gt;Ask each model independently, without showing it the other answers.&lt;br&gt;
Compare the conclusions and list the exact points of disagreement.&lt;br&gt;
Ask each model to critique its own reasoning.&lt;br&gt;
Verify the disputed claim using a reliable source, calculation, or test.&lt;br&gt;
Choose the answer supported by evidence—not the one written with the most confidence.&lt;/p&gt;

&lt;p&gt;For coding questions, the final judge might be a test. For math, it might be recalculating the result step by step. For current facts, it should be a reliable and recent source.&lt;/p&gt;

&lt;p&gt;The important part is that another AI model is not automatically the final authority.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More models can reveal uncertainty, not remove it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I still think using multiple models is valuable. A second model can catch assumptions the first one missed, and disagreement can expose weak points that would otherwise remain hidden.&lt;/p&gt;

&lt;p&gt;But multiple models don’t automatically create a verification system.&lt;/p&gt;

&lt;p&gt;Sometimes they produce consensus. Sometimes they produce a better answer. And sometimes they simply give you several polished versions of uncertainty.&lt;/p&gt;

&lt;p&gt;The real advantage is not that one of them must be right. It’s that their disagreement can show you where verification is actually needed.&lt;/p&gt;

&lt;p&gt;So now, when GPT, Claude, and Gemini give me three different answers, I don’t immediately ask for a fourth.&lt;/p&gt;

&lt;p&gt;First, I ask them why they disagree.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Adding One Limitation Made the AI Recommendation More Confident</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Wed, 22 Jul 2026 07:09:46 +0000</pubDate>
      <link>https://dev.to/sarahpan/adding-one-limitation-made-the-ai-recommendation-more-confident-4pgf</link>
      <guid>https://dev.to/sarahpan/adding-one-limitation-made-the-ai-recommendation-more-confident-4pgf</guid>
      <description>&lt;p&gt;I recently tested two versions of the same product description to see whether a small wording change would affect how an AI evaluated it.&lt;/p&gt;

&lt;p&gt;The first version was simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Gives developers access to multiple AI models through one API.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The second version added one sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;TokenBay gives developers access to multiple AI models through one API. It is most useful for teams that regularly compare or switch models, but may be unnecessary for projects that only use one provider.&lt;br&gt;
I expected the second version to make the product sound less appealing. After all, product copy usually tries to remove doubt, not introduce it. Saying that some people may not need the product felt a little like putting a warning label on the homepage.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model reacted in the opposite way.&lt;/p&gt;

&lt;p&gt;Its recommendation became more specific and, strangely, more confident. Instead of describing the product as a generally useful multi-model platform, it explained that TokenBay made sense for teams testing several providers, comparing outputs, or trying to avoid maintaining separate accounts and integrations. It also noted that a small project committed to one model would probably not benefit as much.&lt;/p&gt;

&lt;p&gt;Nothing about the product had changed. The model simply had a clearer boundary around when the product was useful.&lt;/p&gt;

&lt;p&gt;That made me realize that limitations can give an AI something important: a reason to rule a product out.&lt;/p&gt;

&lt;p&gt;Most product descriptions only explain why someone should use the product. They list features, benefits, and broad claims about who it is for. The result often sounds positive but vague. A platform is “flexible,” “powerful,” and “built for modern teams,” which could describe half the software products currently asking for my email address.&lt;/p&gt;

&lt;p&gt;A limitation narrows the picture. Once the description says that the product may be unnecessary for single-provider projects, the intended user becomes easier to identify. The model can compare the product against a real situation instead of trying to interpret a list of features.&lt;/p&gt;

&lt;p&gt;I do not think the lesson is that every product page should suddenly lead with everything the product cannot do. That would be honest, but perhaps not excellent marketing. The useful part is giving enough context for the product to be evaluated properly.&lt;/p&gt;

&lt;p&gt;For example, these two statements communicate very different levels of information:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Supports multiple AI models.&lt;br&gt;
and:&lt;/p&gt;

&lt;p&gt;Designed for teams that regularly test or switch models and do not want to maintain separate provider integrations.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The first tells you what exists. The second tells you when it matters.&lt;/p&gt;

&lt;p&gt;Adding a limitation makes that distinction even clearer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Probably unnecessary if your application only relies on one provider.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now the product has a recognizable shape. It is no longer trying to be useful to everyone, which makes the recommendation feel less like generic promotion and more like an actual judgment.&lt;/p&gt;

&lt;p&gt;This may matter more as people increasingly ask AI systems to compare tools for them. A model needs enough information to decide not only why a product fits, but also when another option would make more sense. Without that boundary, the safest answer is often vague: the product “could be useful depending on your needs.”&lt;/p&gt;

&lt;p&gt;With a clear limitation, the model can say something more helpful.&lt;/p&gt;

&lt;p&gt;The part I found most interesting was that the limitation did not weaken the recommendation. It made the recommendation easier to justify.&lt;/p&gt;

&lt;p&gt;Maybe good product descriptions should not only help a model rule the product in. They should also help it rule the product out.&lt;/p&gt;

&lt;p&gt;That sounds slightly uncomfortable from a marketing perspective, but it may be what makes the recommendation believable.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>reviews</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Do Students Still Need to Learn Coding the Old Way?</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Tue, 21 Jul 2026 07:24:43 +0000</pubDate>
      <link>https://dev.to/sarahpan/do-students-still-need-to-learn-coding-the-old-way-546g</link>
      <guid>https://dev.to/sarahpan/do-students-still-need-to-learn-coding-the-old-way-546g</guid>
      <description>&lt;p&gt;A friend of mine teaches at a university, and she has been noticing that more students are using AI heavily for coding assignments and projects. Some of them cannot explain most parts of the code they submit, and they struggle to build the same thing from scratch, but their projects can still work surprisingly well. They are able to turn an idea into something people can actually use from making small apps, 2D &amp;amp; 3D games, connect APIs. When they meet some troubles in their code, they first asked their AI to debug or not go to the Office Hour or discuss with friends. A few years ago, students would usually need much more traditional programming knowledge before they could even get to that stage, so this shift makes me wonder what learning to code is supposed to look like now.&lt;/p&gt;

&lt;p&gt;I do not think this means traditional programming skills have suddenly become useless. Students still need enough technical understanding to notice when the AI is producing something unreliable and to understand why a project becomes difficult to maintain when they finding a job. AI can generate a lot of code very quickly, but it can also generate a very large mess with impressive confidence. Without some knowledge of data structures, students may be able to make a project look finished without really knowing whether it is stable.&lt;/p&gt;

&lt;p&gt;At the same time, it may no longer make sense to require students to spend months proving they can build everything without assistance before they are allowed to create something useful. In most real jobs, people already use frameworks, libraries, documentation, search engines, code examples, and now AI assistants. The ability to work with tools has always been part of programming. What has changed is how much those tools can now do, and how quickly someone with limited technical experience can reach a working first version.&lt;/p&gt;

&lt;p&gt;The biggest change may be who gets to build. A design student with a strong visual idea can create an interactive prototype. A film student can make a tool for organizing footage or planning a production. A humanities student who understands a niche community can build something for users that a traditional engineering team might never have considered. In this environment, creativity, taste, subject knowledge, and the ability to describe a problem clearly become much more valuable, because technical execution is no longer the only barrier between an idea and a working product.&lt;/p&gt;

&lt;p&gt;This does not necessarily mean that art or humanities students will become more competitive than engineering students. It may mean that the advantage is shifting toward people who can combine several kinds of knowledge. An engineering student who understands users, design, and business may become far more effective with AI. A creative student who learns enough technical basics to evaluate and improve AI-generated code may also become capable of building things that were previously out of reach.&lt;/p&gt;

&lt;p&gt;Maybe the real question is no longer whether students should learn programming in the traditional way, but which parts of traditional programming still matter most. Should they spend as much time memorizing syntax, or more time understanding systems, testing, debugging, and making good technical decisions? Should assignments measure whether students can write every line alone, or whether they can explain, verify, and improve what AI helped them produce? These are also questions that universities and instructors need to take seriously. &lt;/p&gt;

&lt;p&gt;From what I have seen, many computer science and engineering programs still discourage or completely prohibit the use of AI and vibe coding in coursework. That concern is understandable, especially when teachers need to know whether students actually understand the material, but the curriculum also has to keep pace with the tools students will encounter in the workplace. If schools continue teaching as if these tools do not exist, they may leave students at a disadvantage rather than protect their learning.&lt;/p&gt;

&lt;p&gt;I am curious how other developers and teachers see this. If you were designing a programming curriculum for students who will grow up with AI coding tools, what would you keep, and what would you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>career</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Building a Model-Agnostic Agent Harness: What Breaks When You Switch LLMs</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Mon, 20 Jul 2026 08:01:00 +0000</pubDate>
      <link>https://dev.to/sarahpan/building-a-model-agnostic-agent-harness-what-breaks-when-you-switch-llms-32j1</link>
      <guid>https://dev.to/sarahpan/building-a-model-agnostic-agent-harness-what-breaks-when-you-switch-llms-32j1</guid>
      <description>&lt;p&gt;I’ve seen quite a few “multi-model agents” lately.&lt;/p&gt;

&lt;p&gt;Usually there’s a dropdown in the corner of the screen that allows switching from one model to another. Which is useful. I also enjoy using dropdowns. It makes everything feel very configurable.&lt;/p&gt;

&lt;p&gt;But then, when you click on that dropdown…&lt;/p&gt;

&lt;p&gt;The model may return an answer that is in a different format than you asked for, or which calls into tools that either omit or include one that was previously not called at all.&lt;/p&gt;

&lt;p&gt;The dropdown worked perfectly.&lt;/p&gt;

&lt;p&gt;The agent, unfortunately, has developed a new personality.&lt;/p&gt;

&lt;p&gt;The model changed. Everything around it noticed.&lt;/p&gt;

&lt;p&gt;Imagine a small agent that reads a support ticket, looks up the customer’s account, and returns a suggested reply.&lt;/p&gt;

&lt;p&gt;One model might return exactly what you asked for:&lt;br&gt;
&lt;code&gt;&lt;br&gt;
{&lt;br&gt;
  "priority": "high",&lt;br&gt;
  "action": "refund",&lt;br&gt;
  "reply": "I’ve issued a refund..."&lt;br&gt;
}&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Another may return the same information with a short introduction:&lt;/p&gt;

&lt;p&gt;`Here is the recommended response:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "priority": "high",&lt;br&gt;
  "action": "refund",&lt;br&gt;
  "reply": "I’ve issued a refund..."&lt;br&gt;
}`&lt;/p&gt;

&lt;p&gt;A third may decide recommended_action is a much nicer field name than action.&lt;/p&gt;

&lt;p&gt;To a person, these answers are basically the same. To the next function in the pipeline, they may be three completely different situations.&lt;/p&gt;

&lt;p&gt;This is where “supports multiple models” starts becoming less about access and more about everything wrapped around the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool use is not identical either&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The differences become more obvious when the agent has tools.&lt;/p&gt;

&lt;p&gt;One model may look up the customer account immediately. Another may ask a follow-up question first. A third may write a confident response without checking the account at all, which is efficient in the same way skipping all your tests is efficient.&lt;/p&gt;

&lt;p&gt;Even when models have access to the same tools, they do not always make the same decisions about when to use them.&lt;/p&gt;

&lt;p&gt;That means the harness has to do more than pass messages back and forth. It needs to check whether the required steps actually happened.&lt;/p&gt;

&lt;p&gt;Did the model call the account tool before recommending a refund? Did it receive a valid result? Did it finish the task, or did it simply announce that it was finished?&lt;/p&gt;

&lt;p&gt;Apparently, “done” is also a matter of opinion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A failed agent does not always return an error&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The annoying failures are rarely the dramatic ones.&lt;/p&gt;

&lt;p&gt;A timeout is easy to notice. A rate-limit error is clear. You can retry or switch providers.&lt;/p&gt;

&lt;p&gt;The stranger case is when the API returns 200 OK, the model gives you a polished answer, and the workflow has absolutely no idea what to do next.&lt;/p&gt;

&lt;p&gt;The request succeeded. The task did not.&lt;/p&gt;

&lt;p&gt;So a useful multi-model agent needs its own definition of success. The output may need to pass a schema check. Certain tools may need to appear in the trace. A value may need to fall within an expected range. If those checks fail, the harness can retry, repair the response, or move the task to another model.&lt;/p&gt;

&lt;p&gt;Without that layer, model switching is mostly optimistic copy-pasting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The less photogenic part of multi-model systems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We’ve been running into this while working on TokenBay.&lt;/p&gt;

&lt;p&gt;Connecting several models through one API is the neat part you can show in a demo. The less photogenic part is making sure the rest of the workflow does not panic every time the model changes.&lt;/p&gt;

&lt;p&gt;In practice, that means normalizing outputs, checking tool calls, handling different failure patterns, and keeping the application’s internal contract stable even when the model is not.&lt;/p&gt;

&lt;p&gt;The models do not need to behave identically. They probably never will.&lt;/p&gt;

&lt;p&gt;The workflow just needs to know what to do when they don’t.&lt;/p&gt;

&lt;p&gt;A model dropdown can give users a choice. A real multi-model agent needs enough plumbing behind that choice to make switching boring.&lt;/p&gt;

&lt;p&gt;And boring, in production, is usually a compliment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>50,000 AI Dramas a Month. So Why Have You Never Seen One?</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Fri, 17 Jul 2026 05:48:37 +0000</pubDate>
      <link>https://dev.to/sarahpan/china-made-more-ai-dramas-in-3-months-than-existed-in-the-platforms-entire-history-before-this-year-28nd</link>
      <guid>https://dev.to/sarahpan/china-made-more-ai-dramas-in-3-months-than-existed-in-the-platforms-entire-history-before-this-year-28nd</guid>
      <description>&lt;p&gt;On Douyin (China's TikTok), the total number of AI-generated comic/anime-style dramas that had ever aired before 2026 was around 58,000 titles.&lt;/p&gt;

&lt;p&gt;In Q1 2026 alone that number jumped to roughly 180,000. About 50,000 of those were added in March alone. So, in one quarter, China produced more than double the entire back catalog that took years to build.&lt;/p&gt;

&lt;p&gt;The reason is the cost curve, and it's not subtle:&lt;/p&gt;

&lt;p&gt;Traditional live-action drama: ~$1,400 per minute of footage&lt;br&gt;
AI-generated comic drama: as low as $14 per minute&lt;/p&gt;

&lt;p&gt;That's a 100x difference. A full traditional series costs $40K–170K. A full AI comic-drama series costs $7K–21K.&lt;/p&gt;

&lt;p&gt;Meanwhile, of the ~61,000 comic dramas released in 2025, only 96 crossed 100 million views. A hit rate of 0.16%.&lt;br&gt;
And this isn't staying in China: Southeast Asia (Thailand, Malaysia, Philippines) is picking it up fast, with around a third of internet users there already watching this kind of content monthly. However, the mainstream Western audiences largely haven't joined in. In Spain, the highest-engagement country in Europe, only about 7% of internet users have watched one in the past month.&lt;/p&gt;

&lt;p&gt;So the question I keep coming back to: is this what an industry actually maturing looks like, or is it just what it looks like right before the correction hits?&lt;/p&gt;

&lt;p&gt;Curious if anyone's tracking similar cost-collapse → content-flood patterns in other media categories (stock photography, game asset generation, etc), feeling like the same shape.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>reviews</category>
      <category>news</category>
      <category>discuss</category>
    </item>
    <item>
      <title>The API Call Succeeded. The Workflow Still Broke.</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Thu, 16 Jul 2026 06:24:26 +0000</pubDate>
      <link>https://dev.to/sarahpan/the-api-call-succeeded-the-workflow-still-broke-5cpb</link>
      <guid>https://dev.to/sarahpan/the-api-call-succeeded-the-workflow-still-broke-5cpb</guid>
      <description>&lt;p&gt;We're building TokenBay, a unified API that routes requests across GPT, Claude, Gemini, DeepSeek, and a few dozen other models. A few weeks ago we ran a fairly boring internal test: send the same structured-output prompt to five different models and compare what came back. It turned into one of the more useful debugging sessions we've had this quarter, mostly because nothing about it looked broken at first.&lt;/p&gt;

&lt;p&gt;The task was simple. Extract product, budget, and location from a user request, return JSON. Every request succeeded. Every model returned something a human would call correct. And our downstream workflow — the thing that actually parses the response and does something with it — broke anyway.&lt;/p&gt;

&lt;p&gt;The responses all looked fine&lt;/p&gt;

&lt;p&gt;Here's the input we sent:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;I need a laptop under $800. I'm in California.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;GPT-5.4 came back with exactly what we expected:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;{&lt;br&gt;
  "product": "laptop",&lt;br&gt;
  "budget": 800,&lt;br&gt;
  "location": "California"&lt;br&gt;
}&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Claude Opus 4.8 gave us the right values, but wrapped in a sentence first:&lt;/p&gt;

&lt;p&gt;`Here is the extracted information:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "product": "laptop",&lt;br&gt;
  "budget": "$800",&lt;br&gt;
  "location": "California"&lt;br&gt;
}`&lt;/p&gt;

&lt;p&gt;Still correct, if you're a human reading it. Gemini 3.1 changed the field names entirely:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;{&lt;br&gt;
  "item": "laptop",&lt;br&gt;
  "max_price": 800,&lt;br&gt;
  "region": "California"&lt;br&gt;
}&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Also correct. DeepSeek V4 followed the schema but represented the missing currency differently again. Four models, four slightly different shapes, and our parser — which was expecting exactly product, budget, and location — choked on three out of four.&lt;/p&gt;
&lt;h2&gt;
  
  
  A unified endpoint doesn't mean unified behavior
&lt;/h2&gt;

&lt;p&gt;This is the part that's easy to underestimate when you're evaluating a multi-model API. The pitch is usually about the request layer: one key, one endpoint, swap the model name and you're done. That part is true, and it's genuinely useful. But it quietly implies the models are interchangeable once you can call them the same way, and they aren't.&lt;/p&gt;

&lt;p&gt;They interpret the same instruction differently. They make different assumptions about missing values. They disagree on how strict "return JSON" actually means. Two models can agree completely on the answer and disagree completely on how that answer should be formatted — and that difference is invisible in a demo, because a person is reading the output. In production, the next thing reading it is usually another piece of software, and it's a lot less forgiving than a person skimming a chat window.&lt;/p&gt;
&lt;h2&gt;
  
  
  Tightening the prompt helped, but not enough
&lt;/h2&gt;

&lt;p&gt;Our first move was the obvious one — make the instruction stricter. We went from:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Return the result as JSON.&lt;/code&gt;&lt;br&gt;
to&lt;br&gt;
&lt;code&gt;Return only valid JSON. Do not include any explanation or Markdown.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That fixed the Claude wrapper-text problem immediately. It didn't fix everything. Some models represented a missing field as null, some as an empty string, some just omitted the key entirely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"laptop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"budget"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"location"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"laptop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"budget"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"location"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"product"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"laptop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"budget"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;800&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three are valid JSON. None of them are the same contract. At that point it was clear we weren't asking the models for a format — we needed to ask our own workflow to enforce a schema, because the models were never going to agree on one for us.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual fix was a normalization layer, not a smarter prompt
&lt;/h2&gt;

&lt;p&gt;Prompt tuning bought us a few percentage points, but it's fragile by nature — a model can follow an instruction 99 times out of 100 and still surprise you on the hundredth call. If that hundredth response ends up in a billing flow or an automated action, "usually correct" isn't a bar we were willing to accept.&lt;/p&gt;

&lt;p&gt;So we stopped passing raw model output directly into the app. Everything now goes through a small adapter first: pull JSON out of any surrounding text, map alternate field names onto our internal schema, coerce currency strings into numbers, collapse empty strings to null, and reject anything incomplete outright. A simplified version of what that looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;RawResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;ProductRequest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;product&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;normalizeResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RawResult&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;ProductRequest&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;product&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rawBudget&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;budget&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;max_price&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;location&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;location&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;region&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;product&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;product&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Missing product&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;rawBudget&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;number&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;budget&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rawBudget&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;rawBudget&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;parsed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rawBudget&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;[^&lt;/span&gt;&lt;span class="sr"&gt;0-9.&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isNaN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;budget&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;parsed&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;product&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;product&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="nx"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
      &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;location&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;location&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;
        &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;location&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing here is clever, and that's kind of the point — the real production version handles a longer list of edge cases (arrays instead of strings, nested objects, models that helpfully "correct" your currency to EUR), but the core idea doesn't change: stop trusting the raw shape and normalize before anything downstream touches it.&lt;/p&gt;

&lt;h2&gt;
  
  
  This also changed what we count as a failure
&lt;/h2&gt;

&lt;p&gt;Before this, our fallback logic only triggered on request failure — timeout, rate limit, a 4xx or 5xx from the provider. If a request came back as 200 OK, we treated it as done.&lt;/p&gt;

&lt;p&gt;But a 200 with an unusable payload isn't a success, it's a failure wearing a success status code. So we moved the fallback trigger downstream:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Send request → Receive response → Parse output → Validate schema → Normalize → Continue, or retry with another model&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The question changed from "did the API respond?" to "did the API produce something the application can actually use?" Those aren't the same question, and building fallback logic around the first one instead of the second one will pass every test you write until it doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we test now
&lt;/h2&gt;

&lt;p&gt;We used to eyeball a handful of outputs and call the prompt "good enough." Now the test suite specifically checks for the boring stuff: extra text wrapped around the JSON, renamed fields, numbers returned as strings, omitted optional fields, inconsistent handling of missing values, and whether the normalized result from every model passes the exact same downstream checks. A response can be fluent, accurate, and still completely unusable to a parser — that's its own failure mode, and it needs its own test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Being able to call five models through one API makes them easier to reach. It doesn't make them interchangeable. The call can succeed, the model can understand the task perfectly, the answer can even be right — and the workflow can still break quietly downstream. Routing the request to a different model is a config change. Deciding what the rest of your application is allowed to assume about that model's output is the actual engineering work, and it's the part that doesn't show up in a benchmark chart.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;We ran into this while building TokenBay, a unified API for routing across GPT, Claude, Gemini, DeepSeek and other models. If you've hit similar cross-model quirks, I'd be curious to hear how you handled them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>javascript</category>
      <category>api</category>
    </item>
    <item>
      <title>The Best AI Model Is Usually the Wrong Question</title>
      <dc:creator>Sarah Pan</dc:creator>
      <pubDate>Wed, 15 Jul 2026 06:25:50 +0000</pubDate>
      <link>https://dev.to/sarahpan/the-best-ai-model-is-usually-the-wrong-question-1koh</link>
      <guid>https://dev.to/sarahpan/the-best-ai-model-is-usually-the-wrong-question-1koh</guid>
      <description>&lt;p&gt;The more AI becomes part of ordinary product infrastructure, the less sensible it feels to treat a single model as a permanent foundation.&lt;/p&gt;

&lt;p&gt;Models will keep improving. Prices will keep moving. New providers will keep appearing.&lt;/p&gt;

&lt;p&gt;A good workflow should expect that.&lt;/p&gt;

&lt;p&gt;Why we started building TokenBay&lt;/p&gt;

&lt;p&gt;This is also one of the reasons we started building TokenBay.&lt;/p&gt;

&lt;p&gt;Not because developers need another long list of AI models, but because testing and changing models should not require rebuilding everything around them.&lt;/p&gt;

&lt;p&gt;The idea is simple: one account, one API, and access to different models without having to manage a separate setup for every provider.&lt;/p&gt;

&lt;p&gt;We are still learning where this is most useful, but the problem has become clearer to me:&lt;/p&gt;

&lt;p&gt;The difficult part is not finding a powerful model.&lt;/p&gt;

&lt;p&gt;It is building a workflow that does not become outdated every time a better one arrives.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>programming</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
