<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Marco Reyes</title>
    <description>The latest articles on DEV Community by Marco Reyes (@marcoreyes1).</description>
    <link>https://dev.to/marcoreyes1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4009360%2F43fe6a4e-0aa7-4057-b186-980a417d8302.png</url>
      <title>DEV Community: Marco Reyes</title>
      <link>https://dev.to/marcoreyes1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marcoreyes1"/>
    <language>en</language>
    <item>
      <title>WAIC 2026 Robotics Hardware: Why Dexterous Hands and Tiny Motors Matter</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Mon, 20 Jul 2026 09:52:51 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/waic-2026-robotics-hardware-why-dexterous-hands-and-tiny-motors-matter-58lk</link>
      <guid>https://dev.to/marcoreyes1/waic-2026-robotics-hardware-why-dexterous-hands-and-tiny-motors-matter-58lk</guid>
      <description>&lt;p&gt;The most interesting robotics story from WAIC 2026 was not another humanoid doing a stage trick. It was the hardware that makes the boring work possible: hands, actuators, motor drivers, tactile sensing, and the control layer that sits between them.&lt;/p&gt;

&lt;p&gt;That sounds less exciting than a walking robot demo, but it is where the real engineering problem lives.&lt;/p&gt;

&lt;p&gt;If a robot can walk across a booth but cannot pick up a soft package, hold a tool, recover from a slip, or survive repeated cycles in a warehouse, it is still mostly a demo platform. The hand is where “embodied AI” stops being a model slogan and starts becoming a bill of materials.&lt;/p&gt;

&lt;p&gt;WAIC 2026 made that shift pretty visible. Shanghai’s official conference coverage says the event ran from July 17 to 20, with more than 1,100 companies and over 3,000 exhibits. Several Chinese reports also framed this year’s robotics area around real-world work instead of pure motion demos. That framing matters because the questions change. Instead of asking “Can the robot move?”, buyers and engineers start asking “Can this subsystem ship, integrate, and keep working?”&lt;/p&gt;

&lt;p&gt;For dexterous hands, the first thing to check is not the number of fingers. It is the split between mechanical freedom and controllable freedom.&lt;/p&gt;

&lt;p&gt;A five-finger robotic hand may list 20 or more degrees of freedom, but those DOF can mean different things. Some are actively driven. Some are passively coupled. Some rely on tendons, linkages, direct drive, micro linear actuators, or a mix. That design choice affects weight, backlash, responsiveness, repairability, and how painful the control stack becomes.&lt;/p&gt;

&lt;p&gt;Joyson Electronics used WAIC 2026 to introduce its TeleHand series. The company describes the professional version as a 20-DOF hand using an in-palm hybrid actuation design, combining direct drive, tendon drive, and linkage drive. That is a useful example of where the category is going: not one magic actuator, but a packaging problem where multiple drive types are placed where they make the most mechanical sense.&lt;/p&gt;

&lt;p&gt;Xynova’s Flex 2 is another good reference point. Xynova’s own product page lists a 23-DOF hand, a 400 g palm, repeatability of ±0.1 mm, force control accuracy down to 0.05 N, a 12 kg single-hand grasp load, and a 4 kg rated continuous load. Those are strong headline numbers, but for integration work I would still ask for the less glamorous details: active/passive DOF breakdown, thermal curves, cycle-life test conditions, cable replacement process, calibration steps, and the control API.&lt;/p&gt;

&lt;p&gt;That is where many procurement conversations get soft. “Dexterous” is easy to say. “Here is the latency budget from tactile event to actuator response under load” is harder.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzgwYTc4N2ZiNjM2M2Q3ODM5YTc1ZGJjMTExZjM4MzlfNmE3YWFMMVhpeFNXa1ltRVRKQllBRFg0R2E1cllwUFVfVG9rZW46Wk9WZGJZTTJZb3NMMkd4TkhPV2NmRmQ5blJmXzE3ODQ1NDEwMTQ6MTc4NDU0NDYxNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzgwYTc4N2ZiNjM2M2Q3ODM5YTc1ZGJjMTExZjM4MzlfNmE3YWFMMVhpeFNXa1ltRVRKQllBRFg0R2E1cllwUFVfVG9rZW46Wk9WZGJZTTJZb3NMMkd4TkhPV2NmRmQ5blJmXzE3ODQ1NDEwMTQ6MTc4NDU0NDYxNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from WAIC Official Page" width="1159" height="785"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The motor story is similar.&lt;/p&gt;

&lt;p&gt;Humanoid robots usually divide motor choices by space and load. Larger joints such as shoulders, hips, knees, and elbows often push designers toward high-torque, high-density actuator modules. Fingers and small palm mechanisms have tighter packaging constraints, so coreless motors, micro linear actuators, and compact drive trains show up often.&lt;/p&gt;

&lt;p&gt;Frameless torque motors matter because they strip the motor down to rotor and stator so the robot designer can embed the motor inside a joint structure. That can reduce redundant housing, improve integration, and support compact joint modules. But “frameless torque motor” is not enough information by itself. For real comparison, you need outer diameter, stack length, peak torque, continuous torque, torque density, cogging torque, encoder options, heat dissipation method, and production yield.&lt;/p&gt;

&lt;p&gt;Some Chinese coverage around WAIC highlighted micro frameless torque motors and even used “global smallest” language. I would treat that kind of claim as a prompt to request a data sheet, not as a conclusion. A 10 mm-class or otherwise ultra-small motor may be impressive, but the procurement question is whether it can hold continuous torque in a small enclosed hand, survive heat, and be manufactured consistently at scale.&lt;/p&gt;

&lt;p&gt;A tiny actuator that wins a headline but needs custom tuning for every batch is not really a supply-chain win.&lt;/p&gt;

&lt;p&gt;The driver electronics are easy to overlook, but they decide whether the mechanical design is controllable. Awinic’s WAIC-related Vanex dexterous hand article is useful here because it talks about the less flashy layer: motor driver chips with H-bridge integration, current sensing, current regulation, protection circuits, PWM/GPIO control, and fault protection. Those are not social-media-friendly specs, but they are exactly the things you need when a finger stalls, slips, or hits an unexpected object.&lt;/p&gt;

&lt;p&gt;For me, a practical WAIC 2026 robotics hardware checklist looks like this.&lt;/p&gt;

&lt;p&gt;Dexterous hand:&lt;/p&gt;

&lt;p&gt;Check active DOF, passive DOF, hand weight, rated continuous load, peak load, tactile sensing channels, repeatability, cycle life, and supported communication interfaces.&lt;/p&gt;

&lt;p&gt;Miniature motor or actuator:&lt;/p&gt;

&lt;p&gt;Check diameter, torque curve, efficiency, thermal behavior, gearbox or screw pairing, backlash, noise, service life, and whether the published number is peak or continuous.&lt;/p&gt;

&lt;p&gt;Driver and sensing electronics:&lt;/p&gt;

&lt;p&gt;Check current sensing, stall detection, overcurrent protection, voltage range, control mode, sensor sampling rate, and whether diagnostics are exposed to the host controller.&lt;/p&gt;

&lt;p&gt;Control layer:&lt;/p&gt;

&lt;p&gt;Check whether the hand requires a vendor-specific stack, whether it supports ROS 2 or common industrial buses, how calibration works, and how much task logic lives on the hand versus the main robot controller.&lt;/p&gt;

&lt;p&gt;This is also why the “robotics Android” idea keeps coming back.&lt;/p&gt;

&lt;p&gt;A robotic hand is not useful in isolation. The interesting system is a hardware abstraction layer that lets a model or planner control different hands without rewriting everything from scratch. Xinhua’s coverage of RoboScience at WAIC described a demo where different dexterous hands were swapped and controlled by the same embodied model stack within a short time window. That is not proof that robotics has found its Android moment, but it points at the real need: hardware should not force every manipulation model to become a one-off integration project.&lt;/p&gt;

&lt;p&gt;The industry is still far from clean standardization. Hands differ too much in mechanics, sensors, safety behavior, and control frequency. But the direction is clear. The winners will probably not be the teams with the loudest booth demo. They will be the teams that make the hand, actuator, driver, and control interface feel boring enough for other engineers to build on.&lt;/p&gt;

&lt;p&gt;That is a good thing.&lt;/p&gt;

&lt;p&gt;Robotics hardware needs fewer miracle claims and more boring reliability. WAIC 2026 was interesting because the conversation moved closer to that reality. Dexterous hands are becoming a serious component category. Micro motors and actuators are becoming strategic supply-chain items. Driver chips and tactile sensing are moving into the center of the stack.&lt;/p&gt;

&lt;p&gt;For developers, the takeaway is simple: do not evaluate humanoid robots only from the model layer down. Evaluate them from the contact point up.&lt;/p&gt;

&lt;p&gt;The hand tells you what the robot can actually do.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>sharepoint</category>
      <category>productivity</category>
      <category>automation</category>
    </item>
    <item>
      <title>Sora 2, Seedream 5.0 Pro, and the Shift from AI Video Demos to Production Workflows</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Fri, 17 Jul 2026 09:05:03 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/sora-2-seedream-50-pro-and-the-shift-from-ai-video-demos-to-production-workflows-k51</link>
      <guid>https://dev.to/marcoreyes1/sora-2-seedream-50-pro-and-the-shift-from-ai-video-demos-to-production-workflows-k51</guid>
      <description>&lt;p&gt;I have stopped judging AI video tools by the first clip they generate.&lt;/p&gt;

&lt;p&gt;That sounds a little unfair, because the first clip is usually what gets shared: a clean camera move, a dancer with believable motion, a product floating through a glossy studio scene. It looks great in a feed. But if you have ever had to turn generated visuals into an actual edit, ad concept, product page, or client review deck, you know the harder question comes later.&lt;/p&gt;

&lt;p&gt;Can I revise it without starting over?&lt;/p&gt;

&lt;p&gt;That is where the current wave of multimodal generation updates gets interesting.&lt;/p&gt;

&lt;p&gt;Sora 2, Seedream 5.0 Pro, MiniMax Hailuo, and Google DeepMind’s recent vision work are not all doing the same thing. They are not even all video models. But together they point toward the same shift: AI media tools are moving from “generate something impressive” toward “hold enough control to be useful in production.”&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI’s Sora 2 is the obvious headline name.&lt;/strong&gt; According to OpenAI’s own release post, Sora 2 was introduced as a video and audio generation model with better physical accuracy, realism, controllability, synchronized dialogue, and sound effects. The physics part is not just marketing language. OpenAI specifically describes the model as better at handling failure states: a missed basketball shot rebounds instead of magically correcting itself into a make.&lt;/p&gt;

&lt;p&gt;That is the sort of thing editors notice.&lt;/p&gt;

&lt;p&gt;In a music video or product spot, a model that always “completes” the prompt in the prettiest way can be less useful than one that respects cause and effect. If a skater lands badly, if a liquid spills, if a camera move misses the subject for half a second, the clip may still feel more real than a perfect but slippery hallucination.&lt;/p&gt;

&lt;p&gt;But there is an availability caveat here. OpenAI’s Sora 2 page now notes that the Sora product is no longer available as of April 26, 2026, and OpenAI’s help center says the Sora API has its own discontinuation timeline. So I would be careful with any “Sora 2 review” that sounds like a fresh hands-on recommendation for developers today. The right way to talk about it is as a benchmark moment in video/audio model design, not necessarily as a tool you can just build around this week.&lt;/p&gt;

&lt;p&gt;The “real-time video” story is also not something I would attach to Sora 2 unless OpenAI says it directly. Krea, for example, has its own Realtime Video work, where it describes generating faster than playback and responding to painting, text prompts, webcam input, or screen streams. That is a different product path from Sora 2, and mixing those claims together makes the whole comparison muddy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDkzNWE4ZWI2MzY1MWIxMDA0YWRhMGY0ZTAwZjg0MDhfMzIybTVFN0NNckNybVlOVzBXS2w2VHlJMm4yTVhOSDRfVG9rZW46WHdoN2JIcDBtb2lpVlN4SW1MWGNDaGdFbnZoXzE3ODQyNzg1Nzc6MTc4NDI4MjE3N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDkzNWE4ZWI2MzY1MWIxMDA0YWRhMGY0ZTAwZjg0MDhfMzIybTVFN0NNckNybVlOVzBXS2w2VHlJMm4yTVhOSDRfVG9rZW46WHdoN2JIcDBtb2lpVlN4SW1MWGNDaGdFbnZoXzE3ODQyNzg1Nzc6MTc4NDI4MjE3N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="screenshot from MiniMax Official Page" width="801" height="441"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seedream 5.0 Pro is a different case again.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Krea announced Seedream 5.0 Pro live on Krea on July 8, 2026, calling it ByteDance’s new flagship image generation and editing model. That is important: it is an image generation and editing model, not a text-to-video model. Still, I think it belongs in the same production-workflow conversation because many video projects start as stills: product frames, storyboard panels, look references, architecture renders, pitch decks, thumbnails, and campaign key visuals.&lt;/p&gt;

&lt;p&gt;Krea’s writeup says Seedream 5.0 Pro accepts text, a single image, or multiple images. It can generate single images or grouped outputs, supports streaming output, and is positioned around editing, blending, image sequences, and batch generation. Krea’s product photography examples focus on targeted changes: fixing glare, swapping materials, preserving bottle silhouettes, keeping labels stable, and varying approved product finishes without rebuilding the whole scene.&lt;/p&gt;

&lt;p&gt;That is exactly the sort of control that “AI video generation comparison” posts often skip.&lt;/p&gt;

&lt;p&gt;For product photography, the useful Seedream 5.0 Pro prompt is not “make this look premium.” It is closer to a production brief:&lt;/p&gt;

&lt;p&gt;Product truth: what must stay recognizable.&lt;/p&gt;

&lt;p&gt;Scene: studio, shelf, counter, pedestal, room.&lt;/p&gt;

&lt;p&gt;Lighting: softbox, rim light, window light, hard reflection.&lt;/p&gt;

&lt;p&gt;Camera: crop, lens feel, angle, distance.&lt;/p&gt;

&lt;p&gt;Protected details: label, geometry, logo, material boundaries.&lt;/p&gt;

&lt;p&gt;Revision target: what is allowed to change.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Create a studio product photograph of an amber glass serum bottle on pale stone. Keep the bottle silhouette, label placement, cap shape, and readable logo unchanged. Use soft overhead diffusion with a narrow warm rim light from camera right. Eye-level 70mm product crop, clean negative space above. Change only the reflection strength on the glass so it feels premium but not blown out.&lt;/p&gt;

&lt;p&gt;That is less poetic than most prompt guides. It is also more reviewable.&lt;/p&gt;

&lt;p&gt;The same logic applies to AI architecture visualization. Krea’s architecture guide for Seedream 5.0 Pro leans on references, anchors, and layered edits so a user can preserve massing, materials, camera angle, planting, or daylight while changing one variable at a time. For architecture, that matters more than raw beauty. A render that changes the building every time you adjust the sky is not a design tool. It is a slot machine with nice lighting.&lt;/p&gt;

&lt;p&gt;A better architecture prompt has to separate the fixed design from the variable being tested:&lt;/p&gt;

&lt;p&gt;Render a mass-timber community hall in a regional landscape. Preserve the structural span rhythm, roof pitch, main entrance position, and glazing proportions. Use exposed CLT and glulam, warm interior light, overcast daylight outside, and a wide architectural photograph viewpoint. Generate a revision that changes only the surrounding season from early summer to late autumn while keeping the building geometry unchanged.&lt;/p&gt;

&lt;p&gt;That is the move from prompt art to prompt-controlled workflow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNTJmYTkwYzgzMmJkMTUyYjI0MjdkMDgzZmQ5MmVlYmFfVEplVzFKajdxRmp1U1FOYlpJbmJsbHBrVmZSYmdXOUhfVG9rZW46SnNOUWJKZUU1b3pTeFd4WktYYmNuc2tkbkRnXzE3ODQyNzg1Nzc6MTc4NDI4MjE3N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNTJmYTkwYzgzMmJkMTUyYjI0MjdkMDgzZmQ5MmVlYmFfVEplVzFKajdxRmp1U1FOYlpJbmJsbHBrVmZSYmdXOUhfVG9rZW46SnNOUWJKZUU1b3pTeFd4WktYYmNuc2tkbkRnXzE3ODQyNzg1Nzc6MTc4NDI4MjE3N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from Dreamina Offical Page" width="760" height="415"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;MiniMax is useful to mention here, but I would not call it “MiniMax Code 2.0” unless there is a reliable source for that exact name in this video context. The stronger source is MiniMax Hailuo 02. MiniMax’s own announcement describes Hailuo 02 as a video generation model with native 1080p, stronger instruction following, and “extreme physics” performance. More interestingly, MiniMax explains an architecture called Noise-aware Compute Redistribution, or NCR, claiming 2.5x training and inference efficiency at comparable parameter scale, 3x the parameters of its predecessor, and 4x the training data.&lt;/p&gt;

&lt;p&gt;That is the kind of architecture detail I trust more than a vague leaderboard screenshot.&lt;/p&gt;

&lt;p&gt;In practice, the point is not that Hailuo, Sora, Seedance, Pika, or Veo “wins.” The point is that video generation is now being judged on different axes: motion coherence, physical plausibility, prompt adherence, editability, cost per usable second, and whether the output can survive a revision loop.&lt;/p&gt;

&lt;p&gt;Google DeepMind’s work pushes the discussion even further. Its 2026 paper “Image Generators are Generalist Vision Learners,” with Kaiming He listed among the authors, argues that generative image training can produce strong general visual representations. The paper reframes vision tasks as image generation tasks and presents Vision Banana as evidence that generation can become a general interface for vision.&lt;/p&gt;

&lt;p&gt;That is not the same as saying “video generators are now AGI.” Please do not write that headline.&lt;/p&gt;

&lt;p&gt;But it does suggest why video and image generation models are converging with broader visual intelligence. A model that learns to generate coherent scenes, respect geometry, preserve identity, and simulate motion is also learning a lot about the structure of the visual world.&lt;/p&gt;

&lt;p&gt;For developers and creative technologists, my takeaway is simple: stop asking only which model makes the prettiest clip.&lt;/p&gt;

&lt;p&gt;Ask these instead:&lt;/p&gt;

&lt;p&gt;Can I lock the subject across revisions?&lt;/p&gt;

&lt;p&gt;Can I preserve composition while changing one material?&lt;/p&gt;

&lt;p&gt;Can I control camera movement without fighting random motion?&lt;/p&gt;

&lt;p&gt;Can I generate reference stills, storyboard frames, and final clips in one pipeline?&lt;/p&gt;

&lt;p&gt;Can I export something that a designer, editor, or client can actually review?&lt;/p&gt;

&lt;p&gt;Can I reproduce the result next week?&lt;/p&gt;

&lt;p&gt;That is where the professional market is going. Not just text-to-video. Not just image-to-video. Not just one more impressive demo.&lt;/p&gt;

&lt;p&gt;The useful stack will look more like this:&lt;/p&gt;

&lt;p&gt;Seedream-style image and edit models for controlled key visuals, product shots, and architecture concepts.&lt;/p&gt;

&lt;p&gt;Video models like Sora 2, Veo, Hailuo, Seedance, Pika, or Kling for motion tests and short generated sequences.&lt;/p&gt;

&lt;p&gt;Realtime tools for fast ideation and live visual exploration.&lt;/p&gt;

&lt;p&gt;Traditional editing, compositing, grading, and review systems for finishing.&lt;/p&gt;

&lt;p&gt;The winner will not always be the most cinematic model. It may be the one that lets you keep the bottle label intact, preserve the building massing, repeat the camera move, and get a second version without throwing away the first one.&lt;/p&gt;

&lt;p&gt;That sounds less magical.&lt;/p&gt;

&lt;p&gt;For production work, it is much more useful.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>productivity</category>
      <category>reviews</category>
    </item>
    <item>
      <title>I do not think the next useful AI coding tool is necessarily another coding agent.</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Fri, 10 Jul 2026 03:30:00 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/i-do-not-think-the-next-useful-ai-coding-tool-is-necessarily-another-coding-agent-2hj6</link>
      <guid>https://dev.to/marcoreyes1/i-do-not-think-the-next-useful-ai-coding-tool-is-necessarily-another-coding-agent-2hj6</guid>
      <description>&lt;p&gt;At least for me, the bottleneck has moved somewhere else. I can open Claude Code in one terminal, Codex in another, Gemini CLI somewhere else, and maybe a GitHub Copilot agent session in the browser. The problem is not that I lack agents. The problem is that every agent brings its own queue, workspace, logs, review flow, and mental overhead.&lt;/p&gt;

&lt;p&gt;That is why the “Agent OS” idea is starting to make sense, even if I would not use the term too literally. I do not mean a real operating system. I mean a control layer for planning work, assigning it to different coding agents, reviewing the results, and deciding what actually gets merged.&lt;/p&gt;

&lt;p&gt;One concrete open-source example is Vibe Kanban, from BloopAI. Its GitHub README describes it as a way to get more out of Claude Code, Gemini CLI, Codex, Amp, and other coding agents. The same README says it can switch between more than ten coding agents, including Claude Code, Codex, Gemini CLI, GitHub Copilot, Amp, Cursor, OpenCode, Droid, Claude Code Router, and Qwen Code.&lt;/p&gt;

&lt;p&gt;The official docs have a “Supported Coding Agents” page listing those integrations and noting that each agent still requires its own installation and authentication.&lt;/p&gt;

&lt;p&gt;That last detail matters.&lt;/p&gt;

&lt;p&gt;Vibe Kanban is not a magic way to make paid agents free. If your Claude Code, Codex, or Gemini CLI setup requires an account, subscription, API key, or local authentication, you still need that. The useful part is different: the repository documents a local workbench flow through npx vibe-kanban, so the management layer itself does not have to become another subscription surface.&lt;/p&gt;

&lt;p&gt;More precisely: no additional Vibe Kanban subscription does not mean no underlying agent cost.&lt;/p&gt;

&lt;p&gt;There is also an important status update. On April 10, 2026, Bloop announced the shutdown of the company behind Vibe Kanban. The announcement said the project would continue as open source and community maintained. It also said remote services would remain available for 30 days, after which Vibe Kanban would move toward a fully local architecture. The announcement specifically listed remote kanban issues, comments, projects, and organisations among the services being removed, while saying local workspaces would continue to function.&lt;/p&gt;

&lt;p&gt;That makes Vibe Kanban a slightly strange product story, but a useful architecture story.&lt;/p&gt;

&lt;p&gt;The architecture is what I care about here. Vibe Kanban is not trying to be the smartest agent in the room. It is trying to be the room.&lt;/p&gt;

&lt;p&gt;The basic workflow is familiar if you already work with tickets. You plan work as kanban issues. When you are ready, you create a workspace where a coding agent can execute. According to the README, each workspace gives an agent a branch, a terminal, and a dev server. You can review diffs, leave inline comments, preview the app in a built-in browser, and then create a pull request.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDM4YTYyYTNkNjc4YTJkODRmMWVjNTQ1NDI4MmUwMzRfcDJ1UXhEVVRQc1RYQ0p1d0VQWFhqS0ZIclM3YnJLcThfVG9rZW46UWJramJqZHhob21oTUt4a2dTU2NrdVl1bk1lXzE3ODM2NTM1MjM6MTc4MzY1NzEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDM4YTYyYTNkNjc4YTJkODRmMWVjNTQ1NDI4MmUwMzRfcDJ1UXhEVVRQc1RYQ0p1d0VQWFhqS0ZIclM3YnJLcThfVG9rZW46UWJramJqZHhob21oTUt4a2dTU2NrdVl1bk1lXzE3ODM2NTM1MjM6MTc4MzY1NzEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="760" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That sounds boring in the best possible way.&lt;/p&gt;

&lt;p&gt;A lot of multi-agent demos focus on agents talking to agents. One agent plans, another codes, another reviews, another writes documentation, and then everyone pretends the result is a software team. Maybe that will become useful. For day-to-day engineering, I am more interested in a human-centered version of Multi-Agent Orchestration: let me send a narrow task to the agent that is best suited for it, keep the workspace isolated, review the diff, and decide whether it deserves to move forward.&lt;/p&gt;

&lt;p&gt;That is where an Agent OS-style workbench becomes practical.&lt;/p&gt;

&lt;p&gt;For me, a useful agent management platform needs five things.&lt;/p&gt;

&lt;p&gt;First, it needs task isolation. If I ask Codex to try one approach and Claude Code to try another, I do not want both agents fighting over the same working tree. A branch-per-workspace model is easier to reason about. It also makes cleanup less painful when one attempt goes nowhere.&lt;/p&gt;

&lt;p&gt;Second, it needs visible logs. Coding agents are good at producing a finished diff that looks plausible. That is not enough. I want to see what commands ran, what tests failed, which files changed, and where the agent got stuck. If the management layer hides the process, it becomes just another trust-me interface.&lt;/p&gt;

&lt;p&gt;Third, it needs review as a first-class step. The best part of the Vibe Kanban model is not that it launches agents. Lots of tools can launch agents. The useful part is that it keeps planning, execution, diff review, preview, and pull request creation in one loop. That is closer to how real code gets shipped.&lt;/p&gt;

&lt;p&gt;Fourth, it needs agent choice without workflow lock-in. Claude Code may be better for one kind of repo. Codex may be better for another. Gemini CLI or OpenCode may be enough for a smaller task. A management layer should make those choices easy without forcing every project into one vendor’s interface.&lt;/p&gt;

&lt;p&gt;Fifth, it needs honest boundaries. “No extra subscription” should mean the orchestration product is not adding another SaaS bill. It should not imply that the underlying model usage disappears. The clean model is: bring your own agent auth, use a local or self-hosted workbench where possible, and keep the review process under your control.&lt;/p&gt;

&lt;p&gt;This is also where I would be careful with the phrase “Agent OS.” It is a good search term, and honestly, it captures the mood of the moment. But the useful version is not a grand AI operating system that owns your whole dev environment. It is closer to a process operating layer: task queues, workspaces, permissions, logs, previews, review gates, and Git operations.&lt;/p&gt;

&lt;p&gt;That may sound less exciting. It is also more likely to survive contact with real teams.&lt;/p&gt;

&lt;p&gt;The security side is not optional either. If one UI can launch multiple agents across repositories, it also becomes a place where permissions and blast radius need to be thought through. I would want separate agent profiles, clear repository boundaries, explicit MCP server connections, and a habit of reviewing generated diffs before giving anything merge rights.&lt;/p&gt;

&lt;p&gt;A unified agent platform should make supervision easier, not make automation invisible.&lt;/p&gt;

&lt;p&gt;The bigger trend is clear enough: we are moving from single-agent usage to agent operations. The work is less about asking one assistant for code and more about deciding which agent gets which task, where it runs, how much context it sees, and who approves the output.&lt;/p&gt;

&lt;p&gt;That is why open-source tools like Vibe Kanban are worth watching, even with the messy company shutdown context. They point toward a future where the winning layer may not be the model or the coding agent itself. It may be the workbench that helps developers manage several agents without losing control of the repo.&lt;/p&gt;

&lt;p&gt;I would not call that a full operating system yet.&lt;/p&gt;

&lt;p&gt;But as an Agent OS-style workbench for Codex, Claude Code, Gemini CLI, and the rest of the fast-growing coding agent stack, the idea is practical enough to take seriously.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>GitLost Explained: When a GitHub Issue Becomes an AI Agent Security Boundary</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Thu, 09 Jul 2026 02:00:00 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/gitlost-explained-when-a-github-issue-becomes-an-ai-agent-security-boundary-4iae</link>
      <guid>https://dev.to/marcoreyes1/gitlost-explained-when-a-github-issue-becomes-an-ai-agent-security-boundary-4iae</guid>
      <description>&lt;p&gt;I do not usually get dramatic about GitHub Issues.&lt;/p&gt;

&lt;p&gt;In most teams, an issue is boring infrastructure: bug reports, meeting follow-ups, half-shaped feature requests, and the occasional “can someone check this?” note that sits there longer than anyone wants to admit.&lt;/p&gt;

&lt;p&gt;GitLost is uncomfortable because it turns that boring surface into something much more interesting: an instruction channel for an AI agent.&lt;/p&gt;

&lt;p&gt;On July 6, 2026, Noma Security published a report called “GitLost: How We Tricked GitHub’s AI Agent into Leaking Private Repos.” According to Noma’s PoC, the vulnerable setup involved GitHub Agentic Workflows, where an agent could read a public GitHub Issue, use configured tools, access repository content, and then post a response back to the issue.&lt;/p&gt;

&lt;p&gt;The attacker’s entry point, in the reported case, was not a stolen token or a compromised maintainer account. It was a crafted issue in a public repository that belonged to the same organization as the private repository being targeted.&lt;/p&gt;

&lt;p&gt;That is the part worth sitting with for a minute.&lt;/p&gt;

&lt;p&gt;This was not a classic “someone leaked a secret in CI” story. It was closer to a workflow design failure: untrusted text entered through a public collaboration feature, and an agent with broader access treated that text as something it could act on.&lt;/p&gt;

&lt;p&gt;GitHub’s own Agentic Workflows documentation helps explain why the boundary is subtle. A workflow is written as a Markdown file with frontmatter for things like triggers, permissions, safe outputs, and the AI engine. The Markdown body contains natural-language instructions the agent follows when the workflow runs. As of the documentation checked on July 8, 2026, GitHub lists multiple supported engines, including GitHub Copilot, Anthropic Claude, OpenAI Codex, and Google Gemini.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDEwYzgwNzNlYWE2YWIzNDhkYmIyMmYxNDRiOWE2ZjBfQTN2aGdvYUxydURsbVVyOWNHT2IxRk1oek9Db2dDSTNfVG9rZW46SEJNdmJvMjNYbzUzQ2N4UnJnbmNnbWtZblliXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDEwYzgwNzNlYWE2YWIzNDhkYmIyMmYxNDRiOWE2ZjBfQTN2aGdvYUxydURsbVVyOWNHT2IxRk1oek9Db2dDSTNfVG9rZW46SEJNdmJvMjNYbzUzQ2N4UnJnbmNnbWtZblliXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image about how GitHub sets the boundary" width="759" height="503"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That design is useful. It is also exactly why the line between “workflow instruction” and “issue content being inspected” has to be very clear.&lt;/p&gt;

&lt;p&gt;In the GitLost PoC described by Noma, the workflow was triggered when an issue was assigned. The workflow read the issue title and body, and the agent was allowed to add a comment. Noma says the tested workflow also had read access to other repositories in the organization. The crafted issue asked about the README in the current public repo and then asked about the same file in another repo. The PoC issue referenced by Noma is sasinomalabs/poc issue #153.&lt;/p&gt;

&lt;p&gt;I am intentionally not turning that into a copy-paste attack recipe. The defensive lesson is already clear enough: if an agent can read private repository content and can write into a public issue comment, then a public issue can become both the input path and a potential exfiltration path.&lt;/p&gt;

&lt;p&gt;This is why I do not think GitLost is only a “prompt injection” story. That label is accurate, but it can make the problem sound smaller than it is.&lt;/p&gt;

&lt;p&gt;The agent did not need to break into GitHub in the usual sense. In the reported configuration, it followed instructions inside context it was already allowed to read, then used an output channel it was already allowed to use.&lt;/p&gt;

&lt;p&gt;That is a different review problem from the one many teams are used to.&lt;/p&gt;

&lt;p&gt;For traditional automation, we usually ask:&lt;/p&gt;

&lt;p&gt;What event triggers this workflow?&lt;/p&gt;

&lt;p&gt;What token does it run with?&lt;/p&gt;

&lt;p&gt;Which APIs can it call?&lt;/p&gt;

&lt;p&gt;Can it write to the repo?&lt;/p&gt;

&lt;p&gt;Can it publish comments, artifacts, or pull requests?&lt;/p&gt;

&lt;p&gt;Those questions still matter. But agentic workflows add another one:&lt;/p&gt;

&lt;p&gt;Whose text is allowed to become an instruction?&lt;/p&gt;

&lt;p&gt;A GitHub Issue body is user-controlled content. So is a pull request description. So is a comment thread. So is a Markdown file from a repository the agent is asked to inspect.&lt;/p&gt;

&lt;p&gt;In older automation, those were usually strings passed into scripts. In agentic automation, those strings may be summarized, prioritized, interpreted, mixed with system instructions, or accidentally followed.&lt;/p&gt;

&lt;p&gt;That is the shift.&lt;/p&gt;

&lt;p&gt;The context window is not just memory. It is part of the attack surface.&lt;/p&gt;

&lt;p&gt;If I were reviewing an agentic GitHub workflow tomorrow, I would start with the access map.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNTY1NWRiYmYzMGUyNjM4YTFlMGQyMDkxNjM5MTRmNDlfQ3VrTGM3MFdYcWVvU2pWd084U0RTdGZpTkN2V2tYUVFfVG9rZW46SllwT2JqUU55b0tpZ2d4U3R6VmNWM3ZJblVjXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNTY1NWRiYmYzMGUyNjM4YTFlMGQyMDkxNjM5MTRmNDlfQ3VrTGM3MFdYcWVvU2pWd084U0RTdGZpTkN2V2tYUVFfVG9rZW46SllwT2JqUU55b0tpZ2d4U3R6VmNWM3ZJblVjXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image about the access map" width="760" height="380"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;First, I would list every trigger. From a security review perspective, an issue assignment trigger is very different from a manual dispatch by a maintainer. Anything triggered by public or semi-public user input deserves extra attention.&lt;/p&gt;

&lt;p&gt;Second, I would list every repository the agent can read. GitHub documentation describes permission controls and safe output mechanisms, but “read” should not be treated as harmless by default. Read access can create a potential leakage path when it is combined with broad repository scope and a public output channel.&lt;/p&gt;

&lt;p&gt;Third, I would review safe outputs as data-leak controls, not just write-operation controls. Allowing an agent to add a comment may feel safer than allowing it to push code. In a case like GitLost, a public comment is exactly where sensitive content can escape.&lt;/p&gt;

&lt;p&gt;Fourth, I would separate trusted instructions from untrusted content in the workflow design. The agent should treat issue text as data to inspect, not authority to obey. That sounds obvious until you see prompts that effectively say, “read the issue and do what it asks.”&lt;/p&gt;

&lt;p&gt;Fifth, I would add a human checkpoint before any agent output that includes repository content, cross-repository summaries, private project names, or anything that looks even slightly secrets-adjacent.&lt;/p&gt;

&lt;p&gt;Some practical controls are boring, which is usually a good sign:&lt;/p&gt;

&lt;p&gt;Use the narrowest permissions you can.&lt;/p&gt;

&lt;p&gt;Avoid giving issue-triggered workflows access to unrelated private repositories.&lt;/p&gt;

&lt;p&gt;Use repository allowlists for agent tools.&lt;/p&gt;

&lt;p&gt;Treat public issues, PR descriptions, comments, and artifacts as possible exfiltration sinks.&lt;/p&gt;

&lt;p&gt;Log cross-repository reads by agent workflows.&lt;/p&gt;

&lt;p&gt;Add prompt-injection test cases to workflow review.&lt;/p&gt;

&lt;p&gt;Use harmless canary files in private repos to test whether an agent repeats content into public channels.&lt;/p&gt;

&lt;p&gt;Review imported or third-party agentic workflows before enabling them.&lt;/p&gt;

&lt;p&gt;None of this requires panic. It does require teams to stop treating “the agent only has read access” as a satisfying answer.&lt;/p&gt;

&lt;p&gt;Read access plus public output can still create a leak, depending on the workflow configuration.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDI1MTNiNDhkZTQzOGE3OWQ4ZmFkN2VkN2Q0MWYzZGRfQzkyMmxEQW5Zcno2R09Pb2pBTlRBdFhnNmxUZ3pzSUZfVG9rZW46R2Nrc2IyOExMb2Q0VnJ4RDNCMmM0QmhVblJlXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZDI1MTNiNDhkZTQzOGE3OWQ4ZmFkN2VkN2Q0MWYzZGRfQzkyMmxEQW5Zcno2R09Pb2pBTlRBdFhnNmxUZ3pzSUZfVG9rZW46R2Nrc2IyOExMb2Q0VnJ4RDNCMmM0QmhVblJlXzE3ODM1MDU4NjU6MTc4MzUwOTQ2NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image about the information leak" width="760" height="467"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The useful thing about GitLost is that it makes the failure visible. A lot of AI agent security discussion still floats around abstract words: jailbreaks, autonomy risk, hidden prompts. This case is easier to reason about. A public issue contained untrusted instructions. An agent had access the issue author should not have had. The output channel was public. According to Noma, private README content was exposed through that path.&lt;/p&gt;

&lt;p&gt;That is not a benchmark problem.&lt;/p&gt;

&lt;p&gt;It is a boundary problem.&lt;/p&gt;

&lt;p&gt;As more teams wire agents into GitHub, Jira, Slack, CI, docs, ticket queues, and internal search, this pattern will show up in different shapes. The question will not only be “can the model follow instructions?” It will be “whose instructions are allowed to matter?”&lt;/p&gt;

&lt;p&gt;For me, that is the practical takeaway from GitLost.&lt;/p&gt;

&lt;p&gt;Before giving an agent access to private repositories, decide what public text is allowed to ask it to do. Before letting it comment publicly, decide what private context it is allowed to repeat. And before calling a workflow safe because it cannot push code, look at where it can speak.&lt;/p&gt;

&lt;p&gt;An agent that can read privately and speak publicly may be crossing a trust boundary, depending on how you configure it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>github</category>
      <category>security</category>
    </item>
    <item>
      <title>AI Video Tools Are Getting Better. The Real Shift Is Coding Agents Editing the Timeline</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Mon, 06 Jul 2026 07:00:00 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/ai-video-tools-are-getting-better-the-real-shift-is-coding-agents-editing-the-timeline-25k7</link>
      <guid>https://dev.to/marcoreyes1/ai-video-tools-are-getting-better-the-real-shift-is-coding-agents-editing-the-timeline-25k7</guid>
      <description>&lt;p&gt;I spent part of last weekend doing the thing I always pretend I will automate someday: cleaning up a rough video cut.&lt;/p&gt;

&lt;p&gt;There were three interview takes, two b-roll folders, one music bed that was too loud in the second half, and a bunch of tiny edit notes sitting in a text file. None of this was glamorous. It was the normal middle layer of production: rename files, find the usable lines, trim dead space, add captions, check levels, render, watch it back, fix the weird cut, render again.&lt;/p&gt;

&lt;p&gt;That is why the current AI video tool conversation feels slightly misframed to me.&lt;/p&gt;

&lt;p&gt;Everyone wants to compare Seedance, Sora, Kling, Runway, Veo, and whatever model is trending this week by asking: which one makes the prettiest clip?&lt;/p&gt;

&lt;p&gt;That is a fair question. But it is not the only question anymore.&lt;/p&gt;

&lt;p&gt;The more interesting question for developers is: can an AI system operate the production workflow?&lt;/p&gt;

&lt;p&gt;Because generating a ten-second shot is not the same as finishing a video.&lt;/p&gt;

&lt;p&gt;Quick caveat before getting into Seedance 2.5: I’ve seen the phrase “Seedance 2.5” showing up in search demand and community discussion, but I could not confirm a public official Seedance 2.5 product page from ByteDance at the time of writing. The official Seedance page I checked points to Seedance 2.0 materials.&lt;/p&gt;

&lt;p&gt;So I would treat “Seedance 2.5 review” as a keyword that needs careful handling. If you are writing a buyer’s guide or a comparison table, don’t publish exact claims about Seedance 2.5 pricing, specs, model limits, or commercial rights unless you can verify them from an official product page or a current platform listing.&lt;/p&gt;

&lt;p&gt;That said, the Seedance line itself is still worth watching. According to ByteDance’s public Seedance materials, the product direction is about more controllable video generation: prompts, references, camera motion, visual consistency, and cinematic output. That is the direction AI video generation seems to be moving. Less “make a cool clip” and more “follow this reference, keep this character, use this camera language, and give me motion I can actually cut with.”&lt;/p&gt;

&lt;p&gt;For social clips and concept shots, that matters. If I am making a music video treatment, a mood trailer, or a pitch deck for a small artist, I do not need a perfect final render on the first try. I need to explore visual direction fast enough that the idea survives the meeting.&lt;/p&gt;

&lt;p&gt;That is where Seedance-style tools are useful.&lt;/p&gt;

&lt;p&gt;Sora 2 is the other side of the conversation, mostly because it made the economics impossible to ignore. According to OpenAI’s Sora 2 announcement, Sora 2 was introduced as a video-and-audio generation model designed to handle synchronized dialogue, sound effects, and more physically coherent scenes than the earlier Sora release. OpenAI also launched it alongside a social app experience.&lt;/p&gt;

&lt;p&gt;The part I care about as a developer is not just the model quality. It is the cost shape.&lt;/p&gt;

&lt;p&gt;OpenAI said in its original Sora 2 announcement that the app would initially be free with generous limits, and that its monetization plan at the time was eventually to let users pay for extra generation if compute demand required it. That is a very honest AI video sentence. Video generation is not cheap. Someone always pays for the retry loop.&lt;/p&gt;

&lt;p&gt;OpenAI has also posted a Sora discontinuation notice, saying the Sora web and app experiences were discontinued and that the Sora API has a scheduled discontinuation timeline. I would re-check that page before publishing any current “Sora 2 pricing” guide, because this is exactly the kind of product status that can change or become stale fast.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzMxMDVjMGMwNjIwMWNhMjA5MTNlMTdhYjkzMWRmZWNfdHhia0ljSTFXSUJ3S05DWHFlNEJqYThYRXdIRVV2SGdfVG9rZW46V1Nlb2JSY2htb2l2YnJ4WEpPQ2NObWw2blJlXzE3ODMzMTk4ODM6MTc4MzMyMzQ4M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzMxMDVjMGMwNjIwMWNhMjA5MTNlMTdhYjkzMWRmZWNfdHhia0ljSTFXSUJ3S05DWHFlNEJqYThYRXdIRVV2SGdfVG9rZW46V1Nlb2JSY2htb2l2YnJ4WEpPQ2NObWw2blJlXzE3ODMzMTk4ODM6MTc4MzMyMzQ4M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So I would not write about Sora 2 today as if it were simply another stable tool in a shopping list. I would write about it as a useful case study: video generation is expensive, hard to scale, legally sensitive, and difficult to align with user expectations.&lt;/p&gt;

&lt;p&gt;This is why I’m more interested in the open-source workflow projects showing up around video.&lt;/p&gt;

&lt;p&gt;OpenMontage is the one that caught my eye first. According to its GitHub README, OpenMontage describes itself as an open-source agentic video production system. The project’s pitch is to use a coding assistant as part of a video production workflow: research, scripting, asset sourcing, media generation, scene composition, music, captions, and final rendering.&lt;/p&gt;

&lt;p&gt;I would not treat every README capability as production-proven just because it appears in the repo. But the framing is important. OpenMontage is not only trying to generate clips. It is trying to coordinate the pipeline.&lt;/p&gt;

&lt;p&gt;That feels much closer to how production actually works.&lt;/p&gt;

&lt;p&gt;A finished video is a stack of decisions. What is the hook? What is the script? What footage is available? What has to be generated? What music is licensed? Where do captions go? Does the render have audio pops? Are the subtitles inside the safe area? Did the final file actually export correctly?&lt;/p&gt;

&lt;p&gt;A video model can help with some of that. A coding agent can potentially coordinate more of it.&lt;/p&gt;

&lt;p&gt;OpenMontage also pushes the cost discussion in a useful direction. The repo includes examples and workflow notes that try to make the tool path visible. I would treat any listed costs as examples, not promises, but I like the habit. AI video tools often hide the painful part behind credits, queues, or “try again” loops. A developer workflow should make cost visible before generation starts.&lt;/p&gt;

&lt;p&gt;Then there is browser-use/video-use, which is narrower but maybe more immediately practical.&lt;/p&gt;

&lt;p&gt;According to the repo, video-use experiments with using coding agents to automate parts of video editing. The project describes workflows around raw footage, transcripts, cuts, captions, overlays, audio handling, and rendered output review. I would describe it as experimental, not as a finished replacement for an editor.&lt;/p&gt;

&lt;p&gt;But if you have ever edited talking-head footage, tutorials, interviews, internal launch videos, or creator clips, you can see the appeal immediately. The boring work is the work.&lt;/p&gt;

&lt;p&gt;The design idea is smart too. Instead of asking an LLM to “watch” every frame, the system can use structured transcripts, timestamps, timeline representations, and visual checks when needed. That is the video equivalent of giving a web agent a DOM instead of only a screenshot. Less token waste, more useful structure.&lt;/p&gt;

&lt;p&gt;This is where I think the new AI video stack splits into two layers.&lt;/p&gt;

&lt;p&gt;Layer one is generation.&lt;/p&gt;

&lt;p&gt;That includes Seedance-style models, Kling, Veo, Runway, Higgsfield-style hosted workflows, and whatever comes next. These tools answer: can I create or transform visual material?&lt;/p&gt;

&lt;p&gt;Layer two is orchestration.&lt;/p&gt;

&lt;p&gt;That includes OpenMontage, video-use, Remotion-based renderers, FFmpeg scripts, caption pipelines, audio cleanup, metadata checks, source licensing, and agent review loops. This layer answers: can I turn material into a deliverable?&lt;/p&gt;

&lt;p&gt;For developers, the second layer may be the more durable opportunity.&lt;/p&gt;

&lt;p&gt;Video models will keep changing. Prices will move. Access will shift. Some tools will disappear. Some will get folded into bigger creative suites. But the workflow problems stay weirdly stable: ingest, plan, cut, compose, caption, mix, export, review.&lt;/p&gt;

&lt;p&gt;That is why I would build around the pipeline, not around one model.&lt;/p&gt;

&lt;p&gt;If I were setting up a small AI video production workflow today, I would start with three tracks.&lt;/p&gt;

&lt;p&gt;First, use hosted generators for shots you cannot source. Seedance-style tools may be useful for fast visual exploration, especially when you need camera motion or stylized scene generation. But I would keep every generated clip replaceable.&lt;/p&gt;

&lt;p&gt;Second, use open-source agent workflows for structure. OpenMontage is interesting when you want to experiment with a full concept-to-render pipeline. video-use is interesting when you already have footage and want an agent to help with the editing layer.&lt;/p&gt;

&lt;p&gt;Third, keep old production tools in the loop. FFmpeg, Remotion, subtitles, waveform checks, file naming, render validation. Boring tools are still the backbone. The agent should drive them, not replace them with vibes.&lt;/p&gt;

&lt;p&gt;Prompting also changes in this world.&lt;/p&gt;

&lt;p&gt;For a pure video generator, the prompt is visual direction: camera, subject, motion, lighting, duration, style, continuity.&lt;/p&gt;

&lt;p&gt;For an agentic video workflow, the prompt is production direction: audience, runtime, source assets, pacing, deliverable format, approval points, cost ceiling, licensing constraints, and what must not be generated.&lt;/p&gt;

&lt;p&gt;That second prompt is less glamorous, but it gets you closer to a real finished video.&lt;/p&gt;

&lt;p&gt;My current take is this: Seedance 2.5 might be the keyword people are chasing this week, but I would not write about it as a confirmed official product version unless there is a current official source to cite. Sora 2 is useful as a case study in AI video economics, but I would re-check OpenAI’s current product status before treating it as an active tool recommendation.&lt;/p&gt;

&lt;p&gt;The bigger shift is not one model beating another model.&lt;/p&gt;

&lt;p&gt;The bigger shift is that video production is becoming programmable.&lt;/p&gt;

&lt;p&gt;Not easy. Not fully automatic. Not something I would trust unsupervised on client work yet.&lt;/p&gt;

&lt;p&gt;But programmable.&lt;/p&gt;

&lt;p&gt;And once video becomes something a coding agent can inspect, plan, edit, render, and self-check, the category starts looking less like “AI magic” and more like software.&lt;/p&gt;

&lt;p&gt;That is the part I care about.&lt;/p&gt;

&lt;p&gt;Because I do not need a model to make one impressive clip. I need a system that can take messy inputs, make reasonable decisions, ask before it commits to expensive steps, and give me a file I can actually send.&lt;/p&gt;

&lt;p&gt;That is a much less viral promise.&lt;/p&gt;

&lt;p&gt;It is also a much more useful one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Sonnet 5 vs Fable 5: What I’d Actually Use in Production</title>
      <dc:creator>Marco Reyes</dc:creator>
      <pubDate>Fri, 03 Jul 2026 06:52:05 +0000</pubDate>
      <link>https://dev.to/marcoreyes1/claude-sonnet-5-vs-fable-5-what-id-actually-use-in-production-47mc</link>
      <guid>https://dev.to/marcoreyes1/claude-sonnet-5-vs-fable-5-what-id-actually-use-in-production-47mc</guid>
      <description>&lt;p&gt;I usually test new AI models in a very unglamorous way.&lt;/p&gt;

&lt;p&gt;Not with benchmark prompts. Not with “build me a company in one message.” I throw them at the boring parts of my week: cleaning up a script that renames audio stems, fixing a small FFmpeg wrapper, turning production notes into a tiny dashboard, or explaining why a Node tool works on my laptop but breaks on the studio machine.&lt;/p&gt;

&lt;p&gt;That is the lens I’m using for Claude Sonnet 5 and Fable 5.&lt;/p&gt;

&lt;p&gt;Anthropic announced Claude Sonnet 5 on June 30, 2026, positioning it as the most agentic Sonnet model so far. Around the same window, Fable 5 returned after a temporary restriction and redeployment process tied to safeguards. So the question is not just “which model is smarter?” The better question is: which one would I actually trust inside a working developer workflow?&lt;/p&gt;

&lt;p&gt;The short version: Sonnet 5 is the more interesting daily model.&lt;/p&gt;

&lt;p&gt;Not because it is the absolute strongest Claude model. It is not. Anthropic positions Claude Fable 5 as its highest-capability widely released model, and its own model overview describes Opus 4.8 as a better fit for complex agentic coding and enterprise work. But Sonnet 5 looks like the model many developers will actually use day to day: cheaper, fast enough, stronger than the previous Sonnet line, and designed for the kind of multi-step coding tasks that show up in real projects.&lt;/p&gt;

&lt;p&gt;That last part matters. A lot of models can produce a decent first answer. Fewer models can keep context across a messy repo, notice the test failure they caused, use tools properly, and avoid rewriting half the project just because they saw an opportunity to be “helpful.”&lt;/p&gt;

&lt;p&gt;For me, Sonnet 5 is most interesting in those middle-weight tasks: refactoring a helper module, writing tests, debugging a brittle script, generating small internal tools, or turning vague notes into something I can actually run.&lt;/p&gt;

&lt;p&gt;There are migration details worth paying attention to, though.&lt;/p&gt;

&lt;p&gt;The model name is simple: claude-sonnet-5. But if you are moving from Sonnet 4.6, I would not just swap the model string and call it done.&lt;/p&gt;

&lt;p&gt;According to Anthropic’s documentation, Sonnet 5 supports adaptive thinking and may vary reasoning effort depending on the task. That is usually good for agentic work, but it can also affect latency, token usage, and how predictable the response feels if your old integration was tuned around a previous Sonnet model.&lt;/p&gt;

&lt;p&gt;If your app still sends a fixed thinking budget, it may throw an error depending on API version and request parameters. I would check the current migration notes before touching production. The practical takeaway is simple: do a small compatibility pass before you migrate any real user workflow.&lt;/p&gt;

&lt;p&gt;Anthropic recommends using the newer effort-based approach where supported, rather than relying on older extended-thinking patterns. I would also check whether your app sets custom sampling parameters like temperature, top_p, or top_k. If those are no longer accepted in your request path, move style control into the system prompt instead of treating sampling as the main control surface.&lt;/p&gt;

&lt;p&gt;The quiet migration issue is tokenization. In its Sonnet 5 announcement and related docs, Anthropic notes that higher effort levels and token usage are part of the availability and pricing discussion. So if you have a big prompt template full of repo summaries, tool instructions, and guardrails, measure it again. A prompt that felt cheap and safe on Sonnet 4.6 may behave differently once you move it.&lt;/p&gt;

&lt;p&gt;So where does Fable 5 fit?&lt;/p&gt;

&lt;p&gt;Fable 5 is not really the “cheap lite version” people sometimes imply when discussing restricted models. I also do not think “crippled version” is the most useful mental model. Anthropic describes Fable 5 as a Mythos-class model made safe for general use. Mythos 5 uses the same underlying model family, but with some safeguards lifted for approved trusted-access users.&lt;/p&gt;

&lt;p&gt;That distinction matters. Fable 5 is powerful, but guarded.&lt;/p&gt;

&lt;p&gt;Anthropic has stated that some restricted requests may be routed to Opus 4.8. In the Fable 5 launch explanation, this applies to areas like cybersecurity, biology and chemistry, and distillation when classifiers detect requests that fall into those categories. Anthropic also says users are informed when that fallback happens.&lt;/p&gt;

&lt;p&gt;That does not make Fable useless. It just means “Fable 5 is back” does not mean every capability is available in every context.&lt;/p&gt;

&lt;p&gt;If I had to split usage, I would put it like this:&lt;/p&gt;

&lt;p&gt;Sonnet 5 is for normal developer work. The stuff you do repeatedly. Debugging, code review, tests, repo navigation, small agents, tool use, project scaffolding, and practical automation.&lt;/p&gt;

&lt;p&gt;Opus 4.8 is for harder engineering judgment. Anthropic positions Opus 4.8 as its higher-capability model for more demanding reasoning tasks, especially when the failure cost is higher or the model needs to reason across architecture and trade-offs.&lt;/p&gt;

&lt;p&gt;Fable 5 is for high-value work where deeper capability matters enough to justify the cost and the safeguards. Long planning sessions, complex analysis, creative technical exploration, or tasks where you really want the stronger model and can tolerate more constraints.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2RmOTY3MjkzYTY1MzQ3ZWY1NmVkNjcwNTNlOGRhOGVfZ2J5VDcwbkRPTG5xRUVRNGhsTnNJd29mczlYcm5wZ2RfVG9rZW46SXdWZ2JGVXFvb1lQRlZ4aGZMYWN0TjZ5bmhkXzE3ODMwNjI1Mzg6MTc4MzA2NjEzOF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2RmOTY3MjkzYTY1MzQ3ZWY1NmVkNjcwNTNlOGRhOGVfZ2J5VDcwbkRPTG5xRUVRNGhsTnNJd29mczlYcm5wZ2RfVG9rZW46SXdWZ2JGVXFvb1lQRlZ4aGZMYWN0TjZ5bmhkXzE3ODMwNjI1Mzg6MTc4MzA2NjEzOF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8 (a more generally capable model, for reference). The Claude Sonnet 5 System Card reports a broader set of evaluations in detail.&lt;br&gt;
" width="2600" height="1234"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Price is part of the story here. Sonnet 5 launched with introductory API pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026. After that, Anthropic lists standard pricing at $3 input and $15 output per million tokens. Fable 5 is much more expensive at $10 input and $50 output per million tokens.&lt;/p&gt;

&lt;p&gt;That gap is big enough to affect architecture. I would not point every small request at Fable unless the product economics are very forgiving.&lt;/p&gt;

&lt;p&gt;The better pattern is routing. Use Sonnet 5 by default, escalate to Opus or Fable only when the task deserves it, and log enough metadata to understand when escalation actually helped.&lt;/p&gt;

&lt;p&gt;AWS access is also worth separating clearly.&lt;/p&gt;

&lt;p&gt;There is Claude Platform on AWS, which gives access to Anthropic’s native platform through AWS billing, IAM, and CloudTrail. Then there is Claude in Amazon Bedrock, which is the managed AWS model service path. Anthropic’s model overview currently lists Claude Sonnet 5, Opus 4.8, and Fable 5 with Bedrock IDs, and notes that they are available through Claude in Amazon Bedrock using the Messages-API Bedrock endpoint.&lt;/p&gt;

&lt;p&gt;If your integration is old, check Anthropic and AWS documentation for current API support before assuming the model is unavailable. This is exactly the kind of boring platform detail that can waste an afternoon.&lt;/p&gt;

&lt;p&gt;The China access question is more sensitive, and I would be careful with it.&lt;/p&gt;

&lt;p&gt;When I checked Anthropic’s supported countries and regions page on July 3, 2026, Taiwan appeared in the list for API and Claude.ai access, but I did not find mainland China, Hong Kong, or Macau listed there. That page is the official source I would re-check before making any access decision.&lt;/p&gt;

&lt;p&gt;So if someone asks “Can Claude Sonnet 5 be used in China?” I would phrase the answer this way: not officially via standard Claude.ai or direct Anthropic commercial access, based on the current supported-regions list.&lt;/p&gt;

&lt;p&gt;I would not build anything serious around unofficial account workarounds. They may work for a while, but they are fragile, and they tend to fail at the worst time. If you are building for a team with China-based access needs, check official cloud-provider availability, account eligibility, compliance requirements, and regional constraints. Also keep your model layer abstract enough that you can switch providers if needed.&lt;/p&gt;

&lt;p&gt;My actual take after reading through the release notes is simple: Sonnet 5 is probably the model I would test first.&lt;/p&gt;

&lt;p&gt;It is not the flashiest option. It is not the model I would pick for every hard problem. But it sits in the right place for practical developer work: capable enough to handle real agentic tasks, cheap enough to iterate with, and available across the surfaces many teams already use.&lt;/p&gt;

&lt;p&gt;Fable 5 is more like the expensive session player you bring in when the track really needs it. Very useful, but not for every pass. Sonnet 5 is the working model you keep in the room all day.&lt;/p&gt;

&lt;p&gt;That may sound less dramatic than “this model changes everything,” but honestly, that is usually how useful tools arrive. Not as a lightning strike. More like something you start reaching for without thinking, because it saves you twenty minutes here, an hour there, and maybe one bad deploy on a tired Thursday night.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>productivity</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
