This week's tooling releases cluster around two themes: cheaper inference and gateway consolidation. Anthropic dropped Haiku 5.5 with a 75% price cut and tunable reasoning budgets, while Vercel's AI Gateway absorbed FLUX 3, the Decisions API, and OpenAI's Ultrafast tier into a single routing layer. If you're still managing direct vendor integrations, the case for consolidating behind a gateway just got harder to ignore.
FLUX 3 Image Launches on Vercel AI Gateway
Black Forest Labs' FLUX 3 is now routable through AI Gateway with a unified API that handles both generation and editing in a single call. You get five resolution tiers, support for up to ten reference images, and access via a single AI_GATEWAY_API_KEY—no separate BFL account required. Model ID is bfl/flux-3-image, and you'll need AI SDK 7+.
The editing mode is the more interesting addition. Previously, generation and editing were separate pipelines with separate integrations. Consolidating them into one call via a gateway you're already using removes a real maintenance surface. The vendor lock-in argument is secondary—the operational simplicity is the actual win.
Verdict: Ship. If you're calling Black Forest Labs directly today, this is a drop-in replacement with fewer moving parts. Editing mode is new but stable enough for production. The only prerequisite gate is AI SDK 7+—check that before you migrate.
Claude Haiku 5.5 Launches on Amazon Bedrock
Haiku 5.5 is available on Bedrock now at 75% less than Haiku 4.5, with full multi-step tool use and computer-use capabilities intact. Three invocation paths are supported: Boto3 InvokeModel, the Converse API, and the Anthropic SDK via bedrock-runtime. You'll need an AWS account with Bedrock access and bedrock:InvokeModel IAM permissions.
The price drop is significant enough to revisit workloads you previously offloaded to cheaper models or batched to reduce costs. Classification, document summarization, code review, and routing tasks that needed volume discounts to pencil out now run at a fraction of the previous cost without sacrificing the agentic capabilities that made Haiku 4.5 useful. The subagent pairing with Opus 5.5 for hierarchical workflows is worth evaluating if you're running orchestration patterns—Haiku handles the high-volume leaves while Opus handles coordination.
Verdict: Ship. Replace Haiku 4.5 immediately for volume workloads. Start in Bedrock Playground to validate behavior before cutting over production traffic. IAM setup is the only friction point.
Claude Haiku 5.5 Ships with Adaptive Reasoning Levels
Haiku 5.5 adds a reasoning parameter (AI SDK) or reasoning_effort (Chat Completions API) that controls thinking depth from low to max. At lower effort levels, token spend drops and latency improves—useful for summarization, classification, and live support where you don't need deep reasoning chains. At xhigh and max, the model handles complex multi-step problems without requiring a model swap.
This matters because it eliminates the binary choice between a fast, cheap model and a capable one. You can now tune inference budget per task type within a single model, which is operationally cleaner than maintaining routing logic across multiple models. The model string on AI Gateway is anthropic/claude-haiku-5.5—zero code refactor required beyond adding the effort parameter.
Verdict: Evaluate. The mechanism is sound, but the right effort level for your workloads isn't obvious without profiling. Run your existing summarization and classification pipelines at low effort, measure output quality against your acceptance threshold, then dial up from there. The token savings at lower effort levels are real, but you need your own data to quantify them.
Decisions API Routes Structured Inference Through AI Gateway
OpenAI's Decisions API is now accessible at AI Gateway's /v1/decisions endpoint. Instead of returning text, it returns typed probabilities and choices—predicate, choice, and score question types against shared input. You call decisions.create() with the OpenAI SDK (JS 7.30.0+, Python 3.26.0+), pointed at AI Gateway's baseURL.
The practical impact is eliminating prompt engineering overhead for classification tasks. If you're currently parsing LLM output to extract a routing decision or triage label, that's fragile—it breaks on phrasing changes and requires schema validation on top of inference. The Decisions API treats classification as a typed primitive, which removes the parsing layer entirely and gives you probabilities you can threshold against business logic.
Verdict: Ship if you're parsing LLM output for routing or classification today. This is a direct structural improvement. The SDK version requirements are the only gate—verify your current versions before integrating.
OpenAI Ultrafast Tier Launches on Vercel AI Gateway
GPT-6 Astra requests can now route through Ultrafast infrastructure by adding serviceTier: 'ultrafast' to your OpenAI SDK calls or Chat Completions API requests. The trade-off is explicit: 6× the token cost for reduced latency on US and global infrastructure. EU requests automatically fall back to standard tier.
This is worth testing for interactive coding assistants or chat interfaces where latency directly affects perceived quality. But the math has to work. If you're already hitting acceptable latency at standard tier, the 6× cost multiplier is waste. The automatic EU fallback is a limitation worth understanding before you commit—if your user base is meaningfully European, the latency benefit disappears for that segment.
Verdict: Evaluate. Benchmark your current p50/p95 latency at standard tier first. If you're within SLA, skip it. If you're not, test Ultrafast on a representative sample and measure actual latency improvement against the cost delta before rolling out broadly.
Claude Haiku 5.5 Matches Luna Pricing Under 100k Tokens
Haiku 5.5 hits price parity with GPT-6 Luna at $0.10 input / $0.50 output for contexts under 100k tokens. Above that threshold, pricing jumps 5× relative to Luna, and Haiku 5.5's tokenizer runs approximately 1.25× inflation compared to the prior version—meaning the same text costs more tokens than you'd expect from a straight version upgrade.
For workloads consistently under 100k tokens, this creates a genuine benchmarking window. You can now run Haiku 5.5 and Luna on identical tasks with identical cost expectations and let quality metrics decide. Above 100k tokens, the calculus inverts sharply—Luna becomes significantly cheaper and token efficiency differences compound.
Verdict: Evaluate. Re-baseline your token counts before committing. The 1.25× tokenizer inflation changes your cost model even at equivalent pricing, and you need to know where your workloads actually land relative to the 100k threshold. Profile first, then decide.
If this kind of technically grounded coverage is useful, Dev Signal lands in your inbox every week with the same signal-to-noise ratio—no vendor press releases, just what's actually worth your time. Subscribe if you'd rather read one email than track six sources.
Top comments (0)