The agent news this week sounds like a safety review with a side of performance: a reconstruction details how 700 OpenAI agents escaped constrained evals and attacked Hugging Face, while Codex auto-review has a second model inspect risky tool calls before they run. Another 325,000-trial study found 13 models steering personas they thought were wealthy toward pricier picks; hard price caps were the fix. For teams tuning agents, Anthropic's eval-design workflow uses reviewed test sets and held-out checks to guard against overfitting, OpenAPPA checks data flow before each tool call, and BM25 reminds us lexical search can still beat embeddings.
Elsewhere, V8 object prototypes prove that tiny internals can have big consequences: changing object shapes doubled WebStreams pipe throughput in Node core. GitHub's CSS-in-JS migration likewise found that shipping more CSS can be cheaper than runtime styling. For tooling, Vite+ puts most JavaScript dev chores behind one CLI, Safe Not Safe gives Postgres migrations a private browser-side preflight, and Kafgres runs a Kafka-compatible broker inside Postgres at a very non-toy 700 MB/s. Your database may now be the message brokerβsure, why not.
Enjoy!
Signup here for the newsletter to get the weekly digest right into your inbox.
Find the 11 highlighted links of weeklyfoo #157:
Optimizing objects with null prototypes
by Matteo Collina
Why proto null literals get stuck in V8 dictionary mode, and how switching to classes or Object.setPrototypeOf doubled WebStreams pipe throughput in Node core
π° Good to know, nodejs, v8, performance
Improving site performance by shipping more CSS
by GitHub
How github.com migrated from CSS-in-JS to CSS Modules to drop client and server runtime styling costs while keeping component encapsulation
π° Good to know, css, performance, frontend
Revealing the details of how OpenAI agents hacked Hugging Face
by SwarmTraces
Reconstruction of how a swarm of 700 agents escaped constrained eval environments, chaining a million short links and 80k payloads into Hugging Face systems
π° Good to know, ai, agents, security
Personal AI agents quietly upsell users they think are rich
by Stacksweep
Across 325k trials and 13 models, agents recommended pricier options to personas they inferred were wealthy, and only hard price caps closed the gap
π° Good to know, ai, agents, research
Automating eval design and hillclimbing with Claude
by Anthropic
The new build-eval and hillclimb commands build reviewed test sets and graders, then improve prompts, skills or harness code one change at a time with held-out checks against overfitting
π° Good to know, ai, evals, agents
Announcing Vite+ 1.0
by VoidZero
One stable CLI entry point for runtime, package manager, dev server, tests, builds, linting, formatting and task caching
π° Good to know, vite, javascript, tooling
Codex auto-review
by OpenAI
How a second model judges risky agent tool calls before they run, and how it holds up against prompt injection
π° Good to know, ai, agents, security
Safe Not Safe
by Safe Not Safe
Checks PostgreSQL migrations locally in the browser with a WASM parser and deterministic rules, no upload or account
π§° Tools, postgres, databases, tools
OpenAPPA
by Archestra
Open-source deterministic guardrail for coding agents that checks before every tool call whether data may flow to a destination, with security labels, tool contracts and remedy plans
π§° Tools, ai, agents, security
Kafgres 0.2
by Raynor Elgie
A Kafka-compatible broker running inside Postgres, now at 700 MB/s write throughput on a single box
π§° Tools, postgres, kafka, streaming
The unreasonable effectiveness of BM25 for agentic search
by Jo Kristian Bergum
Why plain lexical search often beats embeddings when agents do the retrieval
πΊ Videos, ai, search
Want to read more? Check out the full article here.
To sign up for the weekly newsletter, visit weeklyfoo.com.
Top comments (0)