Originally published on Medium.
AI fatigue is real. New technologies have always pulled developers in. Ruby on Rails, Go, Rust: we'd write a quick "Hello World," play with it for a while, and over months or years settle on what it was actually good for. With AI, that cycle has shrunk to days. There's a new framework, or a new way of doing things, almost every week, and keeping up is so tiring that many developers stop trying.
Generative AI brought something traditional programming never had: reasoning. Life is never a straight line, and the beauty of machine learning is that it doesn't assume the rules of life are written in advance. Our own rules aren't either. They come from history and practice etched into our brains, the neural patterns of our ancestors, reinterpreted from time to time. AI offers us a glimpse, a shimmer of hope, that this accumulated human knowledge can be codified to automate workflows that until now needed a chain of people in between.
Look closely, though, and most of what we're building with it is the same thing. Those workflows are graphs. Coding agents, support bots, personal assistants like OpenClaw all follow that shape. A human workflow is replaced by a network of agents.
What's being commoditized is the framework for building that graph. LangChain and Spring AI pioneered it, and I have a lot of respect for both. At work we've been using LangChain and LangGraph, and I've come to know them well. When LangChain introduced Deep Agents, a "batteries included" agent harness, I realised it was, underneath, a more opinionated, standardized graph. CrewAI, n8n and the rest each offer their own opinion on how a workflow should be built. As an industry, we're still early in finding the right abstraction.
Meanwhile the work takes several translations. We picture a workflow, describe it to a coding agent in English, and the agent writes Python or Java that we then have to review to check it built what we pictured. Along the way we have to answer the hard questions in framework terms: what should be a node, what should be an agent, when to ask a person, how much is safe to automate. No wonder so many developers give up on frameworks and hand-roll their own loops.
It doesn't have to work this way. What if the coding agent didn't write the implementation at all? What if it wrote the workflow itself, in a language you can read in a minute and see as a graph, and the framework turned that into a running system? You would decide what the workflow is and what it must guarantee, and stop reviewing pages of generated plumbing.
So yes, this is another tool. But it isn't another way to wire graphs. It's a way to stop writing the wiring at all.
Re-introducing Loom
I first wrote about Loom's design earlier this year. That piece covers the compiler side of the story. Since then Loom has grown a security audit, an eval harness, tasks and a CLI. This one is about developer experience.
Say I want an automated workflow where I describe a problem on my Ubuntu laptop, and the workflow researches it on the web and writes a script to fix it. I describe that in English to my coding agent, and using the Loom guide skill (more on that at the end), it drafts this:
// laptop-doctor: describe a laptop problem in English. Agents diagnose it and research it, another writes a fix
// script, a code check blocks anything that deletes data, you approve twice, and only then is the script run.
budget { tokens: 400000 calls: 60 }
tool Inspect {
use: shell
allow: "df, free, uptime, lsblk, ps, lscpu, uname, journalctl, dmesg, systemd-analyze, ip, ss"
timeout: 20s
unattended: true
}
agent Diagnostician { model: "gemini-2.5-flash" prompt: "diagnostician" temperature: 0.1 tools: [Inspect] max_iterations: 10 guard { pii: mask } }
agent Researcher { model: "gemini-2.5-flash" prompt: "researcher" temperature: 0.2 tools: [web_search] }
agent ScriptAuthor { model: "gemini-2.5-flash" prompt: "scriptauthor" temperature: 0.2 tools: [web_search] }
agent Remediator { model: "gemini-2.5-flash" prompt: "remediator" temperature: 0.1 tools: [Inspect] max_iterations: 6 guard { pii: mask } }
workflow Main(laptop_issue) {
// 1. Look at the machine, then research it. The researcher may send the diagnostician back for more (at most 3 extra rounds).
delegate "Problem: {laptop_issue}\nCollect evidence." to Diagnostician -> evidence_text
loop until (research_result.verdict == "ENOUGH") max 4 {
delegate "Round {_loopRound} of 4. Problem: {laptop_issue}\nEvidence:\n{evidence_text}" to Researcher -> research_result expecting { verdict: enum["ENOUGH", "NEED_MORE"], findings: string, request: string }
alt (research_result.verdict == "NEED_MORE") {
// the 4th round is the last: whatever it asks for is not collected (so at most 3 extra rounds of evidence)
alt (_loopRound != "4") {
delegate "Problem: {laptop_issue}\nEarlier evidence:\n{evidence_text}\nExtra checks requested: {research_result.request}" to Diagnostician -> evidence_text
}
}
}
// 2. Write a fix script. If no safe fix is supported, stop here.
delegate "Problem: {laptop_issue}\nResearch findings:\n{research_result.findings}" to ScriptAuthor -> fix_draft expecting { has_script: enum["YES", "NO"], script: string, summary: string }
alt (fix_draft.has_script == "NO") {
note "No safe fix was found: {fix_draft.summary}\n\nWhat the research found:\n{research_result.findings}"
} else {
// 3. The safety check is code, not a model. A blocked script goes back to the author, at most 2 rewrites.
run ValidateScript(script = fix_draft.script) -> check_result
alt (check_result.outcome == "blocked") {
delegate "Rewrite 1 of 2. Your script was blocked by the safety check: {check_result.reason}\nRewrite it so this cannot happen. Problem: {laptop_issue}\nResearch findings:\n{research_result.findings}" to ScriptAuthor -> fix_draft expecting { script: string, summary: string }
run ValidateScript(script = fix_draft.script) -> check_result
alt (check_result.outcome == "blocked") {
delegate "Rewrite 2 of 2. Your script was blocked again by the safety check: {check_result.reason}\nRewrite it so this cannot happen. Problem: {laptop_issue}\nResearch findings:\n{research_result.findings}" to ScriptAuthor -> fix_draft expecting { script: string, summary: string }
run ValidateScript(script = fix_draft.script) -> check_result
}
}
alt (check_result.outcome == "blocked") {
note "Stopped. The safety check blocked every draft: {check_result.reason}\nNothing was saved or run."
} else {
// 4. First approval: the person reads the script. A script the check was unsure about says so.
alt (check_result.outcome == "ask") {
human_prompt "THE SAFETY CHECK COULD NOT TELL if this script is safe: {check_result.reason}\n\n{fix_draft.script}\n\nWhat it does: {fix_draft.summary}\n\nApprove it anyway? (type yes or no)" -> first_approval
} else {
human_prompt "The safety check found nothing that deletes data.\n\n{fix_draft.script}\n\nWhat it does: {fix_draft.summary}\n\nApprove this script? (type yes or no)" -> first_approval
}
alt (first_approval == "yes") {
run SaveScript(script = fix_draft.script, label = "fix") -> saved_file
// 5. Second approval, just before running: the person may keep the saved file and run it later.
human_prompt "Saved to {saved_file.value}\nRun it now? (type yes or no)" -> second_approval
alt (second_approval == "yes") {
// 6. Running is code: only the saved file, checked against the approved checksum, 2 minutes at most.
run RunScript(path = saved_file.value, sha256 = saved_file.sha256) -> run_outcome
// shown at once, so the result is never lost if the spend cap stops the remediator's report
note "The script finished ({run_outcome.outcome}):\n{run_outcome.value}"
delegate "Original problem: {laptop_issue}\nRun result: {run_outcome.outcome}\nScript output:\n{run_outcome.value}" to Remediator -> final_text
note "{final_text}"
} else {
note "Not run. The script is saved at {saved_file.value}; you can read it and run it yourself later."
}
} else {
note "Not approved. Nothing was saved or run."
}
}
}
}
The VS Code extension validates it as you type and draws the network of agents beside the editor:
What each line buys you
budget { tokens: 400000 calls: 60 }
is a hard ceiling on what one run can spend. It isn't a dashboard you check afterwards. When the allowance is gone, the run stops. The same limits can sit on a single agent or a single step.
tool Inspect { use: shell allow: "df, free, uptime, ..." }
gives the agents a shell, but only for the commands you list, all of which look rather than change. Each call has twenty seconds. Because nothing on the list can modify the machine, the agents may use it without asking you each time.
tools: [Inspect]
and
tools: [web_search]
decide who can do what. The Diagnostician and the Remediator can look at the machine. The Researcher and the ScriptAuthor can search the web. No agent gets both, and none of them can run the fix.
guard { pii: mask }
matters more than it looks. System logs are full of usernames, email addresses and IP addresses. They're replaced with placeholders before anything reaches the model.
loop until (research_result.verdict == "ENOUGH") max 4
lets the Researcher send the Diagnostician back for more evidence, at most three extra times. The loop can't run away, and the agents don't decide when to stop. The script does.
expecting { verdict: enum["ENOUGH", "NEED_MORE"], ... }
makes each answer arrive in a fixed shape, so the next step branches on verdict or has_script instead of guessing at free text. If the ScriptAuthor finds no safe fix, it says so with
has_script: "NO"
, and the workflow stops honestly instead of inventing one.
run ValidateScript(...)
is the heart of it. That isn't a model; it's Java code. The model writes the script, and code decides whether it's safe. A blocked script goes back to the author with the reason, at most twice. If the check can't tell, it says so, and your approval prompt tells you plainly. Same input, same answer, no tokens spent, and no clever wording in a log file can talk it into anything.
human_prompt
First you read the script and approve it. Then it's saved, and you're asked again whether to run it now or keep it for later. If nobody is at a console, the workflow pauses without tying up a thread and picks up when the answer arrives, without repeating the work already done.
run RunScript(path = ..., sha256 = ...)
runs only the file you approved, checked against its checksum so nothing can swap it in between, for two minutes at most. The result is shown immediately, so even if the spend cap stops the Remediator's report, you still see what happened.
Around the script there's more you didn't have to write: a security audit that reads this file and flags any agent that can read untrusted text, reach private data and act; an evaluation harness for testing it against your own examples; a journal that lets a stopped run resume instead of starting over; a graph of the workflow drawn straight from the script; and a guide your coding agent can follow to build in the same way.
That's the point of Loom: the workflow is the interesting part, so everything else is built in.
One tool, two ways in: a command line and VS Code
Loom ships as a single jar, the weave command line, and a VS Code extension that carries it. There is no server to stand up and no platform to sign in to.
The command line is how you work in a terminal and how your pipeline works with your workflow:
weave check laptop-doctor.loom # every problem, with its line, without running or spending
weave graph laptop-doctor.loom --format mermaid # the workflow as a diagram
weave explain laptop-doctor.loom # the script described back to you in plain English
weave audit laptop-doctor.loom # a security review; exits 1 on a high finding
weave run laptop-doctor.loom --max-tokens 50000 # run it, with a cap
weave package laptop-doctor.loom --fat # a standalone jar
weave next # what to do next in this project
It exits with the codes your CI expects.
The VS Code extension is where most of the writing happens: syntax highlighting, problems underlined as you save, an outline of your workflows, snippets that expand the common statements, a live graph of the workflow that follows your cursor, a one-click run, and the guide one command away.
Same project, same files, either way in. Check it from the terminal in CI, build it in the editor, and nobody has to translate between the two.
Personal data and guardrails, in one line
A laptop's logs say more about you than you'd think: your username, your email address, the IP addresses you've connected from. Before laptop-doctor lets a model read any of that, the question is what the model gets to see. In Loom, that's a property of the agent, written next to its model:
agent Diagnostician { model: "gemini-2.5-flash" prompt: "diagnostician" temperature: 0.1 tools: [Inspect] max_iterations: 10 guard { pii: mask } }
The guard has three settings, plus a bias check:
mask: emails, phone numbers, social security numbers, card numbers and IP addresses become placeholders like [EMAIL] before anything reaches the model. That covers the task, the context, the agent's memory and the results of its tools, so whatever journalctl prints is masked before the Diagnostician sees it. The answer is masked before it's stored too, so the evidence the Researcher takes to the web carries no trace of you.
block: if a task or an answer contains personal data, the step fails, and the error names the kinds of data, never the values.
warn: it records the finding and carries on.
bias: warn or bias: block checks the answer for bias, with simple rules by default, or a model you name as the judge.
Want to protect just part of a workflow? Wrap those steps in a guardrail and say what happens if it trips:
guardrail (PII) {
delegate "Problem: {laptop_issue}\nEvidence:\n{evidence_text}" to Researcher -> research_result
} on_violation {
note "Personal data reached the research step. Stopped before anything was searched."
}
The audit trail follows the same rule: counts and kinds, never the data. Secrets such as API keys, webhook URLs and passwords can't be written into a script at all. They come from the environment or a secret store, and they're scrubbed from results, errors, traces and logs.
A security review that fits in CI. A built-in audit reads your workflow without running it and shows what each agent can touch. It asks the question that matters most for agents: does any single agent read untrusted content, reach private data and have a way to send or act? That combination is Simon Willison's lethal trifecta, and it's how a hidden instruction in a log file or a web page becomes a leak or a harmful action. Loom flags it, along with unapproved side effects, a missing budget, or personal data with no guard. Every finding is mapped to the OWASP Top 10 for LLM applications and says what to change.
In laptop-doctor, the agents that look at the machine can't search the web, the ones that search can't touch the machine, and none of them can run the fix. The steps that must always happen aren't prompts at all. They're code. That's next.
Write your own tools, and make the critical steps unskippable
Real workflows have to do things: look at a system, call an internal API, run a script. Loom gives you a clear way to think about it.
In laptop-doctor, Inspect is a tool. The Diagnostician decides which commands to run and when. ValidateScript, SaveScript and RunScript are tasks. No agent decides whether the safety check happens. It happens every time a script is written.
Tools. You get a lot for free: web search, a calculator, an OpenAPI tool that exposes an API from its spec, MCP servers, and six general-purpose tools: webhook (Slack, Discord, Teams), email, http, file, shell and sql. They come with safe defaults: https only, no calls to internal or cloud-metadata addresses, the agent picks a path but never the host, credentials only from the environment or a secret store, SQL that only reads, and a shell tool that runs programs directly (no shell syntax) and needs a person's approval. Inspect is that shell tool narrowed to a list of commands that only look, which is why it may run unattended. When you need something that isn't there, a tool is a small Java class:
public class WordCount implements Tool {
@Override public String getName() { return "word_count"; }
@Override public String getDescription() { return "Counts the words in the text argument."; }
@Override public String execute(Map<String, Object> args) {
String text = String.valueOf(args.getOrDefault("text", "")).strip();
return String.valueOf(text.isEmpty() ? 0 : text.split("\\s+").length);
}
}
It's an ordinary class, which makes it exactly the kind of code your coding agent writes well for you.
Tasks. For the steps that must happen, you write a task, plain Java run with a run step:
/** The mandatory safety check. Plain code, no model: outcome is "pass", "ask" (a person decides) or "blocked". */
public class ValidateScript implements Task {
@Override public String getName() { return "ValidateScript"; }
@Override public TaskEffect effect() { return TaskEffect.NONE; }
@Override public TaskResult run(TaskContext context) {
String script = context.requireArg("script", String.class);
ScriptSafety.Verdict verdict = ScriptSafety.check(script);
TaskResult result = TaskResult.outcome(verdict.outcome().name().toLowerCase());
return verdict.reasons().isEmpty() ? result : result.reason(verdict.reasonText());
}
}
And refer to it in the Loom file:
run ValidateScript(script = fix_draft.script) -> check_result
alt (check_result.outcome == "blocked") {
delegate "Your script was blocked by the safety check: {check_result.reason}\nRewrite it so this cannot happen." to ScriptAuthor -> fix_draft expecting { script: string, summary: string }
}
Same input, same output, no tokens spent, and no prompt injection can change what it does. A task that changes the outside world says so, gets an idempotency key so a restarted run doesn't do it twice, and can require a person's approval. That's why SaveScript and RunScript are tasks: if a run stops after your approval and resumes the next day, the script isn't saved twice or run twice. It's also what lets you put a safety check, a refund policy or a payment inside an AI workflow. The model writes, and code decides.
Autonomy, earned
laptop-doctor asks you twice, every time. For a script about to run on your own machine, that's right. But picture a decision your team makes fifty times a day: approve this refund, grant this access, escalate this ticket. Ask a person to confirm every one, and within a week nobody is reading the prompt.
Jairo Cadena makes this point well in Human-in-the-Loop as an Authorization Primitive. Consent fatigue, he argues, isn't people failing. It's what you should expect from a control designed with no regard for how much attention a person actually has. His remedy is to treat each approval as a real grant of authority: narrowly scoped, used up by the action it allows, recorded with who approved and exactly what they saw, and required only where the risk earns it. He also gives a test for a control that has decayed: approval rates near 100% and decisions made in no time at all mean nobody is reading.
Loom's answer to that is earned autonomy. Every organisation asks the same question before it lets an agent act alone: has this agent earned it? Usually that's a gut call made once, before launch. Loom turns it into a record. A workflow declares a decision. The agent proposes, a person decides, and the runtime keeps a ledger of both. The agent starts out watched, moves up as its record earns it, and moves back down when the record says it should.
decision Refund {
proposed by: Triager
choices: approve, reject, escalate
group cases by: tier
remember: amount, reason, customer_since
dangerous mistake: propose approve, person decides reject
ask: support-lead
keep records for: 180 days
when the agent changes: test it on past cases
trust {
start at watch
never go above suggest
to suggest: after 100 cases over 14 days, agreeing at least 90%
to act: after 300 cases over 30 days, agreeing at least 97%, with no dangerous mistakes
judge on the latest 300 cases
check 5% of cases with a person who doesn't see the proposal
always ask a person when amount > 200
always ask a person after 50 cases a day
drop to suggest when 2 dangerous mistakes in 50 cases
drop to suggest when 2 reversals in 100 cases
drop to suggest when agreement falls below 92%
drop to suggest when 5 unusable proposals in 50 cases
moving up needs approval from: risk-owner
}
}
workflow Triage(ticket, tier, amount, reason, customer_since) {
decide Refund -> verdict
alt (verdict == "approve") { note "refund {ticket}" }
}
It reads aloud, and every line is doing something:
start at watch means the agent proposes quietly and a person decides without seeing the proposal. That blindness matters. If people could see the agent's answer, high agreement might only mean they were nodding along. Only blind cases count as evidence.
to suggest: andto act: are the rungs. At suggest, the person sees the proposal and its reasoning, and confirms or overrides it. At act, the agent decides and nobody is asked. Agreement is judged by the lower end of a 95% confidence interval, not the raw rate, so twenty out of twenty doesn't count as proof.
never go above suggest is the default. Reaching act is a sentence someone has to write on purpose, and moving up needs approval from: risk-owner puts a named person behind every promotion.
dangerous mistake: lets you say which disagreement really matters. Approving a refund a person would have refused is worse than the reverse, so it's counted on its own.
always ask a person when amount > 200 and after 50 cases a day are the attention budget written down. Big cases always reach a person. Past a daily volume, people are asked again rather than letting the agent run unwatched.
check 5% of cases with a person who doesn't see the proposal keeps measuring even after the agent is trusted, so the record never goes stale.
drop to suggest when... works without a meeting. Two dangerous mistakes, a run of reversals, or slipping agreement, and the agent loses its autonomy from the next case.
group cases by: tier means trust is earned separately for each kind of case. Being good at gold-tier customers says nothing about basic ones.
When the agent changes, the trust doesn't come along for free. Each case records the identity of the agent behind it: its model, prompt, tools, schema and guard settings. Edit the prompt or change the model, and that's a new agent as far as the ledger is concerned. With when the agent changes: test it on past cases, Loom replays the old agent's blind cases under the new one before it handles a single live case, and lets it keep only the level that replay earns. You can do the same by hand before you ship anything:
weave replay refund.loom --decision Refund --candidate refund-v2.loom --since 14d --limit 200
The report shows both agents' agreement with your people on the same cases, and every case where their answers differ, unsafe ones first. In other words, you find out whether a change is better or worse on your own history, before a single customer sees it.
Set against Cadena's tests, the fit is close. Each case is a record of who decided, what they saw and how long they took. The tiers are written in the script: amounts, daily volumes, kinds of mistake. Questions park and wait for the right person, on their phone if you set it up, instead of whoever happens to be online. And because routine cases stop needing a person once the agent has earned it, the approvals that remain are fewer, and more likely to be read.
It's also clear about what it isn't. It doesn't learn; the ledger grades, it never trains or rewrites a prompt. It doesn't find out on its own that a verdict was wrong; you record that with weave autonomy outcome. And it doesn't certify compliance with anything. It produces evidence. No agent, tool or case can write to the ledger; only the runtime and the commands can.
The idea in one line: don't trust the agent, make it earn it, with your own history as the exam.
Let your coding agent do the rest
I started by asking why a coding agent should write a workflow in Java or Python at all. It's the question Loom answers, and the last piece is what makes the answer practical.
The repository carries a skill for coding agents. One command installs it into your project:
weave guide --install-skill .
That puts the Loom guide and a set of instructions in .claude/skills/llm4j-workflow-guide, so your coding agent knows how to build a Loom workflow the way this article does. It starts from a ready-made project, writes the workflow in Loom, puts the steps that must always happen in Java tasks, adds a budget and the guards, and checks the script before anything spends a cent. You describe what you want in English. The agent writes the workflow you can read in a minute and see as a graph, not pages of plumbing for you to review.
Try it in a few minutes
Clone the repository and run one script:
git clone https://github.com/srijithunni7182/llm4j.git
cd llm4j
scripts/quick-build.sh
That's the whole setup. It builds and tests the libraries, packages the skill as dist/llm4j-workflow-guide.zip, and builds and installs the VS Code extension. If VS Code's code command isn't on your PATH, the script tells you how to install the extension by hand.
Then add the packaged skill to your coding agent, describe a workflow you've been meaning to automate, and let it build your own version, with guardrails, spending caps, an evaluation report and security built in.
srijithunni7182
/
llm4j
A Native Java Ecosystem for AI - Agents, Tools, Harness and Workflows
llm4j
AI agents, written the Java way.
Typed, testable, observable: from a single tool call to autonomous workflows that run for days
Get started · The stack · Examples · Docs · Why ai-agent4j? · Why Loom? · Security
Important
The most complete explainable-AI toolkit for Java agents. Every agent step is recorded as an immutable audit event (traceability)
each run carries a confidence score with shouldEscalateToHuman(), PII is masked before it reaches a model
or a log, and bias monitors can flag or block a response. The four pillars are mapped to GDPR, the EU AI Act and
the NIST AI RMF in xAI: Beyond Black Boxes.
These ship in the box. With Spring AI or LangChain4j you get observability and PII guardrails, but you build confidence scoring, bias monitoring and the audit trail yourself.
Why llm4j
Java runs the systems that can't go down…
Clone it, try it, and if you like the idea, give the repo a star. It's still young, so you'll likely find rough edges and bugs. Tell me about them in the comments or open an issue, and I'll keep improving it with your feedback.
What's the part of your current agent workflow you'd most like to stop hand-wiring? I'd love to hear in the comments.


Top comments (0)