DEV Community

Yuuichi Eguchi
Yuuichi Eguchi

Posted on

What were my AI-written UI tests actually proving? Green XCUITests over a broken macOS terminal

Calyx is a native macOS terminal for running and supervising coding agents in parallel. Most of its code is written by Claude Code subagents, and the tests are no exception: I run a TDD loop where one agent writes a failing test and another writes the implementation. The UI tests are XCUITest, and the CalyxUITests target has more than 130 test methods as I write this.

The suite was green many times while the product was broken. The clearest case was the switch in Settings that enables persistent sessions. Clicking it did nothing, and every persistent-session E2E test passed. On the same day I found out that the UI tests had left 47 shells running against the session daemon on my own Mac.

This post groups those failures by kind. Each one starts with something specific to Calyx, but they all come back to two questions: what does this assertion actually observe, and what does this test leave behind on the developer's machine? It is written for people who have agents writing their tests, and for anyone writing E2E tests for a macOS app.

1. Where can a test observe a terminal?

XCUITest observes an app through accessibility. Buttons and labels show up as elements; the text in a Calyx pane does not. Calyx renders the terminal on the GPU with Ghostty's engine (libghostty), and the glyphs on screen never become accessibility elements.

The first persistent-session E2E test ran straight into this. A persistent session keeps the shell in a daemon outside Calyx, so you can quit the app and come back to the same shell after relaunch. The first version of the test echoed a marker string in a pane and, after relaunch, looked for it with app.staticTexts. That assertion could not pass even with a fully working session (c01be27e0).

The first fix moved the observation into a file. The pane's shell writes its own PID to a file; after the app restarts, the restored pane writes it again. Equal PIDs mean the daemon kept the same shell alive and the restored pane reattached to it.

That version still typed into the pane, and typeText sends key events to whatever app is actually frontmost on the system, not to the app under test. When the app under test loses focus, the keystrokes land in the developer's own terminal. I saw that happen twice, and about an hour later the test was rewritten to type nothing at all (62564941c). Section 4 has the story of the leak.

The rewritten test observes the daemon's ledger. Persistent sessions are owned by a daemon called calyx-session, which records its sessions in a ledger file at $HOME/.calyx/state/sessions.json. The test reads the ledger before and after the restart and asserts that the same session id is still Running with the same PID and creation time. It also asserts that attached_clients is at least 1 on both sides of the restart. The daemon keeps a session alive with zero clients attached, so matching ids and PIDs alone would not prove that the restored pane reattached.

The test reads the ledger as a file rather than asking the CLI, and there is a reason for that too. The earlier version spawned calyx-session ls --all --json from the test runner, and the output was always empty, even while the daemon was healthy (2bb1b24a2). The UI test runner is itself App-Sandboxed; codesign -d --entitlements on it showed the sandbox enabled with only a read-only file exception. A sandboxed process cannot connect to the daemon's Unix domain socket. The same query run outside the sandbox succeeded, so the problem was on the runner's side. Writes are restricted too: a write from the runner to /tmp failed, and because it was written with try?, it failed silently.

For a terminal app, the most reliable thing an E2E test can observe turned out to be the state the product writes to disk. Screen text is invisible to accessibility, the socket is out of reach, and typed keys can leak. A file the runner can read was the one channel none of those constraints touched.

2. Launch arguments and a switch that did nothing

According to the July 7, 2026 commit that fixed them (1a17867c1), the four switches in Settings > Sessions were all dead. They surfaced through user reports and a deliberate sweep for UI that looks functional but is not. The switches had no target and no action, and nothing seeded them from the stored value. Clicking one wrote nothing; reopening Settings showed it off again. There was no way to turn persistent sessions on from the UI.

Every persistent-session E2E test still passed, because every one of them turned the feature on with a launch argument before starting the app:

override var additionalLaunchArguments: [String] {
    ["-calyx.session.persistentSessionsEnabled", "YES"]
}
Enter fullscreen mode Exit fullscreen mode

A -key value launch argument goes into the UserDefaults argument domain, which is read ahead of the app's own persistent domain. What these tests could check was "what happens when persistent sessions are on", not "can a user turn persistent sessions on". The switch's wiring was not on any test's path.

It was also easy to miss by hand. An NSSwitch flips its own on-screen state when clicked whether or not anything is wired to it, so at the moment you click, it looks like it worked.

The E2E test written for the fix clicks the real switch, quits through the app menu, relaunches, and reads the switch back. It deliberately avoids the launch argument: the argument domain would shadow every read of that key for the life of the process, so the test could never tell whether set() had taken effect. The relaunch matters because the Settings window controller is a per-process singleton. Close and reopen the window in the same process and you get the same NSSwitch instance you just clicked, which shows "on" whether or not anything was saved. After a relaunch, a fresh controller builds its state from what is actually on disk, and "looked like it worked" becomes distinguishable from "was saved".

A commit that evening (1cab009aa) closed the gaps found by an audit that classified every assertion in all 20 UI test classes by what it actually proves. Ten of them would have let through a defect a user would notice within a second. A few examples:

  • The group-switching test checked that group rows existed. It did not check that the active group changed, which is to say that the visible set of tabs changed.
  • Of the seven switches in Settings, one had a click, quit, relaunch, and read-back E2E test. The other six had none.
  • The search test checked that the search bar appeared. It now types a real query and asserts the match count and that Find Next moves.
  • The test for pressing Enter in the Compose box checked that the box cleared. It did not check that the text reached the terminal and ran.
  • The restore-after-restart tests looked only at ledger fields. They now also assert the window count and the restored tab's visible title.

Each of these tests existed, was green, and checked something. What they checked was not what the user sees. A red-first TDD loop does not catch this. The launch-argument E2E tests fail without the flag and pass with it, so they go from red to green exactly as the process asks. Red-first guarantees that an assertion reacts to something; it says nothing about whether that something is visible to a user.

3. When the test's actions are not the user's actions

The launch arguments in section 2 were a case of the test reaching a state by a path no user takes. The same gap showed up in the actions themselves, such as dragging and quitting.

After moving to Xcode 27 and macOS 27, press(forDuration:thenDragTo:) stopped delivering any mouse events to the app (516a9afe1). I instrumented mouseDown and mouseDragged on the tab view and added an app-wide NSEvent local monitor; nothing fired. click(forDuration:thenDragTo:), the same click-hold-and-drag gesture through a different XCTest code path, did deliver events.

A test that asserts a reorder happens will fail when the drag never arrives. The problem was the test pointing the other way. test_tapStillWorksAfterDrag asserts that a drag shorter than the 5pt threshold does not reorder anything and the tab stays where it was. With no drag delivered, no reorder could happen, so the test passed without exercising anything:

// `click(forDuration:thenDragTo:)`, not `press(forDuration:
// thenDragTo:)` -- see test_dragTabBarTab_reordersCorrectly's own
// comment. Using the broken `press` variant here made this
// assertion pass vacuously (no drag was ever delivered to assert
// "no reorder" against).
startCoordinate.click(forDuration: 0.1, thenDragTo: nearbyCoordinate)
Enter fullscreen mode Exit fullscreen mode

An assertion that something does not happen also holds when the input never arrived. If you keep a test like that, pair it with a test that drives the same input and asserts that the expected thing does happen. When the pair goes red, you learn that the input is not getting through.

app.terminate() takes a path no user takes either (8d292b778). It sends SIGTERM, which skips AppKit's termination flow entirely. Calyx saves its session snapshot synchronously in applicationWillTerminate, and on a test-driven quit that never ran. After relaunch there was nothing to restore, and the daemon's ledger showed a third, fresh session. This time the product was right and the test's way of quitting was wrong. Mid-test quits now go through the app's own menu, the same path as a real quit. The terminate() in tearDown stays, since cleanup does not need the save.

Two more cases where the test assumed something the real environment did not provide. When a pane appears on screen, the shell the daemon started may not be reading its PTY yet. A keystroke sent in that window is silently dropped, and the test failed against a healthy pane (68b9fcf86). And initializing Sparkle, the auto-update framework, at launch put its first-run "Check for updates automatically?" dialog in front of the app under test, where it swallowed every keystroke XCUITest sent (0b9a60da9). Under the --uitesting launch argument, Calyx now skips updater setup.

4. What the test harness left on my machine

Running the E2E suite kept leaving things on my Mac. Here are three, in the order they happened. None of them showed up in the test results; the damage piled up on the developer's side instead.

Keystrokes in my real terminal

According to a July 5 commit (a3f729d57), an earlier E2E run had crashed partway through, and the echo marker the test was typing went into a real Ghostty window I had open. typeText sends keys to the frontmost app; when the app under test disappears, whatever is behind it receives them.

The first defense had two layers. The first was a separate bundle ID. A dedicated DebugUITesting build configuration gives the Calyx built for UI tests its own bundle identifier, com.calyx.terminal.e2e. According to the comment in project.yml, LaunchServices had previously been able to confuse a locally installed production Calyx with the UI-test build, letting keystrokes leak into whichever one it treated as the real app. The second layer was a guard that re-checks the frontmost app's bundle ID before every keystroke.

Later the same day, the persistent-session test moved to the no-typing design from section 1 anyway. If the test types nothing, nothing can leak.

47 zombie shells

According to the July 7 commit that fixed it (3e1aea2b8), the real daemon under my own ~/.calyx had accumulated 47 shells in a single day. The cause was the UI tests' base class. Its setUp() launched the app with no UserDefaults isolation and no HOME override. The app under test's defaults domain, com.calyx.terminal.e2e, was shared by every run, and during that period it had persistent sessions turned on. Every time a tab-creation or tab-renaming test launched the app, the app read that persistent sessions were enabled and created a session in my real daemon. The bundle-ID split from the previous section separated the test build from the production app; it did nothing to separate one test run from the next.

The fix gives every test its own UserDefaults suite, named with a fresh UUID:

let suiteName = "com.calyx.tests.e2e.CalyxUITestCase-\(UUID().uuidString)"
defaultsSuiteName = suiteName
app.launchEnvironment["CALYX_UITEST_DEFAULTS_SUITE"] = suiteName
Enter fullscreen mode Exit fullscreen mode

When the app sees this environment variable, it uses that suite. After the fix, a run of TabManagementUITests added zero sessions to the real ledger.

You might expect that overriding HOME with a temp directory would be enough. In this project it was not, and I confirmed that in two places:

  • UserDefaults.standard goes through cfprefsd and turned out to be tied to the real macOS account, not to HOME. The first version of the Settings switch test changed HOME for every test and was still reading and writing my real e2e defaults domain.
  • Commands running in a pane did not inherit the HOME the app was launched with. ps aux on a live run showed each pane's command started through login -flp <username> ..., which rebuilds the shell's environment for the real user. A bare calyx-session run in a pane from a test would operate on my daemon, not the test's. So every calyx-session call a test runs in a pane passes flags that name its state directories explicitly.

Even with a per-test suite, some code went around it. Three of the seven switches from section 2 (smooth scrolling and two LSP settings) read and write UserDefaults.standard directly, without going through the isolation suite. Their test saves those three keys before it runs and restores them afterwards.

Moving HOME into a temp directory caused a problem of its own (bd79c9482). The daemon's socket lives at $HOME/.calyx/run/sessiond.sock, and with HOME under NSTemporaryDirectory() that path came to 131 bytes. The path in a macOS sockaddr_un holds at most 104 bytes, so every calyx-session attach failed instantly and the pane just kept reconnecting. Test HOME directories now live at /tmp/cxe2e-<first 8 characters of a UUID>-h, which keeps the socket path at 46 bytes at most.

92,287 lock files

According to the August 20 commit that fixed it (b4e3bf393), 92,287 lock files had accumulated under my ~/Library/Application Support. When Calyx edits an agent CLI's configuration file, it takes an exclusive lock on a lock file named after a hash of the config path. It never deletes those lock files, on purpose. If you delete a flock lock file after use, another process can recreate the path as a new inode, and two processes each end up believing they hold "the" lock. A real user has a handful of config paths, so they end up with a handful of lock files.

The unit tests, though, write to a fresh UUID temp path in every test. Every write minted a new hash and left another file behind in the real directory. The files arrived in bursts that matched test-suite runs, not any configuration activity.

The fix puts the lock directory under the test host's own temp directory whenever the code runs inside the test host. I did not do it as a per-test override, because any write creates a lock file, and a single test that forgot to opt in would bring the leak back. In the current code this redirect lives in one place, CalyxPathRoot.testRoot.

5. What typeText actually sends

Some tests still type into panes, and typeText tripped me up repeatedly.

typeText does not send characters. It presses, under the test runner's keyboard layout, the key that produces each character, and the pane turns that key back into a character under its own layout. When the command-log E2E test typed a marker containing an underscore, the pane received =, and the marker never matched (c883f2fe5).

Quotes broke too. Typing sh '<path> arrived in the pane as sh ;Su/Users/..., and sh, handed a mangled argument, started an interactive shell instead of running the script. With an IME such as Japanese input active, keystrokes become input to the IME's composition, and Return commits the composition instead of reaching the terminal.

Pasting avoids layouts and brings its own problem. On a loaded machine, Return can arrive while the paste is still being delivered and land inside the bracketed paste, and then the command never runs. That happened even with the short one-line sh '<path>'.

The current setup combines three things:

  1. For the duration of each test, switch the input source to ABC (or US if ABC is not installed) and restore the original afterwards (InputSourceGuard).
  2. Write the command itself to a script file and type only sh <path> into the pane with typeText. Before typing, check that the path contains only characters that need no quoting.
  3. When a test types a command directly, route it through a helper that accepts only layout-invariant characters:
func typeIntoPane(_ line: String) {
    let allowed = CharacterSet(charactersIn: "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789 /.;-\n")
    guard line.unicodeScalars.allSatisfy(allowed.contains) else {
        XCTFail("pane text contains a character that is not layout-invariant: \(line.debugDescription)")
        return
    }
    app.typeText(line)
}
Enter fullscreen mode Exit fullscreen mode

The second change broke another test's assumption (9ba33d3e0). The focus-movement test tagged each pane by running export PANE=... and read the tag back with echo $PANE. Once commands ran from a script file inside a child sh (23999ef0a, September 9), the export ended with the child shell and never reached the pane's interactive shell. echo $PANE always came back empty, and the test failed on main until I fixed it on September 26. The tag is now the pane's tty, written to a file and compared there.

6. Rules for letting an agent run the tests

I develop Calyx with Claude Code running inside Calyx. When the agent runs the UI tests, the Calyx I am working in is running on the same machine, next to the app under test. That combination gets two rules in my local setup. Neither is in the public repository; both are my own machine's configuration.

The first is a hook that forbids concurrent xcodebuild test runs. A Claude Code PreToolUse hook inspects every Bash command, and for xcodebuild ... test it creates /tmp/.calyx-test-lock with mkdir. mkdir is atomic, so a failure means another run holds the lock. In that case the hook checks with pgrep whether a test process is really running, and if so, it denies the command. The denial reason tells the agent not to run tests in parallel, not to kill the other run, and to wait for it to finish before re-running the same command. If no test process is running, the hook treats the lock as stale, removes it, and takes it.

The second rule, written into the skill that runs the tests, is that the agent never quits Calyx on its own. The UI suite is best run with no Calyx open, but pgrep cannot tell the Calyx in /Applications from a debug build in DerivedData. A running Calyx is almost certainly the terminal I am working in, and quitting it would destroy every shell in it along with the scrollback. So the agent stops and asks, and I quit it myself with Cmd-Q. This step is never handed to a subagent, because a subagent cannot stop partway and ask a human.

The test code follows the same principle. An instance of the app under test left behind by an aborted run blocks the next run's automation connection, so CalyxUITestCase terminates such instances before launching. It only touches instances whose bundle lives under a /DebugUITesting/ build directory, never the Calyx in /Applications or a plain Debug build.

7. A checklist for the next macOS E2E suite

  • For each assertion, write down what the user should be seeing when it passes. An assertion you cannot describe that way only proves that an element exists.
  • Test settings by operating the real control, not a launch argument, then restart the process and read the value back.
  • Pair every "nothing happens" test with a test that drives the same input and asserts that something does happen.
  • Quit mid-test through the app's menu. app.terminate() is SIGTERM.
  • Terminal contents are not in the accessibility tree, so observe the state the product writes to disk. The UI test runner is sandboxed: it cannot reach your sockets, but it can read files.
  • Do not treat a HOME override as isolation. Isolate UserDefaults with a per-test suite, and isolate commands that run in a pane with explicit flags.
  • If test HOME directories live in a temp directory, do the arithmetic on your Unix socket paths against the 104-byte limit.
  • For any code that deliberately never deletes a file, check where it writes when a test calls it.
  • Give typeText only layout-invariant characters. Put long commands in a file, and switch the input source to ABC for the duration of the test.
  • If an agent runs your tests, enforce the no-concurrent-runs rule with a hook, and write "never quit the developer's app" into the agent's instructions as a step.

Having agents write tests makes the number of tests grow fast. The ten holes from section 2 were sitting inside that growing pile of green tests. What found them was not more tests; it was going back through the existing assertions one at a time and asking what each of them observes.

About Calyx

Calyx is a native macOS terminal (Swift 6, AppKit and SwiftUI, MIT) for running and supervising coding agents in parallel: an approval inbox for Claude Code and Codex permission prompts and for Grok (always-approve mode) and pi tool calls, an agent status sidebar driven by each CLI's own hooks, in-terminal diff review with line comments sent back to the agent, MCP tools that let agents read command output and language-server results, and the persistent sessions that came up throughout this post. It requires macOS 26 or later on Apple Silicon.

brew install --cask calyx
Enter fullscreen mode Exit fullscreen mode

Source: https://github.com/yuuichieguchi/Calyx
Docs: https://help.getcalyx.app

Top comments (0)