In December, Maya's tokenized-equity position is down $3,140 on paper and she does nothing about it. Banking that loss properly means picking which of her 64 tax lots to sell (highest cost basis first, HIFO), then policing a 30-day calendar so a rebuy doesn't trip the wash-sale rule and void the deduction.
I built HarvestBot for the Arbitrum Open House Singapore buildathon: an agent that does that job on Robinhood Chain testnet. The contracts bound it: the lot selection is recomputed on chain, a rebuy inside the window reverts, and an off-map trade gets its bond slashed. The HIFO ledger is a Rust contract compiled to WASM with Arbitrum Stylus.
This post covers the one number I published, and the part of it I had to take back.
Links: live site · verify on chain in your browser, no wallet · 30-second judge path · repo · demo video, 2:39 · bench results
Why HIFO belongs in WASM
HIFO selection is a nested loop. Each pass scans every open lot for the highest basis-per-unit, takes it, and prorates the last lot when the sell quantity runs out partway through one. With 64 lots and 9 picks, that's 9 full scans, and every comparison is a cross-multiplication on 256-bit integers (basis_i · qty_b > basis_b · qty_i, so there's no division and no rounding drift).
Stylus runs WASM next to the EVM on Arbitrum chains, and WASM costs less per instruction than EVM opcodes. A loop full of comparisons seemed like the textbook case for it. Before writing any code, I estimated about 8× cheaper.
What the chain said: 2.87×
I kept an ABI-identical Solidity twin of the ledger. It has the same integer semantics and the same tie-break (lowest lot id wins). It serves as the benchmark control and as the fallback for chains without Stylus. I deployed both to Robinhood Chain testnet (chain 46630), seeded them with the same lots, and measured computeHarvest using Arbitrum's own NodeInterface.gasEstimateComponents. That call splits the L1 calldata share from L2 compute, and L2 compute is where the two engines actually differ.
At 64 lots it came out to 1,456,905 L2 gas for Solidity and 507,993 for Stylus: 2.87×. That's a lot less than 8×, but I reported what the chain measured. The 2.87× went into the README, the demo video, the OG images and an X thread.
The catch: my Solidity was reading storage in the loop
Two weeks later I put the two inner loops side by side. This is the Solidity twin I shipped (contracts/src/TaxLotLedger.sol):
Lot[] storage lots = _lots[owner][asset];
// ...
while (remaining > 0) {
// select the open, untaken lot with the highest basis/qty
uint256 best = type(uint256).max;
for (uint256 i = 0; i < n; i++) {
if (!lots[i].open || taken[i]) continue;
if (best == type(uint256).max) {
best = i;
continue;
}
// lots[i].basis/qty > lots[best].basis/qty ⇔ basis_i·qty_best > basis_best·qty_i
if (lots[i].costBasis * lots[best].qty > lots[best].costBasis * lots[i].qty) {
best = i;
}
}
// ...
}
lots is a storage pointer, so every lots[i].open, .costBasis and .qty is a storage read, repeated on each of the 9 selection passes. The Rust ledger (stylus/ledger/src/lib.rs) reads storage once and runs the rest in memory:
// one pass over storage → memory
let lots_s = self.lots.get(owner);
let lots_s = lots_s.get(asset);
let n = lots_s.len();
let mut mem: Vec<MemLot> = Vec::with_capacity(n);
for i in 0..n {
let lot = lots_s.get(i).expect("index < len");
mem.push(MemLot {
qty: lot.qty.get(),
basis: lot.cost_basis.get(),
open: lot.open.get(),
});
}
// ...
while !remaining.is_zero() {
let mut best: Option<usize> = None;
for (i, l) in mem.iter().enumerate() {
if !l.open || taken[i] {
continue;
}
match best {
None => best = Some(i),
Some(b) => {
// basis_i/qty_i > basis_b/qty_b ⇔ basis_i·qty_b > basis_b·qty_i
if l.basis * mem[b].qty > mem[b].basis * l.qty {
best = Some(i);
}
}
}
}
// ...
}
So I hadn't compared engines. I'd compared two data paths. Part of the 2.87× came from the Rust code having a better structure, and that structure isn't specific to WASM. Solidity can copy to memory too.
Building a fair baseline
I wrote a third ledger, contracts/src/bench/TaxLotLedgerMemCopy.sol, for the benchmark only. It uses the same selection rule, the same integer semantics and the same tie-break. The only change is the data path: one pass from storage into memory, and every later selection pass reads memory:
/// @dev One pass over storage → memory (the Rust `mem: Vec<MemLot>`).
function _load(Lot[] storage lots) internal view returns (MemLot[] memory mem) {
uint256 n = lots.length;
mem = new MemLot[](n);
for (uint256 i = 0; i < n; i++) {
Lot storage s = lots[i];
mem[i] = MemLot({qty: s.qty, basis: s.costBasis, open: s.open});
}
}
A Foundry differential test checks that it picks the same lots and returns the same realized loss as the shipped twin on random books, both before and after realize. I deployed it next to the other two on chain 46630 and reran all three engines together, with 50 eth_call samples per cell.
Three engines, one chain
L2 compute gas for computeHarvest (HIFO, up to 9 picks, 4 at 8 lots), from bench/RESULTS.md:
| Open lots | Solidity shipped | Solidity memory-copy | Stylus | Fair: memory-copy ÷ Stylus | Shipped ÷ Stylus |
|---|---|---|---|---|---|
| 8 | 132,363 | 108,713 | 125,807 | 0.86× | 1.05× |
| 16 | 352,160 | 231,323 | 179,982 | 1.29× | 1.96× |
| 32 | 720,396 | 440,005 | 288,999 | 1.52× | 2.49× |
| 64 | 1,456,891 | 857,559 | 507,979 | 1.69× | 2.87× |
| 128 | 2,929,924 | 1,693,678 | 945,964 | 1.79× | 3.1× |
All three engines return identical lotIds and realizedLoss at every size. The 64-lot row is Maya's scenario, exactly −$3,140.000000 (-3140000000 in 6-dp USD, the same figure forge test and cargo test pin).
The fair engine-to-engine figure is 1.69× at 64 lots, not 2.87×. About 41% of the ratio I first published (2.87× → 1.69×) came from my Solidity baseline re-reading storage, not from WASM. The 2.87× is still accurate for the Solidity twin I shipped. It just isn't the gap between the engines.
The repo's README, JUDGE.md and DEMO.md now give the fair figure next to every mention of 2.87×, and JUDGE.md has a dated correction entry.
Where Stylus still loses, and where it wins
At 8 lots, the memory-copy Solidity ledger beats Stylus (0.86×). The Stylus program is uncached on this testnet, so every call pays a fixed WASM initialisation cost, and at 8 lots there isn't enough loop to make up for it. Caching the program (cargo stylus cache bid) would remove that cost, but Robinhood Chain testnet has no ArbOS cache manager yet: ArbWasmCache.allCacheManagers() returns [] on chain 46630. So these are worst-case Stylus numbers.
Past the fixed cost, the fair ratio goes 1.29× → 1.52× → 1.69× → 1.79× as the book grows. All three engines pay the same cold storage read per lot. Stylus only wins on the comparison loop, and that loop grows with the number of lots times the number of picks. A tax-lot ledger does grow: Maya's 64 lots turn into hundreds after a few years of recurring buys. In this use case the advantage grows with the portfolio. For small books, I'd write it in Solidity.
Reproduce it
Read-only, no wallet. This asks the deployed Stylus ledger for the next HIFO picks:
cast call 0xEff7B46049fC677F58264e0ebb19dF1a39195a21 \
'computeHarvest(address,address,uint256,uint256)(uint64[],int256)' \
0x72cd3cB98A5d9B830b386EeBA7B2340132Ba557b 0x5884aD2f920c162CFBbACc88C9C51AA75eC09E02 638501157098660213 412300000 \
--rpc-url https://rpc.testnet.chain.robinhood.com
The full three-engine bench needs a funded testnet key, because it seeds the three ledgers:
git clone --recurse-submodules https://github.com/edycutjong/harvestbot && cd harvestbot
cd contracts && forge test # 91 tests, incl. the memory-copy differential
cd ../stylus/ledger && cargo test # 6 native tests on the WASM engine
cd ../.. && RPC=... PRIVATE_KEY=... python3 scripts/bench.py
# --components-only refreshes the gas split without re-seeding
The bench ledger addresses are in deployments/46630-bench.json, and the raw numbers, including p50/p95 latency, are in bench/results.json. Latency is eth_call wall-clock through the public RPC (1181.9–1215.2 ms p50 across every engine and size). It's bound by the network, not the engine, so I don't claim anything from it.
Honest limitations
- Testnet only. Robinhood Chain mainnet exists; I didn't deploy there (gas budget, unaudited code).
- Uncached Stylus. Caching would remove the init floor. I haven't measured how much of the 8-lot gap that closes, so I don't claim it.
- Mocks, labeled. The oracle mark, the swap venue and the bond's USDC are MOCK contracts with MOCK in their names. The assets are real Robinhood testnet Stock Tokens from the faucet. The bench ledgers use synthetic asset keys and never touch a token.
- One workload. This measures HIFO selection. A different access pattern (few comparisons, many writes) would give different numbers.
- The benchmark control is mine. Someone better at gas golfing could probably narrow the gap further. The memory-copy ledger is the best Solidity I wrote for it, not the best Solidity possible.
Takeaway
If you benchmark Stylus against Solidity, make sure both sides use the same data path first. My first number compared WASM plus memory against EVM plus repeated storage reads, and it overstated the engine gap by about 41%. The corrected number is smaller, but it measures only the engine.
Live site · /verify · /judge · Repo · Demo video · bench/RESULTS.md
If you've measured Stylus with a cached program, I'd like to see how your numbers compare.
Top comments (0)