Bespoke Labs is looking for a researcher to help design and evaluate RL environments and benchmarks for long-horizon agentic tasks — the kind that take an agent hours, days, or weeks of coherent multi-step reasoning to complete, not single-turn prompts.
What you'll do
Design and build long-horizon RL environments and verifiers grounded in real-world tasks (code, tool-use, or enterprise workflows)
Develop evaluation benchmarks that measure agent coherence, planning, and reliability over extended trajectories
Analyze failure modes in long-horizon rollouts (drift, reward hacking, loss of task state) and propose fixes
Collaborate with the broader team on open datasets and reproducible eval recipes
Must-have (hard requirement)
Demonstrated long-horizon agent/RL experience — this is non-negotiable. You should be able to point to specific work involving multi-step, multi-day, or sequential-reasoning agent systems (e.g., contributions to environments like SWE-bench, Vending-Bench, FrontierSWE, DeepSWE, OpenReward, Gymnasium, or equivalent original research/production work). Applications without concrete long-horizon evidence will not be considered.
Strong Python; comfort with RL training/eval frameworks (e.g., Verifiers, Gymnasium-style APIs, or custom environment tooling)
Track record of publishing or shipping work others can verify (GitHub, papers, benchmarks, or production systems)
Nice to have
Experience with reward-hacking detection or "fuzzy" quality verifiers beyond pass/fail correctness
Background in multi-agent coordination or agent memory systems
Prior contributions to open-source RL environment or agent-eval projects
Logistics
Type: Contract, remote
Location: Remote (any timezone considered; some overlap with US/India hours preferred)
Compensation: Based on experience — happy to discuss
How to apply
Send a short note plus links to your relevant long-horizon work (GitHub, papers, benchmarks, or production systems you've shipped) to [https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5]. No long-horizon evidence, no need to apply — we will reject on this filter first.
Top comments (0)