<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ahmed Khan</title>
    <description>The latest articles on DEV Community by Ahmed Khan (@ahmed_khan_6c5f55092f881b).</description>
    <link>https://dev.to/ahmed_khan_6c5f55092f881b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4061870%2Fa3b11883-6a31-494c-974a-65122647a5fa.jpg</url>
      <title>DEV Community: Ahmed Khan</title>
      <link>https://dev.to/ahmed_khan_6c5f55092f881b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ahmed_khan_6c5f55092f881b"/>
    <language>en</language>
    <item>
      <title>Look for Long Horizon Agents for frontier Labs, located in Mountain View (Remote)</title>
      <dc:creator>Ahmed Khan</dc:creator>
      <pubDate>Tue, 04 Aug 2026 06:58:08 +0000</pubDate>
      <link>https://dev.to/ahmed_khan_6c5f55092f881b/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote-3b5p</link>
      <guid>https://dev.to/ahmed_khan_6c5f55092f881b/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote-3b5p</guid>
      <description>&lt;p&gt;Bespoke Labs is looking for a researcher to help design and evaluate RL environments and benchmarks for long-horizon agentic tasks — the kind that take an agent hours, days, or weeks of coherent multi-step reasoning to complete, not single-turn prompts.&lt;/p&gt;

&lt;p&gt;What you'll do&lt;/p&gt;

&lt;p&gt;Design and build long-horizon RL environments and verifiers grounded in real-world tasks (code, tool-use, or enterprise workflows)&lt;br&gt;
Develop evaluation benchmarks that measure agent coherence, planning, and reliability over extended trajectories&lt;br&gt;
Analyze failure modes in long-horizon rollouts (drift, reward hacking, loss of task state) and propose fixes&lt;br&gt;
Collaborate with the broader team on open datasets and reproducible eval recipes&lt;/p&gt;

&lt;p&gt;Must-have (hard requirement)&lt;/p&gt;

&lt;p&gt;Demonstrated long-horizon agent/RL experience — this is non-negotiable. You should be able to point to specific work involving multi-step, multi-day, or sequential-reasoning agent systems (e.g., contributions to environments like SWE-bench, Vending-Bench, FrontierSWE, DeepSWE, OpenReward, Gymnasium, or equivalent original research/production work). Applications without concrete long-horizon evidence will not be considered.&lt;/p&gt;

&lt;p&gt;Strong Python; comfort with RL training/eval frameworks (e.g., Verifiers, Gymnasium-style APIs, or custom environment tooling)&lt;/p&gt;

&lt;p&gt;Track record of publishing or shipping work others can verify (GitHub, papers, benchmarks, or production systems)&lt;/p&gt;

&lt;p&gt;Nice to have&lt;/p&gt;

&lt;p&gt;Experience with reward-hacking detection or "fuzzy" quality verifiers beyond pass/fail correctness&lt;br&gt;
Background in multi-agent coordination or agent memory systems&lt;br&gt;
Prior contributions to open-source RL environment or agent-eval projects&lt;/p&gt;

&lt;p&gt;Logistics&lt;/p&gt;

&lt;p&gt;Type: Contract, remote&lt;br&gt;
Location: Remote (any timezone considered; some overlap with US/India hours preferred)&lt;br&gt;
Compensation: Based on experience — happy to discuss&lt;/p&gt;

&lt;p&gt;How to apply&lt;/p&gt;

&lt;p&gt;Send a short note plus links to your relevant long-horizon work (GitHub, papers, benchmarks, or production systems you've shipped) to [&lt;a href="https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5" rel="noopener noreferrer"&gt;https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5&lt;/a&gt;]. No long-horizon evidence, no need to apply — we will reject on this filter first.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
