DEV Community

#reinforcementlearning

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I Implemented the Algorithm Behind ChatGPT From Scratch - Day 8 (PPO).

I Implemented the Algorithm Behind ChatGPT From Scratch - Day 8 (PPO).

11
Comments
3 min read
How Early Digital Systems Quietly Shaped the Minds Building Tomorrow

How Early Digital Systems Quietly Shaped the Minds Building Tomorrow

Comments
5 min read
Decoding the Link Between Pretraining and Reinforcement Learning

Decoding the Link Between Pretraining and Reinforcement Learning

Comments
3 min read
Muon Optimizer Boosts Agentic Reinforcement Learning Performance

Muon Optimizer Boosts Agentic Reinforcement Learning Performance

Comments
3 min read
The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)

The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)

Comments
6 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Comments
3 min read
Agent Apprenticeship turns finished agent tasks into reusable experience

Agent Apprenticeship turns finished agent tasks into reusable experience

Comments
3 min read
I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).

I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).

5
Comments
4 min read
AI Agents Are Learning to Build the Worlds They Train In

AI Agents Are Learning to Build the Worlds They Train In

Comments 1
4 min read
Why teaching AI agents to use tools keeps blowing up in training

Why teaching AI agents to use tools keeps blowing up in training

Comments
3 min read
Building a Self-Optimizing Python Trading Bot with Reinforcement Learning and Binance API

Building a Self-Optimizing Python Trading Bot with Reinforcement Learning and Binance API

Comments
4 min read
The Whole Paper Fits in One Sigmoid: Implementing the SDAR Gate

The Whole Paper Fits in One Sigmoid: Implementing the SDAR Gate

Comments 1
5 min read
Four Models in One Training Loop: Architecting SDAR on AWS (Before Renting a Single GPU)

Four Models in One Training Loop: Architecting SDAR on AWS (Before Renting a Single GPU)

Comments
5 min read
How to Add Live Telemetry and Failure Diagnosis to Isaac Lab, MuJoCo, or Gazebo Training in Under 5 Minutes

How to Add Live Telemetry and Failure Diagnosis to Isaac Lab, MuJoCo, or Gazebo Training in Under 5 Minutes

Comments
4 min read
Why robotics RL training pipelines fail at scale

Why robotics RL training pipelines fail at scale

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.