DEV Community

Cover image for MemoraX AI at NeurIPS 2026: Ten Papers on Reliable Learning, Efficient Reasoning, and Long-Term Agent Memory
MemoraX AI
MemoraX AI

Posted on

MemoraX AI at NeurIPS 2026: Ten Papers on Reliable Learning, Efficient Reasoning, and Long-Term Agent Memory

We are excited to share that 10 papers from MemoraX AI have been accepted to NeurIPS 2026.

These papers cover a range of topics, including robust self-training, reinforcement learning, difficulty-aware reasoning, diffusion models, long-term Agent Memory, and the continuous evolution of AI systems.

Although the papers address different technical problems, they share a common research question:

How can AI systems learn from experience, use computation more effectively, and continue improving over time?

Beyond Bigger Models

As AI systems become capable of handling longer and more complex tasks, simply increasing model size is no longer enough.

A capable AI system should also be able to:

  • recognize when a problem requires deeper reasoning;
  • learn from reliable experience rather than noisy feedback;
  • avoid repeating failed strategies;
  • align training improvements with real inference performance;
  • retain and reuse useful information over long-running tasks.

This perspective guides much of MemoraX AI’s research.

Building More Robust Learning Systems

One line of work focuses on improving the reliability of self-training for large language models.

In “Beyond Correctness: Robustness-Driven Evolutionary Self-Training for Large Language Models,” the authors study how correct solutions can be selected and evolved more reliably during training.

The central idea is that a useful solution should not only be correct once. It should also be robust, stable, and transferable to related problems.

The work introduces a robustness-driven evolutionary self-training framework that uses correctness feedback to guide the generation and selection of new solution trajectories. This helps the model learn from high-quality reasoning paths while reducing the risk of overfitting to fragile solutions.

Aligning Training with Real Inference

Another research direction studies the gap between training-time improvements and inference-time behavior.

In “The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning,” the authors examine a common problem in reinforcement learning for language models: improving the training policy does not always lead to better behavior when the model is actually deployed.

The paper introduces the idea of optimizing for monotonic inference policy improvement. Its proposed framework separates training-side updates from inference-side validation, allowing potentially harmful updates to be detected and rejected.

This line of research asks an important question:

Are we optimizing the metric used during training, or are we truly improving the behavior users experience at inference time?

Making Reasoning More Compute-Efficient

AI systems often spend the same amount of computation on problems with very different levels of difficulty.

“Does Your Large Language Model Have an Intuitive Sense of the Difficulty of a Question?” investigates whether a model can estimate problem difficulty before committing to a reasoning strategy.

The work proposes a difficulty-perception mechanism that helps models dynamically select different reasoning intensities. Easier problems can be handled with a faster path, while harder problems receive additional computation.

This direction is especially important for practical AI systems, where accuracy and efficiency must be balanced. Better difficulty awareness can improve token efficiency without simply reducing the quality of responses.

Exploring New Policy Models for Reinforcement Learning

MemoraX AI also studies how diffusion models can be used in online reinforcement learning.

“What Kind of Diffusion Models Do We Need in Online Reinforcement Learning?” explores the challenges of using diffusion policies in high-dimensional action spaces, including computational efficiency, expressiveness, and distribution coverage.

The paper introduces Coupled Flow, a structured approach designed to preserve expressive policy modeling while improving optimization efficiency.

This work contributes to a broader effort to understand how generative models can support more capable decision-making systems in complex environments.

Long-Term Agent Memory

One of MemoraX AI’s central research areas is long-term Agent Memory.

As AI Agents move from short interactions to longer-running tasks, memory becomes more than a convenience. It becomes part of the system’s ability to learn from experience.

An Agent working on a real software project may need to remember:

  • why a particular architecture was chosen;
  • which implementation paths have already failed;
  • how a previous bug was fixed;
  • which project rules are stable;
  • how the team prefers to structure and review code.

This is the motivation behind MemoraX Code, our long-term memory plugin for Coding Agents.

MemoraX Code helps Coding Agents retain and reuse project context, historical decisions, debugging experience, failed approaches, verified fixes, and reusable development workflows.

It is designed to work with Coding Agents such as Codex, Claude Code, OpenCode, Trae, and WorkBuddy. The goal is not to replace a developer’s preferred Agent, but to help the Agent maintain continuity across long-running development work.

From Research to Evaluation

In addition to building memory systems, MemoraX AI is also working on how Agent Memory should be evaluated.

Through the Agent Memory Leaderboard, we are exploring an open and reproducible evaluation framework for comparing memory systems across different scenarios.

These scenarios include:

  • long conversations and cross-session history;
  • coding memory and historical engineering experience;
  • multimodal memory;
  • factual recall and multi-hop reasoning;
  • temporal understanding;
  • personalization and rule following;
  • memory governance and safe usage.

A shared evaluation framework can help researchers and developers move beyond vague claims and better understand what different memory systems can actually do.

Toward Continuously Evolving AI

The ten accepted papers represent different parts of a larger research agenda.

Some focus on how models learn. Some study how they allocate computation. Others examine how policies improve during reinforcement learning or how Agents can retain useful experience over time.

Together, they point toward a broader direction for AI development:

Future AI systems should not only generate answers. They should learn from experience, understand when to reason more deeply, reuse reliable knowledge, and improve continuously through interaction.

At MemoraX AI, we are continuing to explore this direction through both fundamental research and practical systems such as MemoraX Code and Agent Memory Leaderboard.

We look forward to sharing more details about the accepted papers, released benchmarks, and open research projects in the coming weeks.

Learn More

ai #machinelearning #agents #reinforcementlearning #llm #opensource

Top comments (0)