I'm Alina, Head of Web Applications at EXANTE. I'm responsible for the technical strategy of our web projects, delivery efficiency and engineering processes across teams. I also lead several initiatives, including our AI transformation.
When we started working with AI, we didn't want to stop at individual developer tools. We wanted an autonomous pipeline that could take some tasks end to end, and we wanted to measure how much that improves developer productivity and shortens the time from task to finished result. The result is Codey: a multi-agent system of AI coding agents that independently covers much of the path from a Jira task to a working feature.
We started rolling it out six months ago. The numbers so far:
- 250 tasks have gone through the system
- 75% of them were completed successfully
- 7% of all tasks in the teams using Codey are now handled by it
- ~30 minutes is the typical time from the start of coding to a deployed feature environment
The main lesson: a single powerful model is not enough to make AI part of a real development process. You also need high-quality requirements, well-prepared projects, measurable checks and clear rules for how the system operates.
From individual automations to an AI software engineer
We started small: AI code review and analysis of affected modules based on the changed diff. These sped up individual stages after the code was written, but they couldn't complete the full cycle on their own, from analysing a Jira task and exploring the codebase to implementation and checks. I wrote about that stage in my previous article.
Codey was the next step. It explores the code, draws up a plan, writes the implementation and tests, checks the result and fixes the problems it finds, all on its own.
Codey currently works well with Node.js, React and Python projects. We began with local frontend and user interface (UI) tasks, then moved on to full-stack changes that touch both frontend and backend.
Today we use it across all projects in my area: from internal customer relationship management (CRM) and back-office systems to client-facing fintech services, including the client portal, the web trading terminal and the desktop application.
How Codey works: a multi-agent architecture
Architecturally, Codey is an in-house orchestration layer for AI agents. By an AI agent, we mean a large language model (LLM) running in a connected runtime environment, with a defined role and access to the tools it needs. We use three runtimes: Claude Agent SDK, Codex and OpenCode. They let agents work with repositories, files and the terminal.
Codey connects them into a single agentic workflow. It assigns roles and stages, runs checks, sends work back for revision and syncs the result with Jira, GitLab and continuous integration (CI).
How tasks reach Codey
Scenario 1: sprint planning
How this works depends on how far Codey has been introduced in a team. At the start, the team manually picks well-described tasks and labels them for the system.
In teams with an established process, Codey and developers share a common backlog. By default, Codey picks up suitable tasks, and the team applies an AI skip label to anything a person should handle for specific reasons. Over time, we plan to extend this to all connected teams and projects.
Scenario 2: Slack
The requester describes what they need. Codey helps them navigate the system, clarifies the details and creates a Jira task. It then completes the task, opens a merge request (MR) and returns to the same thread: it tags the requester, asks them to review the result and shares a link to the feature environment.
When Codey receives a task from Jira, it identifies the related GitLab repositories. Each project has predefined settings for:
- which repositories should receive the task
- which automated checks to run
- who to assign as reviewer
- how to deploy the feature environment and link the right frontend and backend versions
With these settings, Codey knows how to take a task through a specific project's process and prepare the result for the team's review.
AI agent roles in the system
Six core roles work on each task:
Researcher → Planner → Developer → Reviewer → QA → Verifier
- Researcher studies the requirements and the codebase
- Planner draws up a plan
- Developer writes code and tests
- Reviewer checks the quality and correctness of changes
- QA (quality assurance) runs tests and technical checks
- Verifier compares the result with the mandatory requirements in the plan
Different roles use different models. If Reviewer or the automated checks find a problem, the task goes back to Developer for rework.
Once the result is handed over, the team reviews it. Codey picks up comments on the code and business logic from the MR and the Jira task, makes changes and runs the checks again.
To stop a task from looping through revisions forever, Codey makes no more than three iterations. If comments are still unresolved, the task gets the AI failed status. The outcome is recorded in the statistics and feeds into the system's self-improvement loop. If needed, the run can be restarted manually.
Checks and error handling
Once an MR is created, Codey runs the project's quality gates: tests, linters, security checks and AI code review. If a problem can be fixed, the task goes back for rework. If a check can't run because of the project's configuration, Codey opens a Draft MR and lists the checks that weren't completed.
Most AI failed statuses come down to one of three causes:
- missing context
- changes across several interdependent repositories
- first tasks in a new project or technology stack
Team review of the result
For each MR, Codey identifies the affected modules and notes in Jira which parts of the product need regression testing. Once the checks pass, the frontend and backend are deployed to a separate feature environment (a preview environment), where the requester can check the whole feature end to end.
If the result meets expectations, the requester clicks Confirm. The task then goes through team review and into the release build. Codey works with a human in the loop: product acceptance, the final decision and responsibility for the release stay with people.
Business challenges of adopting AI coding agents
Quality of task definitions
Like any member of a development team, Codey can't do a task well without clear requirements. So we introduced Jira task templates with a mandatory description of the acceptance flow and expected checks, including requirements for unit and integration tests.
Team trust in Codey
We suggest starting with routine tasks or refactoring, which go through specialist review anyway. As successful results build up, teams move from manual selection to a shared backlog for Codey and developers, marking exceptions as AI skip. Trust grows gradually, and the system gets more context along the way.
A shift in workload to review and refinement
Early in adoption, part of the workload moves to these stages. A developer has to understand not only the task but also how Codey interpreted and implemented it. We reduce this with more precise task descriptions, detailed rules for each project and multi-level automated and human review.
Infrastructure costs
Separate backend feature environments cost more to run. To keep this under control, we limit environment lifetime and use reduced test databases.
Technical challenges of a multi-agent system
System versatility
Projects differ in architecture, codebase structure, technology stack and development rules. One set of agent instructions won't give equally good results everywhere.
So we maintain a profile for each project. It covers the architecture, codebase structure, build and test commands and local development rules. For recurring tasks, Codey has skills at three levels:
- global
- language-specific
- project-specific
The system generates new versions of skills from accumulated patterns. They are reviewed first and only then added to the pipeline. This approach to context engineering lets us adapt Codey's behaviour to each project.
Change quality
Reliable changes required projects with a simple, reproducible set-up, clear check commands and mature quality gates. Tests, linters, security checks and AI code review in the CI pipeline stop problematic changes before they add to the team's workload or reach production.
Choosing LLMs for different roles
We run regular benchmarks to evaluate our coding agents: different combinations of models, roles and tools run the same set of general-purpose tasks. We compare three things:
- solve rate: the share of solutions that passed the mandatory checks and received a score comparable to a human review
- speed of execution
- cost of execution
To keep evaluation independent, a result is never reviewed by a model from the same family.
The benchmarks showed that using the single most powerful model at every stage doesn't always give the best result. For some roles, a simpler model works better: it follows the plan more closely, overcomplicates the solution less often and finishes faster. So we regularly review how models are allocated across roles based on test results. After the latest round, we moved some roles to Codex models.
Continuous improvement
To improve Codey systematically, we use a Learning loop. It analyses logs and problematic sessions, identifies recurring errors and forms hypotheses. Each change is tested against the same benchmark and reaches production only after a measurable improvement.
In one cycle, we tested 20 hypotheses and kept nine. We also factor in feedback from developers and testers. In parallel, we optimise how the system works with code and context: RTK reduces the volume of command output, and an abstract syntax tree (AST) index speeds up the search for the right modules.
Next steps for Codey
Here is what we are working on next:
- Automated backlog analysis: the system will propose suitable tasks, and the team will exclude those that can't be handed over to AI
- A stronger Learning loop: continuous analysis of human comments in MRs, with that feedback used to improve the engine automatically
- Wider coverage: connecting new requesters and departments, projects, technology stacks and systems to Codey
- More sources of context: documentation, Jira and work communications, so that Codey can support analytical and product work on tasks, not only development
- Feature-based delivery: so that changes from Codey reach production faster after mandatory checks
Conclusion
Six months in, we have confirmed that an AI developer can be part of the working process. Connecting a powerful model to development is not enough on its own. It takes high-quality requirements, well-prepared projects, measurable checks and clear rules for how the system operates.
Our next step is to hand Codey the full cycle for some tasks, from task definition to production, while keeping the agreed rules and risk controls in place.
If you run coding agents in your own pipeline, I'd be interested to hear how you split roles and models in the comments.
Originally published on EXANTE Technology on Medium.

Top comments (0)