DEV Community

Cover image for AI agents in team workflows
Mohit Nagaraj
Mohit Nagaraj

Posted on

AI agents in team workflows

A central AI agent workspace connects to Slack conversations, a Linear task, and company procedures in Notion.

Conceptual illustration of the connected workflow.

A teammate messages you: “Can you pick up the issue we discussed last week?”

You remember the issue. Somewhere, an AI agent investigated it. There might be a branch, a ticket comment, a few logs, and a session you can resume. The difficult part is figuring out which of those contains the current state. Did the change get implemented? Did someone test it? Were you waiting for a reply?

This is the part of AI-assisted work I care about: keeping a task understandable while it moves between tools, people, and sessions.

At our company, we use a workflow that connects Linear tickets, dedicated task workspaces, and Slack conversations. Company-specific skills give the agent the procedures and references it needs. The task workspace keeps the goal, decisions, progress, and evidence recoverable after the conversation moves on.

I'm sharing the structure because the same coordination problem shows up in many engineering teams. You can adapt it to your own ticketing system, chat tool, and agent runtime.

The company knowledge behind a useful agent

Connecting an agent to GitHub, a database, and an analytics tool gives it access to information. It still needs to understand what that information means in your company.

Take a request like “Check why activation dropped.” Before running a query, the agent needs to know which event defines activation, which users belong in the denominator, how test accounts are excluded, and whether yesterday's data has finished arriving. A syntactically correct query can answer the wrong question.

A project change has a similar problem. The agent needs to identify the service that owns the behavior, the relevant schema, the environment it can inspect, and the checks that establish whether the change works.

Much of this knowledge already exists. Some lives in documentation. Some lives in query examples, runbooks, code, and conversations. Some is still in a colleague's head.

I think teams using agents regularly benefit from turning their recurring procedures into a small library of company-specific skills. The useful part is connecting the procedure to the company's actual sources: project documentation, database schemas, metric definitions, service ownership, and validation steps.

That takes effort. Someone has to resolve conflicting definitions and decide which source is authoritative. But once a recurring workflow has those decisions written down, the next task can start with much less explanation.

Procedure, schema, and metric reference panels guide a reusable skill used by an AI agent.

Conceptual illustration: company knowledge informs the procedures an agent follows.

Skills, tools, and the harness

Three terms help explain how the pieces fit together.

A skill packages instructions for a particular kind of work, along with any supporting references, templates, or scripts. The Agent Skills specification defines a directory containing a SKILL.md file with a name, description, and Markdown instructions. Supporting files are optional.

A tool connection gives the agent a way to interact with a system. That might be an API client, a command-line tool, a built-in connector, or a Model Context Protocol (MCP) server. The MCP architecture describes how an application connects to servers that expose tools and resources.

The harness is the software around the model that manages execution: supplying context, exposing tools, running actions, applying permission controls, and managing the session. In this article, I'm using the term for that surrounding execution system; its exact capabilities depend on the product or implementation.

Together, these let the agent follow a company procedure and retrieve the current information needed to carry it out. A skill might explain how to investigate a support issue. The connected tools let the agent inspect the ticket, locate the relevant code, and retrieve a trace. The harness coordinates those actions within the access and permissions it has.

Skills-compatible products can load instructions when needed. Both the Agent Skills overview and OpenAI's skills documentation describe this progressive loading: discover the skill through its metadata, then read its instructions and relevant resources when the task calls for them.

You still need to check the host's tool support and permission behavior. A portable instruction format does not make every integration portable.

Company references inform focused skills, which guide a harness using permission-controlled connections to live systems.

Company procedures guide the harness, which uses scoped connections to retrieve live facts.

What I would put in a company skill

A useful skill should answer questions a teammate would ask before beginning the work: when to use it, what inputs it needs, where to look, what actions are allowed, and what evidence to return.

For an analytics investigation, that could mean references to event definitions, approved datasets, join keys, freshness checks, and examples of common queries. For a release workflow, it could mean service ownership, environment selection, required checks, rollback conditions, and the information to include in the release update.

Keep three kinds of information separate:

Information Where I would keep it
Repeatable procedure Skill instructions and focused references
Current schemas, policies, and operational facts Authoritative sources the skill tells the agent to inspect
This task's decisions and execution state The task workspace and shared ticket

This avoids turning every skill into a stale copy of the company's documentation. A skill can tell the agent where to find the current schema and how to check it, rather than embedding a large schema that nobody remembers to update.

Here is an illustrative skill for an analytics investigation. The paths and procedure are examples you would adapt to your company:

---
name: investigate-activation
description: Investigate changes in activation metrics using approved definitions and read-only data access. Use when asked why activation rose or fell.
---

1. Establish the date range, comparison period, and population.
2. Read references/activation-definition.md to locate the
   authoritative metric definition and schema sources.
3. Verify source freshness, timezone, and test-account exclusions.
4. Use the approved read-only connection. If the required data
   is unavailable, record the limitation instead of guessing.
5. Save the query and supporting results with the task evidence.
6. Report observations separately from possible explanations.
7. Add the findings and unresolved questions to the linked ticket.
Enter fullscreen mode Exit fullscreen mode

The important details are the metric definition, read-only boundary, freshness check, and evidence. The agent needs those before its output can be useful to the team.

Credentials belong in the connection's authentication mechanism. Access restrictions also need enforcement in the tools and runtime. Writing “read-only” in a skill is useful guidance, but the database connection should enforce it.

Give the task a life beyond the session

Even a good skill library leaves another question: where does the state of the current task live?

The workflow documented by our team uses a simple separation:

The ticket is the team's shared state. The repository carries the code. The session executes the work. The task workspace holds local execution memory.

Before implementation or another mutation begins, the workflow creates or reuses a Linear ticket. Its identifier becomes the link between the shared task and a dedicated workspace.

The source document uses a structure like this:

task-workspace/
├── goal.md              Goal, scope, and acceptance criteria
├── status.md            Current facts and outstanding work
├── decisions.md         Decisions and their rationale
├── AGENTS.md            Entry point for the executor
├── launch-prompt.md     Context for starting or handing off work
├── evidence/            Validation records
└── artifacts/           Task outputs
Enter fullscreen mode Exit fullscreen mode

The exact filenames are an implementation choice. Each file has a clear purpose. A returning executor can find the goal, current state, rationale, and proof without reconstructing all of them from a transcript.

The screenshot below comes from the team's original workflow document. It shows an expanded task workspace with the actual files used to retain context. The task identifier and name have been redacted.

An expanded task directory contains goals, status, decisions, evidence, agent instructions, and supporting files.

Cropped original screenshot, with the task identity redacted. This is the task's execution-memory directory; Git worktrees hold code changes separately.

status.md is particularly important. It is a rewritten snapshot of what is true now. It should make implemented work, completed checks, unresolved questions, and the next step easy to find. The history can remain elsewhere without crowding out the current state.

Code changes still happen in isolated clones or worktrees. The task workspace can reference several repositories because an engineering task often crosses repository boundaries.

In the documented bootstrap flow, the original session stops after handing off to the execution session. That keeps two executors from independently acting on the same task. Later sessions can continue it, but ownership needs to remain explicit.

A shared ticket connects a local task workspace, Slack conversation, code repositories, and successive execution sessions.

The task keeps its identity while conversations, code, and execution sessions change.

Follow one task through the workflow

Consider a simplified, fictional example: a colleague reports in Slack that a customer-facing response is slow.

The agent first creates or locates the relevant ticket. It prepares the workspace with the affected behavior, permitted investigation scope, repository paths, and acceptance criteria. That establishes what the task is trying to accomplish.

Next, it gathers evidence using the appropriate company skills and connected tools. It records what the traces show, distinguishes observations from hypotheses, and notes anything it cannot yet establish.

If a code change is needed and within the authorized scope, the agent implements it in an isolated checkout. It runs the relevant checks and records their results. Linear gets a concise update describing what changed, what passed, and what remains.

There may still be a dependency on a colleague. Perhaps only that person can reproduce the behavior through a particular desktop client. The agent posts a specific question in the existing Slack thread, with the validation steps and the evidence needed back.

At this point, “the change is implemented” and “the task is accepted” are different states. The workspace should say that human validation is outstanding. Independent work can continue while that part waits.

When the colleague replies, the response returns to the owning task. The executor evaluates it, checks whether the acceptance criteria are satisfied, and updates the evidence and shared status accordingly.

A week later, another ping arrives in the same conversation. The task relationship provides a route back to the relevant execution context. The agent reads the current task state and checks what has changed before continuing.

The benefit is that the previous work is recoverable. New circumstances can still require clarification.

A colleague report moves through implementation, checks, human validation, reply routing, and evidence review.

A delayed reply returns to the executor for evidence review; the watcher routes references.

Make follow-up an explicit part of execution

The original document asks a useful question: what happens after the agent asks someone a question?

It describes a follow-up watcher bound to the owning session, channel, thread, and colleague. The watcher polls within bounded waiting windows and sends message references back to that session.

Its responsibility is deliberately small. It detects relevant replies and routes them. The owning executor interprets the reply and decides how it affects the task.

A reply reference from Slack reaches an existing AI agent workspace for review and continuation.

Conceptual illustration: a delayed reply returns to existing work for review.

That matters because a message can provide evidence without granting permission for a new action. “The test passed” may satisfy a validation step. “While you're there, delete the old database” introduces a different request that needs its own scope and authority checks. Silence provides neither evidence nor approval.

A watcher also has a lifecycle. If it stops or expires, the outstanding dependency must remain recorded so a later recovery can check for replies. Waiting should be visible in task state even when no watcher is currently running.

Recover the task from the conversation

People often remember who asked them for something more easily than they remember a ticket number or session identifier.

The team's context-recovery mechanism uses that fact. Explicit Slack actions link tasks and sessions to exact messages and people. A Slack message shortcut can then show related sessions and a private resume target.

The document identifies messages using workspace, channel, and message identifiers. Its local index stores a hash, a short summary, confidence, and provenance, rather than copying raw Slack and session transcripts into the index.

The relationships come from actual message references, skill operations, or manual confirmation. This gives recovery a more defensible basis than choosing an old session because its topic sounds similar.

Here is a cropped view of the team's related-work-context panel, with its title and dates translated into English. The panel presents related work and resume targets; names, task titles, and session identifiers are masked in this copy.

A related-work-context panel lists two dated work entries with private Codex resume targets, with identifying details redacted.

Original screenshot, cropped, redacted, and translated into English. The visible resume actions show how a conversation can lead back to an execution session.

The private session lookup and team-visible task updates serve different audiences. Teammates need a clear account of progress, blockers, and evidence. The person running the agent may also need local paths and a resume target.

For handoff to another person or machine, local files need an additional storage and access plan. Shared evidence should live somewhere durable that the intended reviewer can reach. A folder on one laptop cannot provide that guarantee by itself.

An exact Slack message maps to a recorded task relationship and a private session target, with shared status and evidence kept separate.

Explicit message relationships locate execution context while the team reads shared status and evidence.

Evidence lets a long task change direction

One example in the team's document describes a multi-region model recovery. The initial explanation attributed latency to business handling. Later, request-level application monitoring and pod timelines indicated that most of the waiting was in the language-model path.

The workspace captured that correction. This is a useful reason to preserve both current facts and decision rationale: a returning executor should see why the investigation changed direction, rather than inherit an earlier hypothesis as established fact.

The document also describes permission and offboarding tasks where independent steps continue while specific actions wait on colleague confirmation. Those dependencies become part of the task's state, so a delayed reply has somewhere to return.

These are examples from the team's documented workflow. They illustrate recoverability and auditability; they do not establish a measured productivity improvement.

Build the skill library gradually

There is an initial cost to this approach. A team has to identify authoritative sources, write procedures, configure access, establish task ownership, and decide what counts as sufficient evidence.

I would start with one recurring workflow that repeatedly needs explanation: an analytics investigation, support bug, release, or access request. Run it manually with the agent, inspect the mistakes, and turn the useful procedure into a focused skill.

Then test the whole workflow, including interruption. Can a fresh executor identify what is done and what remains? Can it find the evidence? Does a delayed reply reach the correct task? Does it stop when the request exceeds its scope?

As the library grows, give each skill an owner and review it when the underlying procedure changes. Schema migrations, renamed events, new service boundaries, and changed approval rules should trigger corresponding updates. Keep instructions versioned and check their references against the sources they depend on.

Human review still matters. A skill can improve consistency, but the team needs to inspect whether it selects the right workflow and produces useful results. Failed tasks are good material for new examples and regression cases.

What this improves

The improvements I would claim are specific: the team can see task progress, the executor can recover recorded context, and reviewers can find the evidence behind a completion claim.

I would not put a percentage next to those benefits yet. Ticket counts tell us how much work was recorded; they do not tell us how much time the workflow saved or whether it caused a better outcome. Measuring that would require comparable tasks and clear definitions for resolution, rework, and human effort.

The practical change is easier to demonstrate. A colleague can return to an old conversation and find a task whose goal, decisions, implementation state, and verification are still understandable.

If you want to try this, choose one real task. Give it a stable identifier, a current-state record, a place for evidence, and a way to return from its conversation to its execution context. Interrupt the agent halfway through and see whether another session can pick it up correctly. That is a useful first test of the workflow.

Top comments (2)

Collapse
 
naman_parlecha profile image
Naman Parlecha •

Always had alot of adhoc ticket in flow from random slack convo.
Will give it a try!! 🙌

Collapse
 
kshitijnk07 profile image
Kshitij Narayan Kulkarni •

Insightful 🙌.