DEV Community

Cover image for FAIR Data for Agentic AI
N Chandra Prakash Reddy for AWS Community Builders

Posted on Originally published at devopstour.hashnode.dev

FAIR Data for Agentic AI

I got the wonderful opportunity to attend AWS Community Day Chennai on 7th March 2026. There were a ton of great sessions throughout the day, but one specific session absolutely transformed the way I thought about the future of artificial intelligence. The topic was “FAIR Data for Agentic AI” and the speaker was Naveena Ravi.

Whether you’re an experienced data engineer or just beginning your cloud adventure, knowing how to prepare your data for the next generation of AI is key.

The AI Transformation: From Predictions to Actions

Let's be honest, keeping up with AI has felt a bit like trying to sip from a firehose the past few months. To help us understand exactly where we are headed, Naveena opened her session by breaking down the growth of artificial intelligence into three separate phases.

Traditional AI was mostly about using current data to generate predictions and drive decisions. Think of it like your phone’s weather app, trying to figure out if it’s going to rain tomorrow based on previous data. Then came Generative AI that unleashed the power to generate and create brand new material from our text prompts. It’s like asking a chef to create a unique recipe for you based on your favorite ingredients.

And this is when it gets interesting. We are nearing the age of Agentic AI. Agentic AI is about performing action, rather than predicting an outcome or generating a block of text. These complex algorithms can think for themselves, strategize and really do things on your behalf.

Decoding AI Agents

You may be asking yourself, what does it mean for an AI to be “agentic”? "A clear definition was given by Naveena: AI Agents are semi or fully independent pieces of software. They have the unique ability to reason, plan, and act to achieve certain goals. And they can also work smoothly in digital and physical situations.

Think of your company’s database as a large disorganized library. A typical Generative AI model is like a speed reading assistance, if you give it a book, it will summarize the book for you. But an AI Agent is a proactive researcher. You tell it you need a report, and it scans the shelves, pulls the five most relevant-looking books, extracts the best quotes, turns it into a polished paper, and emails it to your employer.

The Building Blocks of an Agent

To enable Agentic AI to do this independent magic, four main components function in perfect balance:

  • LLM (Large Language Model): This is the main brain of the operation , allowing the system to understand human language and process complex logic .

  • AI Agent: The orchestrator that takes the user’s aim and turns it into steps, doing the thinking and planning.

  • RAG (Retrieval-Augmented Generation): The memory system draws in relevant business facts related to the scenario so the AI does not just guess the responses.

  • MCP (Model Context Protocol): The bridge or communication layer that allows the agent to connect securely to outside tools, applications and environments.

The Next Generation of Amazon SageMaker

These complex, interconnected systems need a very robust base for developers to build on. This is where Amazon SageMaker comes in. “The next generation of SageMaker is built to be the central hub for all of your data, analytics and AI needs.

If you’re developing a startup, stitching together a dozen different tools might be an operational headache. SageMaker fulfills this need with a Unified Studio architecture.

A Unified Architecture

The architectural Naveena gave was a great roadmap for today’s data teams. At the core level you have Open Lakehouse, sitting just underneath a vital layer for Data & AI Governance and on top of this safe base stands the Unified Studio, which divides your workflow into specialized toolsets:

  • SQL Analytics: Powered by tools such as Amazon Redshift and Amazon Athena to query your data.

  • Data Processing: Using Amazon EMR and AWS Glue to clean and process raw data.

  • Model Development: Build and train your algorithms powered by Amazon SageMaker AI.

  • Gen AI App Development: Built securely using Amazon Bedrock.

  • Future Capabilities: The ecosystem is also expanding with Streaming (Amazon MSK, Kinesis), Business Intelligence (Amazon QuickSight) and Search Analytics (Amazon OpenSearch Service) products coming soon.

Transforming Raw Data into AI Wisdom

Here’s the thing: An AI agent is only as smart as the data you feed it. But to really enable these agents to make the right choices, your business data needs to be run through a complete, step-by-step process to be transformed into actionable information.

The Data Pipeline Steps

  1. Data Ingestion: The process of getting raw, unstructured data into your system.

  2. Taxonomy / Ontology: Classifying and organizing the data in a way that clearly establishes linkages and hierarchies.

  3. Data Modelling – Schema Creation: Building the actual blueprints/formats of how the data is saved.

  4. Data Quality Management: Cleaning the data such that the information is correct, clean and free of errors.

  5. Data Catalog: Index everything to make the data easy to find for your teams.

The Knowledge Pyramid

To give a sense of this transformation, Naveena presented a stunning pyramid that showed the journey from raw inputs to complex AI.

At the very bottom base is raw Data, the first Data Ingestion phase. One rung higher, Data Processing turns those raw inputs into useful Information. Then at the Knowledge level, we have the Data Catalog which organizes the data for easy discovery. And last Wisdom where the AI Agents are. These AI Agents use all the underlying layers to act smartly.

The Core: F.A.I.R. Data Principles

To ensure your data can safely and securely reach that peak “Wisdom” stage, it must strictly follow the F.A.I.R. Data Principles. Is that you? F.A.I.R. stands for Findable, Accessible, Interoperable and Reusable.

Let's go over what this means for your data infrastructure.

1. Findable

For data to be useful, it needs to be easily discoverable by human workers and computer systems.

  • Organizations need to have the right governance to tightly control metadata.

  • The metadata must be significant, very detailed and continuously persistent across the whole business.

Consider this as an online e-commerce store. If a new pair of shoes is not labeled with the correct category, color and size description, no customer (or search engine algorithm) will be able to find it.

2. Accessible

When data is found it must be accessible in a safe manner, with no extra barriers.

  • To keep security, the actual data should be strictly accessible to authorized individuals only.

  • But the descriptive metadata needs to be available to relevant people and AI agents so they know what’s there.

  • The protocol for accessing this information must be open or easily recognized for conventional authentication and authorization procedures.

  • Crucially, metadata must be available even if the underlying data is removed or is no longer available.

3. Interoperable

Data should not be separate, disconnected silos.

  • Metadata shall be in standardized, FAIR certified formats.

  • This makes it easy to identify the data across numerous distinct systems and different AIs Agents.

  • The data itself must be able to easily cooperate with apps and workflows, allowing easy storage, processing and analysis.

Simply said, your data should have a common language. If your marketing software speaks French and your sales software speaks Japanese, your AI Agent won’t be able to help you. Compatibility means everyone speaks the same technological language.

4. Reusable

And finally, solid data is an asset that should continue to provide value over time.

  • Metadata must be extremely reusable so engineering teams may design complex, interconnected systems.

  • It should describe both the metadata and the actual data in a thorough and good way.

  • At this level of description, the information is easily copied or merged for fresh new use cases down the line.

Key Takeaways

The secret to building mind-blowingly intelligent, autonomous AI agents isn't picking the latest language model, it's really a data organization challenge.

At the end of the session, Naveena raised a vital question to the audience, "Is your DATA Agent Ready?" To be able to answer yes, with confidence, organizations need to heavily focus on these essential pillars:

  • Data Quality is non-negotiable: unavoidably bad data leads to unavoidably faulty, and even dangerous, agent decisions.

  • Strict Data Governance: You want to make sure security, compliance and adequate control of access are all locked down entirely.

  • Clear Data Lineage: You must know precisely where your data started and how it has evolved throughout its existence.

  • Embrace F.A.I.R. Principles: Only Findable, Accessible, Interoperable, and Reusable data can bridge the gap between fundamental information and actual AI wisdom.

Conclusion

After all, transitioning from standard AI to Agentic AI is a tremendous technological leap. We’re not simply having computers predict the future or write words anymore; we’re trusting them to develop plans and perform real-world activities for us.

To summarize, before unleashing autonomous AI to tackle your complicated business difficulties, you need to make sure your data house is clean, structured, and properly managed. If you’re leading a data team today, I highly recommend checking out the unified design of Amazon SageMaker, and start evaluating your own pipelines to see just how F.A.I.R. your data really is.

About the Author

As an AWS Community Builder, I enjoy sharing the things I've learned through my own experiences and events, and I like to help others on their path. If you found this helpful or have any questions, don't hesitate to get in touch! 🚀

🔗 Connect with me on LinkedIn

References

Event: AWS Community Day Chennai

Topic: FAIR Data for Agentic AI

Date: March 7, 2026

Also Published On

AWS Builder Center

Hashnode

Top comments (0)