Foundry IQ: Inside the Managed Knowledge Layer That Turns RAG Into an Agent Tool Call
Ask any team that shipped a "chat with your docs" bot in 2024 what happened six months later, and you'll hear a familiar story: the chunking strategy needed retuning, the reranker was hand-rolled and brittle, permissions leaked because the vector index didn't respect SharePoint ACLs, and every new agent needed its own copy-pasted retrieval pipeline. Retrieval-Augmented Generation (RAG) was never really the hard part — productionizing RAG as a shared, governed, multi-tenant capability was.
Microsoft Foundry's answer to that problem is Foundry IQ, a managed knowledge layer released as the productized wrapper around Azure AI Search's agentic retrieval engine. It is one of the more quietly significant additions to the Foundry ecosystem this year, because it changes the unit of reuse in enterprise AI from "a RAG pipeline I built" to "a knowledge base I connect to N agents," with permission enforcement baked into the query path instead of bolted on after retrieval.
This article is a deep, implementation-level walkthrough of Foundry IQ: what it actually is, how the agentic retrieval pipeline works internally, how it plugs into Foundry Agent Service over MCP, what the security model really enforces (and doesn't), and where it breaks down at scale.
Table of Contents
- Why This Exists
- Core Concepts: Knowledge Sources, Knowledge Bases, Agentic Retrieval
- Architecture: How a Query Actually Flows
- The MCP Bridge Into Foundry Agent Service
- Implementation Walkthrough
- Permission Enforcement: What's Real and What's Marketing
- Real-World Developer Scenario: HR Policy Assistant Across Three Data Silos
- Production Considerations
- Cost Considerations
- Common Mistakes and Pitfalls
- Alternatives and Trade-offs
- Practical Recommendations
- Conclusion
- References
1. Why This Exists
A Foundry model — even the largest ones you can deploy — has a knowledge cutoff and zero awareness of your tenant's SharePoint sites, blob containers, Fabric lakehouses, or internal wikis. The standard fix is RAG: chunk your documents, embed them, index them in a vector store, retrieve the top-k chunks at query time, and stuff them into the prompt.
The problem isn't the concept — it's everything around it:
- Query complexity. A single dense-vector similarity search handles "what is our parental leave policy" fine. It falls over on "compare our parental leave policy in the US and Germany and tell me which team's leads are most affected by the difference" — a query that actually needs decomposition into sub-questions.
- Fragmentation across sources. Real enterprise knowledge is never in one place. It's in SharePoint, Blob Storage, a Fabric lakehouse, and sometimes it needs to come from the live web. Building one retrieval pipeline per source, per agent, doesn't scale organizationally.
- Permission leakage. If your index doesn't carry ACL metadata and your query path doesn't check it, your RAG bot becomes a permission-escalation vector — the classic "I asked the HR bot and it told me the CEO's salary" failure mode.
- Reuse. Ten different teams building ten different agents against the same underlying corporate knowledge shouldn't mean ten different embedding pipelines, ten different chunking strategies, and ten different bugs.
Foundry IQ addresses this by promoting retrieval from "a pipeline you write" to "a first-class, shareable resource" — the knowledge base — that sits on top of Azure AI Search's agentic retrieval engine and is consumable by any number of Foundry agents (or Microsoft Agent Framework apps, or Copilot Studio agents) via a standard protocol.
2. Core Concepts: Knowledge Sources, Knowledge Bases, Agentic Retrieval
Three objects matter here, and it's worth being precise about the layering because the docs use "Foundry IQ" and "agentic retrieval" almost interchangeably, which causes confusion.
Knowledge source
A knowledge source is a top-level Azure AI Search resource describing where content comes from and how it's queried. Knowledge sources are either:
- Indexed — Azure AI Search ingests the content ahead of time via an indexer pipeline (chunking, embedding generation, metadata extraction). Supported indexed kinds: Search index (wraps an existing index), Azure Blob, Azure SQL (preview), File (preview), OneLake, and Indexed SharePoint (preview).
- Remote — content is fetched live at query time, not pre-indexed. Supported remote kinds: Remote SharePoint (preview, uses the Copilot Retrieval API and enforces SharePoint's own permissions directly), Fabric Data Agent (preview), Fabric Ontology (preview), MCP server (preview — yes, a knowledge source can itself be a proxy to another MCP server), Work IQ (preview), and Web (via Bing).
This indexed-vs-remote split matters architecturally: indexed sources trade freshness for query speed and semantic reranking quality; remote sources trade some query-time latency and reduced reranking control for zero duplication of source-of-truth data and native enforcement of the origin system's permission model.
Knowledge base
A knowledge base is the orchestration object. It references one or more knowledge sources and holds the parameters that control retrieval behavior — most importantly the retrieval reasoning effort (minimal, low, or medium), which determines whether an LLM is used to plan/decompose the query before execution. Multiple agents can point at the same knowledge base. This is the reusable unit: build it once, govern it once, connect N agents to it.
Agentic retrieval
Agentic retrieval is the actual multi-query pipeline that a knowledge base executes when called. It is a genuinely distinct pattern from naive single-query vector search:
-
Query planning (skipped entirely at
minimaleffort): an LLM — an Azure OpenAI deployment you configure on the knowledge base — takes the user's query plus conversation history and decomposes it into a set of focused subqueries. This is where "compare parental leave in the US and Germany" becomes two or three separate, well-formed sub-questions instead of one blurry embedding. - Parallel query execution: every subquery runs concurrently against every configured knowledge source, using keyword, vector, or hybrid search as appropriate to each source.
- Semantic reranking: each subquery's results are reranked with Azure AI Search's L2 semantic reranker to surface the truly relevant matches, not just the nearest-neighbor matches.
- Result synthesis: everything is merged into a unified response. You always get merged extractive content; source references and an execution activity log are optional, and — in preview — full natural-language answer synthesis (an LLM writes the final grounded answer with citations, rather than the caller having to do that step itself).
The important design decision here: agentic retrieval returns grounding data, not necessarily a final answer. Whether you consume it as raw extractive passages (GA path) or ask it to synthesize a natural-language answer (preview path) is your choice, made per knowledge base configuration.
3. Architecture: How a Query Actually Flows
At a component level:
| Component | Owning service | Role |
|---|---|---|
| Knowledge base | Azure AI Search | Orchestrates the pipeline; owns query parameters and reasoning effort |
| Knowledge source(s) | Azure AI Search | Define what content is queried and how |
| Search index | Azure AI Search | Backing store for indexed sources; holds text + vectors + semantic config |
| Semantic ranker | Azure AI Search | L2 reranking of subquery results |
| LLM | Azure OpenAI (via Foundry Models) | Powers query planning, web-result summarization, and answer synthesis |
| Foundry Agent Service | Microsoft Foundry | Consumes the knowledge base as an MCP tool from a PromptAgentDefinition
|
Note what's not in this list: there is no separate "Foundry IQ service" runtime. Foundry IQ is the productized, governed front door — the naming and portal experience layer — over Azure AI Search's agentic retrieval, surfaced inside the Microsoft Foundry portal and consumable through Foundry Agent Service. This matters operationally: your quotas, region availability, and REST API versioning all live under Azure AI Search, not under a separate Foundry billing meter.
Two API generations you must not mix
As of the 2026-04-01 GA REST API, Azure AI Search supports agentic retrieval for GA knowledge source types with minimal reasoning effort only (i.e., no LLM-based query planning, extractive results only). The 2026-08-01-preview REST API version unlocks preview knowledge source types (SharePoint, Fabric, MCP-as-source, Web), non-minimal reasoning effort (LLM query planning), answer synthesis, and multi-turn message arrays.
The Microsoft Foundry portal and Azure portal currently only expose the preview surface — meaning anything you wire up through the portal UI may need a deliberate migration pass before it's a supportable GA production configuration. If you're building for production today, decide explicitly which REST API version you're targeting rather than letting the portal default you into preview schemas you didn't intend to depend on.
4. The MCP Bridge Into Foundry Agent Service
This is the part developers most need to internalize: Foundry Agent Service talks to a Foundry IQ knowledge base exclusively through the Model Context Protocol. The knowledge base itself exposes an MCP endpoint:
{search_service_endpoint}/knowledgebases/{knowledge_base_name}/mcp?api-version=2026-08-01-preview
That endpoint exposes exactly one MCP tool today: knowledge_base_retrieve. Your agent's PromptAgentDefinition gets an MCPTool pointed at that endpoint via a project connection — a RemoteTool connection category with ProjectManagedIdentity auth, which is specific to Foundry project connections and lets the project's system-assigned managed identity authenticate to Azure AI Search without you juggling API keys in agent config.
This design has a consequence worth calling out explicitly: the knowledge base is not a Foundry-native resource — it's a remote tool the agent calls over network protocol. That means:
- Every retrieval is a tool-call round trip with MCP framing overhead, not an in-process function call.
- The agent's own tracing shows it as a tool invocation (
mcp_approval_request/ tool-call spans), which is good for observability but means retrieval latency shows up as tool latency in your traces, not model latency — budget your P95 SLAs accordingly. - Because it's a generic MCP integration, the same knowledge base can be consumed by Microsoft Agent Framework code, a Copilot Studio agent, or any custom MCP client — not just Foundry Agent Service. That's the actual "share one knowledge base across many surfaces" story made concrete.
5. Implementation Walkthrough
Here's the real, end-to-end path: create the connection, then create the agent, then call it. This mirrors the officially supported pattern (Python SDK ≥ 2.0.0, REST API 2026-08-01-preview for the knowledge base MCP endpoint, 2025-10-01-preview for the ARM connection).
5.1 — Create the project connection to the knowledge base's MCP endpoint
# create_kb_connection.py
# Production pattern: creates (or updates) an ARM connection on a Foundry project
# that points at an Azure AI Search knowledge base's MCP endpoint.
import requests
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
credential = DefaultAzureCredential()
# ARM resource ID of the Foundry project (Microsoft.CognitiveServices/accounts/.../projects/...)
project_resource_id = (
"/subscriptions/<sub-id>/resourceGroups/<rg>/providers/"
"Microsoft.CognitiveServices/accounts/<account>/projects/<project>"
)
project_connection_name = "hr-kb-mcp-connection"
# The knowledge base's own MCP surface — this is what the agent will call
mcp_endpoint = (
"https://hr-search-svc.search.windows.net/knowledgebases/hr-policy-kb/mcp"
"?api-version=2026-08-01-preview"
)
# Token scoped to Azure Resource Manager, not to Azure AI Search itself —
# we're calling the ARM control plane to *create* the connection object.
bearer_token_provider = get_bearer_token_provider(
credential, "https://management.azure.com/.default"
)
headers = {"Authorization": f"Bearer {bearer_token_provider()}"}
response = requests.put(
f"https://management.azure.com{project_resource_id}/connections/{project_connection_name}"
"?api-version=2025-10-01-preview",
headers=headers,
json={
"name": project_connection_name,
"type": "Microsoft.MachineLearningServices/workspaces/connections",
"properties": {
# ProjectManagedIdentity + RemoteTool are specific to this scenario:
# the project's own managed identity authenticates to Search at
# call time — no API keys stored in the connection.
"authType": "ProjectManagedIdentity",
"category": "RemoteTool",
"target": mcp_endpoint,
"isSharedToAll": True,
"audience": "https://search.azure.com/",
"metadata": {"ApiType": "Azure"},
},
},
)
response.raise_for_status()
print(f"Connection '{project_connection_name}' created or updated successfully.")
Before this call succeeds in a real tenant, three RBAC assignments have to be in place:
- Foundry Project Manager on the project's parent resource — needed to create the connection.
- Search Index Data Reader (and Search Index Data Contributor if the agent writes back) for the project's managed identity, granted on the Azure AI Search service.
- Cognitive Services User for the search service's own system-assigned managed identity on the Foundry account — only required if the knowledge base has an LLM configured for query planning or answer synthesis, since Search calls back into Azure OpenAI using its own identity.
That third one trips people up: it's Azure AI Search's identity, not the agent's, that needs the model-calling permission, because Search is the thing invoking the LLM mid-pipeline for query planning — the agent never sees that intermediate call.
5.2 — Create the agent with the knowledge base as an MCP tool
# create_hr_agent.py
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import PromptAgentDefinition, MCPTool
from azure.identity import DefaultAzureCredential
credential = DefaultAzureCredential()
project_endpoint = "https://hr-foundry.services.ai.azure.com/api/projects/hr-project"
mcp_endpoint = (
"https://hr-search-svc.search.windows.net/knowledgebases/hr-policy-kb/mcp"
"?api-version=2026-08-01-preview"
)
project_connection_name = "hr-kb-mcp-connection" # from step 5.1
project_client = AIProjectClient(endpoint=project_endpoint, credential=credential)
# Instructions are doing real work here, not boilerplate: they force tool
# invocation on every turn and mandate the citation annotation format,
# which is what makes the retrieved grounding data auditable downstream.
instructions = """
You are a helpful HR assistant that must use the knowledge base to answer all
questions from the user. You must never answer from your own knowledge under
any circumstances.
Every answer must include citations for the knowledge base sources you used,
rendered as: [message_idx:search_idx | source_name]
If the knowledge base does not contain the answer, respond with "I don't know".
"""
mcp_kb_tool = MCPTool(
server_label="knowledge-base",
server_url=mcp_endpoint,
require_approval="never", # skip human-in-the-loop approval per call
allowed_tools=["knowledge_base_retrieve"], # explicit allow-list — the only tool exposed anyway
project_connection_id=project_connection_name,
)
agent = project_client.agents.create_version(
agent_name="hr-policy-assistant",
definition=PromptAgentDefinition(
model="gpt-4.1-mini",
instructions=instructions,
tools=[mcp_kb_tool],
),
)
print(f"Agent '{agent.name}' version {agent.version} created.")
A few implementation details worth flagging:
-
require_approval="never"is a real security decision, not a formality. Because the knowledge base's only exposed tool is a read-only retrieval call, blanket auto-approval is defensible here — but if you later add a knowledge source or MCP-server-as-source configuration that can trigger side effects (e.g., a Fabric Data Agent that can kick off a query job with cost implications), revisit this. -
allowed_tools=["knowledge_base_retrieve"]looks redundant since that's the only tool the endpoint exposes today, but it's cheap insurance against a future server-side addition silently expanding your agent's capability surface. - The instructions block is functionally part of your security and quality posture, not just prompt-craft. Microsoft's own guidance explicitly frames the citation-format instruction as improving MCP tool invocation rates — i.e., a badly worded system prompt measurably reduces how often the agent actually bothers to call the knowledge base instead of hallucinating from parametric memory.
5.3 — Calling the agent (unchanged from any other Foundry agent)
from azure.ai.projects.models import ResponsesHostServer # illustrative import path
response = project_client.agents.responses.create(
agent_name="hr-policy-assistant",
input="How does parental leave differ between our US and Germany offices, "
"and which of our team leads have direct reports in Germany?",
)
print(response.output_text)
Under the hood, this single call triggers: query decomposition into "US parental leave policy" and "Germany parental leave policy" and "team leads with direct reports in Germany" sub-questions (assuming low/medium reasoning effort), three parallel hybrid searches possibly across two different knowledge sources (an HR policy blob index and an org-chart SQL-backed index), semantic reranking of each, and a synthesized, citation-bearing answer handed back to the agent's model to finalize.
6. Permission Enforcement: What's Real and What's Marketing
This is the section worth reading twice before you put Foundry IQ in front of sensitive data.
The permission story has two genuinely different enforcement mechanisms depending on knowledge source type, and conflating them is a common and dangerous mistake:
For indexed sources (Blob, SQL, OneLake, Indexed SharePoint): permission enforcement requires you to have synchronized access control list (ACL) metadata fields into your search index yourself, and to pass the calling user's identity via the x-ms-query-source-authorization header at query time so Search can filter results per-caller. This is query-time RBAC/ACL enforcement, and as of this writing it is itself a preview capability layered on top of the base agentic retrieval feature. If you don't populate those permission fields and don't forward that header, your index has zero awareness of who's asking — it will happily surface a document to anyone whose query is a good semantic match, regardless of whether that person should be able to read it.
For remote SharePoint sources specifically: content isn't indexed into Azure AI Search at all. Instead, the same authorization header is forwarded to SharePoint's own Copilot Retrieval API, and SharePoint enforces its native permission model directly at query time. No duplicated data, no duplicated ACL sync job, no drift between the source system's permissions and a stale index copy.
The practical takeaway: "Foundry IQ enforces permissions" is true only if you built the plumbing for it. The platform gives you the mechanism — header propagation, ACL metadata schema support, Purview sensitivity label honoring for supported sources — but it does not retroactively secure an index you built without permission metadata. Teams migrating an existing, permission-naive Azure AI Search index into a Foundry IQ knowledge source need to treat ACL backfill as a hard blocking prerequisite, not a nice-to-have.
There's a second, more subtle risk: who runs the query planning LLM call, and against what? Query decomposition sends the user's raw query and conversation history to the configured Azure OpenAI model. If your conversation history contains sensitive context from a prior turn (say, a previous answer that quoted a restricted document), that content is now part of the payload sent to the planning model, which may sit in a different resource/network boundary than the eventual retrieval target. Model input/output logging and data residency policy on that Azure OpenAI deployment therefore becomes part of your knowledge base's overall data-handling boundary — not an implementation detail you can ignore because "it's just doing retrieval."
7. Real-World Developer Scenario: HR Policy Assistant Across Three Data Silos
Consider a mid-size enterprise with HR policy PDFs in SharePoint, a structured headcount/org-chart table in Azure SQL, and a benefits FAQ maintained as a Fabric lakehouse table. Before Foundry IQ, three separate retrieval pipelines, three separate embedding refresh jobs, and three separate places for permissions to silently diverge from the source systems.
With Foundry IQ:
- Three knowledge sources are created — an Indexed SharePoint source (preview) for the policy PDFs, an Azure SQL source (preview) for the org chart, and... at time of writing there's no native Fabric lakehouse knowledge source distinct from OneLake, so the benefits FAQ, if it lives in OneLake-backed storage, is added as a OneLake knowledge source; if it needs live Fabric semantics, it becomes a Fabric Data Agent remote source instead.
- A single knowledge base,
hr-policy-kb, references all of them, with reasoning effort set tolowso multi-part questions get decomposed. - Two separate agents — an internal HR-assistant chat agent and a manager-facing "org insights" agent — both connect to the same knowledge base via their own MCP tool + project connection. Neither team re-implements retrieval; both inherit whatever reranking and permission-enforcement improvements the platform team makes to the shared knowledge base later.
- When a manager asks "who on my team is eligible for extended parental leave under the German policy," the pipeline decomposes into a policy-lookup subquery (hits the SharePoint source) and an org-chart subquery (hits the SQL source), executes both in parallel, reranks each, and returns merged, cited grounding data that the agent's model turns into a coherent answer — while the ACL header ensures the manager only sees direct reports they're actually authorized to see in the org data.
This is the actual value proposition in concrete terms: not "better RAG," but organizational reuse of a governed retrieval capability across otherwise-independent agent teams.
8. Production Considerations
-
Pin your REST API version deliberately. Portal-created objects default to preview schemas. Decide up front whether you're shipping against
2026-04-01GA (stability, butminimal-effort/extractive-only, GA source types only) or2026-08-01-preview(LLM query planning, answer synthesis, preview sources) and document the migration path before you have production traffic depending on preview-only behavior. -
Latency budgeting. Agentic retrieval is explicitly slower than a single-query pipeline by design — query planning adds an LLM round trip, and multiple subqueries each get semantically reranked. Measure P95/P99 tool-call latency in your agent traces (this shows up as MCP tool latency, not model latency) and set reasoning effort (
minimalvslowvsmedium) based on your actual SLA, not just answer quality. - Hub-based projects are not supported. Foundry IQ's MCP integration requires a standard (non-hub) Foundry project. If you're still running older hub-based projects, this is a forcing function to migrate, not a footnote.
- Region availability is Azure AI Search's, not a separate Foundry footprint. Agentic retrieval is only available in select Azure AI Search regions — check regional availability before you assume a knowledge base can be co-located with every Foundry project region.
- Managed identity hygiene. Both directions need identities: the project's managed identity needs Search Index Data Reader/Contributor, and Search's own managed identity needs Cognitive Services User on the Foundry account when an LLM is configured on the knowledge base. Missing either produces confusing partial failures — retrieval works but planning/synthesis silently degrades, or vice versa.
- Test the citation contract, not just the answer. Because instructions drive whether citations render correctly and whether the tool gets invoked at every turn, treat your agent instructions as a versioned artifact with regression tests (does it call the tool on ambiguous queries? does it correctly say "I don't know" when the knowledge base returns nothing?), not a one-time prompt you wrote once and forgot.
9. Cost Considerations
Foundry IQ has no independent billing meter — costs roll up through the underlying services it composes:
- Azure AI Search billing for the service tier (which determines available compute for indexing and semantic ranking), storage, and — where applicable — the semantic ranker feature itself.
- Indexer runs for indexed knowledge sources: initial ingestion plus every scheduled incremental refresh consumes indexing compute; chunking, embedding generation, and metadata extraction all happen during these runs.
- Azure OpenAI token consumption for query planning and answer synthesis calls made by the knowledge base — this is in addition to the tokens your agent's own model consumes, and it's easy to undercount because it doesn't show up in the agent's own model deployment metrics; it bills against whatever Azure OpenAI deployment the knowledge base itself is configured to call.
- Remote source query-time cost — e.g., calls into SharePoint's Copilot Retrieval API or a Fabric Data Agent invocation — which may carry their own service-specific charges or throttling limits separate from Azure AI Search.
The practical implication: a single user question against a medium-reasoning-effort knowledge base with three knowledge sources can trigger one planning LLM call, three-plus parallel search queries, three reranking passes, and one synthesis LLM call — all before your agent's own model ever generates a token. Budget and monitor this as a distinct cost center from your agent's model spend, and prefer minimal or low reasoning effort for high-volume, latency- and cost-sensitive endpoints where query complexity doesn't warrant full decomposition.
10. Common Mistakes and Pitfalls
- Assuming ACL sync is automatic for indexed sources. It isn't — you must design your indexer pipeline to write permission metadata fields, and your query path must forward the caller's identity. Skipping this silently turns your "permission-aware" knowledge base into a permission-blind one that merely looks secure in a demo.
-
Mixing GA and preview objects without a migration plan. Building entirely through the portal (preview-only) and then discovering your production REST API pin (
2026-04-01) can't read those objects as configured. -
Forgetting the search service's own identity needs Azure OpenAI access. This produces the confusing failure mode where retrieval works fine but query planning or answer synthesis silently falls back to (or fails on)
minimalbehavior. -
Treating
require_approval="never"as a universal default without re-evaluating it as you add knowledge sources that might have side effects (billable external calls, write-capable remote MCP sources). -
Under-specifying agent instructions, resulting in the model answering from parametric memory instead of invoking
knowledge_base_retrieve, especially on questions the model "thinks" it already knows the answer to — exactly the class of question where a stale or wrong parametric answer is most dangerous. - Ignoring conversation-history leakage into the planning LLM. If earlier turns contain sensitive retrieved content, that content re-enters the pipeline as planning input on every subsequent turn, an easy oversight in multi-turn agents with long-lived sessions.
- Sizing one knowledge base for a single team's schema and then reusing it broadly without validating that unrelated agents' typical queries are well served by the same reasoning effort and source mix — a knowledge base optimized for short internal support queries won't necessarily handle a legal-research agent's much longer, more open-ended prompts well.
11. Alternatives and Trade-offs
| Approach | When it makes sense | Trade-off vs. Foundry IQ |
|---|---|---|
| Hand-rolled RAG pipeline (your own chunker, embedder, vector store, reranker) | Highly specialized domains (e.g., genomics, legal citations) where generic chunking/embedding underperforms, or when you need a non-Azure vector store | Full control, but you own query decomposition, reranking, ACL enforcement, and multi-source orchestration yourself — and you rebuild it per agent unless you invest separately in your own shared-service layer |
| Foundry Agent Optimizer prompt/tool tuning alone (no retrieval layer) | Small, static knowledge bases that fit comfortably in a system prompt or a single small index | Doesn't scale past a few documents; no dynamic multi-source retrieval, no citation infrastructure |
| Direct Azure AI Search agentic retrieval without Foundry IQ framing | Non-Foundry applications, or when you need the GA 2026-04-01 REST API surface without any Foundry-specific portal/MCP conventions |
Same underlying engine, but you lose the Foundry-portal knowledge-base authoring UX and the standardized MCP tool contract for Foundry agents specifically |
| Vendor RAG-as-a-service platforms outside Azure | Multi-cloud strategies or existing investment in another vector database ecosystem | Loses native ACL/Purview integration and native MCP wiring into Foundry Agent Service; you're back to building your own bridge |
The honest framing: Foundry IQ doesn't introduce a fundamentally new retrieval algorithm — multi-query decomposition, parallel hybrid search, and semantic reranking are all patterns you could implement yourself against Azure AI Search directly, or against any vector database. What it buys you is standardization and reuse: one governed object, one permission model, one MCP contract, consumable by every agent in your tenant instead of reinvented per team.
12. Practical Recommendations
- Start with
minimalreasoning effort and a single indexed knowledge source to validate your citation and instruction contract before turning on LLM-based query planning — it's easier to debug decomposition quality once basic retrieval and citation rendering are proven correct. - Treat ACL metadata design as part of your index schema from day one if there is any chance the underlying content is permission-sensitive; retrofitting ACL fields into a live production index is significantly more painful than designing for it upfront.
- Pin a REST API version in code (don't rely on portal defaults) and write an explicit migration note in your repo referencing the official migration guidance before you go to production.
- Instrument tool-call latency for
knowledge_base_retrieveseparately from your agent's model latency in your OpenTelemetry traces — this is the single most useful signal for deciding whether to lower reasoning effort or split an overloaded knowledge base into more targeted ones. - Version your agent instructions alongside your code, and add regression tests that assert the agent invokes the retrieval tool on a fixed set of canary questions — instruction drift is a silent failure mode that won't show up until users notice hallucinated answers.
13. Conclusion
Foundry IQ is best understood not as a new retrieval technology but as Microsoft formalizing agentic retrieval — Azure AI Search's multi-query, LLM-assisted retrieval pipeline — into a reusable, governed, MCP-addressable resource inside the Foundry ecosystem. The technical substance (query planning, parallel hybrid search, semantic reranking, optional answer synthesis) has existed in Azure AI Search independently; what Foundry IQ adds is the organizational contract: one knowledge base, many agents, one place to reason about permissions, cost, and quality.
The permission story is genuinely good when built correctly — ACL sync plus header propagation plus Purview label honoring is a real, defensible security model — but it is opt-in machinery, not a default you get for free by pointing an agent at an index. Teams that skip the ACL and header plumbing get a bot that looks secure and isn't. Teams that respect the preview/GA API boundary, budget for the extra LLM calls query planning and synthesis introduce, and treat instructions as a tested contract rather than a one-off prompt will get real leverage: a knowledge layer that multiple agent teams can build on without each reinventing retrieval from scratch.
If you're currently maintaining more than one hand-rolled RAG pipeline against overlapping enterprise content inside the same tenant, that's the strongest signal that Foundry IQ's reuse model is worth the migration effort.
14. References
- What is Foundry IQ? — Microsoft Learn
- Agentic Retrieval Overview — Azure AI Search
- What is a Knowledge Source? — Azure AI Search
- Connect Agents to Foundry IQ Knowledge Bases — Microsoft Foundry
- Create a Knowledge Base — Azure AI Search
- Migrate Agentic Retrieval Code to the Latest Version
- Query-time ACL and RBAC Enforcement (preview)
- Sample: agentic-retrieval-pipeline-example (Azure-Samples/azure-search-python-samples)
(Note: Some figures and API version references above reflect documentation current as of late September 2026 and reference preview features that may change before general availability — verify against current Microsoft Learn documentation before building production systems.)


Top comments (0)