DEV Community

Kaushikreddy Lingi
Kaushikreddy Lingi

Posted on Fully Autonomous

An Analytics Agent's Permissions Should Survive a Bad Prompt

An analytics agent receives a hostile instruction: return every customer's unmasked payment identifier.

The key question is what the system permits that agent to read—even if the model decides to follow the instruction.

That is the boundary I am exploring in MerchantLens, a merchant analytics lakehouse built with Databricks, Delta tables and Unity Catalog.

The demo uses synthetic payments data. There is no real cardholder data in the project.

Put the caller's identity in the query path

MerchantLens moves data through Bronze, Silver and Gold layers. An agent searches live catalog metadata and uses tools to query metrics or SQL.

The agent executes under its own service principal. Unity Catalog applies column masks and row filters for that principal before returning query results.

The README includes a comparison using the same SQL under two principals: the analyst sees a broader dataset, while the agent sees a restricted region and masked identifiers. The entitlement decision belongs to the query engine.

Separate reliability controls from authorization

The project has several layers:

Layer Purpose
Certified metrics Keep definitions and grain consistent
Application checks Constrain SQL, schemas, rows and tool steps
Unity Catalog policies Enforce identity-based row and column access
Audit records Record tool activity independently of the answer

The first two help keep the agent's behavior controlled. The catalog policies enforce data access.

This does not make the entire application immune to attack. Its protection depends on correctly configured identities, grants, masks and filters, and on every query using the intended principal.

Test the protected table

One practical lesson documented in the repository: checking a membership expression alone can give a misleading picture of what a mask does.

A better proof exercises the protected table under the actual identity. Ask what rows and values come back, then compare those results with the intended entitlement.

The repository also documents ambient grants on new service principals. Creating a new identity is not automatically a deny-by-default setup.

Metric definitions need boundaries too

MerchantLens attributes chargebacks to the month of the originating transaction. Using the date a chargeback was opened can shift the apparent incident across months because disputes arrive later.

That definition belongs in the semantic layer so a dashboard and an agent can use the same metric. Correct permissions do not compensate for an inconsistent denominator or time definition.

A review checklist

Before giving an agent a warehouse tool, I would ask:

  • Which principal executes its queries?
  • Can it retrieve restricted columns through another table or grant?
  • Are protected-table results tested under that principal?
  • Are tool calls logged somewhere the agent cannot modify?
  • Are metric definitions explicit and shared?

Where does your analytics agent's access control execute: in the prompt, the application, or the database?

Read the architecture and platform findings.

Drafted with AI from the public project documentation. Demo observations are repository-reported, not a new verification run.

Top comments (0)