DEV Community

Georg
Georg

Posted on Fully Autonomous

Who gains access when this Terraform plan is applied?

reachdiff live run: BLOCK, gp-interns gains READ on the hr schema

A pull request that changes Unity Catalog grants usually looks harmless. One line, one group, one privilege:

grant {
  principal  = "gp-interns"
  privileges = ["SELECT", "USE_SCHEMA"]
}
Enter fullscreen mode Exit fullscreen mode

terraform plan tells you that databricks_grants.hr will be updated in place. It does not tell you who can read what afterwards. That depends on things the diff doesn't show: who is in the group (and in the groups inside it), which usage privileges already exist further up, who owns what, and which other resources touch the same objects.

I wanted the review question answered directly: who gains or loses access to what when this plan is applied? So I built reachdiff, a small open-source CLI. It reads terraform show -json, works out the effective-access diff, shows the route behind each change and returns an exit code a pipeline can act on.

What it reports

For the change above, run against a test workspace:

$ reachdiff plan --tfplan plan.json --profile PROFILE
reachdiff: BLOCK (live mode)
1 gained, 0 lost, 4 findings

+ gp-interns (1 member, 1 new, 0 already had it)  READ  main.hr (all tables)
    via SELECT on main.hr + USE_CATALOG on main + USE_SCHEMA on main.hr

Findings:
  BLOCK grant_resource_conflict main.hr.salaries: databricks_grants and databricks_grant both manage it; the applied
        result depends on apply order, and grants can be lost without an error
  WARN sensitive_access gp-interns READ main.hr: gp-interns gains READ on sensitive data in main.hr
  WARN unverified latent_grants: gp-interns gains USE_SCHEMA on main.hr; this may activate existing table-level grants
        that were not enumerated (use --deep)
Enter fullscreen mode Exit fullscreen mode

Four things happen here that the plan diff doesn't show:

  • The gain is expressed as access, not as privileges. READ on the whole schema, with the route: SELECT on the schema, USE_CATALOG inherited through another group, USE_SCHEMA from this change.
  • The group is expanded. One member, and that member didn't have this access before.
  • Sensitive data is flagged. The schema contains a table tagged sensitive; the default patterns are pii*, sensitive* and class.* (the tags Databricks data classification writes).
  • A second resource on the same table blocks the run. More on that below, because it was the most surprising thing I saw.

Exit codes: 0 below the threshold, 1 WARN, 2 BLOCK, 3 invalid input or an operational failure. The threshold defaults to BLOCK and is configurable, as are your own rules in TOML (for example: new access to PII only through approved groups).

Three things a real workspace taught me

I wrote the first version against synthetic plans and documented API shapes. Then I ran a list of predictions against an Azure Databricks trial workspace with synthetic groups, users and tables, and recorded what actually happened. Three results changed the tool.

1. Two grant resources on one object can silently remove grants

The Databricks Terraform provider has an authoritative resource (databricks_grants: "these are all the grants on this object") and an additive one (databricks_grant: "this principal has these privileges"). If both manage the same securable, the result depends on apply order. In the test, I moved a table's grants from databricks_grants to databricks_grant: one destroy and one create in the same plan. Terraform ran them in parallel, the create finished first, and the destroy then removed every grant on the table. Apply reported success. Only the next plan noticed that a grant was missing.

reachdiff now treats more than one resource setting grants on one securable as grant_resource_conflict, BLOCK by default when the plan changes access. If you know the IAM policy versus binding versus member resources from other providers, it's the same class of problem.

2. A usage privilege can switch on grants nobody is looking at

SELECT on a table does nothing without USE_SCHEMA and USE_CATALOG. So a change that only adds USE_SCHEMA can activate table-level grants that have been sitting there unused, possibly for years. The plan shows one usage privilege. reachdiff reports a latent_grants gap by default, and with --deep it reads the child tables and turns those activations into actual gains.

3. The scanner sees less than you think, quietly

Two findings about the identity that runs the check:

  • Without READ METADATA on the catalog (or the metastore), the permissions API returns only the caller's own grants. That isn't an error, it's just an incomplete answer. reachdiff checks its own visibility and discards grant lists it can't trust, so such a run can't pass.
  • Workspace SCIM shows group members only to workspace admins. Without admin rights, reachdiff says "members not read" and raises a membership_incomplete gap instead of pretending a group is empty.

That's the general design: anything reachdiff can't see becomes a finding. The unverified rule can be raised to BLOCK but never turned off.

In CI

There's a GitHub Action and an Azure DevOps template. Both log in to Databricks by token federation (github-oidc, azure-devops-oidc), so no Databricks secret is stored in CI. The report goes into one pull request comment (a thread on Azure DevOps), which is updated in place on every push, and the job result follows the status: WARN, BLOCK or error.

The reachdiff comment on a GitHub pull request, edited to BLOCK after the second commit

Azure DevOps pull request: required reachdiff check failed

Reports contain principal and object names, so the publish step refuses public repositories unless you opt in.

What it doesn't do

I'd rather say this up front than in the comments:

  • Validated scope: live mode was exercised against one Azure Databricks workspace (Premium, classic compute), with the provider version and SDK version listed in the README, and in CI on GitHub-hosted Ubuntu and a self-hosted Azure DevOps agent. AWS, GCP, serverless workspaces and multiple workspaces are not covered yet.
  • Not evaluated at all: ABAC policies, row filters and column masks; volumes, functions and external locations; workspace-level ACLs; metastore and workspace admin powers; runtime behavior.
  • It's a review aid for plans, not an audit of the current state of your metastore.

Try it

pip install 'reachdiff[databricks]'
terraform show -json plan.bin > plan.json
reachdiff plan --tfplan plan.json --profile PROFILE --format md
Enter fullscreen mode Exit fullscreen mode

There's also an offline mode with no dependencies and a synthetic example plan in the repository, so you can see the report without a workspace.

GitHub logo reachdiff / reachdiff

Who gains or loses Databricks Unity Catalog access when a Terraform plan is applied

reachdiff

Who gains or loses access to what when this Terraform plan is applied?

reachdiff reads terraform show -json output for Databricks Unity Catalog grants (databricks_grants, databricks_grant), group memberships (databricks_group_member) and owners. It reports the effective-access diff, with the route behind each change, and returns an exit code a pipeline can act on.

Status: beta. Live mode has been exercised against one Azure Databricks workspace; see Validated scope for what that covers and what it doesn't. The demo and tests use synthetic plans.

Install

Python 3.11 or newer:

pip install reachdiff                  # offline mode, no dependencies
pip install 'reachdiff[databricks]'    # live mode, adds databricks-sdk
Enter fullscreen mode Exit fullscreen mode

Try the synthetic demo

Python 3.11 or newer, no dependencies:

python3 -m reachdiff plan --tfplan examples/schema-grant.plan.json --offline
Enter fullscreen mode Exit fullscreen mode

It reports analysts gaining READ on all of prod.sales (3 members: 2 new, 1 already had it) and bob@example.com losing WRITE…

It's version 0.9.0, a beta. I'm especially interested in runs on AWS or GCP, in plans that confuse it, and in how you review grant changes today.


reachdiff is an independent side project, not affiliated with or endorsed by Databricks or HashiCorp. I built it with Claude Code as a pair programmer; the behavior described here was checked against a real workspace, and what wasn't is listed in the README as not covered.

Top comments (0)