DEV Community

Gowtham Potureddi
Gowtham Potureddi

Posted on

ADLS Gen2 for Data Engineers: Hierarchical Namespace, POSIX ACLs & Performance

ADLS Gen2 — Azure Data Lake Storage Gen2 — is not a new storage service you provision alongside Blob storage. It is Blob storage, with one capability switched on at account-creation time: the hierarchical namespace. Flip that one flag and a flat bucket of objects whose names merely happen to contain slashes becomes a real filesystem with true directories, atomic renames, and per-file POSIX permissions. Leave it off and you have ordinary Blob storage. Everything an interviewer will drill you on — why a Spark job commits cleanly, why a mv of a billion-file directory is instant, how you grant one team read on a folder without touching the container role — traces back to that single switch.

That makes ADLS Gen2 a strange thing to explain, because half of it is Blob storage you already know and the other half is a POSIX filesystem grafted on top. This guide walks through the four ideas the interview actually probes — the hierarchical namespace and its atomic directory rename, POSIX ACLs versus Azure RBAC, the partition layout that keeps Spark fast without drowning in small files, and the access tiers plus identity-based security you reach through the abfss:// driver — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.

PipeCode blog header for ADLS Gen2 — bold white headline 'ADLS Gen2' with subtitle 'Hierarchical Namespace · POSIX ACLs · Performance' and a stylised blob-to-directory-tree data-lake scene on a dark gradient with purple, green, orange, and blue accents and a small pipecode.ai attribution.

When you want hands-on reps immediately after reading, design a lake layout on the partitioning practice library →, rehearse permission models on the access-control practice set →, and tune read paths on the optimization practice set →.


On this page


1. Why ADLS Gen2 rethinks the cloud data lake

ADLS Gen2 is Blob storage with a hierarchical namespace switched on — that one fact explains every other behaviour

The one-sentence invariant: ADLS Gen2 is a StorageV2 account with the hierarchical namespace enabled, so the same bytes are reachable both as flat blobs and as files in a real directory tree. Everything that makes ADLS Gen2 attractive for analytics — atomic renames, folder-level ACLs, cheap directory operations — is a consequence of that namespace, not a separate product. You do not "migrate to a data lake service"; you tick a box when the account is created.

Two namespaces, one set of bytes.

  • Flat namespace (plain Blob). A blob named sales/2026/03/part-000.parquet has no directory called sales/. The slashes are just characters in a single flat key. "Folders" are a UI illusion produced by grouping on prefixes.
  • Hierarchical namespace (HNS). The same path is stored as nested directory objects — sales, then 2026, then 03 — each a first-class node with its own metadata and permissions, with the file as a leaf.
  • Set at creation, permanent. HNS is chosen when the account is created and cannot be toggled later without a migration; you decide up front whether an account is an analytics lake or general object storage.

Two endpoints for the same account.

  • The blob endpoint — https://{account}.blob.core.windows.net — speaks the classic Blob REST API. Existing SDKs, lifecycle rules, and tools keep working.
  • The dfs endpoint — https://{account}.dfs.core.windows.net — speaks the Data Lake Storage (ADLS) API with directory and ACL operations. This is what the abfss:// Hadoop driver talks to.
  • Multi-protocol access. With HNS on, a file written through the dfs endpoint is immediately readable through the blob endpoint and vice-versa — same object, two APIs.

Why data engineers care.

  • Containers become filesystems. The top-level container in an HNS account is the "filesystem" in ABFS terminology; below it you get real, cheap directory operations instead of prefix scans.
  • Analytics engines expect a filesystem. Spark, Hive, and Hadoop commit protocols assume rename and list-directory are fast and atomic. HNS makes those assumptions true on object storage; flat Blob does not.
  • Governance gets granular. POSIX ACLs let you grant a folder to one group without inventing a new container or a broad RBAC role.

What interviewers listen for.

  • Do you say "ADLS Gen2 is Blob storage plus a hierarchical namespace" in the first sentence? — senior signal.
  • Do you connect atomic directory rename to Spark commit correctness unprompted? — required framing.
  • Do you distinguish the blob endpoint from the dfs endpoint and know abfss targets dfs? — the detail that separates readers from users.
  • Do you reach for POSIX ACLs for fine-grained access and RBAC for coarse-grained rather than treating them as alternatives? — the whole governance story.

Worked example — the same directory rename, HNS off versus on

Detailed explanation. The cleanest way to feel what HNS buys you is to rename a "directory" both ways. On a flat account, raw/ is a prefix shared by thousands of blobs, so renaming it means the service (or your tool) must copy every blob to a new key and delete the old one — an O(n) operation that is not atomic. On an HNS account, raw is a single directory node, so renaming it flips one pointer — O(1) and atomic.

Question. A folder raw/ holds 1,000,000 blobs. What does rename raw -> bronze cost on a flat Blob account versus an HNS (ADLS Gen2) account, and is it atomic?

Input.

account type raw/ contents operation
flat Blob 1,000,000 blobs sharing prefix raw/ rename prefix to bronze/
HNS (ADLS Gen2) 1 directory object raw with 1,000,000 leaf files rename directory to bronze

Code.

# Flat Blob: no real rename — you copy then delete, per blob
az storage blob copy start-batch \
  --destination-container data --destination-path bronze \
  --source-container data --pattern "raw/*"
az storage blob delete-batch --source data --pattern "raw/*"

# ADLS Gen2 (HNS): one directory metadata operation
az storage fs directory move \
  -f data --new-directory "data/bronze" -n "raw" \
  --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Step-by-step explanation. On the flat account there is no directory to rename, so the tool enumerates every blob under raw/, issues a server-side copy to the bronze/ key, waits for each copy, then deletes each original — a million round trips, and if the job dies halfway the namespace is left half-renamed. On the HNS account, raw is one node; directory move rewrites a single metadata entry, so all million children are reparented instantly and the change is all-or-nothing.

Output.

account type operations atomic? rough time
flat Blob ~1,000,000 copy + delete no minutes to hours
HNS (ADLS Gen2) 1 metadata update yes milliseconds

Rule of thumb. If your workload renames or moves directories — and every Spark commit does — you want the hierarchical namespace on; on flat Blob a "directory rename" is a full copy that can fail halfway.


2. Hierarchical namespace and atomic directory rename

Real directories make rename and delete single metadata operations — and that is exactly what Spark commit protocols need

The hierarchical namespace is the feature, and its payoff is a property most people take for granted on a local disk and lose on object storage: directory operations are atomic and independent of how many files the directory holds. On flat Blob, list, rename, and delete of a "folder" all scale with the number of objects. On HNS they are O(1) metadata operations, which is what lets analytics engines commit output safely.

What the namespace actually stores.

  • Filesystem = container. The abfss:// URI names a container as the filesystem: abfss://<filesystem>@<account>.dfs.core.windows.net/<path>. Below it, directories are real nodes, not prefixes.
  • Directories are objects with metadata. Each directory carries its own owner, group, permission bits, and ACL — none of which exists in the flat model.
  • Atomic rename and move. Renaming or moving a directory rewrites one parent pointer; every descendant is reparented in a single, all-or-nothing operation.
  • Cheap recursive delete. Deleting a directory tree is one metadata operation on HNS, versus enumerate-and-delete-each-blob on flat Blob.

Why Spark and Hadoop depend on this.

  • The commit dance is a rename. The FileOutputCommitter writes task output to a _temporary directory, then on success renames it into the final path. If rename is slow or non-atomic, commits are slow and can corrupt output on failure.
  • _SUCCESS and job atomicity. A partly-renamed output on flat Blob can leave readers seeing half a dataset; atomic directory rename makes "all rows or none" hold.
  • Listing drives planning. Spark lists input directories to plan splits; O(1)-ish directory listing on HNS beats paging through a flat prefix of millions of keys.
  • Use abfss, not wasb. The modern abfss:// driver (ABFS, secured with TLS) is the supported path to the dfs endpoint; the legacy wasb:// Blob driver does not use the hierarchical namespace semantics.

The trade-offs to name.

  • HNS has a small metadata cost. Directory transactions are billed and there is bookkeeping overhead versus a pure flat store, negligible for analytics but worth naming.
  • Not every tool speaks dfs. Some older Blob-only tools target the blob endpoint; multi-protocol access means they still work, but they see the flat view and its rename cost.

Iconographic ADLS Gen2 hierarchical-namespace diagram — a flat blob store where a directory rename copies every blob one by one, versus a hierarchical namespace where the same rename is a single atomic metadata operation feeding a fast Spark commit.

Worked example — a Spark write commit on HNS

Detailed explanation. The everyday place atomic rename earns its keep is a Spark write. Spark writes each task's output into a staging directory and, when the job succeeds, renames staging into the final location. On HNS that final rename is one metadata operation, so the commit is fast and cannot leave a half-written dataset.

Question. A Spark job writes df to abfss://data@mylake.dfs.core.windows.net/bronze/orders. Walk through where bytes land and what the commit does on ADLS Gen2.

Input. A DataFrame with 3 output partitions written by 3 tasks.

Code.

(df.write
   .mode("overwrite")
   .format("parquet")
   .save("abfss://data@mylake.dfs.core.windows.net/bronze/orders"))
Enter fullscreen mode Exit fullscreen mode

Step-by-step explanation. Each of the 3 tasks writes its part-file into a task-scoped path under bronze/orders/_temporary/. As each task succeeds, its files are promoted toward the job-attempt directory. When the whole job succeeds, the committer renames the staged output into bronze/orders/ and writes a _SUCCESS marker — on HNS these renames are atomic metadata operations, so no reader ever sees a partial orders/. If any task had failed, the _temporary tree is discarded and the final path is untouched.

Output.

stage path visible to readers?
task write bronze/orders/_temporary/.../part-*.parquet no
atomic commit bronze/orders/part-*.parquet + _SUCCESS yes (all at once)
on failure _temporary discarded nothing published

Rule of thumb. On HNS a Spark commit is a cheap atomic rename; on flat Blob the same commit copies every part-file, which is why big writes to non-HNS accounts are slow and riskier.

ADLS Gen2 interview question on directory operations

Question. Your nightly job renames a staging directory holding ~2 million small files into its final partition path. On a plain Blob account the rename takes 40 minutes and once left the table half-published after a crash. How does moving to ADLS Gen2 fix both the speed and the correctness problem, and what exactly changes at the storage layer?

Solution Using the hierarchical namespace atomic rename

Code.

# Same logical operation, two very different storage costs.
# Plain Blob (flat): "rename" = server-side copy of every object, then delete.
# ADLS Gen2 (HNS):   rename = one directory metadata operation, atomic.

final = "abfss://data@mylake.dfs.core.windows.net/gold/sales/dt=2026-03-05"
staging = "abfss://data@mylake.dfs.core.windows.net/gold/_staging/sales_20260305"

# On HNS this is a single atomic directory move:
dbutils.fs.mv(staging, final, recurse=True)   # metadata-only reparent
Enter fullscreen mode Exit fullscreen mode

Step-by-step trace.

step flat Blob behaviour ADLS Gen2 (HNS) behaviour
1 enumerate 2,000,000 keys under staging resolve 1 directory node sales_20260305
2 copy each blob to the new prefix rewrite one parent pointer
3 delete each old blob (nothing to delete — same objects, new parent)
4 crash mid-way → half-copied, half-published crash → operation either fully applied or not at all
  1. On flat Blob there is no directory, so "rename" is 2,000,000 copy-then-delete round trips — that is the 40 minutes.
  2. Because those round trips are independent and non-transactional, a crash after copying 1,100,000 objects leaves the table half-published.
  3. On HNS, sales_20260305 is a single directory object; the move rewrites one metadata entry and reparents all 2,000,000 children at once.
  4. That single metadata operation is atomic, so a crash leaves the namespace either fully renamed or unchanged — never half-published.

Output:

metric flat Blob ADLS Gen2 (HNS)
storage operations ~4,000,000 (copy + delete) 1 metadata op
wall-clock ~40 min sub-second
partial-failure state possible impossible

Why this works — concept by concept:

  • Hierarchical namespace — storing directories as real objects means a directory has an identity to rename, so the operation touches one node instead of every leaf.
  • Atomic rename — the reparent is a single metadata transaction, so it is all-or-nothing; there is no window where the table is half-published.
  • O(1) not O(n) — cost is independent of the 2,000,000 children, turning a 40-minute copy into a sub-second pointer flip.
  • Commit correctness — Spark/Hadoop committers rely on rename atomicity; HNS gives object storage the guarantee they were designed around.
  • Cost — the move is O(1) metadata versus O(n) data copies; time and dollar cost both collapse from linear in file count to constant.

Partitioning
Topic — partitioning
Directory-layout and partition-move problems

Practice →

Optimization Topic — optimization Commit-cost and write-path optimization problems

Practice →


3. POSIX ACLs versus Azure RBAC

Two layers guard every object — coarse RBAC roles first, then fine-grained POSIX ACLs, and you must know which fires when

ADLS Gen2 has two authorization systems that both apply, and the interview question is almost always "explain how they interact." Say it in one breath: Azure RBAC is coarse-grained and evaluated first; POSIX ACLs are fine-grained and evaluated only when no RBAC role already grants access. Confusing them — or thinking you pick one — is the classic mistake.

Azure RBAC — the coarse layer.

  • Scope is broad. RBAC roles are assigned at subscription, resource-group, storage-account, or container (filesystem) scope — never per file or per directory.
  • Data-plane roles. Storage Blob Data Reader, Storage Blob Data Contributor, and Storage Blob Data Owner grant read/write/own across everything in their scope.
  • Owner is a superuser. Storage Blob Data Owner bypasses ACL checks entirely — it is the POSIX superuser for that scope, so ACLs never restrict it.

POSIX ACLs — the fine layer.

  • Per file and per directory. Every file and directory has an owning user, an owning group, and an other class, each with read/write/execute (rwx) bits.
  • Named entries. Beyond the three POSIX classes you add named-user and named-group entries — e.g. grant group:analysts r-x on one folder — up to a hard limit of 32 entries per file or directory.
  • Execute means traverse. On a directory, the x bit means "traverse into"; to read /a/b/c a principal needs x on /a and /a/b, then r on c. Forgetting traverse permissions on parents is the #1 ACL bug.

Access ACLs versus default ACLs.

  • Access ACL. Controls access to this object right now. It exists on both files and directories.
  • Default ACL. A template that lives only on directories. When a new child is created, it copies the parent's default ACL as its own access ACL (and, for a child directory, its default ACL too).
  • Inheritance is at creation only. Changing a directory's default ACL does not retroactively rewrite children that already exist — you must re-apply recursively. This surprises people constantly.

The evaluation order (the answer they want).

  • RBAC first. If an assigned role grants the requested action at a covering scope, access is allowed and ACLs are not consulted.
  • ACLs second. If no RBAC role grants it, the POSIX ACL on the target (and traverse on its parents) decides.
  • Principals are Microsoft Entra identities. Users, groups, service principals, and managed identities from Microsoft Entra ID (formerly Azure AD) are what both layers evaluate against.

Iconographic ADLS Gen2 access-control diagram — Azure RBAC evaluated first at container scope, then fine-grained POSIX ACLs per file and directory, with access ACLs on the object and default ACLs on directories flowing to child objects at creation.

Worked example — grant one group read on a folder without a new role

Detailed explanation. The everyday governance task is "let the analysts read curated/finance/ but nothing else, without giving them a container-wide role." That is precisely what a named-group ACL entry plus a default ACL is for: an access ACL grants the existing files, and a default ACL makes future files inherit the grant.

Question. Give group:analysts read access to everything under curated/finance/, existing and future, using ACLs (not RBAC). What entries do you set?

Input.

target current access goal
curated/ (parent) analysts: none traverse only
curated/finance/ analysts: none read + traverse, now and future

Code.

# Traverse (x) on the parent so analysts can reach the folder
az storage fs access set --acl "group:analysts:--x" \
  -p "curated" -f data --account-name mylake

# Read+traverse on the target folder (access ACL = existing objects)
az storage fs access set-recursive \
  --acl "group:analysts:r-x" \
  -p "curated/finance" -f data --account-name mylake

# Default ACL so NEW children inherit the grant at creation
az storage fs access set \
  --acl "default:group:analysts:r-x" \
  -p "curated/finance" -f data --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Step-by-step explanation. The --x on curated lets analysts traverse into the tree without being able to list it. The recursive r-x sets the access ACL on curated/finance and everything already inside it, so existing files become readable. The default:group:analysts:r-x entry is a template on the directory: any file created later copies it as its own access ACL, so new data is readable without re-running the recursive apply. No RBAC role was granted, so analysts see only this subtree.

Output.

principal curated/ curated/finance/* (existing) files added tomorrow
analysts traverse only read read (via default ACL)
everyone else unchanged unchanged unchanged

Rule of thumb. Access ACLs fix today's files; default ACLs fix tomorrow's. Set both, and remember default ACLs never rewrite objects that already exist.

ADLS Gen2 interview question on ACLs versus RBAC

Question. A service principal has the Storage Blob Data Contributor role on the whole storage account, but you also removed its ACL entry from curated/finance/ so it "cannot" read finance data. During an audit it reads a finance file anyway. Explain why, and how the two authorization layers combined to allow it.

Solution Using the RBAC-then-ACL evaluation order

Code.

Request: service principal SP reads
  abfss://data@mylake.dfs.core.windows.net/curated/finance/ledger.parquet

Layer 1 — Azure RBAC (evaluated FIRST):
  SP has "Storage Blob Data Contributor" at account scope
  -> read is granted by the role  -> ACCESS ALLOWED, ACLs not consulted

Layer 2 — POSIX ACL (only reached if RBAC did NOT grant):
  SP removed from curated/finance ACL  -> would deny
  ... but this layer is never evaluated, because RBAC already allowed.
Enter fullscreen mode Exit fullscreen mode

Step-by-step trace.

step check result
1 RBAC: does any role grant read at a covering scope? yes — Data Contributor at account scope
2 RBAC allowed → short-circuit ACL check skipped entirely
3 (unreached) ACL: is SP granted on curated/finance? would be deny, but never evaluated
  1. Authorization checks RBAC first; a data-plane role at account scope covers curated/finance/ledger.parquet.
  2. Because Storage Blob Data Contributor includes read, the request is allowed at layer 1 and evaluation stops.
  3. The ACL layer — where you removed the principal — is only consulted when RBAC does not already grant the action, so your ACL edit was dead code.
  4. To actually restrict SP, you must narrow its RBAC (remove the broad role or scope it to a different container), then let ACLs express the fine-grained grant.

Output:

you changed intended effect real effect
removed SP from ACL deny finance read none — RBAC still allows
(fix) remove broad RBAC role ACLs govern ACL deny now takes effect

Why this works — concept by concept:

  • RBAC-first evaluation — a granting role short-circuits authorization, so ACLs can only ever add access on top of RBAC, never subtract from it.
  • Superuser semantics — Storage Blob Data Owner/broad roles behave like a POSIX superuser and bypass ACLs by design, which is powerful and dangerous.
  • Least privilege at the right layer — coarse restrictions belong in RBAC scope; fine-grained grants belong in ACLs; you cannot fix an over-broad role with an ACL.
  • Auditability — knowing the order tells you where to look first when access is unexpectedly allowed: check the role assignments before the ACLs.
  • Cost — the check is O(1) role lookup then, if needed, O(depth) traverse checks up the directory path; negligible per request.

Access control
Topic — access-control
Permission-model and least-privilege problems

Practice →

Optimization Topic — optimization Access-path and traversal-cost problems

Practice →


4. Partitioning layout and the small-files problem

The directory layout you choose is a performance contract — partition for pruning, size files right, and never let the tiny files pile up

Because ADLS Gen2 gives you a real filesystem, how you lay out directories and how big you make files is the single biggest lever on query cost. Two failure modes dominate interviews: a layout that forces every query to scan everything, and a "small-files problem" where millions of tiny files crush the engine with per-file overhead. Say the principle in one line: partition on the columns you filter, size files at 128 MB–1 GB, and compact before the small files multiply.

Partitioning for pruning.

  • Layout by filter column. Write .../sales/year=2026/month=03/day=05/part-*.parquet. The Hive-style key=value folders let the engine skip whole branches — this is partition pruning.
  • Prune only helps if you filter on it. Partitioning by day speeds WHERE day = '2026-03-05' but does nothing for WHERE customer_id = 42; partition on the predicate you actually run.
  • Pick the right cardinality. A partition column with a handful to a few thousand distinct values (date, region) prunes well; one with millions of distinct values (user id) explodes the directory count.

The small-files problem.

  • What it is. Thousands or millions of tiny files (KBs each) instead of fewer right-sized files. Each file is a separate open, a separate metadata entry, and a separate task — overhead swamps useful work.
  • Where it bites. Listing the directory is slow, the driver plans one split per tiny file, and columnar formats lose compression and row-group efficiency at small sizes.
  • Target size. Aim for 128 MB–1 GB per file (roughly one HDFS/Spark block-ish unit). Below ~a few MB you are firmly in the danger zone.

How over-partitioning creates small files.

  • Too many partition levels. Partitioning by year/month/day/hour/minute on modest volume leaves each leaf with a sliver of data, so every write drops a tiny file.
  • Too many writer tasks. N Spark tasks each writing to M partitions produce up to N×M files per run; without a repartition/coalesce you flood the lake with fragments.
  • Streaming micro-batches. Frequent small appends (one file per micro-batch) are the classic streaming small-files source.

The fixes.

  • Repartition before write. df.repartition("day") (or repartitionByRange) so each partition is written by one task → one file per partition.
  • Compaction / OPTIMIZE. Periodically rewrite many small files into few large ones — Delta Lake's OPTIMIZE (with ZORDER for data-skipping) or a scheduled compaction job.
  • Right-size the partition grain. Coarsen from hour to day (or add bucketing) when partitions are too small; the goal is fewer, fuller files.

Iconographic ADLS Gen2 partitioning diagram — a year/month/day directory layout enabling partition pruning on the left, versus an over-partitioned tree producing many tiny files on the right, with a compaction step merging them into right-sized files.

Worked example — repartition to write one file per partition

Detailed explanation. The most common small-files fix is to align the write parallelism with the partition column. If 200 Spark tasks each write to all 30 day-partitions, you get up to 6,000 files. Repartitioning the DataFrame by the partition column first means each partition is handled by one task, so each day gets one right-sized file.

Question. A job writes 30 days of data partitioned by day using 200 tasks, producing ~6,000 tiny files. Rewrite it so each day-partition lands as a single file, and show the file-count change.

Input.

setting before goal
partition column day (30 distinct) day
writer tasks 200 aligned to partitions
files produced ~6,000 ~30

Code.

# Before: 200 tasks x 30 day-partitions -> up to 6,000 small files
# After: repartition by the partition key so one task writes one day

(df.repartition("day")                       # shuffle so each day is one partition
   .write
   .mode("overwrite")
   .partitionBy("day")                       # Hive-style day=... folders
   .parquet("abfss://data@mylake.dfs.core.windows.net/gold/sales"))
Enter fullscreen mode Exit fullscreen mode

Step-by-step explanation. repartition("day") shuffles rows so all rows for a given day sit in one Spark partition. partitionBy("day") writes the Hive-style day=YYYY-MM-DD/ folders for pruning. Because each day is now one Spark partition handled by one task, each day=... folder receives a single part-file instead of up to 200. The result is ~30 right-sized files that prune cleanly and open fast.

Output.

layout files per-file size query on one day
before ~6,000 ~KBs slow: list + open thousands
after ~30 ~hundreds of MB fast: prune to 1 folder, open 1 file

Rule of thumb. Match write parallelism to partition cardinality: repartition(partitionCol) before partitionBy(partitionCol) turns N×M fragments into one file per partition.

ADLS Gen2 interview question on the small-files problem

Question. A streaming job appends one small Parquet file per minute to events/day=.../, and after three months queries over a day have become painfully slow even though each query only touches one day-partition. Diagnose why, and give a fix that keeps the recent-data latency but restores query speed.

Solution Using a scheduled compaction into right-sized files

Code.

# Symptom: 60 files/hour x 24 x 90 days = ~130,000 tiny files per day-partition.
# Fix: keep streaming for freshness, compact yesterday's partition into few files.

from pyspark.sql import functions as F

day = "2026-03-05"
src = f"abfss://data@mylake.dfs.core.windows.net/events/day={day}"

compacted = (spark.read.parquet(src)
             .repartition(4))                 # 4 right-sized files, not 130k tiny

(compacted.write
   .mode("overwrite")
   .parquet(src))                             # atomic overwrite of the partition

# On Delta Lake this is simply:
#   OPTIMIZE events WHERE day = '2026-03-05'
Enter fullscreen mode Exit fullscreen mode

Step-by-step trace.

step state of day=2026-03-05 files query cost
1 streaming appended all day ~130,000 tiny list + open 130k → slow
2 read partition, repartition(4) in memory, 4 partitions —
3 overwrite partition atomically 4 right-sized files list + open 4 → fast
4 today keeps streaming small files (fresh) acceptable for one live day
  1. The slowness is not scanning too much data — pruning already limits it to one day — it is per-file overhead: 130,000 opens, 130,000 metadata reads, 130,000 tasks.
  2. Reading the partition and repartition(4) collapses those rows into four in-memory partitions.
  3. Overwriting the partition (atomic on HNS) replaces the 130,000 fragments with four right-sized files, so the same query now opens four files.
  4. Only the current day keeps the streaming small files, so freshness is preserved while every closed day is compacted — the standard "hot tail, compacted history" pattern.

Output:

metric before compaction after compaction
files in partition ~130,000 4
files opened per query ~130,000 4
query time minutes seconds

Why this works — concept by concept:

  • Small-files problem — cost was dominated by per-file open/metadata/task overhead, not bytes scanned, so fewer larger files fix it without changing the data.
  • Compaction — rewriting a partition into a handful of right-sized files restores columnar efficiency and cheap listing.
  • Hot tail vs cold history — leaving only the live partition fragmented preserves streaming latency while every completed partition is compacted.
  • Atomic overwrite — HNS makes the partition overwrite all-or-nothing, so readers never catch the partition mid-compaction.
  • Cost — compaction is O(rows in partition) once per day, amortized against every future query that now opens 4 files instead of 130,000.

Partitioning
Topic — partitioning
Partition-grain and file-sizing problems

Practice →

Optimization Topic — optimization Compaction and read-path optimization problems

Practice →


5. Access tiers, abfss and identity-based security

Move cold data down the tiers, reach the lake over abfss, and authenticate with an identity instead of a key

The last cluster the interview covers is operational: how you pay for storage (access tiers), how you connect (the abfss driver), and how you prove who you are (identity-based auth). Get these right and you cut cost without losing durability and connect Spark securely without a secret sprawled across notebooks.

Access tiers — pay for what you touch.

  • Hot. Highest storage price, lowest access price. For data read frequently — active partitions, curated marts.
  • Cool. Lower storage, higher access price; minimum ~30 days. For infrequently accessed data you still read occasionally.
  • Cold. Cheaper still, higher access price; minimum ~90 days. For rarely accessed data that must stay online.
  • Archive. Cheapest storage, offline — you must rehydrate (to Hot/Cool) before reading, which can take hours; minimum ~180 days. For compliance retention you almost never read.
  • Set per blob or as an account default. Tiering is set at the blob level (or by lifecycle-management rules that auto-move blobs by age), and early-deletion fees apply if you delete/move before the minimum.

The abfss:// driver — how engines connect.

  • ABFS = Azure Blob File System. The abfss:// scheme (the s = secure/TLS) is the Hadoop driver that talks to the dfs endpoint and uses hierarchical-namespace semantics.
  • URI shape. abfss://<filesystem>@<account>.dfs.core.windows.net/<path> — filesystem is the container, host is the dfs endpoint.
  • Prefer it over wasbs. The legacy wasbs:// (WABS) Blob driver predates ADLS Gen2 and does not use directory semantics; use abfss for lakes.

Identity-based security — stop pasting keys.

  • Account key / connection string. Full control, no rotation story, easy to leak — avoid in shared code.
  • SAS (Shared Access Signature). A time-boxed, scope-limited token; good for handing narrow, expiring access to an external party or a single job.
  • Service principal. A Microsoft Entra app identity with a client id/secret (or certificate); the classic choice for a scheduled pipeline, governed by RBAC + ACLs.
  • Managed identity. An Entra identity Azure manages for you — no secret to store or rotate. The preferred production auth for Azure-hosted compute (Databricks, Synapse, VMs, Functions): grant the identity a role, and code authenticates with no credential in it.

Choosing auth by workload.

  • Azure-hosted, long-lived pipeline → managed identity (no secret).
  • Off-Azure system or automated job with its own lifecycle → service principal.
  • Narrow, expiring, third-party access → SAS.
  • Anything shared or checked into a repo → never the account key.

Iconographic ADLS Gen2 tiers-and-security diagram — hot, cool, cold and archive storage tiers by access frequency and cost, alongside an abfss dfs-endpoint connection secured by managed identity, service principal, SAS or account key.

Worked example — connect Spark with a managed identity, no secret

Detailed explanation. The production-clean way to read the lake from Azure-hosted Spark is a managed identity: you assign the compute's identity a data-plane role (or ACLs), and the abfss client authenticates with an Entra OAuth token minted at runtime — nothing secret lives in the notebook or config.

Question. Configure Spark to read abfss://data@mylake.dfs.core.windows.net/curated/ using a managed identity, with no account key or client secret in the code.

Input. A managed identity granted Storage Blob Data Reader on the data filesystem.

Code.

acct = "mylake.dfs.core.windows.net"

# Tell ABFS to authenticate with the managed identity (OAuth), not a key.
spark.conf.set(f"fs.azure.account.auth.type.{acct}", "OAuth")
spark.conf.set(f"fs.azure.account.oauth.provider.type.{acct}",
    "org.apache.hadoop.fs.azurebfs.oauth2.MsiTokenProvider")

# No secret anywhere — the platform mints the token for the assigned identity.
df = spark.read.parquet(f"abfss://data@{acct}/curated/")
Enter fullscreen mode Exit fullscreen mode

Step-by-step explanation. Setting the auth type to OAuth with the MsiTokenProvider tells the ABFS driver to request an access token for the compute's managed identity from the Azure instance metadata endpoint at runtime. That token is presented to the dfs endpoint, which checks RBAC (the Storage Blob Data Reader role) and, if needed, ACLs. Because the identity is managed, there is no client secret to store or rotate — the credential never appears in code, config, or a secret store.

Output.

aspect account key managed identity
secret in code/config yes (leakable) none
rotation manual automatic (platform)
authorization all-or-nothing key RBAC + ACLs on the identity

Rule of thumb. On Azure-hosted compute, authenticate to abfss with a managed identity and grant it a scoped role — you get least privilege and zero secrets to rotate.

ADLS Gen2 interview question on tiers and cost

Question. Five years of raw event history sits in the Hot tier and is read maybe once a year for a compliance export, yet it dominates your storage bill. How do you cut the storage cost while keeping the data retrievable, and what operational caveats must you flag to stakeholders?

Solution Using lifecycle rules to tier data to Archive

Code.

{
  "rules": [
    {
      "name": "raw-events-to-archive",
      "type": "Lifecycle",
      "definition": {
        "filters": { "blobTypes": ["blockBlob"],
                     "prefixMatch": ["data/raw/events/"] },
        "actions": {
          "baseBlob": {
            "tierToCool":    { "daysAfterModificationGreaterThan": 30 },
            "tierToArchive": { "daysAfterModificationGreaterThan": 180 }
          }
        }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Step-by-step trace.

age of blob tier storage cost readable?
0–30 days Hot highest instant
30–180 days Cool lower instant (higher read cost)
180+ days Archive lowest only after rehydration (hours)
  1. The lifecycle rule matches only the data/raw/events/ prefix, so hot curated data is untouched.
  2. After 30 days without modification a blob auto-moves to Cool, cutting storage cost for data that is no longer actively read.
  3. After 180 days it moves to Archive — the cheapest tier — which is offline storage.
  4. Reading an archived blob requires a rehydration to Hot/Cool that can take hours, so the once-a-year compliance export must schedule a rehydrate step first; and early deletion before the tier minimum incurs a charge.

Output:

metric all-Hot (before) lifecycle-tiered (after)
storage cost of old data highest ~Archive price (a fraction)
retrieval latency (old) instant hours (rehydrate first)
operational caveat none plan rehydration + min-duration fees

Why this works — concept by concept:

  • Access tiers — storage price falls Hot → Cool → Cold → Archive while retrieval price/latency rises, so matching tier to read frequency minimizes total cost.
  • Lifecycle management — age-based rules move blobs automatically, so tiering is policy, not a manual job someone forgets.
  • Archive is offline — the cheapest tier trades instant reads for a rehydration delay, which is fine for once-a-year data but must be flagged to consumers.
  • Minimum durations — Cool/Cold/Archive carry early-deletion charges, so tiering churny data down and back up can cost more than leaving it Hot.
  • Cost — for rarely read history the storage saving dominates the occasional rehydration fee, so the yearly-export dataset belongs in Archive.

Access control
Topic — access-control
Identity, SAS and least-privilege auth problems

Practice →

Partitioning Topic — partitioning Tiering and lifecycle-layout problems

Practice →


Cheat sheet — ADLS Gen2 recipes

Read over abfss (Spark).

df = spark.read.parquet(
    "abfss://data@mylake.dfs.core.windows.net/curated/orders")
Enter fullscreen mode Exit fullscreen mode

Set an access ACL (existing object).

az storage fs access set --acl "group:analysts:r-x" \
  -p "curated/finance" -f data --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Set a default ACL (inherited by new children).

az storage fs access set --acl "default:group:analysts:r-x" \
  -p "curated/finance" -f data --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Move a directory (atomic on HNS).

az storage fs directory move -f data \
  --new-directory "data/bronze" -n "raw" --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Set a blob's access tier.

az storage blob set-tier --container-name data \
  --name "raw/events/2022/part-000.parquet" --tier Archive \
  --account-name mylake
Enter fullscreen mode Exit fullscreen mode

Repartition to avoid small files, then partition-write.

(df.repartition("day")
   .write.mode("overwrite").partitionBy("day")
   .parquet("abfss://data@mylake.dfs.core.windows.net/gold/sales"))
Enter fullscreen mode Exit fullscreen mode

Auth picker.

Situation Auth
Azure-hosted long-lived pipeline managed identity
Off-Azure / external automated job service principal
Narrow, time-boxed, third-party access SAS token
Shared code or repo never the account key

Frequently asked questions

What is ADLS Gen2?

ADLS Gen2 (Azure Data Lake Storage Gen2) is Azure Blob storage with the hierarchical namespace enabled — a StorageV2 account where the namespace flag is switched on at creation. That gives you real directories, atomic directory renames, and per-file POSIX ACLs on top of Blob storage's durability and tiers, reachable through both the classic blob endpoint and the dfs (Data Lake) endpoint. It is the standard object store for analytics and lakehouse workloads on Azure.

How is ADLS Gen2 different from Blob storage?

They are the same underlying service; ADLS Gen2 is Blob storage with the hierarchical namespace turned on. Plain Blob has a flat namespace where "folders" are just prefixes in a blob name, so renaming or deleting a folder means copying or deleting every object (O(n), non-atomic). ADLS Gen2 stores directories as real objects, so rename, move, and delete are single atomic metadata operations, and it adds POSIX ACLs and the dfs endpoint that the abfss driver uses.

What is the difference between ACLs and RBAC in ADLS Gen2?

Azure RBAC is coarse-grained: roles like Storage Blob Data Reader/Contributor/Owner are assigned at account or container scope. POSIX ACLs are fine-grained: read/write/execute entries on each file and directory, including named-user and named-group grants (up to 32 entries). RBAC is evaluated first — if a role grants the action, access is allowed and ACLs are skipped (Data Owner is effectively a superuser). ACLs are only consulted when no RBAC role already grants access, so ACLs add access on top of RBAC rather than subtracting from it.

Why does the hierarchical namespace matter for Spark?

Spark and Hadoop commit protocols write task output to a temporary directory and then rename it into the final location on success. On a flat Blob account that rename is a full copy of every file and is not atomic, so big writes are slow and a crash can leave a half-published table. With the hierarchical namespace, a directory rename is a single atomic metadata operation, so commits are fast and all-or-nothing — which is why you use the abfss driver against an HNS account for lakehouse workloads.

What is the small-files problem and how do I fix it?

The small-files problem is having thousands or millions of tiny files instead of fewer right-sized ones, so per-file open, metadata, and task overhead dominates query cost even when pruning limits the data scanned. It comes from over-partitioning, too many writer tasks (N×M files), or streaming one file per micro-batch. Fix it by targeting 128 MB–1 GB files: repartition(partitionCol) before partitionBy(...), coarsen the partition grain, and run periodic compaction (or Delta Lake OPTIMIZE) on completed partitions.

What is the abfss driver?

abfss:// is the Azure Blob File System (ABFS) Hadoop driver, secured with TLS, that connects Spark/Hadoop to the ADLS Gen2 dfs endpoint using hierarchical-namespace semantics. The URI is abfss://<filesystem>@<account>.dfs.core.windows.net/<path>, where the filesystem is the container. It replaces the legacy wasbs:// Blob driver, which predates ADLS Gen2 and does not use directory semantics; authenticate it with a managed identity, service principal, SAS, or (avoid) the account key.

Practice on PipeCode

Pipecode.ai is Leetcode for Data Engineering — every ADLS Gen2 idea above, from the atomic directory rename to the RBAC-then-ACL evaluation order and the repartition-then-compact fix for small files, maps to a hands-on practice room where you design the layout and reason about access against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you keep this lake fast and locked down?" holds up under a senior interviewer's depth probes.

Practice partitioning problems now →
Access-control drills →

Top comments (0)