ADLS Gen2 — Azure Data Lake Storage Gen2 — is not a new storage service you provision alongside Blob storage. It is Blob storage, with one capability switched on at account-creation time: the hierarchical namespace. Flip that one flag and a flat bucket of objects whose names merely happen to contain slashes becomes a real filesystem with true directories, atomic renames, and per-file POSIX permissions. Leave it off and you have ordinary Blob storage. Everything an interviewer will drill you on — why a Spark job commits cleanly, why a mv of a billion-file directory is instant, how you grant one team read on a folder without touching the container role — traces back to that single switch.
That makes ADLS Gen2 a strange thing to explain, because half of it is Blob storage you already know and the other half is a POSIX filesystem grafted on top. This guide walks through the four ideas the interview actually probes — the hierarchical namespace and its atomic directory rename, POSIX ACLs versus Azure RBAC, the partition layout that keeps Spark fast without drowning in small files, and the access tiers plus identity-based security you reach through the abfss:// driver — and pairs each with a Solution-Tail interview answer: code, a step-by-step trace, an output table, then a concept-by-concept breakdown of why it works.
When you want hands-on reps immediately after reading, design a lake layout on the partitioning practice library →, rehearse permission models on the access-control practice set →, and tune read paths on the optimization practice set →.
On this page
- Why ADLS Gen2 rethinks the cloud data lake
- Hierarchical namespace and atomic directory rename
- POSIX ACLs versus Azure RBAC
- Partitioning layout and the small-files problem
- Access tiers, abfss and identity-based security
- Cheat sheet — ADLS Gen2 recipes
- Frequently asked questions
- Practice on PipeCode
1. Why ADLS Gen2 rethinks the cloud data lake
ADLS Gen2 is Blob storage with a hierarchical namespace switched on — that one fact explains every other behaviour
The one-sentence invariant: ADLS Gen2 is a StorageV2 account with the hierarchical namespace enabled, so the same bytes are reachable both as flat blobs and as files in a real directory tree. Everything that makes ADLS Gen2 attractive for analytics — atomic renames, folder-level ACLs, cheap directory operations — is a consequence of that namespace, not a separate product. You do not "migrate to a data lake service"; you tick a box when the account is created.
Two namespaces, one set of bytes.
-
Flat namespace (plain Blob). A blob named
sales/2026/03/part-000.parquethas no directory calledsales/. The slashes are just characters in a single flat key. "Folders" are a UI illusion produced by grouping on prefixes. -
Hierarchical namespace (HNS). The same path is stored as nested directory objects —
sales, then2026, then03— each a first-class node with its own metadata and permissions, with the file as a leaf. - Set at creation, permanent. HNS is chosen when the account is created and cannot be toggled later without a migration; you decide up front whether an account is an analytics lake or general object storage.
Two endpoints for the same account.
-
The
blobendpoint —https://{account}.blob.core.windows.net— speaks the classic Blob REST API. Existing SDKs, lifecycle rules, and tools keep working. -
The
dfsendpoint —https://{account}.dfs.core.windows.net— speaks the Data Lake Storage (ADLS) API with directory and ACL operations. This is what theabfss://Hadoop driver talks to. -
Multi-protocol access. With HNS on, a file written through the
dfsendpoint is immediately readable through theblobendpoint and vice-versa — same object, two APIs.
Why data engineers care.
- Containers become filesystems. The top-level container in an HNS account is the "filesystem" in ABFS terminology; below it you get real, cheap directory operations instead of prefix scans.
-
Analytics engines expect a filesystem. Spark, Hive, and Hadoop commit protocols assume
renameandlist-directoryare fast and atomic. HNS makes those assumptions true on object storage; flat Blob does not. - Governance gets granular. POSIX ACLs let you grant a folder to one group without inventing a new container or a broad RBAC role.
What interviewers listen for.
- Do you say "ADLS Gen2 is Blob storage plus a hierarchical namespace" in the first sentence? — senior signal.
- Do you connect atomic directory rename to Spark commit correctness unprompted? — required framing.
- Do you distinguish the
blobendpoint from thedfsendpoint and knowabfsstargetsdfs? — the detail that separates readers from users. - Do you reach for POSIX ACLs for fine-grained access and RBAC for coarse-grained rather than treating them as alternatives? — the whole governance story.
Worked example — the same directory rename, HNS off versus on
Detailed explanation. The cleanest way to feel what HNS buys you is to rename a "directory" both ways. On a flat account, raw/ is a prefix shared by thousands of blobs, so renaming it means the service (or your tool) must copy every blob to a new key and delete the old one — an O(n) operation that is not atomic. On an HNS account, raw is a single directory node, so renaming it flips one pointer — O(1) and atomic.
Question. A folder raw/ holds 1,000,000 blobs. What does rename raw -> bronze cost on a flat Blob account versus an HNS (ADLS Gen2) account, and is it atomic?
Input.
| account type |
raw/ contents |
operation |
|---|---|---|
| flat Blob | 1,000,000 blobs sharing prefix raw/
|
rename prefix to bronze/
|
| HNS (ADLS Gen2) | 1 directory object raw with 1,000,000 leaf files |
rename directory to bronze
|
Code.
# Flat Blob: no real rename — you copy then delete, per blob
az storage blob copy start-batch \
--destination-container data --destination-path bronze \
--source-container data --pattern "raw/*"
az storage blob delete-batch --source data --pattern "raw/*"
# ADLS Gen2 (HNS): one directory metadata operation
az storage fs directory move \
-f data --new-directory "data/bronze" -n "raw" \
--account-name mylake
Step-by-step explanation. On the flat account there is no directory to rename, so the tool enumerates every blob under raw/, issues a server-side copy to the bronze/ key, waits for each copy, then deletes each original — a million round trips, and if the job dies halfway the namespace is left half-renamed. On the HNS account, raw is one node; directory move rewrites a single metadata entry, so all million children are reparented instantly and the change is all-or-nothing.
Output.
| account type | operations | atomic? | rough time |
|---|---|---|---|
| flat Blob | ~1,000,000 copy + delete | no | minutes to hours |
| HNS (ADLS Gen2) | 1 metadata update | yes | milliseconds |
Rule of thumb. If your workload renames or moves directories — and every Spark commit does — you want the hierarchical namespace on; on flat Blob a "directory rename" is a full copy that can fail halfway.
2. Hierarchical namespace and atomic directory rename
Real directories make rename and delete single metadata operations — and that is exactly what Spark commit protocols need
The hierarchical namespace is the feature, and its payoff is a property most people take for granted on a local disk and lose on object storage: directory operations are atomic and independent of how many files the directory holds. On flat Blob, list, rename, and delete of a "folder" all scale with the number of objects. On HNS they are O(1) metadata operations, which is what lets analytics engines commit output safely.
What the namespace actually stores.
-
Filesystem = container. The
abfss://URI names a container as the filesystem:abfss://<filesystem>@<account>.dfs.core.windows.net/<path>. Below it, directories are real nodes, not prefixes. - Directories are objects with metadata. Each directory carries its own owner, group, permission bits, and ACL — none of which exists in the flat model.
- Atomic rename and move. Renaming or moving a directory rewrites one parent pointer; every descendant is reparented in a single, all-or-nothing operation.
- Cheap recursive delete. Deleting a directory tree is one metadata operation on HNS, versus enumerate-and-delete-each-blob on flat Blob.
Why Spark and Hadoop depend on this.
-
The commit dance is a rename. The FileOutputCommitter writes task output to a
_temporarydirectory, then on success renames it into the final path. If rename is slow or non-atomic, commits are slow and can corrupt output on failure. -
_SUCCESSand job atomicity. A partly-renamed output on flat Blob can leave readers seeing half a dataset; atomic directory rename makes "all rows or none" hold. - Listing drives planning. Spark lists input directories to plan splits; O(1)-ish directory listing on HNS beats paging through a flat prefix of millions of keys.
-
Use
abfss, notwasb. The modernabfss://driver (ABFS, secured with TLS) is the supported path to thedfsendpoint; the legacywasb://Blob driver does not use the hierarchical namespace semantics.
The trade-offs to name.
- HNS has a small metadata cost. Directory transactions are billed and there is bookkeeping overhead versus a pure flat store, negligible for analytics but worth naming.
-
Not every tool speaks
dfs. Some older Blob-only tools target theblobendpoint; multi-protocol access means they still work, but they see the flat view and its rename cost.
Worked example — a Spark write commit on HNS
Detailed explanation. The everyday place atomic rename earns its keep is a Spark write. Spark writes each task's output into a staging directory and, when the job succeeds, renames staging into the final location. On HNS that final rename is one metadata operation, so the commit is fast and cannot leave a half-written dataset.
Question. A Spark job writes df to abfss://data@mylake.dfs.core.windows.net/bronze/orders. Walk through where bytes land and what the commit does on ADLS Gen2.
Input. A DataFrame with 3 output partitions written by 3 tasks.
Code.
(df.write
.mode("overwrite")
.format("parquet")
.save("abfss://data@mylake.dfs.core.windows.net/bronze/orders"))
Step-by-step explanation. Each of the 3 tasks writes its part-file into a task-scoped path under bronze/orders/_temporary/. As each task succeeds, its files are promoted toward the job-attempt directory. When the whole job succeeds, the committer renames the staged output into bronze/orders/ and writes a _SUCCESS marker — on HNS these renames are atomic metadata operations, so no reader ever sees a partial orders/. If any task had failed, the _temporary tree is discarded and the final path is untouched.
Output.
| stage | path | visible to readers? |
|---|---|---|
| task write | bronze/orders/_temporary/.../part-*.parquet |
no |
| atomic commit |
bronze/orders/part-*.parquet + _SUCCESS
|
yes (all at once) |
| on failure |
_temporary discarded |
nothing published |
Rule of thumb. On HNS a Spark commit is a cheap atomic rename; on flat Blob the same commit copies every part-file, which is why big writes to non-HNS accounts are slow and riskier.
ADLS Gen2 interview question on directory operations
Question. Your nightly job renames a staging directory holding ~2 million small files into its final partition path. On a plain Blob account the rename takes 40 minutes and once left the table half-published after a crash. How does moving to ADLS Gen2 fix both the speed and the correctness problem, and what exactly changes at the storage layer?
Solution Using the hierarchical namespace atomic rename
Code.
# Same logical operation, two very different storage costs.
# Plain Blob (flat): "rename" = server-side copy of every object, then delete.
# ADLS Gen2 (HNS): rename = one directory metadata operation, atomic.
final = "abfss://data@mylake.dfs.core.windows.net/gold/sales/dt=2026-03-05"
staging = "abfss://data@mylake.dfs.core.windows.net/gold/_staging/sales_20260305"
# On HNS this is a single atomic directory move:
dbutils.fs.mv(staging, final, recurse=True) # metadata-only reparent
Step-by-step trace.
| step | flat Blob behaviour | ADLS Gen2 (HNS) behaviour |
|---|---|---|
| 1 | enumerate 2,000,000 keys under staging | resolve 1 directory node sales_20260305
|
| 2 | copy each blob to the new prefix | rewrite one parent pointer |
| 3 | delete each old blob | (nothing to delete — same objects, new parent) |
| 4 | crash mid-way → half-copied, half-published | crash → operation either fully applied or not at all |
- On flat Blob there is no directory, so "rename" is 2,000,000 copy-then-delete round trips — that is the 40 minutes.
- Because those round trips are independent and non-transactional, a crash after copying 1,100,000 objects leaves the table half-published.
- On HNS,
sales_20260305is a single directory object; the move rewrites one metadata entry and reparents all 2,000,000 children at once. - That single metadata operation is atomic, so a crash leaves the namespace either fully renamed or unchanged — never half-published.
Output:
| metric | flat Blob | ADLS Gen2 (HNS) |
|---|---|---|
| storage operations | ~4,000,000 (copy + delete) | 1 metadata op |
| wall-clock | ~40 min | sub-second |
| partial-failure state | possible | impossible |
Why this works — concept by concept:
- Hierarchical namespace — storing directories as real objects means a directory has an identity to rename, so the operation touches one node instead of every leaf.
- Atomic rename — the reparent is a single metadata transaction, so it is all-or-nothing; there is no window where the table is half-published.
- O(1) not O(n) — cost is independent of the 2,000,000 children, turning a 40-minute copy into a sub-second pointer flip.
- Commit correctness — Spark/Hadoop committers rely on rename atomicity; HNS gives object storage the guarantee they were designed around.
- Cost — the move is O(1) metadata versus O(n) data copies; time and dollar cost both collapse from linear in file count to constant.
Partitioning
Topic — partitioning
Directory-layout and partition-move problems
3. POSIX ACLs versus Azure RBAC
Two layers guard every object — coarse RBAC roles first, then fine-grained POSIX ACLs, and you must know which fires when
ADLS Gen2 has two authorization systems that both apply, and the interview question is almost always "explain how they interact." Say it in one breath: Azure RBAC is coarse-grained and evaluated first; POSIX ACLs are fine-grained and evaluated only when no RBAC role already grants access. Confusing them — or thinking you pick one — is the classic mistake.
Azure RBAC — the coarse layer.
- Scope is broad. RBAC roles are assigned at subscription, resource-group, storage-account, or container (filesystem) scope — never per file or per directory.
-
Data-plane roles.
Storage Blob Data Reader,Storage Blob Data Contributor, andStorage Blob Data Ownergrant read/write/own across everything in their scope. -
Owner is a superuser.
Storage Blob Data Ownerbypasses ACL checks entirely — it is the POSIX superuser for that scope, so ACLs never restrict it.
POSIX ACLs — the fine layer.
-
Per file and per directory. Every file and directory has an owning user, an owning group, and an
otherclass, each with read/write/execute (rwx) bits. -
Named entries. Beyond the three POSIX classes you add named-user and named-group entries — e.g. grant
group:analystsr-xon one folder — up to a hard limit of 32 entries per file or directory. -
Execute means traverse. On a directory, the
xbit means "traverse into"; to read/a/b/ca principal needsxon/aand/a/b, thenronc. Forgetting traverse permissions on parents is the #1 ACL bug.
Access ACLs versus default ACLs.
- Access ACL. Controls access to this object right now. It exists on both files and directories.
- Default ACL. A template that lives only on directories. When a new child is created, it copies the parent's default ACL as its own access ACL (and, for a child directory, its default ACL too).
- Inheritance is at creation only. Changing a directory's default ACL does not retroactively rewrite children that already exist — you must re-apply recursively. This surprises people constantly.
The evaluation order (the answer they want).
- RBAC first. If an assigned role grants the requested action at a covering scope, access is allowed and ACLs are not consulted.
- ACLs second. If no RBAC role grants it, the POSIX ACL on the target (and traverse on its parents) decides.
- Principals are Microsoft Entra identities. Users, groups, service principals, and managed identities from Microsoft Entra ID (formerly Azure AD) are what both layers evaluate against.
Worked example — grant one group read on a folder without a new role
Detailed explanation. The everyday governance task is "let the analysts read curated/finance/ but nothing else, without giving them a container-wide role." That is precisely what a named-group ACL entry plus a default ACL is for: an access ACL grants the existing files, and a default ACL makes future files inherit the grant.
Question. Give group:analysts read access to everything under curated/finance/, existing and future, using ACLs (not RBAC). What entries do you set?
Input.
| target | current access | goal |
|---|---|---|
curated/ (parent) |
analysts: none | traverse only |
curated/finance/ |
analysts: none | read + traverse, now and future |
Code.
# Traverse (x) on the parent so analysts can reach the folder
az storage fs access set --acl "group:analysts:--x" \
-p "curated" -f data --account-name mylake
# Read+traverse on the target folder (access ACL = existing objects)
az storage fs access set-recursive \
--acl "group:analysts:r-x" \
-p "curated/finance" -f data --account-name mylake
# Default ACL so NEW children inherit the grant at creation
az storage fs access set \
--acl "default:group:analysts:r-x" \
-p "curated/finance" -f data --account-name mylake
Step-by-step explanation. The --x on curated lets analysts traverse into the tree without being able to list it. The recursive r-x sets the access ACL on curated/finance and everything already inside it, so existing files become readable. The default:group:analysts:r-x entry is a template on the directory: any file created later copies it as its own access ACL, so new data is readable without re-running the recursive apply. No RBAC role was granted, so analysts see only this subtree.
Output.
| principal | curated/ |
curated/finance/* (existing) |
files added tomorrow |
|---|---|---|---|
| analysts | traverse only | read | read (via default ACL) |
| everyone else | unchanged | unchanged | unchanged |
Rule of thumb. Access ACLs fix today's files; default ACLs fix tomorrow's. Set both, and remember default ACLs never rewrite objects that already exist.
ADLS Gen2 interview question on ACLs versus RBAC
Question. A service principal has the Storage Blob Data Contributor role on the whole storage account, but you also removed its ACL entry from curated/finance/ so it "cannot" read finance data. During an audit it reads a finance file anyway. Explain why, and how the two authorization layers combined to allow it.
Solution Using the RBAC-then-ACL evaluation order
Code.
Request: service principal SP reads
abfss://data@mylake.dfs.core.windows.net/curated/finance/ledger.parquet
Layer 1 — Azure RBAC (evaluated FIRST):
SP has "Storage Blob Data Contributor" at account scope
-> read is granted by the role -> ACCESS ALLOWED, ACLs not consulted
Layer 2 — POSIX ACL (only reached if RBAC did NOT grant):
SP removed from curated/finance ACL -> would deny
... but this layer is never evaluated, because RBAC already allowed.
Step-by-step trace.
| step | check | result |
|---|---|---|
| 1 | RBAC: does any role grant read at a covering scope? |
yes — Data Contributor at account scope |
| 2 | RBAC allowed → short-circuit | ACL check skipped entirely |
| 3 | (unreached) ACL: is SP granted on curated/finance? |
would be deny, but never evaluated |
- Authorization checks RBAC first; a data-plane role at account scope covers
curated/finance/ledger.parquet. - Because
Storage Blob Data Contributorincludes read, the request is allowed at layer 1 and evaluation stops. - The ACL layer — where you removed the principal — is only consulted when RBAC does not already grant the action, so your ACL edit was dead code.
- To actually restrict SP, you must narrow its RBAC (remove the broad role or scope it to a different container), then let ACLs express the fine-grained grant.
Output:
| you changed | intended effect | real effect |
|---|---|---|
| removed SP from ACL | deny finance read | none — RBAC still allows |
| (fix) remove broad RBAC role | ACLs govern | ACL deny now takes effect |
Why this works — concept by concept:
- RBAC-first evaluation — a granting role short-circuits authorization, so ACLs can only ever add access on top of RBAC, never subtract from it.
-
Superuser semantics —
Storage Blob Data Owner/broad roles behave like a POSIX superuser and bypass ACLs by design, which is powerful and dangerous. - Least privilege at the right layer — coarse restrictions belong in RBAC scope; fine-grained grants belong in ACLs; you cannot fix an over-broad role with an ACL.
- Auditability — knowing the order tells you where to look first when access is unexpectedly allowed: check the role assignments before the ACLs.
- Cost — the check is O(1) role lookup then, if needed, O(depth) traverse checks up the directory path; negligible per request.
Access control
Topic — access-control
Permission-model and least-privilege problems
4. Partitioning layout and the small-files problem
The directory layout you choose is a performance contract — partition for pruning, size files right, and never let the tiny files pile up
Because ADLS Gen2 gives you a real filesystem, how you lay out directories and how big you make files is the single biggest lever on query cost. Two failure modes dominate interviews: a layout that forces every query to scan everything, and a "small-files problem" where millions of tiny files crush the engine with per-file overhead. Say the principle in one line: partition on the columns you filter, size files at 128 MB–1 GB, and compact before the small files multiply.
Partitioning for pruning.
-
Layout by filter column. Write
.../sales/year=2026/month=03/day=05/part-*.parquet. The Hive-stylekey=valuefolders let the engine skip whole branches — this is partition pruning. -
Prune only helps if you filter on it. Partitioning by
dayspeedsWHERE day = '2026-03-05'but does nothing forWHERE customer_id = 42; partition on the predicate you actually run. - Pick the right cardinality. A partition column with a handful to a few thousand distinct values (date, region) prunes well; one with millions of distinct values (user id) explodes the directory count.
The small-files problem.
- What it is. Thousands or millions of tiny files (KBs each) instead of fewer right-sized files. Each file is a separate open, a separate metadata entry, and a separate task — overhead swamps useful work.
- Where it bites. Listing the directory is slow, the driver plans one split per tiny file, and columnar formats lose compression and row-group efficiency at small sizes.
- Target size. Aim for 128 MB–1 GB per file (roughly one HDFS/Spark block-ish unit). Below ~a few MB you are firmly in the danger zone.
How over-partitioning creates small files.
-
Too many partition levels. Partitioning by
year/month/day/hour/minuteon modest volume leaves each leaf with a sliver of data, so every write drops a tiny file. - Too many writer tasks. N Spark tasks each writing to M partitions produce up to N×M files per run; without a repartition/coalesce you flood the lake with fragments.
- Streaming micro-batches. Frequent small appends (one file per micro-batch) are the classic streaming small-files source.
The fixes.
-
Repartition before write.
df.repartition("day")(orrepartitionByRange) so each partition is written by one task → one file per partition. -
Compaction / OPTIMIZE. Periodically rewrite many small files into few large ones — Delta Lake's
OPTIMIZE(withZORDERfor data-skipping) or a scheduled compaction job. -
Right-size the partition grain. Coarsen from
hourtoday(or addbucketing) when partitions are too small; the goal is fewer, fuller files.
Worked example — repartition to write one file per partition
Detailed explanation. The most common small-files fix is to align the write parallelism with the partition column. If 200 Spark tasks each write to all 30 day-partitions, you get up to 6,000 files. Repartitioning the DataFrame by the partition column first means each partition is handled by one task, so each day gets one right-sized file.
Question. A job writes 30 days of data partitioned by day using 200 tasks, producing ~6,000 tiny files. Rewrite it so each day-partition lands as a single file, and show the file-count change.
Input.
| setting | before | goal |
|---|---|---|
| partition column |
day (30 distinct) |
day |
| writer tasks | 200 | aligned to partitions |
| files produced | ~6,000 | ~30 |
Code.
# Before: 200 tasks x 30 day-partitions -> up to 6,000 small files
# After: repartition by the partition key so one task writes one day
(df.repartition("day") # shuffle so each day is one partition
.write
.mode("overwrite")
.partitionBy("day") # Hive-style day=... folders
.parquet("abfss://data@mylake.dfs.core.windows.net/gold/sales"))
Step-by-step explanation. repartition("day") shuffles rows so all rows for a given day sit in one Spark partition. partitionBy("day") writes the Hive-style day=YYYY-MM-DD/ folders for pruning. Because each day is now one Spark partition handled by one task, each day=... folder receives a single part-file instead of up to 200. The result is ~30 right-sized files that prune cleanly and open fast.
Output.
| layout | files | per-file size | query on one day |
|---|---|---|---|
| before | ~6,000 | ~KBs | slow: list + open thousands |
| after | ~30 | ~hundreds of MB | fast: prune to 1 folder, open 1 file |
Rule of thumb. Match write parallelism to partition cardinality: repartition(partitionCol) before partitionBy(partitionCol) turns N×M fragments into one file per partition.
ADLS Gen2 interview question on the small-files problem
Question. A streaming job appends one small Parquet file per minute to events/day=.../, and after three months queries over a day have become painfully slow even though each query only touches one day-partition. Diagnose why, and give a fix that keeps the recent-data latency but restores query speed.
Solution Using a scheduled compaction into right-sized files
Code.
# Symptom: 60 files/hour x 24 x 90 days = ~130,000 tiny files per day-partition.
# Fix: keep streaming for freshness, compact yesterday's partition into few files.
from pyspark.sql import functions as F
day = "2026-03-05"
src = f"abfss://data@mylake.dfs.core.windows.net/events/day={day}"
compacted = (spark.read.parquet(src)
.repartition(4)) # 4 right-sized files, not 130k tiny
(compacted.write
.mode("overwrite")
.parquet(src)) # atomic overwrite of the partition
# On Delta Lake this is simply:
# OPTIMIZE events WHERE day = '2026-03-05'
Step-by-step trace.
| step | state of day=2026-03-05
|
files | query cost |
|---|---|---|---|
| 1 | streaming appended all day | ~130,000 tiny | list + open 130k → slow |
| 2 | read partition, repartition(4)
|
in memory, 4 partitions | — |
| 3 | overwrite partition atomically | 4 right-sized files | list + open 4 → fast |
| 4 | today keeps streaming | small files (fresh) | acceptable for one live day |
- The slowness is not scanning too much data — pruning already limits it to one day — it is per-file overhead: 130,000 opens, 130,000 metadata reads, 130,000 tasks.
- Reading the partition and
repartition(4)collapses those rows into four in-memory partitions. - Overwriting the partition (atomic on HNS) replaces the 130,000 fragments with four right-sized files, so the same query now opens four files.
- Only the current day keeps the streaming small files, so freshness is preserved while every closed day is compacted — the standard "hot tail, compacted history" pattern.
Output:
| metric | before compaction | after compaction |
|---|---|---|
| files in partition | ~130,000 | 4 |
| files opened per query | ~130,000 | 4 |
| query time | minutes | seconds |
Why this works — concept by concept:
- Small-files problem — cost was dominated by per-file open/metadata/task overhead, not bytes scanned, so fewer larger files fix it without changing the data.
- Compaction — rewriting a partition into a handful of right-sized files restores columnar efficiency and cheap listing.
- Hot tail vs cold history — leaving only the live partition fragmented preserves streaming latency while every completed partition is compacted.
- Atomic overwrite — HNS makes the partition overwrite all-or-nothing, so readers never catch the partition mid-compaction.
- Cost — compaction is O(rows in partition) once per day, amortized against every future query that now opens 4 files instead of 130,000.
Partitioning
Topic — partitioning
Partition-grain and file-sizing problems
5. Access tiers, abfss and identity-based security
Move cold data down the tiers, reach the lake over abfss, and authenticate with an identity instead of a key
The last cluster the interview covers is operational: how you pay for storage (access tiers), how you connect (the abfss driver), and how you prove who you are (identity-based auth). Get these right and you cut cost without losing durability and connect Spark securely without a secret sprawled across notebooks.
Access tiers — pay for what you touch.
- Hot. Highest storage price, lowest access price. For data read frequently — active partitions, curated marts.
- Cool. Lower storage, higher access price; minimum ~30 days. For infrequently accessed data you still read occasionally.
- Cold. Cheaper still, higher access price; minimum ~90 days. For rarely accessed data that must stay online.
- Archive. Cheapest storage, offline — you must rehydrate (to Hot/Cool) before reading, which can take hours; minimum ~180 days. For compliance retention you almost never read.
- Set per blob or as an account default. Tiering is set at the blob level (or by lifecycle-management rules that auto-move blobs by age), and early-deletion fees apply if you delete/move before the minimum.
The abfss:// driver — how engines connect.
-
ABFS = Azure Blob File System. The
abfss://scheme (thes= secure/TLS) is the Hadoop driver that talks to thedfsendpoint and uses hierarchical-namespace semantics. -
URI shape.
abfss://<filesystem>@<account>.dfs.core.windows.net/<path>— filesystem is the container, host is thedfsendpoint. -
Prefer it over
wasbs. The legacywasbs://(WABS) Blob driver predates ADLS Gen2 and does not use directory semantics; useabfssfor lakes.
Identity-based security — stop pasting keys.
- Account key / connection string. Full control, no rotation story, easy to leak — avoid in shared code.
- SAS (Shared Access Signature). A time-boxed, scope-limited token; good for handing narrow, expiring access to an external party or a single job.
- Service principal. A Microsoft Entra app identity with a client id/secret (or certificate); the classic choice for a scheduled pipeline, governed by RBAC + ACLs.
- Managed identity. An Entra identity Azure manages for you — no secret to store or rotate. The preferred production auth for Azure-hosted compute (Databricks, Synapse, VMs, Functions): grant the identity a role, and code authenticates with no credential in it.
Choosing auth by workload.
- Azure-hosted, long-lived pipeline → managed identity (no secret).
- Off-Azure system or automated job with its own lifecycle → service principal.
- Narrow, expiring, third-party access → SAS.
- Anything shared or checked into a repo → never the account key.
Worked example — connect Spark with a managed identity, no secret
Detailed explanation. The production-clean way to read the lake from Azure-hosted Spark is a managed identity: you assign the compute's identity a data-plane role (or ACLs), and the abfss client authenticates with an Entra OAuth token minted at runtime — nothing secret lives in the notebook or config.
Question. Configure Spark to read abfss://data@mylake.dfs.core.windows.net/curated/ using a managed identity, with no account key or client secret in the code.
Input. A managed identity granted Storage Blob Data Reader on the data filesystem.
Code.
acct = "mylake.dfs.core.windows.net"
# Tell ABFS to authenticate with the managed identity (OAuth), not a key.
spark.conf.set(f"fs.azure.account.auth.type.{acct}", "OAuth")
spark.conf.set(f"fs.azure.account.oauth.provider.type.{acct}",
"org.apache.hadoop.fs.azurebfs.oauth2.MsiTokenProvider")
# No secret anywhere — the platform mints the token for the assigned identity.
df = spark.read.parquet(f"abfss://data@{acct}/curated/")
Step-by-step explanation. Setting the auth type to OAuth with the MsiTokenProvider tells the ABFS driver to request an access token for the compute's managed identity from the Azure instance metadata endpoint at runtime. That token is presented to the dfs endpoint, which checks RBAC (the Storage Blob Data Reader role) and, if needed, ACLs. Because the identity is managed, there is no client secret to store or rotate — the credential never appears in code, config, or a secret store.
Output.
| aspect | account key | managed identity |
|---|---|---|
| secret in code/config | yes (leakable) | none |
| rotation | manual | automatic (platform) |
| authorization | all-or-nothing key | RBAC + ACLs on the identity |
Rule of thumb. On Azure-hosted compute, authenticate to abfss with a managed identity and grant it a scoped role — you get least privilege and zero secrets to rotate.
ADLS Gen2 interview question on tiers and cost
Question. Five years of raw event history sits in the Hot tier and is read maybe once a year for a compliance export, yet it dominates your storage bill. How do you cut the storage cost while keeping the data retrievable, and what operational caveats must you flag to stakeholders?
Solution Using lifecycle rules to tier data to Archive
Code.
{
"rules": [
{
"name": "raw-events-to-archive",
"type": "Lifecycle",
"definition": {
"filters": { "blobTypes": ["blockBlob"],
"prefixMatch": ["data/raw/events/"] },
"actions": {
"baseBlob": {
"tierToCool": { "daysAfterModificationGreaterThan": 30 },
"tierToArchive": { "daysAfterModificationGreaterThan": 180 }
}
}
}
}
]
}
Step-by-step trace.
| age of blob | tier | storage cost | readable? |
|---|---|---|---|
| 0–30 days | Hot | highest | instant |
| 30–180 days | Cool | lower | instant (higher read cost) |
| 180+ days | Archive | lowest | only after rehydration (hours) |
- The lifecycle rule matches only the
data/raw/events/prefix, so hot curated data is untouched. - After 30 days without modification a blob auto-moves to Cool, cutting storage cost for data that is no longer actively read.
- After 180 days it moves to Archive — the cheapest tier — which is offline storage.
- Reading an archived blob requires a rehydration to Hot/Cool that can take hours, so the once-a-year compliance export must schedule a rehydrate step first; and early deletion before the tier minimum incurs a charge.
Output:
| metric | all-Hot (before) | lifecycle-tiered (after) |
|---|---|---|
| storage cost of old data | highest | ~Archive price (a fraction) |
| retrieval latency (old) | instant | hours (rehydrate first) |
| operational caveat | none | plan rehydration + min-duration fees |
Why this works — concept by concept:
- Access tiers — storage price falls Hot → Cool → Cold → Archive while retrieval price/latency rises, so matching tier to read frequency minimizes total cost.
- Lifecycle management — age-based rules move blobs automatically, so tiering is policy, not a manual job someone forgets.
- Archive is offline — the cheapest tier trades instant reads for a rehydration delay, which is fine for once-a-year data but must be flagged to consumers.
- Minimum durations — Cool/Cold/Archive carry early-deletion charges, so tiering churny data down and back up can cost more than leaving it Hot.
- Cost — for rarely read history the storage saving dominates the occasional rehydration fee, so the yearly-export dataset belongs in Archive.
Access control
Topic — access-control
Identity, SAS and least-privilege auth problems
Cheat sheet — ADLS Gen2 recipes
Read over abfss (Spark).
df = spark.read.parquet(
"abfss://data@mylake.dfs.core.windows.net/curated/orders")
Set an access ACL (existing object).
az storage fs access set --acl "group:analysts:r-x" \
-p "curated/finance" -f data --account-name mylake
Set a default ACL (inherited by new children).
az storage fs access set --acl "default:group:analysts:r-x" \
-p "curated/finance" -f data --account-name mylake
Move a directory (atomic on HNS).
az storage fs directory move -f data \
--new-directory "data/bronze" -n "raw" --account-name mylake
Set a blob's access tier.
az storage blob set-tier --container-name data \
--name "raw/events/2022/part-000.parquet" --tier Archive \
--account-name mylake
Repartition to avoid small files, then partition-write.
(df.repartition("day")
.write.mode("overwrite").partitionBy("day")
.parquet("abfss://data@mylake.dfs.core.windows.net/gold/sales"))
Auth picker.
| Situation | Auth |
|---|---|
| Azure-hosted long-lived pipeline | managed identity |
| Off-Azure / external automated job | service principal |
| Narrow, time-boxed, third-party access | SAS token |
| Shared code or repo | never the account key |
Frequently asked questions
What is ADLS Gen2?
ADLS Gen2 (Azure Data Lake Storage Gen2) is Azure Blob storage with the hierarchical namespace enabled — a StorageV2 account where the namespace flag is switched on at creation. That gives you real directories, atomic directory renames, and per-file POSIX ACLs on top of Blob storage's durability and tiers, reachable through both the classic blob endpoint and the dfs (Data Lake) endpoint. It is the standard object store for analytics and lakehouse workloads on Azure.
How is ADLS Gen2 different from Blob storage?
They are the same underlying service; ADLS Gen2 is Blob storage with the hierarchical namespace turned on. Plain Blob has a flat namespace where "folders" are just prefixes in a blob name, so renaming or deleting a folder means copying or deleting every object (O(n), non-atomic). ADLS Gen2 stores directories as real objects, so rename, move, and delete are single atomic metadata operations, and it adds POSIX ACLs and the dfs endpoint that the abfss driver uses.
What is the difference between ACLs and RBAC in ADLS Gen2?
Azure RBAC is coarse-grained: roles like Storage Blob Data Reader/Contributor/Owner are assigned at account or container scope. POSIX ACLs are fine-grained: read/write/execute entries on each file and directory, including named-user and named-group grants (up to 32 entries). RBAC is evaluated first — if a role grants the action, access is allowed and ACLs are skipped (Data Owner is effectively a superuser). ACLs are only consulted when no RBAC role already grants access, so ACLs add access on top of RBAC rather than subtracting from it.
Why does the hierarchical namespace matter for Spark?
Spark and Hadoop commit protocols write task output to a temporary directory and then rename it into the final location on success. On a flat Blob account that rename is a full copy of every file and is not atomic, so big writes are slow and a crash can leave a half-published table. With the hierarchical namespace, a directory rename is a single atomic metadata operation, so commits are fast and all-or-nothing — which is why you use the abfss driver against an HNS account for lakehouse workloads.
What is the small-files problem and how do I fix it?
The small-files problem is having thousands or millions of tiny files instead of fewer right-sized ones, so per-file open, metadata, and task overhead dominates query cost even when pruning limits the data scanned. It comes from over-partitioning, too many writer tasks (N×M files), or streaming one file per micro-batch. Fix it by targeting 128 MB–1 GB files: repartition(partitionCol) before partitionBy(...), coarsen the partition grain, and run periodic compaction (or Delta Lake OPTIMIZE) on completed partitions.
What is the abfss driver?
abfss:// is the Azure Blob File System (ABFS) Hadoop driver, secured with TLS, that connects Spark/Hadoop to the ADLS Gen2 dfs endpoint using hierarchical-namespace semantics. The URI is abfss://<filesystem>@<account>.dfs.core.windows.net/<path>, where the filesystem is the container. It replaces the legacy wasbs:// Blob driver, which predates ADLS Gen2 and does not use directory semantics; authenticate it with a managed identity, service principal, SAS, or (avoid) the account key.
Practice on PipeCode
Pipecode.ai is Leetcode for Data Engineering — every ADLS Gen2 idea above, from the atomic directory rename to the RBAC-then-ACL evaluation order and the repartition-then-compact fix for small files, maps to a hands-on practice room where you design the layout and reason about access against real graded inputs. PipeCode pairs each reading with 450+ DE-focused problems and a real-time scoring engine, so your answer to "how would you keep this lake fast and locked down?" holds up under a senior interviewer's depth probes.
Practice partitioning problems now →
Access-control drills →





Top comments (0)