DEV Community

Cover image for Fabric Runtime 2.0 is becoming the default. Here is what actually breaks.
Firat Celik
Firat Celik

Posted on

Fabric Runtime 2.0 is becoming the default. Here is what actually breaks.

Fabric Runtime 2.0 is becoming the default. Here is what actually breaks.

Late September 2026 is when Fabric Runtime 2.0 stops being an opt-in experiment and becomes the path of least resistance for new workspaces and environment items. If you treat that as a release note, you will discover the breakages in production.

I work as a Lead Data Engineer on Azure, Fabric, Databricks, and Spark. This is the migration checklist I wish teams ran before the default flipped under them.

Docs to keep open while you read:

1) Hook: the default is the migration

Microsoft's Runtime 2.0 page is blunt: GA and ready for production, but not the default yet. You opt in. The stated plan is to make Runtime 2.0 the default UX selection and the default for new workspaces and new environment items in late September 2026.

Today is 28 September 2026. That window is now.

Existing workspaces, jobs, and Environment items stay on whatever you configured. The trap is quieter. Every new workspace for a pilot, every Environment someone creates because "defaults are fine," starts landing on Spark 4.1 / Delta 4.2 / Python 3.13 once the plan executes.

Treat the flip as an incident rehearsal date. Pin what you care about. Validate 2.0 on a scoped Environment. Do not discover Python wheel breakage through a failed Monday job.

Runtime 1.3 still has runway. Secondary writeups put Microsoft's LTS window for 1.3 roughly from early October 2026 through March 2027. Verify against current lifecycle docs before you commit a program date. LTS means supported, not "ignore migration."

2) Wrong assumption: GA means notebooks just run

Runtime 1.3 to Runtime 2.0 is not a patch bump. It is a stack jump on nearly every component:

Component Runtime 1.3 (typical) Runtime 2.0
Apache Spark 3.5 4.1
Delta Lake 3.2 4.2
Python 3.11 3.13
Java 11 21
Scala 2.12 2.13
OS (prior Fabric OS) Azure Linux 3.0 (Mariner 3.0)

Industry writeups also call out Delta protocol defaults moving higher (reader/writer versions, Deletion Vectors in the 2.0 path). Treat protocol as a first-class migration risk, not a footnote. See Microsoft's Delta Lake table format interoperability guidance for what cross-experience readers can actually open.

Practical fallout I care about:

  • Python 3.11 -> 3.13. Environment-level libraries may migrate automatically. Validation does not. Custom wheels and pinned packages are your problem. Session-scoped %pip install and notebook Resources are a different blast radius than Environment items.
  • Scala 2.12 -> 2.13 / Java 11 -> 21. JARs and connectors that "worked forever" on 1.3 deserve a dual-run, not hope.
  • SparkR is deprecated in Spark 4.x and may be removed later. If you still have SparkR notebooks, the clock is already running.

GA means Microsoft ships it for production use. It does not mean your Environment's library graph survives the jump without you watching.

3) Wrong assumption: NEE = free 6x

Runtime 2.0 includes the Native Execution Engine (NEE). Microsoft's docs describe it as a vectorized path based on Apache Gluten (offload middle layer) and Velox (C++ engine from Meta). Supported operators leave the JVM and run columnar with SIMD-style acceleration on Parquet, Delta, and (now) CSV. Microsoft cites up to six times faster vs open-source Spark on TPC-DS. That is a vendor benchmark on a specific workload family. It is a ceiling story, not your median job.

How you enable it:

  • Environment -> Spark compute -> Acceleration -> enable native execution engine (preferred UX path), or
  • Spark conf: spark.native.enabled (also usable per notebook / SJD via %%configure)

Hard limits from the docs that matter in production:

  • Structured streaming is not supported. Use streaming? Plan for JVM (or keep NEE off for those jobs).
  • Unsupported ops fall back to classic Spark. Fallback is automatic. Your job keeps running. Your SLA may not.
  • Every boundary between native columnar and JVM row-oriented execution pays a conversion tax. That is the part marketing slides skip.
  • Fabric Spark Advisor can surface fallback alerts in notebook cell output. Use them. Do not ignore green "job succeeded" when Advisor is screaming about offload misses.

NEE also preserves AQE, cost-based rewrites, column pruning, and predicate pushdown when operators do offload. Good. That does not help if half the plan is bouncing in and out of Velox.

Blunt take: a vectorized engine that falls back on half your plan is not acceleration. It is a tax you opted into because a dashboard said "up to 6x."

4) How to see the truth in the plan

Do not argue with anecdotes. Read the plan.

In Spark UI / History Server, look for node names ending in:

  • *Transformer (for example ProjectExecTransformer, BroadcastHashJoinExecTransformer, RollUpHashAggregateTransformer)
  • *NativeFileScan
  • VeloxColumnarToRowExec

Those suffixes mean native offload happened for that node. Light-blue JVM nodes next to green native nodes (Microsoft's Gluten SQL / DataFrame tab uses that color language) are your fallback map. Every VeloxColumnarToRowExec (and the reverse path) is a place where you paid for format conversion.

Also useful: df.explain(), Fabric Spark Advisor fallback alerts, and the Gluten SQL / DataFrame tab's native-vs-JVM node counts.

Surgical disable for one cell (unsupported op, correctness edge, or A/B):

spark.conf.set('spark.native.enabled', 'false')
Enter fullscreen mode Exit fullscreen mode

Re-enable afterward if later cells should use NEE again. Spark runs cells in order. Conf sticks.

5) Delta 4.x trap for multi-workload tables

This is the silent one.

Microsoft is explicit: Delta Lake 4.2-specific features are experimental and only work on Spark experiences (notebooks and Spark Job Definitions). If the same Delta tables must be used across multiple Fabric workloads, do not enable those features.

Why it hurts: SQL analytics endpoint, Power BI Direct Lake, and other non-Spark readers do not all speak the same protocol / feature set. Protocol upgrades (higher reader/writer versions, Deletion Vectors, experimental checkpoints) are effectively one-way doors. Recreate the table is the recovery path, not "uncheck the box." Shared bronze/silver tables that feed Spark and Direct Lake and SQL endpoint are exactly where someone enables a shiny 4.x feature on a Friday and breaks Monday's semantic model.

Rule I use: default to the most constrained reader in the estate. If Direct Lake or the SQL endpoint must read it, keep features inside what those readers support. Save experimental Delta 4.x toys for Spark-only sandboxes.

6) Practical migration pattern

This is the runbook shape I actually use.

1. Inventory

  • Workspaces and Environment items: which runtime are they on?
  • Jobs that create new Environments without pinning
  • Libraries: Environment wheels vs session %pip
  • Tables read by more than one Fabric experience
  • Streaming jobs (NEE: leave off or isolate)
  • SparkR notebooks (debt clock)

2. Baseline on 1.3

  • Duration, shuffle bytes, spill, failure rate, data quality checks
  • Capture plans for critical queries before you change anything
  • Note current Delta protocol / table features on shared tables

3. Dual-run on 2.0

  • Scoped Environment item first (surgical), not whole workspace
  • Same input snapshot, same checks, compare plans and SLA
  • Turn NEE on only after the JVM path is green on 2.0
  • Kill vanity actions that force row materialization (yes, that includes habit .show() in hot paths during timing runs)

4. Rollback in the runbook before you need it

  • Runtime: pin Environment / workspace back to 1.3 while LTS covers you
  • NEE only:
spark.conf.set('spark.native.enabled', 'false')
Enter fullscreen mode Exit fullscreen mode

Put both in the on-call doc. A rollback you invent at 02:00 is not a rollback.

7) CI/CD angle: pin the runtime like a toolchain

You would not let python float to whatever the image shipped last Tuesday without a lockfile. Treat Fabric runtime the same way.

  • Pin runtime version on Environment items that production notebooks and SJDs attach to
  • Fail the pipeline if an Environment publishes without an explicit runtime
  • Gate promotion: 1.3 baseline metrics vs 2.0 dual-run metrics must be compared, not vibes
  • Separate "NEE enabled" as its own promotion flag. Runtime upgrade and native offload are two changes. Ship them as two changes.
  • Use Fabric runtime release channels (preview) to rehearse upcoming defaults on custom pools. Do not park production there.

If CI deploys notebook JSON and never asserts runtime + Acceleration settings, you are continuous-delivering entropy.

8) Close: benchmarks sell the ceiling, production pays the median

Workload type Enable NEE? Watch for...
Large Parquet/Delta batch ETL, heavy agg/join Yes, after dual-run Fallback density, conversion boundaries, Advisor alerts
Interactive lakehouse SQL / notebook exploration Often yes Date type mismatches on filters, wide UDF usage
Structured streaming No (unsupported) Accidental enable via Environment inheritance
JSON/XML heavy paths Limited / often no Format falls back to JVM; measure before celebrating
Shared Delta tables (Spark + SQL endpoint / Direct Lake) NEE optional; Delta 4.x experimental features: no Protocol / feature lockout of non-Spark readers
SparkR or exotic Scala libs Maybe later Language deprecation and binary incompat first

Takeaways

  1. Pin Runtime 1.3 on anything you cannot afford to surprise, then opt into 2.0 on a scoped Environment before the default teaches your newest workspace the hard way.
  2. Treat NEE as a plan shape problem, not a toggle. *Transformer and VeloxColumnarToRowExec tell the truth; "up to 6x" does not.
  3. Never enable experimental Delta 4.x features on multi-experience tables. Shared tables inherit the weakest reader's constraints.

Benchmarks sell the ceiling. Production pays the median. Migrate like you believe that.

What is your scarier unpaid bill right now: Python 3.13 libraries, NEE fallback thrash, or a Delta feature someone enabled on a table Direct Lake still needs to read?

Sources

Top comments (2)

Collapse
 
firfircelik profile image
Firat Celik •

Thanks Dev Supports, I'm gonna believe you and will share my credit card info into that link <3