Capture Spark runtime lineage from Microsoft Fabric notebooks so Atlan can catalog notebooks and Spark jobs with table- and column-level lineage.
Set up Spark runtime lineage Private Preview
Capture lineage from the Spark jobs that your Microsoft Fabric notebooks run, so Atlan can catalog notebooks and Spark jobs and link them to the tables they read and write, down to column level.
When a notebook runs, Fabric Spark emits an OpenLineage event for each Spark job. A Runtime Lineage item in the workspace stores these events as JSON files in OneLake. During each Microsoft Fabric crawl, Atlan reads the new event files and builds lineage from them. Atlan only reads the events: you enable and configure capture in Fabric by following this guide.
Runtime Lineage is a Microsoft Fabric private preview:
- Microsoft must enable the preview for your Fabric tenant. During the preview, Microsoft also pins each capture workspace to the Spark 4.1 validation image that provides the lineage transport. Contact Microsoft to enroll your tenant and workspaces.
- A Runtime Lineage item captures events only from the workspace it lives in, and each workspace supports one active capture.
- The event subscription API uses a private-preview route (
__private), which Microsoft expects to remove when the feature reaches public preview.
Prerequisites
Before you begin, make sure you have:
- Completed Set up Microsoft Fabric with Service Principal authentication. Atlan reads runtime lineage events from OneLake with an Azure Storage token issued to the service principal, which APIM Managed Identity authentication doesn't provide.
- Fabric Administrator privileges, to change tenant settings.
- The Admin, Member, or Contributor role on each workspace whose notebooks you want lineage for.
- Workspaces on a Fabric capacity. A Fabric Trial capacity works.
- The Azure CLI, signed in with
az loginas a user who holds one of the workspace roles listed. The API calls in this guide use it to get access tokens.
Enable tenant settings
Runtime Lineage is off until a Fabric administrator enables it for the tenant.
- Log in to the Fabric admin portal.
- Click the Settings icon on the top panel.
- Click Admin Portal under the Governance and insights section.
- Select Tenant Settings from the sidebar.
- Enable Users can create a Runtime Lineage item (preview). This setting makes the Runtime Lineage item type available. Apply it to the entire organization or to the security groups whose members set up capture, and click Apply.
- Enable Workspace admins can turn on Runtime Lineage monitoring for their workspaces (preview). This setting lets workspace administrators turn on capture. Apply it to the same scope and click Apply.
- Under OneLake settings, confirm that Users can access data stored in OneLake with apps external to Fabric is enabled. Atlan reads the event files from outside Fabric through the OneLake ADLS Gen2 APIs, and these reads fail when this setting is off.
Tenant setting changes can take 15–30 minutes to take effect.
Create Runtime Lineage item
Create one Runtime Lineage item in each workspace whose notebooks you want lineage for. The item is the OneLake destination for the workspace's lineage events. Spark doesn't create it automatically.
-
Open the Microsoft Fabric homepage and go to the workspace.
-
Copy the workspace ID from the browser address bar. It's the value after
/groups/in the URL. -
Click + New item, select Runtime Lineage, enter a name, and click Create.
-
Get the ID of the new item:
export FABRIC_TOKEN=$(az account get-access-token \--resource https://api.fabric.microsoft.com \--query accessToken -o tsv)curl -s -H "Authorization: Bearer $FABRIC_TOKEN" \"https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/items?type=RuntimeLineage"Copy the
idvalue from the response. This is the Runtime Lineage item ID.
Keep exactly one Runtime Lineage item per workspace. When a workspace contains more than one, Atlan can't tell which item holds the capture, and the crawl fails.
To automate setup across many workspaces, create the item with the Fabric Create Item API instead. Send "type": "RuntimeLineage" and a displayName in the request body. The response contains the item ID.
Turn on workspace capture
Subscribe the Runtime Lineage item to its workspace. After the subscription exists, Spark in that workspace sends lineage events to the item. Creating a subscription requires write permission on the item, which the Admin, Member, and Contributor workspace roles include.
-
Create the subscription. Use the same workspace ID in the URL and in the request body:
curl -s -X POST \-H "Authorization: Bearer $FABRIC_TOKEN" \-H "Content-Type: application/json" \-d '{"type": "Workspace", "publisher": {"workspaceId": "<workspace-id>"}}' \"https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/RuntimeLineages/<runtime-lineage-item-id>/__private/eventSubscriptions"A successful request returns
201 Createdand the subscription ID. A409response means the workspace already has an active capture. -
Confirm the subscription is listed:
curl -s -H "Authorization: Bearer $FABRIC_TOKEN" \"https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/RuntimeLineages/<runtime-lineage-item-id>/__private/eventSubscriptions"
To stop capture later, send a DELETE request to .../eventSubscriptions/<subscription-id>.
Configure Spark
Notebooks emit lineage events only when they run on Runtime 2.0 with the following Spark properties set:
| Spark property | Value | Purpose |
|---|---|---|
spark.openlineage.transport.type | sparkcore | Sends events through the Fabric Spark transport to the Runtime Lineage item |
spark.openlineage.disabled | false | Turns on OpenLineage event emission |
spark.fabric.pools.skipStarterPools | true | Applies the pool behavior that lineage capture requires |
spark.computeConf.runtime.releaseChannel | earlyAccess | Uses the early-access runtime channel |
Setting the release channel property doesn't select Runtime 2.0. Select the runtime separately as described in the following steps.
Set the properties in one place for each notebook: an Environment or the notebook's first cell. Using a single source avoids conflicting values and makes troubleshooting easier.
- Environment
- Notebook
Use an Environment to apply the same configuration to many notebooks, including scheduled runs.
- In the workspace, click + New item and select Environment. Enter a name and click Create, or open an existing Environment.
- On the Home tab, open the Runtime dropdown and select Runtime 2.0.
- Open Spark compute > Spark properties and add the four properties from the preceding table.
- Click Save, then click Publish > Publish all. Wait for publishing to complete.
- Apply the Environment to your notebooks in one of these ways:
- All notebooks in the workspace: In Workspace settings > Data Engineering/Science > Spark settings, select the Environment tab. Turn on Set default environment, select the Environment, and save. Notebooks that use Workspace default inherit its configuration.
- Individual notebooks: In the notebook, open the Environment dropdown, select Change environment, choose the Environment, and click Confirm.
A changed Environment takes effect from the next Spark session, so restart any session that's already running.
Use notebook-scoped configuration for a single notebook or a quick validation run.
-
Open the notebook and select Runtime 2.0 in the notebook's runtime selector.
-
Add the following cell as the notebook's first code cell, and run it before any other code:
%%configure -f{"conf": {"spark.openlineage.transport.type": "sparkcore","spark.openlineage.disabled": "false","spark.fabric.pools.skipStarterPools": "true","spark.computeConf.runtime.releaseChannel": "earlyAccess"}}If the first cell already contains a
%%configureblock, merge these four entries into itsconfobject and keep the existing entries. The-fflag restarts the Spark session, so don't rerun the cell in the middle of a workload.
Grant Atlan access
Atlan reads the event files with the connector's service principal. Assign the service principal, or the security group it belongs to, a workspace role in each workspace that the crawl includes:
- Workspaces with a Runtime Lineage item: Contributor or higher. The Viewer role doesn't include read access to data in OneLake, so Atlan can't read the event files with it.
- All other included workspaces: Viewer or higher. Atlan lists the items in every included workspace to catalog its notebooks.
These roles are required even when Enable Scanner API Access is on in the crawler. If the service principal has no role on a workspace, the crawl fails. To keep such workspaces out of the crawl, use the Exclude Workspaces filter or list the workspaces you want in Include Workspaces.
To assign a role:
- Open the Microsoft Fabric homepage.
- Navigate to Workspaces and select the workspace.
- Click Manage Access.
- Click Add people or groups.
- Enter the name of your service principal or its security group.
- Choose the role and click Add.
Verify capture
Confirm that the workspace captures events before you run the crawler.
-
In a notebook configured as described in Configure Spark, attach a Lakehouse and mark it as the notebook's default Lakehouse. In the notebook Explorer, click Lakehouses > Add, choose the Lakehouse, and pin it as the default.
-
Run the following cell to confirm the session settings:
for key in ("spark.openlineage.transport.type","spark.openlineage.disabled","spark.fabric.pools.skipStarterPools","spark.computeConf.runtime.releaseChannel",):print(key, "=", spark.conf.get(key, "<unset>"))The output must show
sparkcore,false,true, andearlyAccess. If a value is<unset>or the transport isfile, confirm the runtime and properties and restart the session. -
Run a transformation that reads one Lakehouse table and writes another. For example:
from pyspark.sql.functions import col, splitspark.createDataFrame([(1, "Jane Smith", "Seattle"), (2, "John Doe", "Redmond")],"id INT, fullname STRING, city STRING",).write.format("delta").mode("overwrite").saveAsTable("rtl_check_source")(spark.table("rtl_check_source").withColumn("firstname", split(col("fullname"), " ").getItem(0)).withColumn("lastname", split(col("fullname"), " ").getItem(1)).select("id", "firstname", "lastname", "city").write.format("delta").mode("overwrite").saveAsTable("rtl_check_target")) -
Wait about three minutes, or stop the Spark session. Fabric sends events in batches, and stopping the session flushes pending events immediately.
-
List the event files for the current UTC date:
export ONELAKE_TOKEN=$(az account get-access-token \--resource https://storage.azure.com \--query accessToken -o tsv)curl -s -H "Authorization: Bearer $ONELAKE_TOKEN" -H "x-ms-version: 2021-06-08" \"https://onelake.dfs.fabric.microsoft.com/<workspace-id>?resource=filesystem&recursive=true&directory=<runtime-lineage-item-id>/RuntimeLineage/V1.0/<yyyy-MM-dd>"Event files are stored under
RuntimeLineage/V1.0/<yyyy-MM-dd>/T<HH>-00/<workspace-id>/Notebook/<notebook-id>/, where the date and hour are in UTC. For example,T14-00holds events from 14:00 to 14:59 UTC. If the run crossed an hour boundary, check both hours. -
Open an event file for the notebook and confirm that:
eventTypeisCOMPLETE.- The
inputsandoutputscontain the source and target tables. - The target's
columnLineagefacet listsfullnameas an input offirstnameandlastname.
Drop the rtl_check_source and rtl_check_target tables when you're done.
Next steps
- Crawl Microsoft Fabric: Turn on Extract Spark runtime lineage in the crawler to catalog notebooks, Spark jobs, and their lineage