Skip to main content
Community Hub
TL;DR

Capture Spark runtime lineage from Microsoft Fabric notebooks so Atlan can catalog notebooks and Spark jobs with table- and column-level lineage.

Your AI can read this via Docs MCPcurl -fsSL "https://docs.atlan.com/install-docs-mcp" | bashConnect

Set up Spark runtime lineage Private Preview

Capture lineage from the Spark jobs that your Microsoft Fabric notebooks run, so Atlan can catalog notebooks and Spark jobs and link them to the tables they read and write, down to column level.

When a notebook runs, Fabric Spark emits an OpenLineage event for each Spark job. A Runtime Lineage item in the workspace stores these events as JSON files in OneLake. During each Microsoft Fabric crawl, Atlan reads the new event files and builds lineage from them. Atlan only reads the events: you enable and configure capture in Fabric by following this guide.

Preview

Runtime Lineage is a Microsoft Fabric private preview:

  • Microsoft must enable the preview for your Fabric tenant. During the preview, Microsoft also pins each capture workspace to the Spark 4.1 validation image that provides the lineage transport. Contact Microsoft to enroll your tenant and workspaces.
  • A Runtime Lineage item captures events only from the workspace it lives in, and each workspace supports one active capture.
  • The event subscription API uses a private-preview route (__private), which Microsoft expects to remove when the feature reaches public preview.

Prerequisites​

Before you begin, make sure you have:

  • Completed Set up Microsoft Fabric with Service Principal authentication. Atlan reads runtime lineage events from OneLake with an Azure Storage token issued to the service principal, which APIM Managed Identity authentication doesn't provide.
  • Fabric Administrator privileges, to change tenant settings.
  • The Admin, Member, or Contributor role on each workspace whose notebooks you want lineage for.
  • Workspaces on a Fabric capacity. A Fabric Trial capacity works.
  • The Azure CLI, signed in with az login as a user who holds one of the workspace roles listed. The API calls in this guide use it to get access tokens.

Enable tenant settings​

Runtime Lineage is off until a Fabric administrator enables it for the tenant.

  1. Log in to the Fabric admin portal.
  2. Click the Settings icon on the top panel.
  3. Click Admin Portal under the Governance and insights section.
  4. Select Tenant Settings from the sidebar.
  5. Enable Users can create a Runtime Lineage item (preview). This setting makes the Runtime Lineage item type available. Apply it to the entire organization or to the security groups whose members set up capture, and click Apply.
  6. Enable Workspace admins can turn on Runtime Lineage monitoring for their workspaces (preview). This setting lets workspace administrators turn on capture. Apply it to the same scope and click Apply.
  7. Under OneLake settings, confirm that Users can access data stored in OneLake with apps external to Fabric is enabled. Atlan reads the event files from outside Fabric through the OneLake ADLS Gen2 APIs, and these reads fail when this setting is off.

Tenant setting changes can take 15–30 minutes to take effect.

Create Runtime Lineage item​

Create one Runtime Lineage item in each workspace whose notebooks you want lineage for. The item is the OneLake destination for the workspace's lineage events. Spark doesn't create it automatically.

  1. Open the Microsoft Fabric homepage and go to the workspace.

  2. Copy the workspace ID from the browser address bar. It's the value after /groups/ in the URL.

  3. Click + New item, select Runtime Lineage, enter a name, and click Create.

  4. Get the ID of the new item:

    export FABRIC_TOKEN=$(az account get-access-token \
    --resource https://api.fabric.microsoft.com \
    --query accessToken -o tsv)

    curl -s -H "Authorization: Bearer $FABRIC_TOKEN" \
    "https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/items?type=RuntimeLineage"

    Copy the id value from the response. This is the Runtime Lineage item ID.

Keep exactly one Runtime Lineage item per workspace. When a workspace contains more than one, Atlan can't tell which item holds the capture, and the crawl fails.

To automate setup across many workspaces, create the item with the Fabric Create Item API instead. Send "type": "RuntimeLineage" and a displayName in the request body. The response contains the item ID.

Turn on workspace capture​

Subscribe the Runtime Lineage item to its workspace. After the subscription exists, Spark in that workspace sends lineage events to the item. Creating a subscription requires write permission on the item, which the Admin, Member, and Contributor workspace roles include.

  1. Create the subscription. Use the same workspace ID in the URL and in the request body:

    curl -s -X POST \
    -H "Authorization: Bearer $FABRIC_TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"type": "Workspace", "publisher": {"workspaceId": "<workspace-id>"}}' \
    "https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/RuntimeLineages/<runtime-lineage-item-id>/__private/eventSubscriptions"

    A successful request returns 201 Created and the subscription ID. A 409 response means the workspace already has an active capture.

  2. Confirm the subscription is listed:

    curl -s -H "Authorization: Bearer $FABRIC_TOKEN" \
    "https://api.fabric.microsoft.com/v1/workspaces/<workspace-id>/RuntimeLineages/<runtime-lineage-item-id>/__private/eventSubscriptions"

To stop capture later, send a DELETE request to .../eventSubscriptions/<subscription-id>.

Configure Spark​

Notebooks emit lineage events only when they run on Runtime 2.0 with the following Spark properties set:

Spark propertyValuePurpose
spark.openlineage.transport.typesparkcoreSends events through the Fabric Spark transport to the Runtime Lineage item
spark.openlineage.disabledfalseTurns on OpenLineage event emission
spark.fabric.pools.skipStarterPoolstrueApplies the pool behavior that lineage capture requires
spark.computeConf.runtime.releaseChannelearlyAccessUses the early-access runtime channel

Setting the release channel property doesn't select Runtime 2.0. Select the runtime separately as described in the following steps.

Set the properties in one place for each notebook: an Environment or the notebook's first cell. Using a single source avoids conflicting values and makes troubleshooting easier.

Use an Environment to apply the same configuration to many notebooks, including scheduled runs.

  1. In the workspace, click + New item and select Environment. Enter a name and click Create, or open an existing Environment.
  2. On the Home tab, open the Runtime dropdown and select Runtime 2.0.
  3. Open Spark compute > Spark properties and add the four properties from the preceding table.
  4. Click Save, then click Publish > Publish all. Wait for publishing to complete.
  5. Apply the Environment to your notebooks in one of these ways:
    • All notebooks in the workspace: In Workspace settings > Data Engineering/Science > Spark settings, select the Environment tab. Turn on Set default environment, select the Environment, and save. Notebooks that use Workspace default inherit its configuration.
    • Individual notebooks: In the notebook, open the Environment dropdown, select Change environment, choose the Environment, and click Confirm.

A changed Environment takes effect from the next Spark session, so restart any session that's already running.

Grant Atlan access​

Atlan reads the event files with the connector's service principal. Assign the service principal, or the security group it belongs to, a workspace role in each workspace that the crawl includes:

  • Workspaces with a Runtime Lineage item: Contributor or higher. The Viewer role doesn't include read access to data in OneLake, so Atlan can't read the event files with it.
  • All other included workspaces: Viewer or higher. Atlan lists the items in every included workspace to catalog its notebooks.

These roles are required even when Enable Scanner API Access is on in the crawler. If the service principal has no role on a workspace, the crawl fails. To keep such workspaces out of the crawl, use the Exclude Workspaces filter or list the workspaces you want in Include Workspaces.

To assign a role:

  1. Open the Microsoft Fabric homepage.
  2. Navigate to Workspaces and select the workspace.
  3. Click Manage Access.
  4. Click Add people or groups.
  5. Enter the name of your service principal or its security group.
  6. Choose the role and click Add.

Verify capture​

Confirm that the workspace captures events before you run the crawler.

  1. In a notebook configured as described in Configure Spark, attach a Lakehouse and mark it as the notebook's default Lakehouse. In the notebook Explorer, click Lakehouses > Add, choose the Lakehouse, and pin it as the default.

  2. Run the following cell to confirm the session settings:

    for key in (
    "spark.openlineage.transport.type",
    "spark.openlineage.disabled",
    "spark.fabric.pools.skipStarterPools",
    "spark.computeConf.runtime.releaseChannel",
    ):
    print(key, "=", spark.conf.get(key, "<unset>"))

    The output must show sparkcore, false, true, and earlyAccess. If a value is <unset> or the transport is file, confirm the runtime and properties and restart the session.

  3. Run a transformation that reads one Lakehouse table and writes another. For example:

    from pyspark.sql.functions import col, split

    spark.createDataFrame(
    [(1, "Jane Smith", "Seattle"), (2, "John Doe", "Redmond")],
    "id INT, fullname STRING, city STRING",
    ).write.format("delta").mode("overwrite").saveAsTable("rtl_check_source")

    (
    spark.table("rtl_check_source")
    .withColumn("firstname", split(col("fullname"), " ").getItem(0))
    .withColumn("lastname", split(col("fullname"), " ").getItem(1))
    .select("id", "firstname", "lastname", "city")
    .write.format("delta").mode("overwrite").saveAsTable("rtl_check_target")
    )
  4. Wait about three minutes, or stop the Spark session. Fabric sends events in batches, and stopping the session flushes pending events immediately.

  5. List the event files for the current UTC date:

    export ONELAKE_TOKEN=$(az account get-access-token \
    --resource https://storage.azure.com \
    --query accessToken -o tsv)

    curl -s -H "Authorization: Bearer $ONELAKE_TOKEN" -H "x-ms-version: 2021-06-08" \
    "https://onelake.dfs.fabric.microsoft.com/<workspace-id>?resource=filesystem&recursive=true&directory=<runtime-lineage-item-id>/RuntimeLineage/V1.0/<yyyy-MM-dd>"

    Event files are stored under RuntimeLineage/V1.0/<yyyy-MM-dd>/T<HH>-00/<workspace-id>/Notebook/<notebook-id>/, where the date and hour are in UTC. For example, T14-00 holds events from 14:00 to 14:59 UTC. If the run crossed an hour boundary, check both hours.

  6. Open an event file for the notebook and confirm that:

    • eventType is COMPLETE.
    • The inputs and outputs contain the source and target tables.
    • The target's columnLineage facet lists fullname as an input of firstname and lastname.

Drop the rtl_check_source and rtl_check_target tables when you're done.

Next steps​

  • Crawl Microsoft Fabric: Turn on Extract Spark runtime lineage in the crawler to catalog notebooks, Spark jobs, and their lineage