Troubleshooting Spark runtime lineage
Resolve crawl failures and missing notebook lineage when extracting Spark runtime lineage from Microsoft Fabric.
Resolve crawl failures and missing lineage when Extract Spark runtime lineage is enabled in the Microsoft Fabric crawler.
Multiple Runtime Lineage items in workspace
Multiple RuntimeLineage items in a workspace; cannot choose a capture safely
Cause
A crawled workspace contains more than one Runtime Lineage item, so Atlan can't tell which item holds the workspace's capture.
Solution
- Open the workspace in Fabric and find its Runtime Lineage items.
- Keep the item that has the active event subscription and delete the others.
- Rerun the crawler.
Authentication method doesn't support runtime lineage
Runtime lineage requires a storage-scoped service-principal token; APIM token routing is not configured
Cause
The crawler uses APIM Managed Identity authentication. Atlan reads runtime lineage events from OneLake with an Azure Storage token issued to a service principal, which this authentication method doesn't provide.
Solution
- Configure the crawler with Service Principal authentication, as described in Set up Microsoft Fabric.
- Alternatively, turn off Extract Spark runtime lineage to crawl without runtime lineage.
Service principal can't access workspace
Fabric API permission denied
Cause
With Extract Spark runtime lineage on, Atlan lists the items in every crawled workspace and reads event files from each Runtime Lineage item. The service principal lacks a role on one of these workspaces, or has Viewer on a workspace with a Runtime Lineage item. The Viewer role doesn't include read access to OneLake data. Both requirements apply even when Enable Scanner API Access is on.
Solution
- Assign the service principal Contributor or higher on each workspace with a Runtime Lineage item, and Viewer or higher on every other crawled workspace. See Grant Atlan access.
- To leave a workspace out instead, add it to Exclude Workspaces in the crawler.
- In the Fabric admin portal, confirm that Users can access data stored in OneLake with apps external to Fabric is enabled under OneLake settings.
- Rerun the crawler.
Notebook has no lineage after crawl
The crawl succeeds and the notebook appears in Atlan, but it has no Spark jobs or lineage.
Cause
Atlan received no usable events for the notebook. Common reasons:
- The notebook ran during the current UTC hour. Atlan reads an hour's events only after the hour ends.
- Capture isn't active for the notebook's workspace, or the notebook didn't run with the required runtime and Spark properties.
- The tables the notebook reads or writes aren't cataloged in Atlan, so the job's datasets don't match any asset. Atlan creates lineage only when at least one input and one output match a cataloged asset.
Solution
- Rerun the crawler after the UTC hour of the notebook run has ended.
- Confirm capture and Spark configuration with the steps in Verify capture. If the transport reads
file, check the requirements in the note at the start of Set up Spark runtime lineage. - Confirm that the notebook, its Runtime Lineage item, and the event subscription are in the same workspace.
- Confirm that the source and target Lakehouse tables appear in Atlan under the same Microsoft Fabric connection. Newly created tables are cataloged after they appear in the Lakehouse SQL analytics endpoint and the crawler runs again.
- For tables cataloged by another connection, run that connection's crawler before the Microsoft Fabric crawler.
Lineage from earlier notebook run is missing
After a notebook ran again, lineage from its earlier run changed or disappeared.
Cause
Atlan keeps the latest execution of each Spark job and replaces the job's lineage when the job runs again. Spark jobs are identified by their job name. When several write operations in a notebook produce the same Spark job name, they share one Spark job asset in Atlan, and only the latest execution's lineage is kept.
Solution
- Open the notebook in Atlan and check its Spark jobs. Each job shows the run ID and status of its latest execution, and its lineage reflects that execution.
- Jobs that didn't run again keep the lineage of their earlier execution. No action is needed for them.
Earlier events aren't picked up
Event files that arrived late for an hour that Atlan already read, or history from before the connection's first crawl, don't appear in Atlan.
Cause
The first crawl reads events from the previous 30 days. Each later crawl continues from where the previous crawl stopped and doesn't reread earlier hours.
Solution
- To read the previous 30 days again, create a new Microsoft Fabric connection with Extract Spark runtime lineage turned on. Rerunning the existing connection doesn't reread earlier hours.
Event subscription request returns 409
409 Conflict
Cause
The workspace already has an active capture. Each workspace supports one active capture.
Solution
- List the existing subscriptions of the workspace's Runtime Lineage item, as described in Turn on workspace capture.
- Keep the existing subscription, or delete it before creating a new one.
See also
- Set up Spark runtime lineage: Enable capture and configure Spark in Fabric
- What lineage does Atlan extract from Microsoft Fabric: How Atlan builds notebook and Spark job lineage
Need help
If you need assistance after trying the steps, contact Atlan support: Submit a request.