What lineage does Atlan extract from Microsoft Fabric
Learn about the lineage relationships that Atlan can map from Microsoft Fabric. The Microsoft Fabric connector is generally available, and Atlan is available in the Microsoft Marketplace and is Microsoft Azure co-sell eligible.
Atlan extracts lineage from Microsoft Fabric in two modes depending on whether Scanner API Access is enabled in the crawler workflow. The mode affects both the depth of lineage and which assets are part of the lineage graph.
Lineage by mode
| Lineage capability | Scanner API enabled | Scanner API disabled |
|---|---|---|
| Lakehouse/Warehouse Table → Semantic Model Table | ✅ Available | ✅ Available |
| Lakehouse/Warehouse Column → Semantic Model Table Column | ✅ Available | ✅ Available |
| Upstream lineage (external SQL sources → Lakehouse/Warehouse) | ✅ Available | ✅ Available |
| Semantic Model Tables → Semantic Model | ✅ Available | ✅ Available |
| Semantic Model → Report | ✅ Available | ✅ Available |
| Dataflow entity column → Semantic Model Table Column | ✅ Available | ✅ Available |
| Source table / column → Dataflow entity column | ✅ Available | ✅ Available |
| Dataflow Gen2 (CI/CD) → destination lakehouse or warehouse table / column ² | ✅ Available | ✅ Available |
| Semantic Model Table Column → Semantic Model Measure | ✅ Available | ✅ Available |
| Semantic Model Measure → Report Page / Visual | ❌ Not available | ✅ Available |
| Semantic Model Table / Column → Report Page | ❌ Not available | ✅ Available |
| Semantic Model Table Column → Report Visual | ❌ Not available | ✅ Available |
| Pipeline Copy Activity (Table → Table) | ❌ Not available | ✅ Available |
| Notebook Spark job: table → table ¹ | ✅ Available | ✅ Available |
| Notebook Spark job: column → column ¹ | ✅ Available | ✅ Available |
¹ Requires Extract Spark runtime lineage in the crawler and Spark runtime lineage setup in Fabric. See Spark runtime lineage.
² Requires the Contributor role or higher on each workspace that contains Dataflow Gen2 (CI/CD) items, in both modes. See Dataflow Gen2 (CI/CD) permissions.
Lineage doesn't yet follow references between queries in the same dataflow, or queries combined with Table.Combine.
For Dataflow Gen2 (CI/CD), Atlan builds lineage from the dataflow definition. If the crawl identity has less than the Contributor role on the workspace, Atlan catalogs the dataflow but extracts no entity columns or lineage for it. This applies when Scanner API Access is enabled too, because no admin API returns Dataflow Gen2 (CI/CD) definitions.
Scanner API enabled
When Scanner API Access is enabled, Atlan uses the Power BI Admin Scanner APIs to derive lineage. The following lineage paths are available:
- External sources → Semantic Models: Upstream lineage from Lakehouse and Warehouse tables and columns to their corresponding Semantic Model Tables and Semantic Model Table Columns, including column-level lineage.
- Upstream SQL lineage: Table- and column-level lineage from external SQL warehouse assets to Fabric assets.
Report Pages, Report Visuals, and Pipeline Copy Activities aren't part of the lineage graph in this mode because those assets aren't cataloged when Scanner API Access is enabled.
Scanner API disabled
When Scanner API Access is disabled, Atlan uses both admin and non-admin APIs, enabling the full lineage graph:
- All lineage paths available in scanner mode (following)
- Semantic Model → Report → Page: Full downstream lineage from Semantic Model Tables and Columns through to Report Pages and Visuals
- Pipeline Copy Activities: Table-level lineage for Copy Activity source and sink tables within Data Pipelines
Spark runtime lineage Private Preview
When Extract Spark runtime lineage is enabled, Atlan builds lineage from the OpenLineage events that Fabric Spark emits while notebooks run. This lineage covers transformations in notebook code, which metadata extraction alone can't see.
How Atlan models runtime lineage
- Notebook: Each notebook in a crawled workspace becomes a
FlowControlOperationasset of typeNotebook. - Spark job: Each distinct Spark job that a notebook runs becomes a
FlowControlOperationasset of typeSpark Job, linked to its notebook. - Table-level lineage: A
Processlinks the job's input tables to its output tables and is linked to the Spark job that orchestrates it. - Column-level lineage: A
ColumnProcessfor each output column links it to the input columns it depends on. Each column process belongs to its job's table-level process.
For the properties Atlan maps on notebooks and Spark jobs, see What does Atlan crawl from Microsoft Fabric.
Which datasets Atlan links
Atlan links a job's input and output datasets to assets that are already cataloged in Atlan:
- Lakehouse tables must be cataloged by the same Microsoft Fabric connection. Atlan matches them by the workspace and Lakehouse IDs in the event, and column names must match the cataloged columns exactly, including case.
- Tables cataloged by another connection must be crawled before the Microsoft Fabric crawler runs. See order of operations.
Datasets that don't match a cataloged asset, such as files under a Lakehouse's Files folder, are skipped. Atlan doesn't create placeholder assets for them. A job gets a table-level process only when at least one input and one output match, and a column process only when its output column and at least one input column match. A Spark job with no matching datasets still appears under its notebook, without lineage.
Column-level dependencies
Column-level lineage includes input columns that feed an output column's value, whether copied, transformed, or aggregated, and input columns used in a conditional expression that produces the value, such as a CASE WHEN condition. Columns used only to filter, join, group, sort, or window rows aren't linked to output columns.
Which run Atlan shows
Atlan keeps the latest execution of each Spark job, chosen by event time, and tracks each job independently:
- When a job runs again, its new execution replaces the earlier lineage.
- When a later notebook run doesn't include a job, that job keeps the lineage of its last execution.
- Spark jobs are identified by their job name within the notebook. Write operations that produce the same Spark job name share one Spark job asset, so only the latest of them is kept.
When a notebook is deleted in Fabric or falls outside the crawl's workspace filters, Atlan stops extracting the notebook, its Spark jobs, and their lineage.
When events appear in Atlan
Fabric stores events in hourly UTC folders, and Atlan reads an hour only after it ends. For example, a crawl at 14:35 UTC reads events up to 14:00 UTC; events from 14:00–14:59 UTC are read by the next crawl after 15:00 UTC.
- The first crawl of a connection reads events from the previous 30 days.
- Each later crawl continues from where the previous crawl stopped, and Atlan keeps the lineage it already extracted between crawls.
- Event files that arrive late for an hour that Atlan already read aren't picked up. To read the previous 30 days again, create a new connection.
See also
- Crawl Microsoft Fabric — Scanner API Access: Configure the Scanner API toggle when setting up the crawler
- Capabilities & limitations: Full comparison of catalog coverage and lineage by mode
- Troubleshooting Spark runtime lineage: Resolve missing notebook lineage and crawl failures