Skip to main content
Community Hub

What lineage does Atlan extract from Microsoft Fabric

TL;DR

Learn about the lineage relationships that Atlan can map from Microsoft Fabric. The Microsoft Fabric connector is generally available, and Atlan is available in the Microsoft Marketplace and is Microsoft Azure co-sell eligible.

Your AI can read this via Docs MCPcurl -fsSL "https://docs.atlan.com/install-docs-mcp" | bashConnect

Atlan extracts lineage from Microsoft Fabric in two modes depending on whether Scanner API Access is enabled in the crawler workflow. The mode affects both the depth of lineage and which assets are part of the lineage graph.

Lineage by mode​

Lineage capabilityScanner API enabledScanner API disabled
Lakehouse/Warehouse Table → Semantic Model Table✅ Available✅ Available
Lakehouse/Warehouse Column → Semantic Model Table Column✅ Available✅ Available
Upstream lineage (external SQL sources → Lakehouse/Warehouse)✅ Available✅ Available
Semantic Model Tables → Semantic Model✅ Available✅ Available
Semantic Model → Report✅ Available✅ Available
Dataflow entity column → Semantic Model Table Column✅ Available✅ Available
Source table / column → Dataflow entity column✅ Available✅ Available
Dataflow Gen2 (CI/CD) → destination lakehouse or warehouse table / column ²✅ Available✅ Available
Semantic Model Table Column → Semantic Model Measure✅ Available✅ Available
Semantic Model Measure → Report Page / Visual❌ Not available✅ Available
Semantic Model Table / Column → Report Page❌ Not available✅ Available
Semantic Model Table Column → Report Visual❌ Not available✅ Available
Pipeline Copy Activity (Table → Table)❌ Not available✅ Available
Notebook Spark job: table → table ¹✅ Available✅ Available
Notebook Spark job: column → column ¹✅ Available✅ Available

¹ Requires Extract Spark runtime lineage in the crawler and Spark runtime lineage setup in Fabric. See Spark runtime lineage.

² Requires the Contributor role or higher on each workspace that contains Dataflow Gen2 (CI/CD) items, in both modes. See Dataflow Gen2 (CI/CD) permissions.

Dataflow lineage

Lineage doesn't yet follow references between queries in the same dataflow, or queries combined with Table.Combine.

For Dataflow Gen2 (CI/CD), Atlan builds lineage from the dataflow definition. If the crawl identity has less than the Contributor role on the workspace, Atlan catalogs the dataflow but extracts no entity columns or lineage for it. This applies when Scanner API Access is enabled too, because no admin API returns Dataflow Gen2 (CI/CD) definitions.

Scanner API enabled​

When Scanner API Access is enabled, Atlan uses the Power BI Admin Scanner APIs to derive lineage. The following lineage paths are available:

  • External sources → Semantic Models: Upstream lineage from Lakehouse and Warehouse tables and columns to their corresponding Semantic Model Tables and Semantic Model Table Columns, including column-level lineage.
  • Upstream SQL lineage: Table- and column-level lineage from external SQL warehouse assets to Fabric assets.

Report Pages, Report Visuals, and Pipeline Copy Activities aren't part of the lineage graph in this mode because those assets aren't cataloged when Scanner API Access is enabled.

Scanner API disabled​

When Scanner API Access is disabled, Atlan uses both admin and non-admin APIs, enabling the full lineage graph:

  • All lineage paths available in scanner mode (following)
  • Semantic Model → Report → Page: Full downstream lineage from Semantic Model Tables and Columns through to Report Pages and Visuals
  • Pipeline Copy Activities: Table-level lineage for Copy Activity source and sink tables within Data Pipelines

Spark runtime lineage Private Preview​

When Extract Spark runtime lineage is enabled, Atlan builds lineage from the OpenLineage events that Fabric Spark emits while notebooks run. This lineage covers transformations in notebook code, which metadata extraction alone can't see.

How Atlan models runtime lineage​

  • Notebook: Each notebook in a crawled workspace becomes a FlowControlOperation asset of type Notebook.
  • Spark job: Each distinct Spark job that a notebook runs becomes a FlowControlOperation asset of type Spark Job, linked to its notebook.
  • Table-level lineage: A Process links the job's input tables to its output tables and is linked to the Spark job that orchestrates it.
  • Column-level lineage: A ColumnProcess for each output column links it to the input columns it depends on. Each column process belongs to its job's table-level process.

For the properties Atlan maps on notebooks and Spark jobs, see What does Atlan crawl from Microsoft Fabric.

Atlan links a job's input and output datasets to assets that are already cataloged in Atlan:

  • Lakehouse tables must be cataloged by the same Microsoft Fabric connection. Atlan matches them by the workspace and Lakehouse IDs in the event, and column names must match the cataloged columns exactly, including case.
  • Tables cataloged by another connection must be crawled before the Microsoft Fabric crawler runs. See order of operations.

Datasets that don't match a cataloged asset, such as files under a Lakehouse's Files folder, are skipped. Atlan doesn't create placeholder assets for them. A job gets a table-level process only when at least one input and one output match, and a column process only when its output column and at least one input column match. A Spark job with no matching datasets still appears under its notebook, without lineage.

Column-level dependencies​

Column-level lineage includes input columns that feed an output column's value, whether copied, transformed, or aggregated, and input columns used in a conditional expression that produces the value, such as a CASE WHEN condition. Columns used only to filter, join, group, sort, or window rows aren't linked to output columns.

Which run Atlan shows​

Atlan keeps the latest execution of each Spark job, chosen by event time, and tracks each job independently:

  • When a job runs again, its new execution replaces the earlier lineage.
  • When a later notebook run doesn't include a job, that job keeps the lineage of its last execution.
  • Spark jobs are identified by their job name within the notebook. Write operations that produce the same Spark job name share one Spark job asset, so only the latest of them is kept.

When a notebook is deleted in Fabric or falls outside the crawl's workspace filters, Atlan stops extracting the notebook, its Spark jobs, and their lineage.

When events appear in Atlan​

Fabric stores events in hourly UTC folders, and Atlan reads an hour only after it ends. For example, a crawl at 14:35 UTC reads events up to 14:00 UTC; events from 14:00–14:59 UTC are read by the next crawl after 15:00 UTC.

  • The first crawl of a connection reads events from the previous 30 days.
  • Each later crawl continues from where the previous crawl stopped, and Atlan keeps the lineage it already extracted between crawls.
  • Event files that arrive late for an hour that Atlan already read aren't picked up. To read the previous 30 days again, create a new connection.

See also​