Crawl Cyera
Crawl classification metadata from Cyera to enrich assets with data sensitivity and privacy information. Extract data sensitivity classifications and security issues.
Configure the Atlan Cyera workflow to crawl data classification metadata from your Cyera instance and enrich assets in Atlan. After completing the prerequisite setup, you can extract data sensitivity classifications, privacy labels, and security findings. Review the order of operations for metadata enrichment workflows before starting.
Prerequisites
Before you begin, make sure you have:
- Created a Cyera API token and obtained your Client ID and Client Secret. If not, follow the Set up Cyera guide.
- Generated an Atlan API token for the Cyera App to write custom metadata structures into your workspace, and to read your connections when you use automatic datastore detection. See Generate Atlan API token.
- The required permissions to configure and run the workflow:
- Atlan: Admin or Workflow Admin permissions
- Cyera: API token with read access to datastores, classifications, and issues
- Reviewed the order of operations for workflow execution.
Create crawler workflow
Create a new Cyera crawler workflow in Atlan by selecting the Cyera App in the Atlan Marketplace, then completing the three configuration steps.
- In the top right of any screen, navigate to New > New Workflow.
- From the list of packages, select Cyera and click Setup Workflow.
Configure credentials
Enter the Cyera credentials Atlan uses to authenticate.
-
In the Credential section:
- Cyera API URL: The Cyera API endpoint, pre-filled as
api.cyera.io. Don't change this value unless instructed by Cyera support. - Client ID: Enter the Client ID from your Cyera API token. See Set up Cyera.
- Client Secret: Enter the Client Secret from your Cyera API token.
- Cyera API URL: The Cyera API endpoint, pre-filled as
-
Expand Advanced Settings and enter the Atlan API Token you generated in Set up Cyera. The Cyera App uses this token to create the Cyera, CyeraIssues, and CyeraIdentities custom metadata structures in your Atlan workspace. Without it, Cyera-tagged assets in Atlan won't receive enriched custom metadata values.
-
Select an Extraction Method:
- Direct (default): Atlan calls Cyera using credentials stored in Atlan.
- Self-Deployed Runtime (SDR): The runtime in your environment holds the credential and pushes metadata to Atlan. See Self-Deployed Runtime for setup requirements.
-
Click Test Authentication to confirm connectivity to Cyera.
-
When the test is successful, click Next.
Configure connection
Set up the Atlan connection and specify who can manage it.
-
Enter a Connection name that represents your Cyera environment. For example:
production,cyera-prod, oranalytics. -
To change who can manage this connection, update the users or groups listed under Connection Admins. By default, your user and all admins are included.
-
Click Next to proceed.
Map datastores
Map Cyera datastores to Atlan connections, configure optional settings, and run the crawler. A Cyera datastore holds the classifications; the mapping tells Atlan which of your connections or assets those classifications belong to. Nothing is enriched until at least one datastore is mapped.
Three approaches are available, and you can combine them. Manual mappings always win: any datastore you map in the repeater or the CSV field keeps that mapping, and automatic detection fills in only the rest.
- Automatic detection
- Bulk connection mapping (CSV)
- Snowflake mapping (UI)
Set Auto-Detect Datastore Mappings to have the app match your Cyera datastores to Atlan connections for you, instead of listing every pair by hand. This is the broadest option, and for object storage it's the only one that works.
| Option | What happens |
|---|---|
| Off (default) | No detection. Only the mappings you enter by hand are used. |
| Suggest (log proposed mappings only) | Detection runs and writes the proposed mappings to the workflow log. Nothing is applied, and no assets change. |
| On (auto-apply high-confidence mappings) | Detection runs and applies its high-confidence matches. |
Matching compares connector type, host and account identifiers, and names. The connector-type gate is strict: a Cyera datastore is only ever compared against Atlan connections that can plausibly hold the same assets, so a Snowflake datastore is never matched to a Databricks connection however similar the names are.
Start with Suggest on your first run, read the proposed mappings in the workflow log, then switch to On once they look right.
If the Bulk Connection Mapping (CSV) field contains any mappings, detection is skipped entirely and only your CSV rows are used. Clear that field to let detection run.
Detection reads your Atlan connections, so it needs the Atlan API Token in the credential step's Advanced Settings. Without it, detection can't run.
Supported sources
Detection covers the sources below. Cyera datastores on any other infrastructure need a manual mapping.
| Cyera infrastructure | Matched Atlan connector |
|---|---|
| Snowflake | Snowflake |
| Databricks | Databricks |
| BigQuery | BigQuery |
| Redshift | Redshift |
| RDS | Postgres, MySQL, MariaDB, Microsoft SQL Server, Oracle, Aurora |
| Cloud SQL, AlloyDB | Cloud SQL Postgres, AlloyDB Postgres, Postgres, MySQL, Microsoft SQL Server |
| Azure SQL Database | Microsoft SQL Server, Synapse |
| Azure SQL Managed Instance, SQL on Azure VM | Microsoft SQL Server |
| Azure Database Server | Postgres, MySQL, MariaDB |
| Oracle Autonomous Database | Oracle |
| MongoDB Atlas cluster | MongoDB |
| Cosmos DB | Azure Cosmos DB, MongoDB |
| DynamoDB | DynamoDB |
| CockroachDB | Postgres |
| Salesforce | Salesforce |
| Amazon S3 | S3 |
| Google Cloud Storage | GCS |
| Azure Blob Storage, Azure File Share | ADLS |
Enter mappings as a comma-separated, semicolon-delimited list in the Bulk Connection Mapping (CSV) field. Use this approach when you have several mappings to configure at once and want to state each one explicitly.
-
Format:
atlan_connection_qualified_name,cyera_datastore_name_or_uid -
Use semicolons to separate multiple mappings; don't add a trailing semicolon
-
The Cyera datastore value can be either the datastore name (case-insensitive) or the datastore UID
-
Example (mixing a Snowflake and a Databricks connection):
default/snowflake/123,my-snowflake-db;default/databricks/456,my-databricks-datastore
This field maps a datastore to a connection, so it fits warehouse and database sources. It doesn't fit object storage. See the callout below.
To build your mapping list from a Cyera export:
- In Cyera, navigate to the Datastores page and export your datastore list to CSV.
- Open the export. Column B contains the Datastore Name. These are the values for the right side of each mapping pair.
- For each datastore you want to map, find the corresponding Atlan connection's qualified name:
- In Atlan, open the connection asset and select the Properties tab in the right panel.
- Copy the Qualified name (for example,
default/snowflake/1234567890for a Snowflake connection ordefault/databricks/1234567890for a Databricks connection).
- Combine them as
atlan_qualified_name,cyera_datastore_nameand join multiple pairs with;. Don't add a trailing semicolon.
Use the Snowflake Mapping repeater to add mappings one at a time.
The UI mapper supports Snowflake datastores only. For any other source, use Auto-Detect Datastore Mappings or Bulk Connection Mapping (CSV).
- In the Snowflake Mapping section, select the Atlan Connection from the dropdown.
- Select the Cyera Datastore (Snowflake) from the dropdown. If the list appears empty, click the refresh icon, then click into the dropdown field to reveal the loaded values.
- Click Add to Snowflake Mapping to save the row.
- Repeat to add more mappings.
Map object storage
Cyera datastores on object storage behave differently from warehouse and database datastores. An S3 bucket, a GCS bucket, or an Azure Blob or File Share container is a single asset in Atlan rather than a whole connection, so the mapping has to point at that asset:
| Cyera infrastructure | Atlan asset it maps to |
|---|---|
| Amazon S3 | S3Bucket |
| Google Cloud Storage | GCSBucket |
| Azure Blob Storage, Azure File Share | ADLSContainer |
Auto-Detect Datastore Mappings is the only option that maps an object-storage datastore to its bucket or container. The Bulk Connection Mapping (CSV) field and the Snowflake Mapping repeater both map to a connection, and mapping an object-storage datastore that way discards the bucket the classifications belong to.
The symptom is quiet: the workflow still succeeds, but it reports zero tables and enriches nothing. If you're cataloging object storage and an apparently successful run leaves your buckets untouched, set Auto-Detect Datastore Mappings to On and clear the Bulk Connection Mapping (CSV) field.
Optional settings: To clean up legacy duplicate metadata attributes from earlier app versions, configure Legacy Metadata Cleanup. Leave as Off (default) unless advised otherwise. Set to Dry run to preview what gets removed, or Apply to delete legacy attributes.
Run preflight checks and start crawler
-
In the Preflight Check section, click Check to run a quick test for necessary permissions before the workflow runs. Click Show details to review the results.
-
Choose your run option:
- To run the crawler immediately, click Run.
- To schedule the crawler to run on a recurring basis, click Schedule & Run and configure the schedule.
Once the crawler completes, you can view the enriched assets on Atlan's asset page.
Need help
- Contact Atlan support: For issues related to the Atlan integration, submit a support request.
See also
- What does Atlan crawl from Cyera: Learn what metadata Atlan extracts from Cyera.