
## MongoDB assets package

URL: https://docs.atlan.com/apps/connectors/database/mongodb/sdk/references/package-reference

> Programmatically crawl and manage MongoDB assets in Atlan using the Python SDK (pyatlan). Reference for the MongoDB assets package.

:::warning[Deprecated—use the app reference]
This package is deprecated. Use the [app reference](app-reference) for the new `client.app` builder, or see [Manage apps](https://docs.atlan.com/llms/platform/python/manage-apps/llms.txt).
:::

# MongoDB assets package

The [MongoDB assets package](https://ask.atlan.com/hc/en-us/articles/7841931946639-How-to-crawl-MongoDB)
crawls MongoDB assets and publishes them to Atlan for discovery.

## Direct extraction

:::warning[Will create a new connection]
This should only be used to create the workflow the first time. Each time you run this method
it will create a new connection and new assets within that connection — which could lead to duplicate
assets if you run the workflow this way multiple times with the same settings.

Instead, when you want to re-crawl assets, re-run the existing workflow
(see [Re-run existing workflow](#re-run-existing-workflow) below).
:::
To crawl assets directly from MongoDB:

### Java

:::warning[Coming soon]
:::

### Python

```python showLineNumbers title="Direct extraction from MongoDB"
from pyatlan.client.atlan import AtlanClient
from pyatlan.model.packages import MongoDBCrawler

client = AtlanClient()

crawler = (
 MongoDBCrawler( # (1)
 client=client, # (2)
 connection_name="production", # (3)
 admin_roles=[client.role_cache.get_id_for_name("$admin")], # (4)
 admin_groups=None,
 admin_users=None,
 row_limit=10000, # (5)
 allow_query=True, # (6)
 allow_query_preview=True, # (7)
 )
 .direct(hostname="test.host.mongodb.net", port=27017) # (8)
 .basic_auth( # (9)
 username="test-user",
 password="test-pass",
 native_host="test.native.mongodb.net",
 default_db="test-default-db",
 auth_db="test-auth-db",
 is_ssl=True,
 )
 .include(assets=["test-asset-1", "test-asset-2"]) # (10)
 .exclude(assets=["test-asset-1", "test-asset-2"]) # (11)
 .exclude_regex(regex="TEST*") # (12)
 .to_workflow() # (13)
)
response = client.workflow.run(crawler) # (14)
```

1. Base configuration for a new MongoDB crawler.
2. You must provide a client instance.
3. You must provide a name for the connection that the MongoDB assets will exist within.
4. You must specify at **least one connection admin**, either:

 - everyone in a role (in this example, all `$admin` users).
 - a list of groups (names) that will be connection admins.
 - a list of users (names) that will be connection admins.
5. You can specify a maximum number of rows that can be accessed for any asset in the connection. defaults to `10000`.
6. You can specify whether you want to allow queries to this connection
(default: `True`, as in this example) or deny all query access to the connection (`False`).
7. You can specify whether you want to allow data previews on this connection
(default: `True`, as in this example) or deny all sample data previews to the connection (`False`).
8. When crawling assets directly from MongoDB,
you are required to provide the following information:

 - hostname of the Atlas SQL connection.
 - port number of the Atlas SQL connection. default: `27017`.

9. When using basic authentication,
you are required to provide the following information:

 - username through which to access Atlas SQL connection.
 - password through which to access Atlas SQL connection.
 - native host address for the MongoDB connection.
 - default database to connect to.
 - authentication database to use (default: `"admin"`).
 - whether to use SSL for the connection (default: `True`).

10. You can also optionally specify the list of assets to include in crawling.
For MongoDB assets, this should be specified as a list of database names.
If set to `[]`, all databases will be crawled.
11. You can also optionally specify the list of assets to exclude from crawling.
For MongoDB assets, this should be specified as a list of database GUIDs.
If set to `[]`, no databases will be excluded.
12. You can also optionally specify the exclude regex for the crawler
to ignore collections based on a naming convention.
13. Now, you can convert the package into a `Workflow` object.
14. Run the workflow by invoking the `run()` method on the workflow client, passing the created object.

 :::warning[Workflows run asynchronously]
Remember that workflows run asynchronously. See the [packages and workflows introduction](https://docs.atlan.com/llms/platform/python/packages/llms.txt)
for details on how you can check the status and wait until the workflow has been completed.
 :::

### Kotlin

:::warning[Coming soon]
:::

### Raw REST API

:::tip[Create the workflow via UI only]
We recommend creating the workflow only via the UI. To rerun an existing workflow, see the steps below.
:::

## Re-run existing workflow

To re-run an existing workflow for MongoDB assets:

### Java

:::warning[Coming soon]
:::

### Python

```python showLineNumbers title="Re-run existing MongoDB workflow"
from pyatlan.client.atlan import AtlanClient
from pyatlan.model.enums import WorkflowPackage

client = AtlanClient()

existing = client.workflow.find_by_type( # (1)
 prefix=WorkflowPackage.MONGODB, max_results=5
)

# Determine which MongoDB workflow (n)

# from the list of results you want to re-run.

response = client.workflow.rerun(existing[n]) # (2)
```

1. You can find workflows by their type using the workflow client `find_by_type()`
method and providing the **prefix** for one of the packages.
In this example, we do so for the `MongoDBCrawler`. (You can also specify
the **maximum number of resulting workflows** you want to retrieve as results.)
2. Once you've found the workflow you want to re-run,
you can simply call the workflow client `rerun()` method.

 - Optionally, you can use `rerun(idempotent=True)` to avoid re-running a workflow that is already in running or in a pending state.
 This will return details of the already running workflow if found, and by default, it is set to `False`.

 :::warning[Workflows run asynchronously]
Remember that workflows run asynchronously. See the [packages and workflows introduction](https://docs.atlan.com/llms/platform/python/packages/llms.txt)
for details on how you can check the status and wait until the workflow has been completed.
 :::

### Kotlin

:::warning[Coming soon]
:::

### Raw REST API

:::warning[Requires multiple steps through the raw REST API]
1. Find the existing workflow.
2. Send through the resulting re-run request.
:::
```json showLineNumbers title="POST /api/service/workflows/indexsearch"
{
 "from": 0,
 "size": 5,
 "query": {
 "bool": {
 "filter": [
 {
 "nested": {
 "path": "metadata",
 "query": {
 "prefix": {
 "metadata.name.keyword": {
 "value": "atlan-mongodb" // (1)
 }
 }
 }
 }
 }
 ]
 }
 },
 "sort": [
 {
 "metadata.creationTimestamp": {
 "nested": {
 "path": "metadata"
 },
 "order": "desc"
 }
 }
 ],
 "track_total_hits": true
}
```

1. Searching by the `atlan-mongodb` prefix will ensure you only find existing MongoDB assets workflows.

 :::tip[Name of the workflow]
The name of the workflow will be nested within the `_source.metadata.name` property of the response object.
(Remember since this is a search, there could be multiple results, so you may want to use the other
details in each result to determine which workflow you really want.)
 :::
```json title="POST /api/service/workflows/submit"
{
 "namespace": "default",
 "resourceKind": "WorkflowTemplate",
 "resourceName": "atlan-mongodb-1684500411" // (1)
}
```

1. Send the name of the workflow as the `resourceName` to rerun it.

---
