> For the complete documentation index, see [llms.txt](https://docs.bdb.ai/pre-sales/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.bdb.ai/pre-sales/technical-faqs/data-pipeline.md).

# Data Pipeline

### Does the BDB Platform provide R script?

Yes, there is an option for R Scripting in the BDB Data Pipeline module.

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FW8YW6CHjNVvafjjjoWyT%2Fimage.png?alt=media&amp;token=6830b94e-2ac5-4d38-a712-98181319bd0c" alt=""><figcaption></figcaption></figure></div>

### Does the BDB Platform enable to read the data from AWS S3, Desktop, Folder on server?

Yes, the BDB Data Pipeline enables to read XML and JSON data from the AWS S3, Desktop, Folder on a server.

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2F2obUbfmGMI8ErlcpdUlI%2Fimage.png?alt=media&amp;token=d2bfcfdd-80f4-4bde-b67b-71681479f6a8" alt=""><figcaption></figcaption></figure></div>

### Does the BDB Platform support Manage Memory Computing?

BDB Pipeline enables data engineers to configure the resource allocation to different components based on the load and operation. It allows the creation of multiple instances to run in parallel to complete the operation in the desired time frame. Also, the BDB Data Pipeline module has an integrated data pipeline monitoring section that showcases important performance statistics such as memory & CPU utilization, no. of records processed, allocated CPU & memory, last processed count, total no. of records processed, last processed record size, no. of instances with each component names associated with the selected pipeline. Along with visualizing the key parameters, there is a component log also available.

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FedN5r17rgAqEIdIeRcaR%2Fimage.png?alt=media&amp;token=34dc9409-a391-4aff-96f0-90df0090b1f0" alt=""><figcaption></figcaption></figure></div>

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FP209F438B0umyutfJNYt%2Fimage.png?alt=media&amp;token=79995ef6-71b6-4bf8-8e03-466e1150c770" alt=""><figcaption></figcaption></figure></div>

### Does the BDB Platform support Incremental Extraction?

BDB Data Pipeline handles incremental data extraction to the BDB Data Store. The pipeline has a built-in scheduler component to run at the scheduled interval to capture the increments. It also enables capturing the changed steam.

&#x20;

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FoYTtKtcE4fV0Ht651RXo%2Fimage.png?alt=media&amp;token=b18c7bd6-ad27-489a-8eb3-cc471314904b" alt=""><figcaption></figcaption></figure></div>

### Does the BDB Platform support Scheduler base Data?

BDB Data Pipeline has a built-in scheduler component to schedule the data refresh.

<div align="left"><figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FfTB49vgXQ98Bs21bxfxd%2Fimage.png?alt=media&amp;token=e5a9a4c2-7979-4fd8-977d-47b9b2c94314" alt=""><figcaption></figcaption></figure></div>

### What type of performance tuning options are available for this type of implementation?

The performance tuning options offered by the BDB Platform are as follows:\
**Data Pipeline:** The users can monitor data workflows through the monitoring feature provided under the BDB Data Pipeline modules. It provides them with a variety of performance-tuning options:

* **Optimizing resource utilization**
* **Detecting and preventing errors**
* **Improving the efficiency of pipelines**
* **Ensuring data quality**
* **Increasing transparency**

E.g., If a user is writes MongoDB queries within the BDB Data Set and the query response time is prolonged, the user can implement performance tuning options such as aggregated views, indexing, sharding, etc. to enhance the query execution performance.

<figure><img src="https://4072512490-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FlLFg37sm5677zgZ4m8et%2Fuploads%2FsvckRjmAvyHgjzoEDaWE%2Fimage.png?alt=media&amp;token=5460afd8-ed57-4d17-b444-465bb23d3864" alt=""><figcaption></figcaption></figure>
