> For the complete documentation index, see [llms.txt](https://docs.bdb.ai/data-pipeline-5/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.bdb.ai/data-pipeline-5/getting-started/homepage/create/creating-a-new-job/pyspark-job.md).

# PySpark Job

Write PySpark scripts and run them flawlessly in the Jobs.

This feature allows users to write their own PySpark script and run their script in the ***Jobs*** section of ***Data Pipeline*** module.&#x20;

Before creating the ***PySpark Job***, the user has to create a project in the ***Data Science Lab*** module under ***PySpark Environment***. Please refer the below image for reference:&#x20;

{% hint style="info" %}
***Please go through the below given demonstration to create and configure a PySpark Job.***
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fc39ZvXI46qjXzpN3rYAg%2Fuploads%2FgTrRLGKZpGmf0dHfvGxN%2F2023-11-14-17-28-57%20(online-video-cutter.com).mp4?alt=media&token=2b859dbd-80ff-4ad7-a3f5-a9b912f22df4>" %}
***Configuring PySpark Job***
{% endembed %}

## **Creating a PySpark job**

* Open the pipeline homepage and click on the **Create** option.

<figure><img src="https://content.gitbook.com/content/6ZP8JhQPMmuTgMyLMcCU/blobs/bcy5wEQPyYAQDfN22lcO/image.png" alt=""><figcaption><p><em><strong>Accessing the Create Job option from the Pipeline Homepage</strong></em></p></figcaption></figure>

* The new panel opens from right hand side. Click on ***Create*** button in Job option.&#x20;
* Provide the following information:
  * **Enter name**: Enter the name for the job.&#x20;
  * **Job Descriptio**n: Add description of the Job (It is an optional field).
  * **Job Baseinfo:** Select ***PySpark Job*** option from the drop in Job Base Information.
  * **Trigger By:** The PySpark Job can be triggered by another Job or PySpark Job. The PySpark Job can be triggered in two scenarios from another jobs:
    * **On Success:** Select a job from drop-down. Once the selected job is run successfully, it will trigger the PySpark Job.
    * **On Failure:** Select a job from drop-down. Once the selected job gets failed, it will trigger the PySpark Job.
  * **Is Schedule:** Put a checkmark in the given box to schedule the new Job.&#x20;
  * **Spark config:** Select resource for the job.
  * Click on ***Save*** option to save the Job.&#x20;

<figure><img src="https://2386645923-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6ZP8JhQPMmuTgMyLMcCU%2Fuploads%2FUdAUFNdHqfCba8Te2VeM%2Fimage.png?alt=media&amp;token=21e0034e-c452-441e-81e9-2fb6eeee4ea2" alt=""><figcaption><p><em><strong>Creating a PySpark Job</strong></em></p></figcaption></figure>

* The ***PySpark*** Job gets saved and it will redirect the user to the Job workspace.

<figure><img src="https://content.gitbook.com/content/6ZP8JhQPMmuTgMyLMcCU/blobs/hM28S4fc2QZE3EO3ahwt/image.png" alt=""><figcaption><p><em><strong>PySpark Job workspace</strong></em></p></figcaption></figure>

## **Configuring a PySpark Job:**

Once the ***PySpark*** Job is created, follow the below given steps to configure the ***Meta Information*** tab of the ***PySpark*** Job.

* **Project Name:** Select the same Project using the drop-down menu where the concerned Notebook has been created.
* **Script Name:** This field will list the exported Notebook names which are exported from the Data Science Lab module to Data Pipeline.

{% hint style="info" %}
*<mark style="color:green;">Please Note:</mark> The script written in DS Lab module should be inside a function. Refer the* [*Export to Pipeline*](https://docs.bdb.ai/data-science-lab/project/tabs-for-a-data-science-lab-project/tabs-for-pyspark-environment/notebook/notebook-list-page/export/export-to-pipeline) *page for more details on how to export a PySpark script to the Data Pipeline module.*
{% endhint %}

* **External Library:** If any external libraries are used in the script the user can mention it here. The user can mention multiple libraries by giving comma(,) in between the names.
* **Start Function:** Here, all the function names used in the script will be listed. Select the start function name to execute the python script.
* **Script:** The Exported script appears under this space.
* **Input Data:** If any parameter has been given in the function, then the name of the parameter is provided as **Key** and value of the parameters has to be provided as **value** in this field.

<figure><img src="https://content.gitbook.com/content/6ZP8JhQPMmuTgMyLMcCU/blobs/63CrPWplYrkXGKDPdHT9/image.png" alt=""><figcaption><p><em><strong>Configuring Meta information of PySpark Job</strong></em></p></figcaption></figure>

{% hint style="info" %}
*<mark style="color:green;">Please note:</mark>* ***We are currently supporting following JDBC connector:***

* MySQL&#x20;
* MSSQL&#x20;
* Oracle&#x20;
* MongoDB&#x20;
* PostgreSQL
* ClickHouse
  {% endhint %}
