> For the complete documentation index, see [llms.txt](https://docs.bdb.ai/data-center-6/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.bdb.ai/data-center-6/data-center/data-preparation/data-preparation-workspace/transforms/advanced.md).

# Advanced

## &#x20;Cluster & Edit

Find out the clusters based on the pronunciation sound and edit the bulk data in a single click.

{% hint style="success" %}
*Check out the illustration on the Cluster & Edit transform.*
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2F2RLk5fgePDXvliZJZk35%2FCluster%20and%20Edit.mp4?alt=media&token=35e9d575-b6e4-4125-bc16-e13473e25ea3>" %}
***Cluster & Edit Transform***
{% endembed %}

Steps to perform the ***Cluster & Edit*** transform.

* Select a column from the given dataset.
* Open the ***Transforms*** tab.
* Select the ***Cluster and Edit*** transform from the Advanced category.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2Fu5rBr73lSsW3fa4prBRQ%2Fimage.png?alt=media&amp;token=a13fa0fd-2d17-466c-a8d4-bf57de143aa3" alt=""><figcaption></figcaption></figure>
* The ***Cluster & Edit*** window opens.
* The ***Method*** drop-down uses the Soundex phonetic algorithm for indexing names by sound as pronounced in English.
* The ***Values found*** column lists number of values found from the data set related to a specific sound. E.g., In the given image the Values found display 5 categories.
* The Replace Value column lists the anticipated replacements of the values found.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FW3S5KlYKqjlXDvlL1pjb%2Fimage.png?alt=media&amp;token=2da117dd-9304-47e5-854a-5f56657caa64" alt=""><figcaption></figcaption></figure>
* Select a value by using the checkbox that needs to be modified or changed. E.g., the '***tesi***', 'test', and 'test3' are selected in the given example.
* Search for a replace value or enter a value that you wish to be used as a replace value using the drop-down menu from the ***Replace Value*** column. For Example, &#x20;
* Click the ***Submit*** option.&#x20;

<figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FeOiVKRUq4JSLp02OKCqe%2Fimage.png?alt=media&amp;token=21367ba1-974f-4637-aeb0-352b3f604307" alt=""><figcaption></figcaption></figure>

* The selected values from the column get modified in the data set.

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FguDFPe3BjpQWwqR3P1tU%2Fimage.png?alt=media&amp;token=9382f750-b7be-43fd-bc7d-a937895d4bb8" alt=""><figcaption></figcaption></figure>

## Expression Editor

This transform helps to execute expressions.

{% hint style="success" %}
*Check out the given illustration to understand the Expression Editor transform.*
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FOZWQp0lezjPIpYnv23dR%2FExpression%20Editor.mp4?alt=media&token=ea75dfdb-44de-4642-ae1c-c1f522d9f4c2>" %}
Expression Editor Transform
{% endembed %}

Steps to perform the ***Expression Editor*** transform:

* Navigate to a Dataset within the ***Data Preparation*** framework.
* Navigate to the ***Transforms*** tab.
* Open the ***Expression Editor*** from the ***Advanced*** transforms.

<figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2Fuq284QV0SugyPlgZLyfR%2Fimage.png?alt=media&amp;token=e44069e2-2122-4229-88fe-34a96e8c035a" alt=""><figcaption></figcaption></figure>

* The ***Expression Edito***&#x72; window opens displaying the following columns:
  * Functions: The first column contains **functions** for the user to search for a function. By using the double clicks on a function, it gets added to the given space provided for creating a formula.
  * Columns: The second column lists all the column names available in the selected dataset.
* The ***Formula*** space is provided to create and execute various formulas/ executions. Click on a Formula name to get the expression on the right side
* Use either of the following ways to consume the created expression or formula in the dataset.
  * Update a selected column by using the ***Update column*** option. The selected column will be updated with the chosen expression.
  * Create a new column with the created expression by selecting the Create a New Column option.
    * &#x20;Provide a column name for the ***New Column***.
* Click the ***Submit*** option to either add a new column or update the selected column based on the executed formula/ expression.

<figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2F4eg8q632U5SxTQzbFhSx%2Fimage.png?alt=media&amp;token=ceb1eabe-14fa-4f27-a828-92a97171b005" alt=""><figcaption></figcaption></figure>

* The recently created or updated column with Formula gets added to the dataset.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FaX20DbtAESC1c9Q2NwLk%2Fimage.png?alt=media&amp;token=33466d52-9918-4fdd-9b66-733f8789cd51" alt=""><figcaption></figcaption></figure>

## Find Anomaly

Anomaly detection is used to identify any anomaly present in the data. i.e., Outlier.  Instead of looking for usual points in the data, it looks for any anomaly. It uses the ***Isolation Forest*** algorithm.

{% hint style="success" %}
*Check out the given walk-through on the Find Anomaly transform.*
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FCzdHw62uYiiv5haRPoDI%2FFind%20Anomaly.mp4?alt=media&token=4ddf4b82-9fad-4962-8a10-4455a3900f04>" %}
Find Anomaly
{% endembed %}

Steps to perform the ***Find Anomaly*** transform:

* Select a dataset within the ***Data Preparation*** framework.&#x20;
* Navigate to the ***Transforms*** tab.
* Select the ***Find Anomaly*** transform from the ***ADVANCED*** category.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FQp6f8YZNdWxwd6epvWlj%2Fimage.png?alt=media&amp;token=27c5c634-eea2-4d48-a4a1-8423e3e6684f" alt=""><figcaption></figcaption></figure>
* The ***Find Anomaly*** window opens.
* Configure the following information:
  * Select ***Feature Columns***: Select one or more columns where you want to find the anomaly.
  * ***Maximum Sample Size***: The ***Isolation Forest*** algorithm takes the training data of a given sample size to find the normal value in the dataset.
  * ***Contamination*** (%): It is the percentage of observations we believe to be outliers. It varies from 0 to 1 (both inclusive).
  * ***Anomaly Flag Name***: The result is either -1 or 1. 1 means the data is standard, and -1 means the data is an outlier. This information gets stored in the new column with the anomaly flag name.
* Click the ***Submit*** option after the required details are provided.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FNeJpxYYLMCv0nTGG8IhW%2Fimage.png?alt=media&amp;token=c33e5a07-1674-42b4-aeeb-e0451c32ecb0" alt=""><figcaption></figcaption></figure>
* The anomaly gets flagged under the column that has been named using the ***Anomaly Flag Name*** option.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2F4kCbNf0b2urmCFKjMPN0%2Fimage.png?alt=media&amp;token=3107ec85-77e0-4372-814f-d66c841040f0" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
*<mark style="color:green;">Please Note:</mark> The other needed parameters such as **Estimators** and **seed values** are considered based on their default values to run the **Isolation Forest logic** on the selected dataset sample.*
{% endhint %}

## SQL Transform

This transform helps to perform SQL queries. The user can customize the data with the&#x20;

{% hint style="success" %}
*Check out the illustration on the SQL Transform.*
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2F7KPOnRVDVmfDCCl61r3e%2FSQL%20Transform.mp4?alt=media&token=4aff4063-12ca-4f96-a43b-0915f4ad291b>" %}
***SQL Transform***
{% endembed %}

Steps to perform the ***SQL*** transform:

* Select a dataset within the Data Preparation framework.
* Navigate to the ***Transforms*** tab.
* Open the ***SQL Transform*** from the ***Advanced*** transforms.

<figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2F7oJTI12WNCshVvHG8kbz%2Fimage.png?alt=media&amp;token=ca94f814-98ad-4b64-95f2-bb1b8aeeb227" alt=""><figcaption></figcaption></figure>

* The ***SQL Editor*** page opens displaying the ***Functions*** and ***Columns*** from the selected dataset.
* Search a function and click it to get the default syntax suggestion in the text space and add it to the text space provided for writing a query.
* Select the columns from the Columns list.
* Write a query sentence with valid syntax by selecting an SQL function and related columns from the dataset. The displayed example may help to suggest a valid SQL query syntax.
* Click the ***Submit*** option.

<figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FyxgtJWK3AD5HnWnNuWIE%2Fimage.png?alt=media&amp;token=15d7eba7-41e2-4343-9422-a446fdd1d3f7" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
*<mark style="color:green;">Please Note:</mark> Function syntax and small examples are displayed at the bottom of the window with the double-clicks on the function name.*
{% endhint %}

* The SQL query gets applied to the Data Set and based on the query the displayed dataset will be customized.&#x20;

  <figure><img src="https://2657181281-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FKg5pfnNkTs1b1YNYX7rD%2Fuploads%2FDv03Ty4tnGrVm38qFXIt%2Fimage.png?alt=media&amp;token=5848208e-4d91-4094-8bda-6b2872aae227" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
*<mark style="color:green;">Please Note:</mark> The **SQL Transform** & **Expression Editor** supports only **Pandas SQL Queries.***
{% endhint %}
