> For the complete documentation index, see [llms.txt](https://docs.rhinofcp.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.rhinofcp.com/creating-and-running-code-objects/creating-and-running-a-new-generalized-compute-code-object.md).

# Creating and Running a New Generalized Compute Code Object

## What is a Generalized Compute Code Object?

In the Rhino Federated Computing Platform (FCP), the Generalized Compute (GC) Code Object represents a versatile and powerful way to execute pre-built container images within the FCP environment. This Code type enables you to run custom code, computations, or processes that are encapsulated within container images. With GC, you can harness the full potential of distributed computing while tailoring your computations to suit your specific needs.

#### Key Features of Generalized Compute

* **Flexible Container Execution:** GC empowers you to execute container images containing diverse computational tasks, irrespective of programming language or complexity. This allows you to leverage existing tools, libraries, and applications within the FCP ecosystem.
* **Batch Processing:** Ideal for tasks that do not require real-time interactivity, GC excels at batch processing scenarios. You can execute computationally intensive tasks, data transformations, or complex analyses efficiently across distributed datasets.
* **Secure Data Access:** GC can only access data provided via the FCP, with no external internet connectivity. GC accesses input Datasets located under the `/input` directory and generates output Datasets under the `/output` directory. This ensures secure data access and maintains the privacy of sensitive healthcare information.

#### Use Cases for Generalized Compute

* **Data Transformation:** Execute data transformation and normalization tasks across distributed healthcare datasets.
* **Custom Analytics:** Run complex statistical analyses, machine learning processes, or custom algorithms on distributed data.
* **Scientific Simulations:** Perform large-scale scientific simulations or simulations requiring substantial computational resources.
* **Resource-Intensive Tasks:** Process resource-intensive tasks, such as image processing, signal analysis, or simulations.

#### Benefits of Generalized Compute

* **Adaptability:** Leverage various programming languages and tools within your container, ensuring flexibility in your computational tasks.
* **Distributed Processing:** Distribute computation tasks across multiple sites, enhancing efficiency and reducing processing time.
* **Customization:** Craft computations tailored to your project's unique requirements, enhancing the depth of insights you can derive.

In summary, Generalized Compute empowers you to harness the versatility and power of containerized computations within the Rhino FCP. By executing custom code across distributed datasets, you can streamline tasks, gain valuable insights, and contribute to advanced healthcare research in a secure and privacy-preserving manner

{% hint style="info" %}
If you want to create a new Generalized Compute Code Object using the SDK, see [Creating a Generalized Compute Code Object Using the Rhino SDK.](/rhino-sdk/running-a-generalized-compute-code-object-using-the-rhino-sdk/creating-a-generalized-compute-code-object-using-the-rhino-sdk.md)
{% endhint %}

{% hint style="info" %}
If your dataset contains sensitive data, then by default, all output datasets are considered to be sensitive. For more information on this, including how to mark which data is sensitive and which is not, please see [Sensitive Datasets](/datasets/sensitive-datasets.md).
{% endhint %}

In the Rhino Federated Computing Platform (FCP), the Generalized Compute (GC) Code Object represents a versatile and powerful way to execute pre-built container images within the FCP environment. This Code type enables you to run custom code, computations, or processes that are encapsulated within container images. With GC, you can harness the full potential of distributed computing while tailoring your computations to suit your specific needs.

#### Key Features of Generalized Compute

* **Flexible Container Execution:** GC empowers you to execute container images containing diverse computational tasks, irrespective of programming language or complexity. This allows you to leverage existing tools, libraries, and applications within the FCP ecosystem.
* **Batch Processing:** Ideal for tasks that do not require real-time interactivity, GC excels at batch processing scenarios. You can execute computationally intensive tasks, data transformations, or complex analyses efficiently across distributed datasets.
* **Secure Data Access:** GC can only access data provided via the FCP, with no external internet connectivity. GC accesses input Datasets located under the `/input` directory and generates output Datasets under the `/output` directory. This ensures secure data access and maintains the privacy of sensitive healthcare information.

#### Use Cases for Generalized Compute

* **Data Transformation:** Execute data transformation and normalization tasks across distributed healthcare datasets.
* **Custom Analytics:** Run complex statistical analyses, machine learning processes, or custom algorithms on distributed data.
* **Scientific Simulations:** Perform large-scale scientific simulations or simulations requiring substantial computational resources.
* **Resource-Intensive Tasks:** Process resource-intensive tasks, such as image processing, signal analysis, or simulations.

#### Benefits of Generalized Compute

* **Adaptability:** Leverage various programming languages and tools within your container, ensuring flexibility in your computational tasks.
* **Distributed Processing:** Distribute computation tasks across multiple sites, enhancing efficiency and reducing processing time.
* **Customization:** Craft computations tailored to your project's unique requirements, enhancing the depth of insights you can derive.

In summary, Generalized Compute empowers you to harness the versatility and power of containerized computations within the Rhino FCP. By executing custom code across distributed datasets, you can streamline tasks, gain valuable insights, and contribute to advanced healthcare research in a secure and privacy-preserving manner

## Creating a New Generalized Compute Code Object

Follow these steps to use the UI to create a new Generalized Compute Code Object.

{% stepper %}
{% step %}

#### Go to the main project's page and select your project.

{% endstep %}

{% step %}

#### Select **Code** from the menu on the left to open the **Code Objects** page.

![](/files/d8b07d4a8da67195913ddf3cedf7217e52b26d74)
{% endstep %}

{% step %}

#### Select the **Create New Code Object** button to open the **Create New Code Object** page.

![](/files/7e7a57a43593e86249e0f7b150f687518d966954)
{% endstep %}

{% step %}

#### Select **Generalized Compute** as the **Code Object Type**.

The page changes to show the fields needed for this object type.

![](/files/cf7777dce8a1aba4909383164c6ee64e572036cc)
{% endstep %}

{% step %}

#### Enter information in the following fields.

*Before you enter information in the fields (see table below), read the following about input and outputs*

*Multiple input and output datasets are supported in Python, Generalized, and Interactive Container Code Objects. While NVFlare Code Objects do not support multiple datasets, they can still be used as before for learning and inference across multiple sites.*

*Code objects usually require that you know the number of inputs and outputs your code needs. Those inputs and outputs are provided as data schemas. For example, if your code splits one dataset into two (like a train/test split), you will have to create a code object with one input schema and two output schemas. For more detail into how to structure your code to run on FCP, see* [*Accessing your datasets in Code Objects*](/creating-and-running-code-objects/accessing-your-datasets-in-code-objects.md)*.*

{% hint style="info" %}
**Pro-tip:** If you don't need this level of control, you can:

* **\[NEW]** set any input or output as optional, which will allow to provide or generate no datasets respectively
* **\[NEW]** set any input or output as a list, which will allow to select 1 or more datasets
* **\[NEW]** select "ANY" as input schema, which will allow code object re-use across any input dataset schema.
* let the system automatically determine the output schemas with [Auto-Generated Data Schemas](/data-schemas/auto-generated-data-schemas.md).
  {% endhint %}

Enter your information in the following fields.

| Field/Button                           | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Name                                   | Indicates the name you want to provide for the code object.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| Description                            | Provides a brief summary of what the code object does.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| Input (including three-dot menu items) | <p>Indicates the file(s) needed for the code object to run properly. Your selection here will affect which Dataset(s) can be selected as input when triggering a Code Run with this Code Object You can enter one or more input files, such as schemas, configuration files, or others. File(s) can be in DICOM, JSON, and or text formats.<br>- If your input doesn't require a schema, select Any.<br>- Select the Add button to the right of the input entry to add another input. You can add as many as you like.<br>- If your input includes more than a single dataset, select the three dot menu to the right of the Input field, then select Input is a list.<br>- If the input is optional, select the three-dot menu to the right of the Input field, then select Optional Input.</p>                                                                                                  |
| Output                                 | <p>Indicates the output file(s) produced when the code object run. You can enter one or more output files. Files can be in DICOM, JSON, and or text formats.<br>- You can also select the option to <em>\[</em> <em>Auto-generate Data Schema from Data</em> <em>]</em>. For more information about Auto-generating Data S <em>chema</em> s, please refer to <a href="/spaces/ydySyCBmy6F7NnGn4bPa/pages/ef1a475d9abe6e933667e507fc5e0b5df2354cb9">Auto-Generated Data <em>Schema</em> s</a>.<br>- Select the Add button to the right of the output entry to add another output. You can add as many as you like.<br>- If your output includes more than a single dataset, select the three dot menu to the right of the Output field, then select Output is a list.<br>- If the output is optional, select the three-dot menu to the right of the Output field, then select Optional Output.</p> |
| Container                              | Specifies the name of the Docker container you would like to execute when running this Code Object. This will either be a container image created by someone from your workgroup that is available on your ECR or an image provided by Rhino. If you need a refresher on pushing containers to your ECR, please refer to [Pushing Containers to the ECR](/getting-started/quick-start-guide/pushing-containers-to-the-ecr.md).                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| {% endstep %}                          |                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |

{% step %}

#### After you've completed your details, you can choose to set required compute resources.

If you want to do this, go to the Setting Required Compute Resources (Auto-scaling) section on this page.
{% endstep %}

{% step %}

#### When complete, select the **Create New Code** button to create the new **Code Object**.

{% endstep %}

{% step %}

#### The code object appears on the **Code Object** page.

{% endstep %}
{% endstepper %}

### Setting Required Compute Resources (Autoscaling)

{% hint style="info" %}
NOTE: This feature is only available if you are running on the Google Cloud Platform (GCP). *For NVFlare code objects and code runs, you can adjust autoscaling methods for both the NVFlare Server and the NVFlare Client. For the NVFlare Server, you can only specify the CPU and RAM.*
{% endhint %}

Autoscaling allows Rhino to run Code Runs on additional compute resources when the default Rhino Client capacity is not sufficient. This capability is designed to support resource‑intensive, "bursty"(processing tasks occur in short, intense or irregular spurts), or GPU‑based workloads while improving isolation, reliability, and cost efficiency. With autoscaling enabled, Rhino provisions temporary cloud VMs on demand, executes workloads on those VMs, and automatically terminates them when execution completes. This can help you make more efficient use of your computing resources and save on costs.

In this section of the screen there are three options: Use Client Resources, Set Minimum, and Use VM Pool.

<figure><img src="/files/GFBA6RaRsa2CcbnjNlEN" alt=""><figcaption></figcaption></figure>

* **Use Client Resources** - Uses the available client settings and resources.
* **Set Minimum** - Allows you to set the minimum number of resources that must be available. If you choose the Set Minimum option, you can indicate the minimum number of CPUs, amount of RAM, the disk size (VRAM), as well as the type and number of GPUs. You can adjust these settings when you create a code object and also when you run your code. Note that adjustments do not persist and will need to be reset with each run.
* **Use VM Pool** - Allows you to select a Pre-allocated Virtual Machine Pool. Select a VM pool to indicate which dedicated compute resource you want to use for your code run, ensuring guaranteed availability for high-demand workloads such as GPU-intensive AI training and inference. This feature provides faster job startup times by using pre-warmed VMs and offers predictable costs through fixed capacity reservations.

Each option is detailed below.

#### *Use Client Resources*

To use client resources, complete the following steps.

1. Select the Use Client Resources button. It should be selected by default.
2. When complete, select Create New Code Object or start a Code Run.

#### *Set Minimum*

To set the minimum required compute resources, complete the following steps.

1. In the Required Compute Resources section, select Set Minimum.
2. Next, make changes to the settings as needed. Note that these setting changes do not persist the next time you run the code object or code run. The settings are explained below.

|           |                                                                                                     |                                                                                                                                                                                                                           |
| --------- | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Setting   | Description                                                                                         | Usage Notes                                                                                                                                                                                                               |
| CPU MIN   | The minimum number of CPU cores on which your code will be run.                                     | If the CPU MIN is too low, you might see slow code execution or throttling.                                                                                                                                               |
| RAM MIN   | The minimum amount of RAM (main memory) that you want to allocate to running your code.             | If this is too low, you might see an out-of-memory crash or slowness in execution due to disk swapping.                                                                                                                   |
| DISK SIZE | Amount of storage allocated to the instance.                                                        | Increasing this setting increases capacity (not necessarily speed)                                                                                                                                                        |
| GPU TYPE  | The model of GPU hardware. The number of GPUs indicate the number of GPUs attached to the instance. | If your code isn’t written to use multiple GPUs, extra GPUs might not be used. Note that GPU availability depends on the selected cloud region, Project‑level GPU quotas, and the current capacity at the cloud provider. |

#### *Use VM Pool*

To select a VM Pool, complete the following steps.

1. If you choose Use VM Pool, available pools are shown. For each pool you will see a name, as well as the current resources available. For each pool you will see the following information.

* Number of VMs (if any)
* Number of vCPUs per VM
* Amount of RAM per VM
* Disk size per VM
* Whether there are GPUs available, and how many

2. Select the pool of your choice.
3. When complete, select Create New Code Object or start a Code Run.

## Viewing a Code Object's Configuration

To get an object's configuration, make sure you are on the page containing the original object you would like to view the configuration of. In the box where your original object is, there should be a row for each version. Navigate to the version you would like to view the configuration for and select the three-dot menu button, shown below:

![](/files/903f8d8b02de3537a785b7355db65dd33f4bf44c)

The menu button is on the right-hand side of the row, click it to view a new menu. The new menu should have an option to show that object's configuration, click on that to see the object's configuration.

For Code Objects and Code Runs, if you chose the Use VM Pool option, you will use the pool that was indicated during the Code Object or Code Run configurations.  Selecting this option shows the resources, including the number of VMs and vCPUs that will be allocated, as well as the amount of RAM, the size of the disk, the number of GPUs, the GPU type/mode&#x6C;*,* and the amount of VRAM.

## Creating a New Generalized Compute Code Object Version

To create a new version, complete the following steps.

{% stepper %}
{% step %}

#### Go to the page for the original object you want to version.

{% endstep %}

{% step %}

#### In the upper right corner of the box where your original object is, select the **+ New Version** button.

{% endstep %}

{% step %}

#### A new window appears that allows you to create a new version of your object.

{% endstep %}
{% endstepper %}

## Removing a Code Object

To remove a code object, complete the following steps.

{% stepper %}
{% step %}

### Open the Code Objects page

Go to the page that lists your code objects.
{% endstep %}

{% step %}

### Find the code object version

If needed, use the search fields at the top right.

* Use **Search Name** to search by code object name.
* Use **Search Description** to search by code object description.

Each code object appears with a row for each version. Find the version you want to delete.
{% endstep %}

{% step %}

### Open the menu for the version

Select the three-dot menu on the right side of the version row, shown below.

![](/files/e30ab6c50f33a01039109108a719df8152c5cb07)
{% endstep %}

{% step %}

### Remove the code object version

Select **Remove code object**.

Repeat these steps for each version you want to remove. When you delete the last version, the code object is fully deleted.
{% endstep %}
{% endstepper %}

## Running a Generalized Compute Code Object

{% hint style="info" %}
Note: To learn how to use the Rhino SDK to run a generalized compute code object see [Running a Generalized Compute Code Object using the Rhino SDK.](/rhino-sdk/running-a-generalized-compute-code-object-using-the-rhino-sdk.md)
{% endhint %}

Create a Generalized Compute code object before you complete these steps.

{% stepper %}
{% step %}

#### Go to your project and open Code Objects

Go to the main project's page and select your project.

Select **Code** from the menu on the left to open the **Code Objects** page.

![](/files/5e8b00865bc09fb8494d115155d47a546eaf6c2f)
{% endstep %}

{% step %}

#### Select the code object to run

Select the Run button for the code object you'd like to run.
{% endstep %}

{% step %}

#### Review the Run Settings page

In the Run Settings page, do the following.
{% endstep %}

{% step %}

#### Fill in the Code Run fields

Fill in the following fields using the table below for reference.

<table><thead><tr><th width="238">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>Input Dataset(s)</strong></td><td>One or more Datasets to be used as input to your Code Run. If multiple Datasets are selected, each will be run separately. If a Dataset happens to be a collaborator's Dataset, the code will be run on that collaborator's Rhino Client. In other words, the data never moves outside your collaborator's Rhino Client.</td></tr><tr><td><strong>Output Dataset Name Suffix</strong></td><td>A suffix that is appended onto the name of each input Dataset to define the name of the output Dataset that will be created during your Code Run.</td></tr><tr><td><strong>Timeout in Seconds</strong></td><td>The number of seconds that must elapse before a Code Run is automatically halted. This is to avoid zombie tasks that run perpetually within a Rhino client.</td></tr><tr><td><strong>Run Parameters (Optional)</strong></td><td>Here you can paste a JSON object that defines additional parameters to be provided to your Code Run. This can be used for specifying hyper-parameters for <em>model</em> validation or any parameter provided to the code. <em>Run parameter</em> s are made available to the container code in a file located at <code>/input/run_params.json</code></td></tr></tbody></table>

{% hint style="info" %}
NOTE: Files that you want to run at run-time should be uploaded to the Rhino Orchestrator.
{% endhint %}

* To add run time files to your code run, select the run time files from the list of the workgroup run files that appears. You will want to only include the ones you need for the code run, like in the example that follows.

![](/files/773ce4c9f454f74ec0eb0ec3c00609ab2d40121d)

{% hint style="info" %}
NOTE: Your workgroup run time files include both published and unpublished files. Files from other workgroups only include the files that they have published and have made visible to you and others in your workgroup that share the same persona. For collaborators to have access to files, they must be published. See [Managing Run Time Files](/creating-and-running-code-objects/managing-run-time-files.md) to learn how to publish run time files.)
{% endhint %}
{% endstep %}

{% step %}

#### Set the required compute resources (if needed).

When you have completed adding all your Code Run details, either set the required compute resources by following the instruction in [Setting Required Compute Resources (Autoscaling)**, or click the Run button to run your code.**](#h_01KJ5WJVMYYGRHFH1JTBR732ZC), or click the **Run** button to run your code.
{% endstep %}
{% endstepper %}

## Reviewing Code Run Logs

Logs that are produced during a Code Run can be viewed by clicking on the Status link in the Code Runs view. If you need to search for a specific code run, use the text boxes located at the top right corner of the page. To search by name enter the full name or part of the name in the Search Name text box. To search by description, enter part of the code run's description in the Search Description text box.

![](/files/b9eac61eb5020042cd8f173ca83e2714a1f39575)

Clicking the Status link will open the Logs view for a particular run:

![](/files/913042c4d818bc668909f24d9376bea0c453f1c8)

Logs can be accessed while the Code Run is active and after the Code Run has completed. While the code is running, an auto-refresh option is available.

This Logs view provides the following information:

* General Info - information about the Code Run, including
  * Code Object name and type
  * Code Object version
  * Code Object description
  * Input and Output Datasets
  * Code Object configuration
  * Code Run configuration
* System Logs
  * Errors or warnings that occurred from the FCP side, e.g. Dataset import errors.
* FL server logs (for NVIDIA FLARE runs)
* FL client logs (for NVIDIA FLARE runs)
  * Each federated client's logs will appear separately
* Client-specific logs (for all other run types)

### Troubleshooting Code Runs

**Error Message that Usage Limit Has Been Exceeded**

If you get an error message indicating that the usage limit has been exceeded, your administrator has set usage limits on the number of input rows a workgroup can process per model within a certain timeframe. Getting this message means that the entire run has been cancelled. Please contact your administrator for more details and help with resolving the issue.

## Viewing a Code Run's Configuration

To get an object's configuration, make sure you are on the page containing the original object you would like to view the configuration of. In the box where your original object is, there should be a row for each version. Navigate to the version you would like to view the configuration for and select the three-dot menu button, shown below:

![](/files/903f8d8b02de3537a785b7355db65dd33f4bf44c)

The menu button is on the right-hand side of the row, click it to view a new menu. The new menu should have an option to show that object's configuration, click on that to see the object's configuration.

For Code Objects and Code Runs, if you chose the Use VM Pool option, you will use the pool that was indicated during the Code Object or Code Run configurations.  Selecting this option shows the resources, including the number of VMs and vCPUs that will be allocated, as well as the amount of RAM, the size of the disk, the number of GPUs, the GPU type/mode&#x6C;*,* and the amount of VRAM.

## Deleting a Code Run

Follow these steps to delete a code run.

{% stepper %}
{% step %}

### Open the Code Runs page

Go to the page that lists your code runs.
{% endstep %}

{% step %}

### Find the code run

If needed, use the search fields at the top right.

* Use **Search Name** to search by code run name.
* Use **Search Description** to search by code run description.

Find the code run you want to delete.
{% endstep %}

{% step %}

### Open the menu for the code run

Select the three-dot menu on the right side of the row, shown below.

![](/files/d6d3efbded53ff3765ccd54d866cfef8630e6c3c)
{% endstep %}

{% step %}

### Remove the code run

Select **Remove code run**.
{% endstep %}
{% endstepper %}

## Adapting your code to run as a Generalized Code Object

For your code to successfully run on the FCP do the following things.

#### Modify Where You Are Reading and Writing From in Your Code

{% stepper %}
{% step %}

#### All inputs should be read from the `/input` folder.

Datasets will be zero-indexed within the `/input` folder (i.e. Input Dataset 1 - `/input/0/`, Input Dataset 2 - `/input/1/`, etc.)

* For Tabular Data: `/input/0/dataset.csv`. If you would like to work with the data as a Pandas DataFrame you can load the Dataset using the following example:

```python
df = pd.read_csv(“/input/0/dataset.csv”)
```

* For file data (data that has been defined in the Data Schema as type Filename): If there is any file data associated with the Dataset, it will be under the `/input/file_data/` directory, structured in the same subdirectory structure as was specified when importing the Dataset. Example:

`/input/0/file_data/`

* For DICOM data (data that has been defined in the Data Schema as type DicomStudyUID, DicomSeriesUID, or DicomInstanceUID.): If there is any DICOM data associated with the Dataset, it will be under the `/input/dicom_data/` directory, structured in the same subdirectory structure as was specified when importing the Dataset. Example:

`/input/0/dicom_data/`
{% endstep %}

{% step %}

#### All output should be written to the `/output` mount.

* For Tabular Data: `/output/0/dataset.csv`. To write your data as a Pandas DataFrame you can use the following example:

```python
df.to_csv(“/output/0/dataset.csv”, index=False)
```

* For file data: `/output/0/file_data/`
* For DICOM data: `/output/0/dicom_data/`
  {% endstep %}
  {% endstepper %}

#### \[If applicable] Modify where you are reading your runtime parameters

Any run parameters that have been specified shall be in a file named `/input/run_params.json`

#### Create the Dockerfile

Create a Dockerfile with the configuration necessary for your code. You may use the example provided in the user-resources github repository as a template and make changes as needed where there are comments beginning with **"# !! EDIT THIS:"**

#### Ensure any files that might be downloaded at runtime are pre-downloaded and copied to the container in the Dockerfile

Copy files that might need to be downloaded at runtime as pre-downloaded copy statements within the Dockerfile

#### Optional recommendations

We recommend that you build and test your container locally. There is a utility script called `docker-run.sh` in the `gc-docker-utils` folder to assist with this. It can be used by running:

```bash
./docker-run.sh path/to/input path/to/output
```
