Skip to main content

Configure AI Runtime bundle tasks

Use this reference to customize an ai_runtime_task in Declarative Automation Bundles. For runnable examples that schedule a GPU workload and add preprocessing tasks, start with Schedule GPU workloads and compose tasks.

For task fields and their definitions, see the AI Runtime task reference. Set retries, timeouts, permissions, and environment_key on the surrounding task or job.

Ship your training code​

The bundle syncs files such as command.sh on deployment. A command-only workload can point command_path at ${workspace.file_path}/command.sh and omit code_source_path.

For a separate directory of training code, set code_source_path to a packaged artifact or a workspace or volume path.

Package a local directory​

Declare a tgz artifact to package your code on databricks bundle deploy. This configuration fragment packages src/ and points the task at the resulting archive:

YAML
artifacts:
code:
type: tgz
path: .
include: [src]
files:
- source: ./dist/code.tgz

resources:
jobs:
train:
tasks:
- task_key: train
ai_runtime_task:
code_source_path: ./dist/code.tgz

The runtime extracts the archive and sets CODE_SOURCE_PATH to its top-level code directory. With path: . and include: [src], the archive contains src/train.py and CODE_SOURCE_PATH is /databricks/code_source/src. After changing to that directory, run python train.py:

Bash
#!/usr/bin/env bash
set -euo pipefail
cd "$CODE_SOURCE_PATH"
python train.py

For artifact settings, including path, include, git, and files, see the bundle artifacts reference.

Use uploaded code​

Set code_source_path to a /Workspace/… or /Volumes/… path for a code archive that is already uploaded. The bundle uses that path without packaging your local directory.

Configure compute and environments​

Set accelerator_count to a multiple of the GPUs per node encoded in accelerator_type. For example, GPU_8xH100 with accelerator_count: 16 runs on two nodes. See Hardware options for hardware guidance.

For a multi-node workload, AI Runtime runs the command on every node. The task environment includes NUM_NODES, WORLD_SIZE, LOCAL_WORLD_SIZE, MASTER_ADDR, and MASTER_PORT. Read these variables in your distributed training command.

Declare dependencies in the job's environments block, and select the environment with the task's environment_key. The following configuration fragment installs PyTorch in the training environment:

YAML
resources:
jobs:
train:
tasks:
- task_key: train
environment_key: training
environments:
- environment_key: training
spec:
environment_version: '6'
dependencies:
- torch

Use environment_version for a Standard environment. To use a Databricks AI environment, set base_environment instead, for example workspace-base-environments/databricks_ai_v6. See Set up your environment.

For a custom Docker image, set docker_image_url on ai_runtime_task and follow Use custom Docker images with the legacy Python CLI.

Pass configuration and data to tasks​

Define a job-level environment_variables entry and select it with the task's environment_variables_key. Put plain values in variables and use {{secrets/scope/key}} for secret references. This configuration requires the environment variables preview described in Configure environment variables for serverless jobs.

To pass parameters through HYPERPARAMETERS_PATH, save a hyperparameters.yaml file alongside the script referenced by command_path. The bundle syncs both files, and the runtime sets HYPERPARAMETERS_PATH to the parameter file for your training command. For conversion from an existing workload YAML, see Convert an AI Runtime workload to a bundle.

note

An ai_runtime_task does not support job task values ({{tasks.<task_key>.values.<name>}} or dbutils.jobs.taskValues). To pass data between tasks, write it to a shared location, such as a Unity Catalog volume, and use the same path in both tasks.

Additional resources​