Schedule GPU workloads and compose tasks
This feature is in Public Preview. To use it, a workspace admin must enable the AI Runtime preview from the Previews page. See Manage Databricks previews.
Use Declarative Automation Bundles to schedule an AI Runtime GPU workload and combine it with other work, such as upstream data preparation. Define the workload and its schedule in YAML, then deploy them as a job. For a training pipeline, run preprocessing on CPU compute and start GPU training after the data is ready.
This guide starts with a workload that prints "Hello world", then builds a scheduled preprocessing and training pipeline. If you already have an AI Runtime workload YAML, see Convert an AI Runtime workload to a bundle.
How it works
- A bundle contains your code and a
databricks.ymlconfiguration. - A job groups tasks and defines when they run.
- An
ai_runtime_taskruns your command on serverless GPU compute.
The depends_on field connects tasks into a directed acyclic graph (DAG), and the schedule field schedules your job to run on a cadence.
databricks bundle deploy uploads the code and creates or updates the job. databricks bundle run starts a job run immediately. On each run, AI Runtime provisions the GPU compute, runs your command, and records the run in the named MLflow experiment.
Requirements
- A workspace in a supported region. See Requirements.
- The latest Databricks CLI, authenticated to your workspace. Update an existing installation before using these examples.
- For the preprocessing example, an existing Unity Catalog volume. The job's run identity needs
USE CATALOG,USE SCHEMA,READ VOLUME, andWRITE VOLUMEprivileges for that volume. See Privileges for Unity Catalog volumes.
Hello world example with ai_runtime_task
This example runs a shell command on one A10 GPU.
-
Create a directory for the bundle. In that directory, create
command.shwith the following contents:Bash#!/usr/bin/env bash
set -euo pipefail
echo "Hello world" -
Create
databricks.ymlin the same directory:YAMLbundle:
name: hello-ai-runtime
resources:
jobs:
hello:
name: hello-ai-runtime
tasks:
- task_key: hello
environment_key: default
ai_runtime_task:
experiment: hello-ai-runtime
deployments:
- command_path: ${workspace.file_path}/command.sh
compute:
accelerator_type: GPU_1xA10
accelerator_count: 1
environments:
- environment_key: default
spec:
environment_version: '6'
targets:
dev:
mode: development
default: true -
From the bundle directory, validate, deploy, and run the job:
Bashdatabricks bundle validate --target dev
databricks bundle deploy --target dev
databricks bundle run hello --target dev -
Open the run URL printed by the CLI and view the
hellotask's output. It containsHello world. The run also appears in thehello-ai-runtimeMLflow experiment.
More complex example: Schedule a preprocessing and training pipeline
This example prepares a small dataset on CPU compute, then trains a linear model on one A10 GPU. Both tasks use the same volume file to pass data between them. The job is scheduled for 09:00 UTC daily, with the schedule paused until you test it.
Create the project
-
Create a separate directory with the following layout:
Textscheduled-training/
├── databricks.yml
├── prep.py
├── command.sh
└── src/
└── train.py -
Create
prep.pywith the following contents. Replace<catalog>,<schema>, and<volume>with your volume's names. This source notebook normalizes the input values and writes the prepared data to the volume:Python# Databricks notebook source
import json
from pathlib import Path
data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
inputs = [value / 10 for value in range(10)]
data = {"x": inputs, "y": [2 * value + 1 for value in inputs]}
data_path.parent.mkdir(parents=True, exist_ok=True)
data_path.write_text(json.dumps(data))
print(f"Prepared {len(inputs)} rows at {data_path}") -
Create
src/train.pywith the following contents. Use the same volume path as inprep.py:Pythonimport json
from pathlib import Path
import torch
data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
data = json.loads(data_path.read_text())
x = torch.tensor(data["x"], device="cuda").reshape(-1, 1)
y = torch.tensor(data["y"], device="cuda").reshape(-1, 1)
model = torch.nn.Linear(1, 1).to("cuda")
optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
for _ in range(200):
optimizer.zero_grad()
loss = torch.nn.functional.mse_loss(model(x), y)
loss.backward()
optimizer.step()
print(f"Trained on {x.device}: loss={loss.item():.4f}") -
Create
command.shto run the training code. The bundle below packagessrc/, soCODE_SOURCE_PATHpoints to the extractedsrcdirectory:Bash#!/usr/bin/env bash
set -euo pipefail
cd "$CODE_SOURCE_PATH"
python train.py -
Create
databricks.ymlwith the following contents. The bundle packagessrc/for the GPU task. Theprepnotebook runs on serverless CPU compute, anddepends_onstartstrainonly afterprepsucceeds.max_concurrent_runs: 1prevents runs of this job from overwriting each other's input file:YAMLbundle:
name: scheduled-training
artifacts:
code:
type: tgz
path: .
include: [src]
files:
- source: ./dist/code.tgz
resources:
jobs:
train_pipeline:
name: scheduled-training
max_concurrent_runs: 1
schedule:
quartz_cron_expression: '0 0 9 * * ?'
timezone_id: UTC
pause_status: PAUSED
tasks:
- task_key: prep
notebook_task:
notebook_path: ./prep.py
- task_key: train
depends_on:
- task_key: prep
environment_key: training
ai_runtime_task:
experiment: scheduled-training
code_source_path: ./dist/code.tgz
deployments:
- command_path: ${workspace.file_path}/command.sh
compute:
accelerator_type: GPU_1xA10
accelerator_count: 1
environments:
- environment_key: training
spec:
environment_version: '6'
dependencies:
- torch
targets:
dev:
mode: development
default: true
Test and enable the schedule
-
From
scheduled-training/, deploy the bundle and start a manual run:Bashdatabricks bundle validate --target dev
databricks bundle deploy --target dev
databricks bundle run train_pipeline --target dev -
Open the run URL printed by the CLI. Confirm that
prepsucceeds beforetrainstarts. Theprepoutput reports 10 prepared rows, and thetrainoutput reportsTrained on cuda:0and the loss. -
In
databricks.yml, change the schedule'spause_statusfromPAUSEDtoUNPAUSED. The schedule block is now:YAMLschedule:
quartz_cron_expression: '0 0 9 * * ?'
timezone_id: UTC
pause_status: UNPAUSED -
Deploy the change to activate the daily schedule:
Bashdatabricks bundle deploy --target devSetting
pause_status: UNPAUSEDexplicitly enables this schedule even for a development target. In Jobs & Pipelines, open the deployed job and confirm that its schedule is active for 09:00 UTC. See Run jobs on a schedule.
To pause scheduled runs, set pause_status: PAUSED and deploy again. To remove the example job and the files uploaded by the bundle, run the following from its bundle directory:
databricks bundle destroy --target dev
The prepared data in your volume remains. Delete the scheduled-training data directory when you no longer need it.
Additional resources
- Convert an existing AI Runtime workload YAML into a bundle. See Convert an AI Runtime workload to a bundle.
- Configure code packaging, compute, environments, and task parameters. See Configure AI Runtime bundle tasks.
- Schedule an existing notebook on GPU. See Schedule with the Jobs API and Declarative Automation Bundles.
- Track training runs. See Experiment tracking and observability.
- Checkpoint model, optimizer, and data pipeline state. See Improve training performance and resiliency on AI Runtime.