Skip to main content

Schedule GPU workloads and compose tasks

Preview

This feature is in Public Preview. To use it, a workspace admin must enable the AI Runtime preview from the Previews page. See Manage Databricks previews.

Use Declarative Automation Bundles to schedule an AI Runtime GPU workload and combine it with other work, such as upstream data preparation. Define the workload and its schedule in YAML, then deploy them as a job. For a training pipeline, run preprocessing on CPU compute and start GPU training after the data is ready.

This guide starts with a workload that prints "Hello world", then builds a scheduled preprocessing and training pipeline. If you already have an AI Runtime workload YAML, see Convert an AI Runtime workload to a bundle.

How it works​

  • A bundle contains your code and a databricks.yml configuration.
  • A job groups tasks and defines when they run.
  • An ai_runtime_task runs your command on serverless GPU compute.

The depends_on field connects tasks into a directed acyclic graph (DAG), and the schedule field schedules your job to run on a cadence.

databricks bundle deploy uploads the code and creates or updates the job. databricks bundle run starts a job run immediately. On each run, AI Runtime provisions the GPU compute, runs your command, and records the run in the named MLflow experiment.

Requirements​

Hello world example with ai_runtime_task​

This example runs a shell command on one A10 GPU.

  1. Create a directory for the bundle. In that directory, create command.sh with the following contents:

    Bash
    #!/usr/bin/env bash
    set -euo pipefail
    echo "Hello world"
  2. Create databricks.yml in the same directory:

    YAML
    bundle:
    name: hello-ai-runtime

    resources:
    jobs:
    hello:
    name: hello-ai-runtime
    tasks:
    - task_key: hello
    environment_key: default
    ai_runtime_task:
    experiment: hello-ai-runtime
    deployments:
    - command_path: ${workspace.file_path}/command.sh
    compute:
    accelerator_type: GPU_1xA10
    accelerator_count: 1
    environments:
    - environment_key: default
    spec:
    environment_version: '6'

    targets:
    dev:
    mode: development
    default: true
  3. From the bundle directory, validate, deploy, and run the job:

    Bash
    databricks bundle validate --target dev
    databricks bundle deploy --target dev
    databricks bundle run hello --target dev
  4. Open the run URL printed by the CLI and view the hello task's output. It contains Hello world. The run also appears in the hello-ai-runtime MLflow experiment.

More complex example: Schedule a preprocessing and training pipeline​

This example prepares a small dataset on CPU compute, then trains a linear model on one A10 GPU. Both tasks use the same volume file to pass data between them. The job is scheduled for 09:00 UTC daily, with the schedule paused until you test it.

Create the project​

  1. Create a separate directory with the following layout:

    Text
    scheduled-training/
    ├── databricks.yml
    ├── prep.py
    ├── command.sh
    └── src/
    └── train.py
  2. Create prep.py with the following contents. Replace <catalog>, <schema>, and <volume> with your volume's names. This source notebook normalizes the input values and writes the prepared data to the volume:

    Python
    # Databricks notebook source
    import json
    from pathlib import Path

    data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
    inputs = [value / 10 for value in range(10)]
    data = {"x": inputs, "y": [2 * value + 1 for value in inputs]}
    data_path.parent.mkdir(parents=True, exist_ok=True)
    data_path.write_text(json.dumps(data))
    print(f"Prepared {len(inputs)} rows at {data_path}")
  3. Create src/train.py with the following contents. Use the same volume path as in prep.py:

    Python
    import json
    from pathlib import Path

    import torch

    data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
    data = json.loads(data_path.read_text())
    x = torch.tensor(data["x"], device="cuda").reshape(-1, 1)
    y = torch.tensor(data["y"], device="cuda").reshape(-1, 1)
    model = torch.nn.Linear(1, 1).to("cuda")
    optimizer = torch.optim.SGD(model.parameters(), lr=0.1)

    for _ in range(200):
    optimizer.zero_grad()
    loss = torch.nn.functional.mse_loss(model(x), y)
    loss.backward()
    optimizer.step()

    print(f"Trained on {x.device}: loss={loss.item():.4f}")
  4. Create command.sh to run the training code. The bundle below packages src/, so CODE_SOURCE_PATH points to the extracted src directory:

    Bash
    #!/usr/bin/env bash
    set -euo pipefail
    cd "$CODE_SOURCE_PATH"
    python train.py
  5. Create databricks.yml with the following contents. The bundle packages src/ for the GPU task. The prep notebook runs on serverless CPU compute, and depends_on starts train only after prep succeeds. max_concurrent_runs: 1 prevents runs of this job from overwriting each other's input file:

    YAML
    bundle:
    name: scheduled-training

    artifacts:
    code:
    type: tgz
    path: .
    include: [src]
    files:
    - source: ./dist/code.tgz

    resources:
    jobs:
    train_pipeline:
    name: scheduled-training
    max_concurrent_runs: 1
    schedule:
    quartz_cron_expression: '0 0 9 * * ?'
    timezone_id: UTC
    pause_status: PAUSED
    tasks:
    - task_key: prep
    notebook_task:
    notebook_path: ./prep.py
    - task_key: train
    depends_on:
    - task_key: prep
    environment_key: training
    ai_runtime_task:
    experiment: scheduled-training
    code_source_path: ./dist/code.tgz
    deployments:
    - command_path: ${workspace.file_path}/command.sh
    compute:
    accelerator_type: GPU_1xA10
    accelerator_count: 1
    environments:
    - environment_key: training
    spec:
    environment_version: '6'
    dependencies:
    - torch

    targets:
    dev:
    mode: development
    default: true

Test and enable the schedule​

  1. From scheduled-training/, deploy the bundle and start a manual run:

    Bash
    databricks bundle validate --target dev
    databricks bundle deploy --target dev
    databricks bundle run train_pipeline --target dev
  2. Open the run URL printed by the CLI. Confirm that prep succeeds before train starts. The prep output reports 10 prepared rows, and the train output reports Trained on cuda:0 and the loss.

  3. In databricks.yml, change the schedule's pause_status from PAUSED to UNPAUSED. The schedule block is now:

    YAML
    schedule:
    quartz_cron_expression: '0 0 9 * * ?'
    timezone_id: UTC
    pause_status: UNPAUSED
  4. Deploy the change to activate the daily schedule:

    Bash
    databricks bundle deploy --target dev

    Setting pause_status: UNPAUSED explicitly enables this schedule even for a development target. In Jobs & Pipelines, open the deployed job and confirm that its schedule is active for 09:00 UTC. See Run jobs on a schedule.

To pause scheduled runs, set pause_status: PAUSED and deploy again. To remove the example job and the files uploaded by the bundle, run the following from its bundle directory:

Bash
databricks bundle destroy --target dev

The prepared data in your volume remains. Delete the scheduled-training data directory when you no longer need it.

Additional resources​