Skip to main content

Deploy agents on Databricks

To run an agent in production, you deploy its code to managed compute that serves requests from your users and applications. A deployed agent has three layers: your framework or harness, an agent server, and an agent runtime. Agent Bricks provides options at each layer, from DurableAgentServer to Agent Runtime on Databricks Apps.

The agent compute stack​

A deployed agent has three layers. Each layer wraps the one above it.

The agent compute stack: your framework or harness, wrapped by an agent server that exposes the invocation API, running on Agent Runtime.

Layer

What it does

Options on Databricks

Framework or harness

Runs the agent loop: calls models and tools and decides what to do next.

Any framework or harness, such as LangGraph or the OpenAI Agents SDK, or a meta-harness such as Omnigent.

Agent server

Wraps the agent loop in an HTTP server. The agent server exposes the invocation API, manages client connections, and handles durability.

DurableAgentServer (recommended), legacy agent servers, or your own server.

Agent runtime

Runs the agent server on managed compute and handles hosting, identity, and scaling.

Agent Runtime on Databricks Apps.

Layer

What it does

Options on Databricks

Framework or harness

Runs the agent loop: calls models and tools and decides what to do next.

Any framework or harness, such as LangGraph or the OpenAI Agents SDK, or a meta-harness such as Omnigent.

Agent server

Wraps the agent loop in an HTTP server. The agent server exposes the invocation API, manages client connections, and handles durability.

DurableAgentServer (recommended), legacy agent servers, or your own server.

Agent runtime

Runs the agent server on managed compute and handles hosting, identity, and scaling.

Agent Runtime on Databricks Apps.

The following terms describe how the layers work together:

Term

What it means

Invocation API

The HTTP API that clients call to run the agent. See Query agents deployed on Databricks.

Serve

The agent runtime runs the agent server, which serves your agent through the invocation API.

Deploy

Put your agent code onto the agent runtime.

Term

What it means

Invocation API

The HTTP API that clients call to run the agent. See Query agents deployed on Databricks.

Serve

The agent runtime runs the agent server, which serves your agent through the invocation API.

Deploy

Put your agent code onto the agent runtime.

Where sandboxes fit

Databricks Sandbox isn't a layer of the compute stack. It's a tool that your agent calls to run code in an isolated environment, separate from the agent runtime, with scoped access to your governed data. Use a sandbox when your agent writes and runs code, such as data analysis scripts. To give a CLI project a sandbox tool, run agentbricks tools add sandbox.

Framework or harness​

The framework or harness is your agent code. It runs the loop that reasons with a model, calls tools, and decides when to respond. Agent Bricks doesn't require a specific framework.

Framework or harness

When to use

Templates

Starting a new agent. The Agent Bricks CLI scaffolds projects for LangGraph and the OpenAI Agents SDK. The templates keep your framework code separate from the code that connects it to the agent server.

Existing agents

Bringing an agent that you already built with LangGraph or the OpenAI Agents SDK. Run agentbricks init --existing in its directory. See Bring an existing agent.

Meta-harness

Composing multiple harnesses, such as coding agents, into one agent with Omnigent.

Framework or harness

When to use

Templates

Starting a new agent. The Agent Bricks CLI scaffolds projects for LangGraph and the OpenAI Agents SDK. The templates keep your framework code separate from the code that connects it to the agent server.

Existing agents

Bringing an agent that you already built with LangGraph or the OpenAI Agents SDK. Run agentbricks init --existing in its directory. See Bring an existing agent.

Meta-harness

Composing multiple harnesses, such as coding agents, into one agent with Omnigent.

Agent server​

The agent server is a library that wraps your agent loop and turns it into a service. It defines the API that clients call, keeps track of each run, and determines what happens when a run is interrupted.

Agent server

When to use

DurableAgentServer (recommended)

New agents built with the Agent Bricks CLI.

Legacy agent servers

Maintaining agents built from the app templates. To compare DurableAgentServer with the legacy agent servers, see Agent servers on Databricks.

Your own server

Agents that need custom endpoints, request formats, or protocols. Run agentbricks init --server custom, or keep your existing server.

Agent server

When to use

DurableAgentServer (recommended)

New agents built with the Agent Bricks CLI.

Legacy agent servers

Maintaining agents built from the app templates. To compare DurableAgentServer with the legacy agent servers, see Agent servers on Databricks.

Your own server

Agents that need custom endpoints, request formats, or protocols. Run agentbricks init --server custom, or keep your existing server.

Agent runtime​

The agent runtime is the managed compute that runs your agent server. You deploy code to it instead of provisioning servers yourself.

Agent runtime

When to use

Agent Runtime

New agents. Runs your agent on Databricks Apps. agentbricks deploy provisions the resources your agent needs, grants the agent access to them, and deploys it to a stable, authenticated endpoint.

Model Serving (legacy)

Earlier agents deployed to Model Serving endpoints. See Deploy an agent for AI applications (Model Serving). To move them to Databricks Apps, see Migrate an agent from Model Serving to Databricks Apps.

Agent runtime

When to use

Agent Runtime

New agents. Runs your agent on Databricks Apps. agentbricks deploy provisions the resources your agent needs, grants the agent access to them, and deploys it to a stable, authenticated endpoint.

Model Serving (legacy)

Earlier agents deployed to Model Serving endpoints. See Deploy an agent for AI applications (Model Serving). To move them to Databricks Apps, see Migrate an agent from Model Serving to Databricks Apps.

Deploy with the Agent Bricks CLI​

The Agent Bricks CLI connects the three layers. It scaffolds your framework code, runs it on DurableAgentServer locally, and deploys it to Agent Runtime:

Bash
agentbricks init my-agent --framework langgraph
cd my-agent
agentbricks dev
agentbricks deploy my-agent

To build and deploy your first agent, see Agent Bricks quickstart.

Permissions to deploy an agent​

The user or service principal that runs agentbricks deploy needs permission to do the following:

  • Create Databricks Apps, or manage the existing app for a redeployment. See Configure permissions for a Databricks app.
  • Create the memory and session stores declared in agent.toml, or manage the existing stores, so that Agent Bricks can grant the app's service principal access to them.
  • Create the MLflow experiment that the project uses for tracing, or edit it if it already exists.
  • Grant the app's service principal access to resources that tools declare with --auth app, such as EXECUTE on a Unity Catalog function, CAN RUN on a Genie Agent, or SELECT on a table in a sandbox scope.

Agent Bricks grants access only to resources that agent.toml declares directly. If a tool uses other resources, such as the tables behind a Genie Agent or the objects that a Unity Catalog function calls, grant the app's service principal access to them yourself. If deploy can't apply a required grant, it stops before it uploads your code and leaves the current deployment running.

Additional resources​