Deploy agents on Databricks
To run an agent in production, you deploy its code to managed compute that serves requests from your users and applications. A deployed agent has three layers: your framework or harness, an agent server, and an agent runtime. Agent Bricks provides options at each layer, from DurableAgentServer to Agent Runtime on Databricks Apps.
The agent compute stack
A deployed agent has three layers. Each layer wraps the one above it.
Layer | What it does | Options on Databricks |
|---|---|---|
Runs the agent loop: calls models and tools and decides what to do next. | Any framework or harness, such as LangGraph or the OpenAI Agents SDK, or a meta-harness such as Omnigent. | |
Wraps the agent loop in an HTTP server. The agent server exposes the invocation API, manages client connections, and handles durability. |
| |
Runs the agent server on managed compute and handles hosting, identity, and scaling. | Agent Runtime on Databricks Apps. |
The following terms describe how the layers work together:
Term | What it means |
|---|---|
Invocation API | The HTTP API that clients call to run the agent. See Query agents deployed on Databricks. |
Serve | The agent runtime runs the agent server, which serves your agent through the invocation API. |
Deploy | Put your agent code onto the agent runtime. |
Databricks Sandbox isn't a layer of the compute stack. It's a tool that your agent calls to run code in an isolated environment, separate from the agent runtime, with scoped access to your governed data. Use a sandbox when your agent writes and runs code, such as data analysis scripts. To give a CLI project a sandbox tool, run agentbricks tools add sandbox.
Framework or harness
The framework or harness is your agent code. It runs the loop that reasons with a model, calls tools, and decides when to respond. Agent Bricks doesn't require a specific framework.
Framework or harness | When to use |
|---|---|
Templates | Starting a new agent. The Agent Bricks CLI scaffolds projects for LangGraph and the OpenAI Agents SDK. The templates keep your framework code separate from the code that connects it to the agent server. |
Existing agents | Bringing an agent that you already built with LangGraph or the OpenAI Agents SDK. Run |
Meta-harness | Composing multiple harnesses, such as coding agents, into one agent with Omnigent. |
Agent server
The agent server is a library that wraps your agent loop and turns it into a service. It defines the API that clients call, keeps track of each run, and determines what happens when a run is interrupted.
Agent server | When to use |
|---|---|
| New agents built with the Agent Bricks CLI. |
Legacy agent servers | Maintaining agents built from the app templates. To compare |
Your own server | Agents that need custom endpoints, request formats, or protocols. Run |
Agent runtime
The agent runtime is the managed compute that runs your agent server. You deploy code to it instead of provisioning servers yourself.
Agent runtime | When to use |
|---|---|
New agents. Runs your agent on Databricks Apps. | |
Model Serving (legacy) | Earlier agents deployed to Model Serving endpoints. See Deploy an agent for AI applications (Model Serving). To move them to Databricks Apps, see Migrate an agent from Model Serving to Databricks Apps. |
Deploy with the Agent Bricks CLI
The Agent Bricks CLI connects the three layers. It scaffolds your framework code, runs it on DurableAgentServer locally, and deploys it to Agent Runtime:
agentbricks init my-agent --framework langgraph
cd my-agent
agentbricks dev
agentbricks deploy my-agent
To build and deploy your first agent, see Agent Bricks quickstart.
Permissions to deploy an agent
The user or service principal that runs agentbricks deploy needs permission to do the following:
- Create Databricks Apps, or manage the existing app for a redeployment. See Configure permissions for a Databricks app.
- Create the memory and session stores declared in
agent.toml, or manage the existing stores, so that Agent Bricks can grant the app's service principal access to them. - Create the MLflow experiment that the project uses for tracing, or edit it if it already exists.
- Grant the app's service principal access to resources that tools declare with
--auth app, such asEXECUTEon a Unity Catalog function,CAN RUNon a Genie Agent, orSELECTon a table in a sandbox scope.
Agent Bricks grants access only to resources that agent.toml declares directly. If a tool uses other resources, such as the tables behind a Genie Agent or the objects that a Unity Catalog function calls, grant the app's service principal access to them yourself. If deploy can't apply a required grant, it stops before it uploads your code and leaves the current deployment running.