Skip to main content

What is Agent Bricks?

Agent Bricks is the Databricks agent developer platform. Use it to build and deploy the agents that power your agentic products and workflows. Agent Bricks is built for developers who write code. You bring an agent built with any framework or harness and choose any model.

Agent Bricks supports the core workflows for building a production agent:

  • Deploy your agent: Agent Runtime hosts agents built with any framework or harness, stateful or stateless. DurableAgentServer, part of the AgentKit library, adds durable execution so that runs survive restarts and crashes.
  • Run code safely: Agents run the code they write in Databricks Sandbox, an isolated environment whose home directory persists across sessions.
  • Give your agent context: Managed agent memory lets your agent recall what it learned in earlier conversations, so it can personalize its responses.
  • Debug and test your agent: MLflow Tracing records each step that your agent takes, and you can store and govern traces at scale in Unity Catalog.

The Agent Bricks CLI scaffolds agent code that connects to these services. You can iterate locally, including with coding agents, and then deploy at scale from the declarative agent.toml file that records your agent's resources.

How Agent Bricks relates to other Databricks offerings:

  • Genie Agents: Genie Agents is for creating agents without code, such as agents that answer questions over your tables and documents. Agent Bricks is for developers who write their own agent code. A custom agent can call Genie Agents as tools.
  • Genie One: Genie One is for business users who want to ask questions about data, view dashboards, and use apps without writing code. Agent Bricks is for developers who build the agents themselves. A custom agent can call Genie One as a tool to draw on your business context.
  • Omnigent: Omnigent is a meta-harness that composes and coordinates multiple agents, such as coding agents. You can use Omnigent to develop a custom agent that you deploy and run on Agent Bricks.

Components of an agent​

An agent is an AI system that perceives a request, decides what to do, and acts to achieve a goal. Unlike a single model call, an agent runs a loop. It reasons about the request with a model, calls tools to gather information or take actions, observes the results, and repeats until it can respond.

Components of an agent: requests from users, apps, and agents reach the agent runtime, which runs the agent server, which keeps durable run state and wraps the agent loop built with a framework or harness. The loop reasons with a model, calls tools, and reads and writes agent state, with observability and governance around it.

A production agent typically combines the following components:

  • Framework or harness: Your code. The agent loop, built with a framework or harness such as LangGraph or the OpenAI Agents SDK, plus the instructions and orchestration logic that define how the agent behaves.
  • Model: The large language model (LLM) that reasons about each step and decides which action to take next.
  • Tools: Functions that the agent calls to retrieve data, including your business data, or to change the state of another system. Examples include Model Context Protocol (MCP) servers, APIs, and sandboxes that run the code the agent writes.
  • Agent state: What the agent keeps track of. Agent state includes the current session, which is the conversation or task in progress, and memory, which is what the agent learns and recalls across sessions.
  • Agent server: The library that wraps the agent loop in an HTTP server. It exposes the API that clients call, manages client connections, and handles durability by keeping each run's state in a Runtime Store.
  • Agent runtime: The managed compute that runs the agent server and handles hosting, identity, and scaling.
  • Observability: Traces that record each step the agent takes, and evaluations that measure whether its output is good.
  • Governance: Identity, permissions, and policies that control which models, tools, and data the agent can use.

The framework or harness, the agent server, and the agent runtime together form the agent compute stack. See Deploy agents on Databricks.

Challenges of building agents in production​

A prototype agent can run on a laptop in an afternoon. Running agents reliably for real users is harder. Teams that build agents on their own often spend most of their time assembling and operating infrastructure instead of improving the agent. The common challenges fall into five areas:

  • Model access: Teams need to choose among frontier and open models, switch or upgrade models without rewriting code, and get enough capacity at production scale, all while controlling cost and access.
  • Compute: Agents need compute that scales with traffic, keeps long-running work alive through restarts and failures, and isolates each session. Agents that write and run code need a secure place to run it.
  • Context: Agents give better answers when they have the right data at the right time. That requires durable conversation state, memory that persists across sessions, and secure access to business data and tools without copying data or sharing credentials.
  • Quality: Agents produce open-ended, non-deterministic output, so "good" is hard to define and measure. Without tracing and evaluation, teams can't debug failures, prove quality to stakeholders, or catch regressions in production.
  • Governance and control: Agents act on data and call external services. Without unified permissions, policies, and usage tracking, organizations risk data leakage, compliance gaps, unauthorized use, and unexpected costs. Teams also want to avoid locking their agents into one model provider or cloud.

How Agent Bricks addresses these challenges​

Agent Bricks provides a managed building block for each layer of the stack. The components share one identity and governance model and work together, but you can also use each one on its own. You keep your framework, harness, and model choices.

Agent Bricks platform: your agent runs on Agent Runtime with durable execution and crash recovery. Unity Gateway connects it to frontier and open-weight models and governs its MCP servers and skills, Sandbox runs code in isolation, MCP servers provide tools, managed memory stores context on Lakebase, and MLflow provides tracing and evaluation.

Challenge

Agent Bricks components

Model access

Unity Gateway and Foundation Model APIs

Compute

Agent Server (DurableAgentServer), Agent Runtime, and Databricks Sandbox

Context

Managed agent memory and managed agent sessions, MCP servers and agent tools, AI Search, and Genie Agents

Quality

MLflow tracing, evaluation, and monitoring

Governance and control

Unity Catalog and Unity Gateway policies, plus support for any framework, harness, and model

Challenge

Agent Bricks components

Model access

Unity Gateway and Foundation Model APIs

Compute

Agent Server (DurableAgentServer), Agent Runtime, and Databricks Sandbox

Context

Managed agent memory and managed agent sessions, MCP servers and agent tools, AI Search, and Genie Agents

Quality

MLflow tracing, evaluation, and monitoring

Governance and control

Unity Catalog and Unity Gateway policies, plus support for any framework, harness, and model

Get started with the Agent Bricks CLI​

The Agent Bricks CLI provides a guided path from an empty directory to a deployed agent. It scaffolds a project from a framework template, runs the agent locally, and deploys it with model access, memory, sessions, tools, and tracing already configured. A declarative agent.toml file records the managed resources the agent depends on. To build your first agent, see Agent Bricks quickstart.

Models​

Unity Gateway gives your agent one API for frontier and open models, including models that Databricks hosts through Foundation Model APIs and models from external providers. You can switch models without changing agent code. Unity Gateway also governs the MCP servers and skills that your agents use, and applies rate limits, guardrails, and usage tracking to the requests that pass through it.

Compute​

  • Agent server: DurableAgentServer wraps your agent loop and serves the invocation API, with synchronous, streaming, and background requests, persistent run state, and crash recovery. Agents that you create with the Agent Bricks CLI use it by default.
  • Agent runtime: Agent Runtime runs the agent server on Databricks Apps at a stable, authenticated endpoint. agentbricks deploy provisions the resources that your agent needs and grants it access to them. To query a deployed agent, see Query agents deployed on Databricks.
  • Databricks Sandbox: Databricks Sandbox is a tool that your agent calls to run the code it writes in an isolated environment, with scoped access to your governed data.

To learn how these layers fit together, see Deploy agents on Databricks.

Context​

  • Managed memory and sessions: Sessions store the state of one conversation or task, and memory stores the facts and preferences that an agent recalls in later conversations. Both are managed stores backed by Lakebase that work with any framework. See Agent memory and sessions.
  • Combined business context with Genie One: The Genie One MCP server gives your agent your organization's combined business context. The agent asks a question in natural language, and Genie answers it across your governed data, using Genie Ontology to resolve business terms, metric definitions, and table relationships. Unity Catalog permissions apply, so the agent can only query data that it's allowed to access. To add it to an Agent Bricks CLI project, run agentbricks tools add genie-one.
  • Tools and MCP servers: Connect agents to Databricks-managed MCP servers for Genie Agents, AI Search, Unity Catalog functions, and SQL, or to custom and external MCP servers. See MCPs and agent tools.
  • AI Search: Retrieve relevant unstructured content with Databricks AI Search, including for retrieval-augmented generation (RAG) applications.
  • Business data: Ground agents in the tables, files, and Genie Agents that your organization already governs in Unity Catalog.

Observability and quality​

Databricks-managed MLflow provides tracing, evaluation, and monitoring for agents. MLflow Tracing records each step that an agent takes, in development and in production. Evaluation uses built-in LLM judges, custom scorers, and expert feedback to measure quality, and production monitoring catches regressions. You can store traces in Unity Catalog and analyze them at scale. See What is agent observability and quality?.

Governance and control​

  • Unified governance: Unity Catalog governs the data, tools, models, and agents that your agents use, with permissions and lineage.
  • Policies on models, MCP servers, and skills: Unity Gateway governs access to models, MCP servers, and skills, and enforces guardrails, rate limits, and usage tracking on the calls that it routes.
  • Choice: Use any framework or harness and any model, and switch models without rewriting your agent.

Agent development lifecycle​

Building a production agent is iterative. The following steps describe a typical lifecycle.

  1. Define the use case and success criteria: Agree on what the agent should do, who uses it, and how you measure quality, cost, and latency. Collect example requests and expected answers early.
  2. Build an initial agent: Start with the least complex design that works. Connect the agent to models, MCP servers, and skills through Unity Gateway, and run it locally.
  3. Evaluate and iterate on quality: Use traces to debug behavior, build an evaluation dataset from real requests and expert feedback, and run evaluations to measure the effect of each change to prompts, tools, or models.
  4. Deploy: Deploy the agent server to the agent runtime, with the stores, tools, and permissions that the agent needs.
  5. Monitor and improve in production: Monitor quality with production traces and scorers, collect user feedback, and feed what you learn back into evaluation.

Agent Bricks and other agent offerings on Databricks​

  • Genie One: The Databricks experience for business users to ask questions about data in natural language, view dashboards, and use apps, without writing code. A custom agent built on Agent Bricks can call Genie One as a tool. To add it to an Agent Bricks CLI project, run agentbricks tools add genie-one.
  • Genie Agents: A low-code way to build agents that answer questions over your tables and documents. Use Genie Agents when you don't need custom agent code. Your custom agents can also call Genie Agents as tools.
  • Omnigent: A meta-harness for composing and collaborating with agents, such as combining coding agents. You can use Omnigent to develop custom agents that run on Agent Bricks.
  • Legacy agent builders: Earlier products, such as Knowledge Assistant and Supervisor Agent, are no longer recommended for new agents. See Legacy agent offerings.

Additional resources​