Agent Server
An agent server is the library that turns your agent code into a service. It wraps the agent loop in an HTTP server, defines the API that clients call to run the agent, manages client connections, and determines what happens when a run is interrupted. The agent server runs on the agent runtime. To learn how the layers fit together, see Deploy agents on Databricks.
Agent servers on Databricks
Databricks provides three agent servers. For new agents, Databricks recommends DurableAgentServer.
Agent server | Package | Client API | Durable execution | Used by |
|---|---|---|---|---|
|
| Invocation API at | A Runtime Store that | Projects that you create with the Agent Bricks CLI |
|
| OpenAI Responses API at | Run state in a Lakebase database that you configure. After a crash, a new attempt continues the run from the interrupted attempt's event log. | The |
MLflow |
| OpenAI Responses API at | None | The base app templates, such as |
LongRunningAgentServer extends the MLflow AgentServer, and both serve agents that implement the MLflow ResponsesAgent interface. To deploy and maintain an agent that uses one of them, see Run agents on Databricks Apps using the legacy agent server. To query an agent on any of these servers, see Query agents deployed on Databricks.
DurableAgentServer
DurableAgentServer is the Agent Bricks agent server. It wraps your agent loop in an HTTP server that serves the invocation API, tracks each run, and recovers runs that a crash or restart interrupts. Agents that you create with the Agent Bricks CLI use DurableAgentServer by default.
DurableAgentServer provides:
- One API for every request mode: Synchronous, streaming, and background invocations, plus stream reconnection, all served by the same handler.
- Idempotent invocations: A client-generated invocation ID ensures that a retried request doesn't start a duplicate run.
- Ordered sessions: Invocations in the same session run one at a time, in order.
- Persistent run state: When deployed, run status, events, and results survive worker restarts.
- Crash recovery: The server detects interrupted runs and starts a replacement attempt.
- Request-user authorization: Tools can act with the permissions of the user who sent the request.
- Custom endpoints:
DurableAgentServeris a FastAPI application, so you can add your own routes.
Requirements
DurableAgentServer has the following requirements:
- Python 3.10 and above.
- The
databricks-agentbrickspackage, which includes thedatabricks_agentkitlibrary. Projects that you create withagentbricks initdeclare it as a dependency.
Register your agent
When you create a project with agentbricks init, the CLI does this for you. The generated runtime/main.py creates the server and registers the template's invoke and recovery handlers, so you only edit the agent code in agent/. Follow the steps in this section to bring an existing agent or to write your own handler.
Create a DurableAgentServer and register an async invocation handler with @app.invoke. The handler receives the request's input and an invocation context, and returns a JSON-serializable result. Publish progress as events with context.emit.
from databricks_agentkit import DurableAgentServer, InvocationContext
app = DurableAgentServer()
@app.invoke
async def invoke(input, context: InvocationContext) -> dict:
await context.emit({"type": "status", "message": "Looking that up"})
answer = await run_my_agent(input, session_id=context.session_id)
return {"answer": answer}
You can register one invocation handler, and the server doesn't start without one. The handler serves every request mode: the client chooses whether to wait for the result, stream events, or run in the background.
To run the server locally, start it with agentbricks dev. Projects that you create with agentbricks init include an entrypoint that runs the server with Uvicorn, and an app.yaml file that starts the same entrypoint after you deploy.
Invocation context
The handler's second argument is an InvocationContext:
Attribute | Description |
|---|---|
| The ID that the client sent for this invocation. |
| The session that the invocation belongs to, or |
| The attempt number. The first attempt is |
|
|
| Stores a JSON event, delivers it to streaming clients, and returns the event's position in the stream. |
| The request-user credential resolver, when the agent requires request-user authorization. Otherwise, |
Invocation API
DurableAgentServer serves the invocation API at /api/invocations:
POST /api/invocationsstarts an invocation. By default, the request waits for the result. Setstreamto receive events as Server-Sent Events, orbackgroundto return immediately with a status URL.GET /api/invocations/<id>returns an invocation's status and, after it completes, its output.GET /api/invocations/<id>/events?after=<event-id>streams stored events, so a client can reconnect after a dropped connection.
For request fields, examples, and response formats, see Query agents deployed on Databricks.
Idempotency
Clients send a UUID id with every invocation. The server treats the ID as an idempotency key while it retains the invocation record: resending the same request returns the existing invocation instead of running the agent again. Reusing an ID for a different request returns a 409 error.
Sessions
Clients can send a session_id to group invocations into one conversation. The server stores the session ID separately from input, passes it to your handler as context.session_id, and runs invocations that share a session ID one at a time, in order. The server doesn't infer a session from the invocation ID or the input. Without a session ID, an invocation is sessionless.
Run state
DurableAgentServer stores each invocation's request, status, heartbeats, events, and result in a Runtime Store.
- Local development:
agentbricks devuses an in-process Runtime Store. The invocation API behaves the same way, but run state is lost when the process stops, and the server doesn't restart interrupted work. - Deployed agents:
agentbricks deployprovisions a dedicated database for each deployment's Runtime Store in a Databricks-managed Lakebase project, and reuses it when you redeploy. You can't use your own Lakebase project for the Runtime Store, and you don't create or bind it yourself. Results and events survive worker restarts, and any instance of the agent can serve status and reconnection requests.agentbricks deployments deleteremoves the Runtime Store with the deployment.
The Runtime Store holds the server's execution state. It's separate from the session and memory stores that your agent uses for conversation history and long-term memory.
Crash recovery
To recover runs that a worker crash or restart interrupts, register a recovery handler with @app.recover. When a deployed server detects that a run's heartbeats have stopped, it starts a replacement attempt on an available worker and calls the recovery handler with the original input.
@app.recover
async def recover(input, context: InvocationContext) -> dict:
# Resume from the agent's last checkpoint in the session store,
# or replay the input if that's safe for your agent.
return await resume_my_agent(input, session_id=context.session_id)
If you don't register a recovery handler, automatic recovery is turned off, and the server logs a warning when it starts.
Recovery works as follows:
- When recovery starts: Each running attempt sends a heartbeat every few seconds. If the heartbeats stop, for example because the worker crashes, restarts, or is replaced during a redeployment, the server detects the stale run within seconds and starts a replacement attempt.
- When recovery doesn't start: If your handler raises an exception, the invocation fails and the server doesn't retry it. Recovery covers interrupted workers, not errors in your agent code.
- Number of attempts: The server doesn't limit the number of recovery attempts. Each replacement attempt increases
context.attemptby one. To stop after a number of attempts, checkcontext.attemptin your recovery handler and raise an error. - Manual recovery: You can't trigger recovery manually. Resending a request with the same invocation ID returns the existing invocation instead of starting a new attempt.
Recovery can run your agent code more than one time for the same invocation. An interrupted attempt might have already called external systems before the replacement attempt starts, so make those calls idempotent.
AgentKit library
DurableAgentServer is part of the AgentKit library, databricks_agentkit, which the databricks-agentbricks package includes. Projects that you create with agentbricks init import from it. The library exports the following helpers:
Export | Description |
|---|---|
| The agent server and the context that it passes to your invoke and recovery handlers. |
| A client for managed memory and session stores. It creates and gets stores, and exposes the stores' memories and sessions as |
| Set up MLflow tracing for the agent, and start a trace around a unit of work. |
| Create an authenticated Databricks SDK |
| List the model services that the agent can call through Unity Gateway. |
The library also includes framework helpers in databricks_agentkit.langgraph and databricks_agentkit.openai, which the generated templates use to connect each framework to the session store. For the memory and session APIs, see Managed agent memory and Managed agent sessions.
Request-user authorization
By default, your agent's tools run with the permissions of the app's service principal. To run a tool with the permissions of the user who sent the request, declare user authorization in agent.toml:
-
For a managed tool, set
auth = "user"on the tool entry. Theagentbricks tools addcommands for MCP servers, sandboxes, and Genie Agents writeauth = "user"by default. Pass--auth appto use the app's identity instead. -
For a tool that you write in code, declare the requirement and any API scopes that Agent Bricks can't infer:
Toml[auth.user]
required = true
additional_api_scopes = ["sql"]
When an agent requires user authorization, DurableAgentServer reads the user's credential from the trusted Databricks Apps request headers and keeps it in memory for the active attempt only. The Runtime Store doesn't store the credential. In your handler, get a workspace client for the user from context.request_auth:
@app.invoke
async def invoke(input, context: InvocationContext) -> dict:
user_client = context.request_auth.client_for("user")
me = user_client.current_user.me()
return {"answer": f"Hello, {me.user_name}"}
client_for("app") returns a client that uses the app's service principal. The resolver closes when the attempt ends, so call it inside the handler instead of storing the client. When you run the agent locally with agentbricks dev, client_for("user") uses your local credentials.
When you deploy, agentbricks deploy requests the Databricks Apps user scopes that your tools need. To add missing scopes to an existing app, pass --allow-user-scope-update. See Configure authorization in a Databricks app.
Request-user invocations use the same synchronous, streaming, background, and reconnection APIs. Because the server doesn't store the user's credential, it can't recover an interrupted request-user invocation. The replacement attempt fails with the MCP_USER_AUTH_RECOVERY_UNSUPPORTED error before your handlers run.
Add custom endpoints
DurableAgentServer is a FastAPI application. Add routes alongside the invocation API the same way you add them to any FastAPI app:
@app.get("/status")
async def status() -> dict:
return {"ready": True}
Framework templates
agentbricks init generates two directories:
agent/contains your framework code: the model, prompts, and tools.runtime/contains the adapter that connects the framework toDurableAgentServer, and the entrypoint that registers the adapter's invoke and recovery handlers.
The adapter translates each invocation into a call to the framework's agent loop, and translates the framework's output into events and a result. Both templates register a recovery handler. The LangGraph template resumes from its last checkpoint in the session store, and the OpenAI Agents SDK template replays the request in the same session. To bring an existing agent, add an adapter and a DurableAgentServer entrypoint, and set server = "agentbricks" in the [agent] section of agent.toml.
Limitations
- You can't change the agent server of an existing deployment. To switch between
DurableAgentServerand your own server, create a new project with theagentbricks init --serveroption that you want, and deploy it under a new name. - Changing the
serverfield inagent.tomldoesn't convert existing server code intoDurableAgentServer. - Request-user authorization requires
server = "agentbricks".