Skip to main content

Query agents deployed on Databricks

How you query an agent depends on the agent server that serves it. Find your agent in the following table, then follow the matching section. To learn about agent servers, see Agent Server.

Your agent

Hosted on

How to query

Uses DurableAgentServer

Agent Runtime on Databricks Apps

Invocation API at /api/invocations

Uses the MLflow AgentServer or LongRunningAgentServer (legacy)

Databricks Apps

Databricks OpenAI client or the OpenAI Responses API at /responses

Deployed to Model Serving (legacy)

Model Serving endpoint

Databricks OpenAI client, REST API, ai_query, or AI Playground

Your agent

Hosted on

How to query

Uses DurableAgentServer

Agent Runtime on Databricks Apps

Invocation API at /api/invocations

Uses the MLflow AgentServer or LongRunningAgentServer (legacy)

Databricks Apps

Databricks OpenAI client or the OpenAI Responses API at /responses

Deployed to Model Serving (legacy)

Model Serving endpoint

Databricks OpenAI client, REST API, ai_query, or AI Playground

Agents hosted on Databricks Apps require a Databricks OAuth token. Personal access tokens don't work for Databricks Apps. To generate OAuth tokens from a script, a service principal, another app, or a notebook, see Connect to an API Databricks app using token authentication.

Query an agent that uses DurableAgentServer​

Agents that use DurableAgentServer serve the invocation API. Agents that you create with the Agent Bricks CLI use DurableAgentServer, and agentbricks deploy deploys them to Agent Runtime as an app named agent-bricks-<name>. Each request to the API starts one run of your agent, called an invocation.

Endpoint

Description

POST /api/invocations

Starts an invocation. By default, the request waits and returns the result. Set stream to receive events as they happen, or background to return immediately.

GET /api/invocations/<id>

Returns the status of an invocation and, after it completes, its output.

GET /api/invocations/<id>/events?after=<event-id>

Streams the stored events that come after <event-id>. Use this endpoint to reconnect to a stream.

Endpoint

Description

POST /api/invocations

Starts an invocation. By default, the request waits and returns the result. Set stream to receive events as they happen, or background to return immediately.

GET /api/invocations/<id>

Returns the status of an invocation and, after it completes, its output.

GET /api/invocations/<id>/events?after=<event-id>

Streams the stored events that come after <event-id>. Use this endpoint to reconnect to a stream.

Request body​

The request body for POST /api/invocations accepts the following fields. The server rejects requests that contain other fields.

Field

Description

id

Required. A UUID that you generate for each invocation. The server treats the ID as an idempotency key: resending the same request with the same ID returns the existing invocation instead of running the agent again. Reusing an ID for a different request returns a 409 error.

session_id

The conversation that the invocation belongs to. Invocations that share a session ID run one at a time, in order. Agents generated from the CLI templates require this field.

input

The input for your agent. Agents generated from the CLI templates accept a list of messages, or an object with a messages list.

stream

Set to true to receive events as Server-Sent Events (SSE).

background

Set to true to return a 202 response immediately with a status URL, and then poll for the result.

Field

Description

id

Required. A UUID that you generate for each invocation. The server treats the ID as an idempotency key: resending the same request with the same ID returns the existing invocation instead of running the agent again. Reusing an ID for a different request returns a 409 error.

session_id

The conversation that the invocation belongs to. Invocations that share a session ID run one at a time, in order. Agents generated from the CLI templates require this field.

input

The input for your agent. Agents generated from the CLI templates accept a list of messages, or an object with a messages list.

stream

Set to true to receive events as Server-Sent Events (SSE).

background

Set to true to return a 202 response immediately with a status URL, and then poll for the result.

Your agent's handler defines the structure of input. Agents generated from the CLI templates read the following fields when input is an object:

Field

Description

messages

The conversation turns to send to the agent.

actor

The identity whose long-term memory the agent reads and writes. If you don't pass an actor, the agent uses the session ID, so memories don't carry over to a new session. Set actor from your application's signed-in user, not from text that the user types.

model

The model to use for this invocation, instead of the model set in the agent code.

resume

The response to an agent that paused for human input, such as approval of a tool call. When an agent pauses, the invocation's status is interrupted. Send resume in a new invocation with the same session_id to continue.

Field

Description

messages

The conversation turns to send to the agent.

actor

The identity whose long-term memory the agent reads and writes. If you don't pass an actor, the agent uses the session ID, so memories don't carry over to a new session. Set actor from your application's signed-in user, not from text that the user types.

model

The model to use for this invocation, instead of the model set in the agent code.

resume

The response to an agent that paused for human input, such as approval of a tool call. When an agent pauses, the invocation's status is interrupted. Send resume in a new invocation with the same session_id to continue.

To let the agent recall what it learned about a user across sessions, pass the user's ID as actor:

JSON
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"session_id": "support-case-123",
"input": {
"messages": [{ "role": "user", "content": "What does Databricks do?" }],
"actor": "user-42"
}
}

To test a deployed agent from your terminal, use agentbricks endpoint invoke. The command looks up the app and authenticates with your CLI profile.

Bash
agentbricks --profile <profile> endpoint invoke agent-bricks-<name> \
--path /api/invocations \
--json "{\"id\":\"$(uuidgen)\",\"session_id\":\"$(uuidgen)\",\"input\":[{\"role\":\"user\",\"content\":\"Hello\"}]}"

To stream the response, add "stream":true to the JSON body and pass --sse. To test the agent while it runs locally with agentbricks dev, replace the app name with --url http://localhost:8000.

Stream, run in the background, and reconnect​

  • Stream: Set "stream": true. The response is an SSE stream that includes run.started and run.completed (or run.failed) events, plus the events your agent emits, such as delta events with streamed text. Each event has an ID.
  • Run in the background: Set "background": true. The server returns a 202 response with a status_url. Poll GET /api/invocations/<id> until the status is completed. If you also set "stream": true, the response includes an events_url that you can read events from.
  • Reconnect: If a stream disconnects, call GET /api/invocations/<id>/events?after=<event-id> with the ID of the last event that you received.

The following Python example streams a response:

Python
with requests.post(
f"{app_url}/api/invocations",
headers=w.config.authenticate(),
json={
"id": str(uuid.uuid4()),
"session_id": session_id,
"input": [{"role": "user", "content": "Summarize our last conversation."}],
"stream": True,
},
stream=True,
) as response:
response.raise_for_status()
for line in response.iter_lines(decode_unicode=True):
if line.startswith("data: "):
print(line[len("data: "):])

If you deploy the agent with more than one instance, send the session ID in an X-Routing-Key header to route all requests in a session to the same instance.

Query an agent that uses the legacy MLflow AgentServer​

Use this section for agents that you deploy on Databricks Apps with the legacy agent server: the MLflow AgentServer or LongRunningAgentServer, with the ResponsesAgent interface. These agents serve the OpenAI Responses API at /responses.

LongRunningAgentServer serves the same API, so the following examples also apply to it. It also supports background runs: set background to true in the request, and then retrieve the response with GET /responses/<response-id>?stream=true&starting_after=<sequence-number>, which streams the events after that sequence number.

Databricks recommends the Databricks OpenAI client for these agents. Include the apps/ prefix in the model name.

Python
from databricks.sdk import WorkspaceClient
from databricks_openai import DatabricksOpenAI

input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
app_name = "<agent-app-name>"

# The WorkspaceClient must use OAuth authentication.
w = WorkspaceClient()
client = DatabricksOpenAI(workspace_client=w)

# Non-streaming request
response = client.responses.create(model=f"apps/{app_name}", input=input_msgs)
print(response)

# Streaming request
streaming_response = client.responses.create(
model=f"apps/{app_name}", input=input_msgs, stream=True
)
for chunk in streaming_response:
print(chunk)

To pass custom_inputs, use the extra_body parameter:

Python
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_body={"custom_inputs": {"id": 5}},
)

To get the trace ID for a request, include the x-mlflow-return-trace-id header. Then use MLflow get_trace to retrieve the full trace.

Python
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_headers={"x-mlflow-return-trace-id": "true"},
)
trace_id = response.metadata["trace_id"]
trace = client.get_trace(trace_id)

Query a legacy agent on Model Serving​

Use this section for legacy agents deployed to Model Serving endpoints. You can authenticate with a Databricks OAuth token or a personal access token. To move these agents to Databricks Apps, see Migrate an agent from Model Serving to Databricks Apps.

For agents that use the ResponsesAgent interface, call responses.create with the endpoint name as the model:

Python
from databricks_openai import DatabricksOpenAI

input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"

client = DatabricksOpenAI()

# Non-streaming request. Calls predict.
response = client.responses.create(model=endpoint, input=input_msgs)
print(response)

# Streaming request. Calls predict_stream.
streaming_response = client.responses.create(model=endpoint, input=input_msgs, stream=True)
for chunk in streaming_response:
print(chunk)

For agents that use the legacy ChatAgent or ChatModel interfaces, use the chat completions client:

Python
from databricks.sdk import WorkspaceClient

messages = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"

client = WorkspaceClient().serving_endpoints.get_open_ai_client()
response = client.chat.completions.create(model=endpoint, messages=messages)
print(response)

With either client, pass custom_inputs or databricks_options through the extra_body parameter. For example, extra_body={"databricks_options": {"return_trace": True}} returns the trace with the response.

Additional resources​