Skip to main content

Serve custom LLMs with Custom Model Serving

Preview

This feature is in Public Preview. Workspace admins can control access to this feature from the Previews page. See Manage Databricks previews.

This page shows how to serve your own LLM on a Model Serving GPU endpoint. Any model that vLLM runs becomes a production endpoint with an OpenAI-compatible API. This is how it works:

  • You run the inference server, such as vLLM, and choose its version and settings.
  • You test the server in a serverless GPU notebook. A wrong flag or an out-of-memory error shows up in seconds instead of after a deployment.
  • You deploy the same command and environment that you built and know works to a production serving endpoint.

Quickstart​

Import the following notebook into your workspace and click Run all on AI Runtime with an A10 GPU. In about 15 minutes, you have a Qwen3.5-4B endpoint that answers OpenAI-compatible chat requests.

Serve Qwen3.5-4B with vLLM

When to use custom LLM serving​

Use custom LLM serving in the following cases:

  • You fine-tuned an LLM on AI Runtime and want to serve it. See the section below for an example.
  • You want to serve an open model that Foundation Model APIs (FMAPI) doesn't support, like a new LLM, a speech-to-text model, or an embedding model.
  • You want to change the runtime or architecture of an LLM and need control of the runtime and environment.

Don't use custom LLM serving in the following cases:

  • FMAPI serves the model you need, unchanged. FMAPI is simpler to use and highly optimized.
  • The model doesn't fit on a single GPU with about 80 GB of memory. For example, Qwen3.8-27B and gpt-oss-120b fit, but Kimi K3 and GLM 5.3 don't.
  • You need frontier-level throughput or price per token without tuning. Performance matches open source vLLM on the same GPU.

Requirements​

Custom LLM serving has the following requirements:

  • Your workspace must have serverless GPU compute.
  • You must log the model from a serverless GPU notebook. A model logged from a CPU environment packages CPU dependencies, and the GPU endpoint fails to start.
  • You must have permission to create models in a Unity Catalog schema.
  • You must register the model with env_pack="databricks_model_serving". Custom LLM serving is built on express deployments, which package the notebook's environment with the model.
  • You must use MLflow 3.12 or above and databricks-sdk 0.102.0 or above. The AI v6 environment includes both.

Starter notebooks​

Each notebook takes one model from Hugging Face to a queried endpoint, like the quickstart. Run one as is, or start from one to serve your own model.

Model

Task

Environment

Notebook GPU

Endpoint GPU

Qwen3.5-4B

Chat

AI v6 or Standard v6

A10

A10G (AWS), A100 (Azure)

Qwen3.8-27B

Chat with tool calls

AI v6

H100

H100 (AWS), A100 (Azure)

Gemma 4 26B-A4B

Chat with tool calls

AI v6

H100

H100 (AWS), A100 (Azure)

Muse Glimmer 30B

Chat

Standard v6

H100

H100 (AWS), A100 (Azure)

Whisper large-v3-turbo

Speech to text

AI v6

A10

A10G (AWS), A100 (Azure)

Qwen3-Embedding-0.6B

Embeddings

AI v6

A10

A10G (AWS), A100 (Azure)

Model

Task

Environment

Notebook GPU

Endpoint GPU

Qwen3.5-4B

Chat

AI v6 or Standard v6

A10

A10G (AWS), A100 (Azure)

Qwen3.8-27B

Chat with tool calls

AI v6

H100

H100 (AWS), A100 (Azure)

Gemma 4 26B-A4B

Chat with tool calls

AI v6

H100

H100 (AWS), A100 (Azure)

Muse Glimmer 30B

Chat

Standard v6

H100

H100 (AWS), A100 (Azure)

Whisper large-v3-turbo

Speech to text

AI v6

A10

A10G (AWS), A100 (Azure)

Qwen3-Embedding-0.6B

Embeddings

AI v6

A10

A10G (AWS), A100 (Azure)

AI v6 and Standard v6 are AI Runtime environments. AI v6 comes with vLLM, PyTorch, Transformers, and other common machine learning packages pre-installed and ready for use. See the AI v6 package list. A Standard v6 notebook installs vLLM itself, so you choose its version. Use a Standard v6 notebook when your model needs a newer vLLM or transformers than AI v6 includes, as with Muse Glimmer.

Serve your own model​

To serve another model, pick the starter notebook with the same task and the GPU your model needs. Change MODEL_REPO_ID and the flags in vllm_command, then run the notebook. Every notebook runs the following steps:

  1. Download the model weights from Hugging Face, or get them from a training checkpoint.
  2. Start vLLM in the notebook and query it.
  3. Log the model with the vLLM command as its entrypoint, and register it to Unity Catalog as an express deployment.
  4. Create a serving endpoint, which starts the same command.
  5. Query the endpoint with the OpenAI client, the Databricks SDK, or SQL ai_query.

In the following example, the model's metadata holds the task and the entrypoint:

Python
import mlflow
from mlflow.pyfunc.model import ChatCompletionResponse, ChatModel

# The endpoint runs the entrypoint and never calls predict, but MLflow needs a model class to log.
class Placeholder(ChatModel):
def predict(self, context, messages, params):
return ChatCompletionResponse.from_dict({"choices": []})

model_info = mlflow.pyfunc.log_model(
name="my-model",
python_model=Placeholder(),
artifacts={"model_dir": "my-model"}, # the weights folder, which --model names
metadata={
"task": "llm/v1/chat",
"entrypoint": (
"python -u -m vllm.entrypoints.openai.api_server "
"--model my-model --served-model-name my-model "
"--host 0.0.0.0 --port 8080 --max-model-len 16384"
),
},
)
mlflow.register_model(model_info.model_uri, "<catalog>.<schema>.my_model", env_pack="databricks_model_serving")

Serve a fine-tuned model​

Fine-tune a model on AI Runtime and serve it from the same notebook. For an example, see Supervised fine-tuning (Full) and serving of Qwen3.5-0.8B. The tutorial fine-tunes Qwen3.5-0.8B on a single H100, compares answers before and after training, and serves the fine-tuned model.

Supported tasks​

The task in the model's metadata sets the API the endpoint serves. Your server must expose the OpenAI-compatible API for that task. Other tasks, such as llm/v1/completions, aren't supported. The following table lists the supported tasks.

task

Model type

Query with

llm/v1/chat

Chat models, including vision-language models

chat.completions

llm/v1/embeddings

Embedding models

embeddings

llm/v1/audio/transcriptions

Speech-to-text models

audio.transcriptions

llm/v1/audio/translations

Speech-to-English translation models

audio.translations

task

Model type

Query with

llm/v1/chat

Chat models, including vision-language models

chat.completions

llm/v1/embeddings

Embedding models

embeddings

llm/v1/audio/transcriptions

Speech-to-text models

audio.transcriptions

llm/v1/audio/translations

Speech-to-English translation models

audio.translations

Choose a GPU​

Databricks recommends developing on the same GPU that you serve on, so the settings you test are the settings you deploy. The following table lists the GPUs for custom LLM serving.

workload_type

GPU

Notes

GPU_SMALL

1x T4 (16 GB)

Small models.

GPU_MEDIUM

1x A10G (24 GB)

Models up to about 10B parameters. Generally available.

GPU_LARGE

1x L40S (48 GB)

Models up to about 15B parameters. Available in us-east-1, us-east-2, us-west-2, eu-central-1, ap-northeast-1, and ap-northeast-2.

GPU_XLARGE

1x H100 (80 GB)

Recommended for large LLMs. Available in us-west-2 only, with enrollment through your Databricks account team. Doesn't support scale-to-zero.

GPU_LARGE_RTX

1x RTX PRO 6000 (96 GB)

More memory than H100, but less optimized for LLM serving. Available in us-west-2, us-east-1, us-east-2, ap-northeast-1, and ap-northeast-2.

workload_type

GPU

Notes

GPU_SMALL

1x T4 (16 GB)

Small models.

GPU_MEDIUM

1x A10G (24 GB)

Models up to about 10B parameters. Generally available.

GPU_LARGE

1x L40S (48 GB)

Models up to about 15B parameters. Available in us-east-1, us-east-2, us-west-2, eu-central-1, ap-northeast-1, and ap-northeast-2.

GPU_XLARGE

1x H100 (80 GB)

Recommended for large LLMs. Available in us-west-2 only, with enrollment through your Databricks account team. Doesn't support scale-to-zero.

GPU_LARGE_RTX

1x RTX PRO 6000 (96 GB)

More memory than H100, but less optimized for LLM serving. Available in us-west-2, us-east-1, us-east-2, ap-northeast-1, and ap-northeast-2.

Common issues​

The following issues are the most common when you change a notebook:

  • The download fails in /Workspace, which doesn't accept multi-GB files. Download the weights to local disk, as the starter notebooks do.
  • The local server doesn't start. Serverless GPU notebooks allow only ports 3000 to 3999, so test on one of those. Only the entrypoint uses port 8080, and it must otherwise match the command you tested.
  • The endpoint can't find the weights. The entrypoint runs in the model's artifacts folder, so --model must name the weights folder you logged.
  • An embedding model doesn't serve embeddings. Start vLLM with --runner pooling.
  • Registration fails with TimeoutError('Timed out after 0:05:00') while uploading model_version.tar or model_environment.tar. Upgrade to databricks-sdk 0.102.0 or above and register the model again.

Create an endpoint​

Create the endpoint from the Serving UI or with the Databricks SDK, as the starter notebooks do. workload_type picks the GPU, and workload_size (Small, Medium, or Large) sets the number of replicas.

Python
ServedEntityInput(
entity_name="<catalog>.<schema>.<model>",
entity_version="<version>",
workload_type=ServingModelWorkloadType.GPU_MEDIUM,
workload_size="Small",
scale_to_zero_enabled=False,
)

Scale-to-zero and capacity​

Currently custom llm endpoints don't autoscale dynamically, so size workload_size for your peak traffic.

With scale-to-zero, an idle endpoint stops all replicas. The next request waits one to several minutes while vLLM loads the model again and all replicas start up. Databricks recommends turning off scale-to-zero for production traffic.

warning

Scale-up capacity is not guaranteed. Whenever Databricks needs to acquire a new GPU for your endpoint, such as on creation, on a workload_size increase, or when an endpoint wakes up from zero, the request can fail if the cloud provider has no GPU capacity in your region. Databricks mitigates this with warm pools and prereservation, which keep GPU capacity available and ready.

GPU_XLARGE (1xH100) endpoints do not support scale_to_zero_enabled=True currently.

Query your endpoint​

A ready chat endpoint appears in the AI Playground. The following examples query an endpoint with the Databricks SDK, the OpenAI client, and REST.

Python
from databricks.sdk import WorkspaceClient
from databricks.sdk.service.serving import ChatMessage, ChatMessageRole

w = WorkspaceClient()
w.serving_endpoints.query(
name="<endpoint-name>",
messages=[ChatMessage(role=ChatMessageRole.USER, content="Hello")],
)

Some embedding models expect a prefix on each input, such as search_query: or search_document:. Check the model card.

Monitor your endpoint​

The endpoint's Logs tab shows your server's stdout and stderr live. The logs API returns the same output.

Databricks forwards your vLLM server's metrics and charts them on the endpoint's Metrics tab:

  • Latency: time to first token, time per output token, request latency, and queue time.
  • Load: requests running and waiting.
  • KV cache: usage and hit rate.
  • Throughput: prompt and generation tokens per second.

The export metrics API returns the same metrics in Prometheus format, so you can scrape them into Prometheus or Datadog.

With telemetry turned on, Databricks also saves the server logs and the vLLM Prometheus metrics to Unity Catalog tables, so you can query them over longer periods. See Persist custom model serving data to Unity Catalog.

Pricing​

You pay per GPU instance hour, the same as for other GPU custom model serving. See Model Serving pricing.

Limitations and region availability​

You log a custom LLM from AI Runtime, so custom LLM serving is available in the same regions as serverless GPU compute. Those are US regions on AWS and Azure. GCP isn't supported.

The following features are coming soon:

  • LoRA adapters.
  • Cross-region serving.
  • Logging models outside AI Runtime, for example from Databricks Runtime or CPU serverless.

The following features aren't supported:

  • KV-cache-aware routing.
  • Route optimization.