Serving concepts
This page defines key concepts used across online serving on Databricks, including model serving endpoints, feature serving endpoints, and Online Feature Stores.
Model serving endpoint
REST API that exposes one or more served models for inference. Model serving endpoints optionally support automatic feature lookup from an Online Feature Store.
Feature serving endpoint
REST API for looking up feature values grouped as a FeatureSpec, which is useful for passing feature data to a model that runs outside of Databricks. For more information, see Serve Feature Views.
Route-optimized endpoint
Endpoint property that enables an improved network path with faster, more direct communication between the client and the endpoint during inference, which lowers overhead latency and increases throughput. Available for both custom model serving endpoints and feature serving endpoints. For more information, see Route optimization on serving endpoints.
Provisioned concurrency
Endpoint property that specifies the maximum number of parallel requests an endpoint can handle. Estimate the required concurrency using the formula: provisioned concurrency = queries per second (QPS) × model execution time (s).
Scale to zero
Endpoint property that automatically reduces resource consumption to zero when endpoint is not in-use. Scale to zero is recommended for testing and development. However, scale to zero is not recommended for production endpoints, as latency is greater and capacity is not guaranteed when scaled to zero.
Served entity
Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.
Online Feature Store
Data store that serves feature values to real-time applications and serving endpoints with low latency. Databricks Online Feature Stores are powered by Lakebase Autoscaling and sync feature data from Feature Views or feature tables. For more information, see Databricks Online Feature Stores.
Traffic configuration
Specification for what percentage of traffic to an endpoint should go to each model. Traffic configuration is required for endpoints with more than one served model.
The following is an example where the endpoint named multi-pt-model hosts version 2 of meta_llama_v3_1_70b_instruct which gets 60% of the endpoint traffic, and also hosts version 3 of meta_llama_v3_1_8b_instruct which gets 40% of the endpoint traffic. For more information, see Serve multiple models to a model serving endpoint.
POST /api/2.0/serving-endpoints
{
"name":"multi-pt-model"
"config":
{
"served_entities":
[
{
"name":"meta_llama_v3_1_70b_instruct",
"entity_name":"system.ai.meta_llama_v3_1_70b_instruct",
"entity_version":"4",
"min_provisioned_throughput":0,
"max_provisioned_throughput":2400
},
{
"name":"meta_llama_v3_1_8b_instruct",
"entity_name":"system.ai.meta_llama_v3_1_8b_instruct",
"entity_version":"4",
"min_provisioned_throughput":0,
"max_provisioned_throughput":1240
}
],
"traffic_config":
{
"routes":
[
{
"served_model_name":"meta_llama_v3_1_8b_instruct",
"traffic_percentage":"60"
},
{
"served_model_name":"meta_llama_v3_1_70b_instruct",
"traffic_percentage":"40"
}
]
}
}
}