Skip to main content

AI models maintenance policy

This page describes the model maintenance policy for the Foundation Model APIs pay-per-token, Foundation Model APIs provisioned throughput, and Batch Inference with ai_query offerings.

To continue supporting the most state-of-the-art models, Databricks manages models through a lifecycle that progresses from legacy status to deprecation to eventual retirement.

  • Legacy: When a model is marked legacy, it is no longer recommended for new workloads, but remains available in workspaces with existing usage of the model. Workspaces that are not using the model when it enters legacy status no longer have access to the model. Customers using such a model should evaluate newer models and consider migration.
  • Deprecation: When a model is deprecated, a retirement date of 30 days or 90 days in the future is announced. It is not recommended for new workloads, but remains available in workspaces with existing usage of the model. Workspaces that are not using the model at the time of deprecation do not have access to the model. Customers using such a model should evaluate newer models and plan to migrate their workloads before the retirement date.
  • Retirement: When a model is retired, it is no longer accessible and support for the model is fully discontinued. Any workload using the model no longer works.

Model lifecycle policy​

The following section summarizes the model lifecycle policy. See Deprecated models for a list of deprecated models and retirement dates.

Many factors are considered when determining the retirement timeline, including model popularity, suitable replacement model, and amount of production usage. Databricks announces retirement dates for models when they are deprecated, following the notification timelines outlined below.

Foundation Model APIs​

The following table summarizes the model lifecycle policy for Foundation Model APIs products: pay-per-token, provisioned throughput, and ai_query (batch inference).

Legacy

Deprecation notification and transition to retirement

On the retirement date

Databricks takes the following steps to inform customers that a model is marked legacy:

  • In the Databricks UI, the model is marked as legacy.
  • The applicable documentation indicates that the model is legacy.

When a model is marked legacy:

  • The model continues to be available only for workspaces with existing workloads that use this model.
  • The legacy model is not accessible from workspaces that were not actively using it before being marked legacy.
  • Customers with existing workloads should evaluate newer models and consider migrating their workloads.

Databricks takes the following steps to notify customers about a model deprecation:

  • In the Databricks UI, a warning message indicates that the model is deprecated.
  • The applicable documentation contains a notice that indicates the model is deprecated, along with a retirement date.

When a model is deprecated, Databricks announces a retirement date 30 days or 90 days in the future. During this transition period:

  • Customers with existing workloads should migrate to the recommended replacement model or sunset affected workloads.

The model is no longer available for use and is removed from the product. Any existing workloads using the model stop working. Applicable documentation is updated to indicate that the model is no longer available, and recommends use of a replacement model.

Legacy

Deprecation notification and transition to retirement

On the retirement date

Databricks takes the following steps to inform customers that a model is marked legacy:

  • In the Databricks UI, the model is marked as legacy.
  • The applicable documentation indicates that the model is legacy.

When a model is marked legacy:

  • The model continues to be available only for workspaces with existing workloads that use this model.
  • The legacy model is not accessible from workspaces that were not actively using it before being marked legacy.
  • Customers with existing workloads should evaluate newer models and consider migrating their workloads.

Databricks takes the following steps to notify customers about a model deprecation:

  • In the Databricks UI, a warning message indicates that the model is deprecated.
  • The applicable documentation contains a notice that indicates the model is deprecated, along with a retirement date.

When a model is deprecated, Databricks announces a retirement date 30 days or 90 days in the future. During this transition period:

  • Customers with existing workloads should migrate to the recommended replacement model or sunset affected workloads.

The model is no longer available for use and is removed from the product. Any existing workloads using the model stop working. Applicable documentation is updated to indicate that the model is no longer available, and recommends use of a replacement model.

Partner model retirement policy​

Partner models are models that third-party partners — specifically OpenAI, Anthropic, and Google — provide through Foundation Model APIs. For these partner models, Databricks generally follows the same deprecation timelines and policies described above.

However, partners might provide retirement dates shorter than the one-month transition period that Databricks publishes. In these cases, Databricks attempts to bridge the gap by temporarily redirecting models to a similar version, so customers receive the full transition time.

For example, if a partner model retirement is announced with two weeks' lead time instead of one month, Databricks redirects the model for an additional two weeks to prevent immediate breakage and allow time for migration. Queries fail at the end of the full one-month period.

note

This redirection can only occur if the replacement model has the same price and is backwards compatible. The replacement model is usually an incremental model version, like 3.0 versus 3.1.

Deprecated and retired models​

The following sections list models that are deprecated (no longer recommended for new workloads) or retired (end-of-life and no longer available). Retirement dates for deprecated models are announced at least one month in advance.

Foundation Model APIs retirements​

The following table shows model retirements, their retirement dates, and recommended replacement models to use for Foundation Model APIs pay-per-token and provisioned throughput serving workloads. Databricks recommends that you migrate your applications to use replacement models before the indicated retirement date.

Partner model

Retirement date

Recommended replacement model

Anthropic Claude Sonnet 4

Pay-per-token: October 9, 2026

Claude Sonnet 4.6

OpenAI GPT-5.1 Codex Max

Pay-per-token: July 16, 2026

OpenAI GPT-5.5

OpenAI GPT-5.1 Codex Mini

Pay-per-token: July 16, 2026

OpenAI GPT-5.4 Codex Mini

OpenAI GPT-5.2 Codex

Pay-per-token: July 16, 2026

OpenAI GPT-5.5

Anthropic Claude 3.7 Sonnet

Pay-per-token: April 12, 2026

Use the latest Claude Sonnet model

Gemini 2.5 Flash

Pay-per-token: October 2, 2026

Provisioned throughput: October 2, 2026

Gemini 3.1 Pro or Gemini 3.5 Flash

Gemini 2.5 Pro

Provisioned throughput: October 2, 2026

Gemini 3.1 Pro or Gemini 3.5 Flash

Gemini 3 Pro

Provisioned throughput: March 26, 2026

Gemini 3.1 Pro. To allow more time for migration, between March 26, 2026 and June 7, 2026, API calls to Gemini 3 Pro will be temporarily redirected to Gemini 3.1 Pro. The pricing for both models is identical.

Partner model

Retirement date

Recommended replacement model

Anthropic Claude Sonnet 4

Pay-per-token: October 9, 2026

Claude Sonnet 4.6

OpenAI GPT-5.1 Codex Max

Pay-per-token: July 16, 2026

OpenAI GPT-5.5

OpenAI GPT-5.1 Codex Mini

Pay-per-token: July 16, 2026

OpenAI GPT-5.4 Codex Mini

OpenAI GPT-5.2 Codex

Pay-per-token: July 16, 2026

OpenAI GPT-5.5

Anthropic Claude 3.7 Sonnet

Pay-per-token: April 12, 2026

Use the latest Claude Sonnet model

Gemini 2.5 Flash

Pay-per-token: October 2, 2026

Provisioned throughput: October 2, 2026

Gemini 3.1 Pro or Gemini 3.5 Flash

Gemini 2.5 Pro

Provisioned throughput: October 2, 2026

Gemini 3.1 Pro or Gemini 3.5 Flash

Gemini 3 Pro

Provisioned throughput: March 26, 2026

Gemini 3.1 Pro. To allow more time for migration, between March 26, 2026 and June 7, 2026, API calls to Gemini 3 Pro will be temporarily redirected to Gemini 3.1 Pro. The pricing for both models is identical.

Open model

Retirement date

Recommended replacement model

DeepSeek V4 Flash (0731)

Pay-per-token: November 5, 2026

DeepSeek V4.1 Flash

Thinking Machine Labs Inkling

Pay-per-token: October 30, 2026

GLM 5.3 or Kimi K3

DeepSeek V4 Pro (0813)

Pay-per-token: October 30, 2026

DeepSeek V4.1 Flash

Kimi K2.7

Pay-per-token: October 30, 2026

Kimi K3

Meta Llama 3.1 405B

Pay-per-token: February 15, 2026

Provisioned throughput: May 15, 2026

OpenAI GPT OSS 120B

DBRX / DBRX Instruct

Pay-per-token: April 30, 2025

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Mixtral 8x7B / Mixtral-8x7B Instruct

Pay-per-token: April 30, 2025

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 3 (70B)

Pay-per-token: July 23, 2024 (Meta-Llama-3-70B-Instruct); December 11, 2024 (Meta-Llama-3.1-70B-Instruct)

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 3 8B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 70B / Meta-Llama-2-70B-Chat

Pay-per-token: October 30, 2024

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 13B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 7B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Mistral 7B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

MPT 30B / MPT 30B Instruct

Pay-per-token: August 30, 2024

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

MPT 7B / MPT 7B Instruct

Pay-per-token: August 30, 2024

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Open model

Retirement date

Recommended replacement model

DeepSeek V4 Flash (0731)

Pay-per-token: November 5, 2026

DeepSeek V4.1 Flash

Thinking Machine Labs Inkling

Pay-per-token: October 30, 2026

GLM 5.3 or Kimi K3

DeepSeek V4 Pro (0813)

Pay-per-token: October 30, 2026

DeepSeek V4.1 Flash

Kimi K2.7

Pay-per-token: October 30, 2026

Kimi K3

Meta Llama 3.1 405B

Pay-per-token: February 15, 2026

Provisioned throughput: May 15, 2026

OpenAI GPT OSS 120B

DBRX / DBRX Instruct

Pay-per-token: April 30, 2025

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Mixtral 8x7B / Mixtral-8x7B Instruct

Pay-per-token: April 30, 2025

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 3 (70B)

Pay-per-token: July 23, 2024 (Meta-Llama-3-70B-Instruct); December 11, 2024 (Meta-Llama-3.1-70B-Instruct)

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 3 8B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 70B / Meta-Llama-2-70B-Chat

Pay-per-token: October 30, 2024

Provisioned throughput: February 27, 2026

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 13B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Meta Llama 2 7B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

Mistral 7B

Provisioned throughput: February 27, 2026

Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

MPT 30B / MPT 30B Instruct

Pay-per-token: August 30, 2024

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

MPT 7B / MPT 7B Instruct

Pay-per-token: August 30, 2024

Provisioned throughput: December 19, 2025

Pay-per-token: Meta-Llama-4-Maverick

Provisioned throughput: Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size.

If you require long-term support for a specific model version, Databricks recommends using Foundation Model APIs provisioned throughput for your serving workloads.

Find workloads that use retired models​

Use the following query to find workloads that are using deprecated models and identify their owners.

SQL
SELECT
eu.requester,
se.endpoint_name,
se.entity_name,
COUNT(*) AS request_count,
SUM(eu.input_token_count) AS total_input_tokens,
SUM(eu.output_token_count) AS total_output_tokens,
MIN(eu.request_time) AS first_request,
MAX(eu.request_time) AS last_request
FROM system.serving.endpoint_usage eu
JOIN system.serving.served_entities se
ON eu.served_entity_id = se.served_entity_id
WHERE LOWER(se.entity_name) LIKE '%<retired-model-name>%'
GROUP BY eu.requester, se.endpoint_name, se.entity_name
ORDER BY request_count DESC