Skip to main content

Stream

View as Markdown

Stream object​

A Stream is a governed UC entity representing an external streaming data source. The source_config oneof determines the streaming platform source (e.g. Kafka, Kinesis, etc.).

namestringBetaIDImmutable

Full three-part (catalog.schema.stream) name of the stream.

descriptionstringBeta

User-provided description.

source_configobjectBeta

Source-specific configuration. Determines the streaming platform source.

Show child attributesHide child attributes
kafka_stream_configobjectBeta

Configuration for Apache Kafka streams.

Show child attributesHide child attributes
subscription_modeobjectBeta

Options to configure which Kafka topics to pull data from.

Show child attributesHide child attributes
assignstringBeta

A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.

subscribestringBeta

A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.

subscribe_patternstringBeta

A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.

extra_optionsobjectBeta

Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:

  • maxOffsetsPerTrigger
  • startingOffsets
  • includeHeaders
  • kafka.request.timeout.ms
  • kafka.session.timeout.ms
  • kafka.max.partition.fetch.bytes The following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:
  • maxOffsetsPerTrigger
  • startingOffsets Auth and connection details belong on the parent Stream's connection_config, not here.
kinesis_stream_configobjectBeta

Configuration for AWS Kinesis Data Streams.

Show child attributesHide child attributes
stream_namesobjectBeta

Kinesis stream names to read from.

Show child attributesHide child attributes
namesarray of stringBeta

Kinesis stream names to read from.

stream_arnsobjectBeta

Kinesis stream ARNs to read from.

Show child attributesHide child attributes
arnsarray of stringBeta

Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.

extra_optionsobjectBeta

Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:

  • consumerMode
  • consumerNamePrefix
  • maxFetchRate
  • minFetchPeriod
  • maxFetchDuration
  • maxRecordsPerFetch
  • shardsPerTask
  • fetchBufferSize
  • shardFetchInterval consumerMode must be efo or polling (case-insensitive). maxRecordsPerFetch applies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream's connection_config, not here.
connection_configobjectBeta

Specifies how to connect and authenticate to the stream platform.

Show child attributesHide child attributes
uc_connection_namestringBeta

Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).

direct_mtls_configobjectBeta

Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.

Show child attributesHide child attributes
bootstrap_serversstringBeta

A comma-separated list of host:port pairs for the Kafka bootstrap servers.

mtls_configobjectBeta

Mutual-TLS authentication configuration.

Show child attributesHide child attributes
keystore_locationstringBeta

Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.

keystore_password_refobjectBeta

Secret-scope reference for the JKS keystore password.

Show child attributesHide child attributes
scopestringBeta

The <Databricks> secret scope name.

keystringBeta

The key within the scope.

key_password_refobjectBeta

Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).

Show child attributesHide child attributes
scopestringBeta

The <Databricks> secret scope name.

keystringBeta

The key within the scope.

truststore_locationstringBeta

Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".

truststore_password_refobjectBeta

Secret-scope reference for the JKS truststore password.

Show child attributesHide child attributes
scopestringBeta

The <Databricks> secret scope name.

keystringBeta

The key within the scope.

disable_hostname_verificationbooleanBeta

Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.

See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification

schema_configobjectBeta

Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.

Show child attributesHide child attributes
direct_schemasobjectBeta

Schema definitions provided directly on the Stream.

Show child attributesHide child attributes
payload_schemaobjectBeta

Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.

Show child attributesHide child attributes
json_schemastringBeta

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastringBeta

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobjectBeta

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringBeta

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringBeta

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

key_schemaobjectBeta

Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.

Show child attributesHide child attributes
json_schemastringBeta

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastringBeta

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobjectBeta

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringBeta

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringBeta

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

schema_registry_configobjectBeta

Resolve schemas from an external schema registry.

Show child attributesHide child attributes
uc_connectionstringBeta

A Schema Registry UC Connection object.

api_secret_refobjectBeta

Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.

Show child attributesHide child attributes
scopestringBeta

The <Databricks> secret scope name.

keystringBeta

The key within the scope.

payload_schema_locatorobjectBeta

Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobjectBeta

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringBeta

The Confluent schema registry subject name.

formatstringBeta

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
key_schema_locatorobjectBeta

Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobjectBeta

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringBeta

The Confluent schema registry subject name.

formatstringBeta

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
ingestion_configobjectBeta

Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.

Show child attributesHide child attributes
ingestion_destinationobjectBeta

Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.

Show child attributesHide child attributes
delta_table_namestringBeta

The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.

backfill_sourceobjectBeta

A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.

Show child attributesHide child attributes
delta_table_namestringBeta

The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.

deduplication_columnsarray of stringBeta

Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g. value.user_id). Empty list means every column is compared.

ingestion_pipeline_idstringBetaOutput only

The ID of the SDP pipeline that continuously copies new events from the streaming source into the ingestion Delta table.

ingestion_job_idint64BetaOutput only

The ID of the Databricks Job that performs the forward-fill ingestion.

backfill_job_idint64BetaOutput only

The ID of the Databricks Job that performs the historical backfill of the ingestion Delta table.

record_type_filterstringBeta

Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".

excluded_columnsarray of stringBeta

Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.

create_timestringBetaOutput only

Time at which this Stream was created.

created_bystringBetaOutput only

Username of the Stream creator.

update_timestringBetaOutput only

Time at which this Stream was last modified.

updated_bystringBetaOutput only

Username of user who last modified the Stream.

browse_onlybooleanBetaOutput only

Indicates whether the principal is limited to retrieving metadata for the associated object through the BROWSE privilege when include_browse is enabled in the request.

Get a Stream Beta​

GET /api/2.0/feature-engineering/streams/{name}

Get a Stream by its full three-part name (catalog.schema.stream).

API scopes: mlflow

Parameters​

namestringRequiredpath

Full three-part name (catalog.schema.stream) of the Stream to get.

Response​

Returns the Stream object.

List Streams Beta​

GET /api/2.0/feature-engineering/streams

List Streams under a given catalog.schema parent.

API scopes: mlflow

Parameters​

parentstringquery

Two-part name (catalog.schema) of the parent under which to list Streams.

page_sizeint32query

The maximum number of results to return.

page_tokenstringquery

Pagination token to go to the next page based on a previous query.

Response​

Returns a list of Stream objects.

Create a Stream Beta​

POST /api/2.0/feature-engineering/streams

Create a Stream, a governed UC entity representing an external streaming data source.

API scopes: mlflow

Request body​

The Stream to create.

namestringRequiredIDImmutable

Full three-part (catalog.schema.stream) name of the stream.

descriptionstring

User-provided description.

source_configobjectRequired

Source-specific configuration. Determines the streaming platform source.

Show child attributesHide child attributes
kafka_stream_configobject

Configuration for Apache Kafka streams.

Show child attributesHide child attributes
subscription_modeobjectRequired

Options to configure which Kafka topics to pull data from.

Show child attributesHide child attributes
assignstring

A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.

subscribestring

A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.

subscribe_patternstring

A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.

extra_optionsobject

Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:

  • maxOffsetsPerTrigger
  • startingOffsets
  • includeHeaders
  • kafka.request.timeout.ms
  • kafka.session.timeout.ms
  • kafka.max.partition.fetch.bytes The following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:
  • maxOffsetsPerTrigger
  • startingOffsets Auth and connection details belong on the parent Stream's connection_config, not here.
kinesis_stream_configobject

Configuration for AWS Kinesis Data Streams.

Show child attributesHide child attributes
stream_namesobject

Kinesis stream names to read from.

Show child attributesHide child attributes
namesarray of string

Kinesis stream names to read from.

stream_arnsobject

Kinesis stream ARNs to read from.

Show child attributesHide child attributes
arnsarray of string

Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.

extra_optionsobject

Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:

  • consumerMode
  • consumerNamePrefix
  • maxFetchRate
  • minFetchPeriod
  • maxFetchDuration
  • maxRecordsPerFetch
  • shardsPerTask
  • fetchBufferSize
  • shardFetchInterval consumerMode must be efo or polling (case-insensitive). maxRecordsPerFetch applies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream's connection_config, not here.
connection_configobjectRequired

Specifies how to connect and authenticate to the stream platform.

Show child attributesHide child attributes
uc_connection_namestring

Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).

direct_mtls_configobject

Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.

Show child attributesHide child attributes
bootstrap_serversstringRequired

A comma-separated list of host:port pairs for the Kafka bootstrap servers.

mtls_configobjectRequired

Mutual-TLS authentication configuration.

Show child attributesHide child attributes
keystore_locationstringRequired

Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.

keystore_password_refobjectRequired

Secret-scope reference for the JKS keystore password.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

key_password_refobjectRequired

Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

truststore_locationstringRequired

Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".

truststore_password_refobjectRequired

Secret-scope reference for the JKS truststore password.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

disable_hostname_verificationboolean

Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.

See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification

schema_configobjectRequired

Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.

Show child attributesHide child attributes
direct_schemasobject

Schema definitions provided directly on the Stream.

Show child attributesHide child attributes
payload_schemaobject

Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.

Show child attributesHide child attributes
json_schemastring

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastring

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobject

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringRequired

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringRequired

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

key_schemaobject

Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.

Show child attributesHide child attributes
json_schemastring

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastring

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobject

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringRequired

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringRequired

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

schema_registry_configobject

Resolve schemas from an external schema registry.

Show child attributesHide child attributes
uc_connectionstring

A Schema Registry UC Connection object.

api_secret_refobject

Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

payload_schema_locatorobject

Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobject

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringRequired

The Confluent schema registry subject name.

formatstringRequired

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
key_schema_locatorobject

Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobject

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringRequired

The Confluent schema registry subject name.

formatstringRequired

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
ingestion_configobjectRequired

Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.

Show child attributesHide child attributes
ingestion_destinationobjectRequired

Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.

Show child attributesHide child attributes
delta_table_namestring

The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.

backfill_sourceobject

A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.

Show child attributesHide child attributes
delta_table_namestring

The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.

deduplication_columnsarray of string

Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g. value.user_id). Empty list means every column is compared.

record_type_filterstring

Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".

excluded_columnsarray of string

Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.

Response​

Returns the Stream object.

Update a Stream Beta​

PATCH /api/2.0/feature-engineering/streams/{name}

Update a Stream. Only fields listed in update_mask are mutated.

API scopes: mlflow

Parameters​

namestringRequiredIDImmutablepath

Full three-part (catalog.schema.stream) name of the stream.

update_maskstringRequiredquery

The list of fields to update.

Request body​

The Stream to update.

descriptionstring

User-provided description.

source_configobjectRequired

Source-specific configuration. Determines the streaming platform source.

Show child attributesHide child attributes
kafka_stream_configobject

Configuration for Apache Kafka streams.

Show child attributesHide child attributes
subscription_modeobjectRequired

Options to configure which Kafka topics to pull data from.

Show child attributesHide child attributes
assignstring

A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.

subscribestring

A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.

subscribe_patternstring

A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.

extra_optionsobject

Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:

  • maxOffsetsPerTrigger
  • startingOffsets
  • includeHeaders
  • kafka.request.timeout.ms
  • kafka.session.timeout.ms
  • kafka.max.partition.fetch.bytes The following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:
  • maxOffsetsPerTrigger
  • startingOffsets Auth and connection details belong on the parent Stream's connection_config, not here.
kinesis_stream_configobject

Configuration for AWS Kinesis Data Streams.

Show child attributesHide child attributes
stream_namesobject

Kinesis stream names to read from.

Show child attributesHide child attributes
namesarray of string

Kinesis stream names to read from.

stream_arnsobject

Kinesis stream ARNs to read from.

Show child attributesHide child attributes
arnsarray of string

Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.

extra_optionsobject

Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:

  • consumerMode
  • consumerNamePrefix
  • maxFetchRate
  • minFetchPeriod
  • maxFetchDuration
  • maxRecordsPerFetch
  • shardsPerTask
  • fetchBufferSize
  • shardFetchInterval consumerMode must be efo or polling (case-insensitive). maxRecordsPerFetch applies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream's connection_config, not here.
connection_configobjectRequired

Specifies how to connect and authenticate to the stream platform.

Show child attributesHide child attributes
uc_connection_namestring

Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).

direct_mtls_configobject

Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.

Show child attributesHide child attributes
bootstrap_serversstringRequired

A comma-separated list of host:port pairs for the Kafka bootstrap servers.

mtls_configobjectRequired

Mutual-TLS authentication configuration.

Show child attributesHide child attributes
keystore_locationstringRequired

Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.

keystore_password_refobjectRequired

Secret-scope reference for the JKS keystore password.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

key_password_refobjectRequired

Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

truststore_locationstringRequired

Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".

truststore_password_refobjectRequired

Secret-scope reference for the JKS truststore password.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

disable_hostname_verificationboolean

Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.

See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification

schema_configobjectRequired

Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.

Show child attributesHide child attributes
direct_schemasobject

Schema definitions provided directly on the Stream.

Show child attributesHide child attributes
payload_schemaobject

Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.

Show child attributesHide child attributes
json_schemastring

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastring

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobject

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringRequired

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringRequired

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

key_schemaobject

Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.

Show child attributesHide child attributes
json_schemastring

Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).

avro_schemastring

Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).

proto_schemaobject

Protocol Buffer schema with its payload message name.

Show child attributesHide child attributes
schema_textstringRequired

The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).

message_namestringRequired

The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.

schema_registry_configobject

Resolve schemas from an external schema registry.

Show child attributesHide child attributes
uc_connectionstring

A Schema Registry UC Connection object.

api_secret_refobject

Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.

Show child attributesHide child attributes
scopestringRequired

The <Databricks> secret scope name.

keystringRequired

The key within the scope.

payload_schema_locatorobject

Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobject

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringRequired

The Confluent schema registry subject name.

formatstringRequired

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
key_schema_locatorobject

Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.

Show child attributesHide child attributes
confluent_schemaobject

Confluent Schema Registry schema locator.

Show child attributesHide child attributes
subjectstringRequired

The Confluent schema registry subject name.

formatstringRequired

Serialization format for this schema.

Values:

  • FORMAT_UNSPECIFIED
  • FORMAT_AVRO
  • FORMAT_PROTOBUF
  • FORMAT_JSON
ingestion_configobjectRequired

Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.

Show child attributesHide child attributes
ingestion_destinationobjectRequired

Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.

Show child attributesHide child attributes
delta_table_namestring

The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.

backfill_sourceobject

A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.

Show child attributesHide child attributes
delta_table_namestring

The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.

deduplication_columnsarray of string

Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g. value.user_id). Empty list means every column is compared.

record_type_filterstring

Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".

excluded_columnsarray of string

Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.

Response​

Returns the Stream object.

Delete a Stream Beta​

DELETE /api/2.0/feature-engineering/streams/{name}

Delete a Stream by its full three-part name (catalog.schema.stream).

API scopes: mlflow

Parameters​

namestringRequiredpath

Full three-part name (catalog.schema.stream) of the Stream to delete.