Stream
A Stream is a governed UC entity representing an external streaming data source. The source_config oneof determines the streaming platform source (e.g. Kafka, Kinesis, etc.).
Stream object
A Stream is a governed UC entity representing an external streaming data source. The source_config oneof determines the streaming platform source (e.g. Kafka, Kinesis, etc.).
- namestringBetaIDImmutable
Full three-part (catalog.schema.stream) name of the stream.
- descriptionstringBeta
User-provided description.
- source_configobjectBeta
Source-specific configuration. Determines the streaming platform source.
Show child attributesHide child attributes
- kafka_stream_configobjectBeta
Configuration for Apache Kafka streams.
Show child attributesHide child attributes
- subscription_modeobjectBeta
Options to configure which Kafka topics to pull data from.
Show child attributesHide child attributes
- assignstringBeta
A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.
- subscribestringBeta
A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.
- subscribe_patternstringBeta
A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.
- extra_optionsobjectBeta
Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:
maxOffsetsPerTriggerstartingOffsetsincludeHeaderskafka.request.timeout.mskafka.session.timeout.mskafka.max.partition.fetch.bytesThe following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:maxOffsetsPerTriggerstartingOffsetsAuth and connection details belong on the parent Stream'sconnection_config, not here.
- kinesis_stream_configobjectBeta
Configuration for AWS Kinesis Data Streams.
Show child attributesHide child attributes
- stream_namesobjectBeta
Kinesis stream names to read from.
Show child attributesHide child attributes
- namesarray of stringBeta
Kinesis stream names to read from.
- stream_arnsobjectBeta
Kinesis stream ARNs to read from.
Show child attributesHide child attributes
- arnsarray of stringBeta
Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.
- extra_optionsobjectBeta
Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:
consumerModeconsumerNamePrefixmaxFetchRateminFetchPeriodmaxFetchDurationmaxRecordsPerFetchshardsPerTaskfetchBufferSizeshardFetchIntervalconsumerModemust beefoorpolling(case-insensitive).maxRecordsPerFetchapplies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream'sconnection_config, not here.
- connection_configobjectBeta
Specifies how to connect and authenticate to the stream platform.
Show child attributesHide child attributes
- uc_connection_namestringBeta
Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).
- direct_mtls_configobjectBeta
Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.
Show child attributesHide child attributes
- bootstrap_serversstringBeta
A comma-separated list of host:port pairs for the Kafka bootstrap servers.
- mtls_configobjectBeta
Mutual-TLS authentication configuration.
Show child attributesHide child attributes
- keystore_locationstringBeta
Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.
- keystore_password_refobjectBeta
Secret-scope reference for the JKS keystore password.
Show child attributesHide child attributes
- scopestringBeta
The <Databricks> secret scope name.
- keystringBeta
The key within the scope.
- key_password_refobjectBeta
Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).
Show child attributesHide child attributes
- scopestringBeta
The <Databricks> secret scope name.
- keystringBeta
The key within the scope.
- truststore_locationstringBeta
Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".
- truststore_password_refobjectBeta
Secret-scope reference for the JKS truststore password.
Show child attributesHide child attributes
- scopestringBeta
The <Databricks> secret scope name.
- keystringBeta
The key within the scope.
- disable_hostname_verificationbooleanBeta
Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.
See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification
- schema_configobjectBeta
Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.
Show child attributesHide child attributes
- direct_schemasobjectBeta
Schema definitions provided directly on the Stream.
Show child attributesHide child attributes
- payload_schemaobjectBeta
Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.
Show child attributesHide child attributes
- json_schemastringBeta
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastringBeta
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobjectBeta
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringBeta
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringBeta
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- key_schemaobjectBeta
Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.
Show child attributesHide child attributes
- json_schemastringBeta
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastringBeta
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobjectBeta
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringBeta
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringBeta
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- schema_registry_configobjectBeta
Resolve schemas from an external schema registry.
Show child attributesHide child attributes
- uc_connectionstringBeta
A Schema Registry UC Connection object.
- api_secret_refobjectBeta
Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.
Show child attributesHide child attributes
- scopestringBeta
The <Databricks> secret scope name.
- keystringBeta
The key within the scope.
- payload_schema_locatorobjectBeta
Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobjectBeta
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringBeta
The Confluent schema registry subject name.
- formatstringBeta
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- key_schema_locatorobjectBeta
Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobjectBeta
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringBeta
The Confluent schema registry subject name.
- formatstringBeta
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- ingestion_configobjectBeta
Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.
Show child attributesHide child attributes
- ingestion_destinationobjectBeta
Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.
Show child attributesHide child attributes
- delta_table_namestringBeta
The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.
- backfill_sourceobjectBeta
A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.
Show child attributesHide child attributes
- delta_table_namestringBeta
The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.
- deduplication_columnsarray of stringBeta
Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g.
value.user_id). Empty list means every column is compared.
- ingestion_pipeline_idstringBetaOutput only
The ID of the SDP pipeline that continuously copies new events from the streaming source into the ingestion Delta table.
- ingestion_job_idint64BetaOutput only
The ID of the Databricks Job that performs the forward-fill ingestion.
- backfill_job_idint64BetaOutput only
The ID of the Databricks Job that performs the historical backfill of the ingestion Delta table.
- record_type_filterstringBeta
Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".
- excluded_columnsarray of stringBeta
Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.
- create_timestringBetaOutput only
Time at which this Stream was created.
- created_bystringBetaOutput only
Username of the Stream creator.
- update_timestringBetaOutput only
Time at which this Stream was last modified.
- updated_bystringBetaOutput only
Username of user who last modified the Stream.
- browse_onlybooleanBetaOutput only
Indicates whether the principal is limited to retrieving metadata for the associated object through the BROWSE privilege when include_browse is enabled in the request.
Get a Stream Beta
GET
Get a Stream by its full three-part name (catalog.schema.stream).
API scopes: mlflow
Parameters
- namestringRequiredpath
Full three-part name (catalog.schema.stream) of the Stream to get.
Response
Returns the Stream object.
List Streams Beta
GET
List Streams under a given catalog.schema parent.
API scopes: mlflow
Parameters
- parentstringquery
Two-part name (catalog.schema) of the parent under which to list Streams.
- page_sizeint32query
The maximum number of results to return.
- page_tokenstringquery
Pagination token to go to the next page based on a previous query.
Response
Returns a list of Stream objects.
Create a Stream Beta
POST
Create a Stream, a governed UC entity representing an external streaming data source.
API scopes: mlflow
Request body
The Stream to create.
- namestringRequiredIDImmutable
Full three-part (catalog.schema.stream) name of the stream.
- descriptionstring
User-provided description.
- source_configobjectRequired
Source-specific configuration. Determines the streaming platform source.
Show child attributesHide child attributes
- kafka_stream_configobject
Configuration for Apache Kafka streams.
Show child attributesHide child attributes
- subscription_modeobjectRequired
Options to configure which Kafka topics to pull data from.
Show child attributesHide child attributes
- assignstring
A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.
- subscribestring
A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.
- subscribe_patternstring
A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.
- extra_optionsobject
Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:
maxOffsetsPerTriggerstartingOffsetsincludeHeaderskafka.request.timeout.mskafka.session.timeout.mskafka.max.partition.fetch.bytesThe following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:maxOffsetsPerTriggerstartingOffsetsAuth and connection details belong on the parent Stream'sconnection_config, not here.
- kinesis_stream_configobject
Configuration for AWS Kinesis Data Streams.
Show child attributesHide child attributes
- stream_namesobject
Kinesis stream names to read from.
Show child attributesHide child attributes
- namesarray of string
Kinesis stream names to read from.
- stream_arnsobject
Kinesis stream ARNs to read from.
Show child attributesHide child attributes
- arnsarray of string
Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.
- extra_optionsobject
Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:
consumerModeconsumerNamePrefixmaxFetchRateminFetchPeriodmaxFetchDurationmaxRecordsPerFetchshardsPerTaskfetchBufferSizeshardFetchIntervalconsumerModemust beefoorpolling(case-insensitive).maxRecordsPerFetchapplies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream'sconnection_config, not here.
- connection_configobjectRequired
Specifies how to connect and authenticate to the stream platform.
Show child attributesHide child attributes
- uc_connection_namestring
Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).
- direct_mtls_configobject
Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.
Show child attributesHide child attributes
- bootstrap_serversstringRequired
A comma-separated list of host:port pairs for the Kafka bootstrap servers.
- mtls_configobjectRequired
Mutual-TLS authentication configuration.
Show child attributesHide child attributes
- keystore_locationstringRequired
Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.
- keystore_password_refobjectRequired
Secret-scope reference for the JKS keystore password.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- key_password_refobjectRequired
Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- truststore_locationstringRequired
Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".
- truststore_password_refobjectRequired
Secret-scope reference for the JKS truststore password.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- disable_hostname_verificationboolean
Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.
See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification
- schema_configobjectRequired
Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.
Show child attributesHide child attributes
- direct_schemasobject
Schema definitions provided directly on the Stream.
Show child attributesHide child attributes
- payload_schemaobject
Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.
Show child attributesHide child attributes
- json_schemastring
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastring
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobject
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringRequired
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringRequired
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- key_schemaobject
Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.
Show child attributesHide child attributes
- json_schemastring
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastring
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobject
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringRequired
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringRequired
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- schema_registry_configobject
Resolve schemas from an external schema registry.
Show child attributesHide child attributes
- uc_connectionstring
A Schema Registry UC Connection object.
- api_secret_refobject
Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- payload_schema_locatorobject
Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobject
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringRequired
The Confluent schema registry subject name.
- formatstringRequired
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- key_schema_locatorobject
Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobject
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringRequired
The Confluent schema registry subject name.
- formatstringRequired
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- ingestion_configobjectRequired
Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.
Show child attributesHide child attributes
- ingestion_destinationobjectRequired
Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.
Show child attributesHide child attributes
- delta_table_namestring
The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.
- backfill_sourceobject
A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.
Show child attributesHide child attributes
- delta_table_namestring
The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.
- deduplication_columnsarray of string
Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g.
value.user_id). Empty list means every column is compared.
- record_type_filterstring
Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".
- excluded_columnsarray of string
Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.
Response
Returns the Stream object.
Update a Stream Beta
PATCH
Update a Stream. Only fields listed in update_mask are mutated.
API scopes: mlflow
Parameters
- namestringRequiredIDImmutablepath
Full three-part (catalog.schema.stream) name of the stream.
- update_maskstringRequiredquery
The list of fields to update.
Request body
The Stream to update.
- descriptionstring
User-provided description.
- source_configobjectRequired
Source-specific configuration. Determines the streaming platform source.
Show child attributesHide child attributes
- kafka_stream_configobject
Configuration for Apache Kafka streams.
Show child attributesHide child attributes
- subscription_modeobjectRequired
Options to configure which Kafka topics to pull data from.
Show child attributesHide child attributes
- assignstring
A JSON string that contains the specific topic-partitions to consume from. For example, for '{"topicA":[0,1],"topicB":[2,4]}', topicA's 0'th and 1st partitions will be consumed from.
- subscribestring
A comma-separated list of Kafka topics to read from. For example, 'topicA,topicB,topicC'.
- subscribe_patternstring
A regular expression matching topics to subscribe to. For example, 'topic.*' will subscribe to all topics starting with 'topic'.
- extra_optionsobject
Optional Kafka source or consumer options, validated against a server-side allowlist at request time. Allowed keys:
maxOffsetsPerTriggerstartingOffsetsincludeHeaderskafka.request.timeout.mskafka.session.timeout.mskafka.max.partition.fetch.bytesThe following keys are ingestion-only and are stripped before being forwarded to the materialization pipeline:maxOffsetsPerTriggerstartingOffsetsAuth and connection details belong on the parent Stream'sconnection_config, not here.
- kinesis_stream_configobject
Configuration for AWS Kinesis Data Streams.
Show child attributesHide child attributes
- stream_namesobject
Kinesis stream names to read from.
Show child attributesHide child attributes
- namesarray of string
Kinesis stream names to read from.
- stream_arnsobject
Kinesis stream ARNs to read from.
Show child attributesHide child attributes
- arnsarray of string
Kinesis stream ARNs to read from. For example, 'arn:aws:kinesis:us-west-2:111122223333:stream/stream-a'.
- extra_optionsobject
Optional Kinesis source options, validated against a server-side allowlist at request time. Allowed keys:
consumerModeconsumerNamePrefixmaxFetchRateminFetchPeriodmaxFetchDurationmaxRecordsPerFetchshardsPerTaskfetchBufferSizeshardFetchIntervalconsumerModemust beefoorpolling(case-insensitive).maxRecordsPerFetchapplies only during ingestion and does not affect the materialization pipeline. Auth and connection details belong on the parent Stream'sconnection_config, not here.
- connection_configobjectRequired
Specifies how to connect and authenticate to the stream platform.
Show child attributesHide child attributes
- uc_connection_namestring
Name of an existing UC Connection for stream platform access. Must be the correct type for the streaming platform (e.g. a Kafka Connection for a Kafka Stream, or a Kinesis Connection for a Kinesis Stream).
- direct_mtls_configobject
Direct mTLS configuration for stream platform access. This is only used in the short term until UC Kafka Connections support mTLS . Once UC Kafka Connections support mTLS, this will be deprecated.
Show child attributesHide child attributes
- bootstrap_serversstringRequired
A comma-separated list of host:port pairs for the Kafka bootstrap servers.
- mtls_configobjectRequired
Mutual-TLS authentication configuration.
Show child attributesHide child attributes
- keystore_locationstringRequired
Unity Catalog volume path to the JKS keystore file containing the client certificate and private key. e.g. "/Volumes/<catalog>/<schema>/<volume>/client.jks". The materialization compute must have read permission on this volume.
- keystore_password_refobjectRequired
Secret-scope reference for the JKS keystore password.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- key_password_refobjectRequired
Secret-scope reference for the private key password. Often the same value as the keystore password (keytool's default), but provided as a separate field because Apache Kafka requires it as a distinct option (kafka.ssl.key.password).
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- truststore_locationstringRequired
Unity Catalog volume path to the JKS truststore file containing the CA certificate(s) trusted to verify the Kafka broker's server certificate. e.g. "/Volumes/<catalog>/<schema>/<volume>/truststore.jks".
- truststore_password_refobjectRequired
Secret-scope reference for the JKS truststore password.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- disable_hostname_verificationboolean
Set to true only when the broker certificate's SAN intentionally does not match the connection endpoint — for example when reaching the cluster through a PrivateLink endpoint whose DNS name is not in the broker certificate. Skipping the hostname check removes a defense against man-in-the-middle attacks; do not enable casually. mTLS client authentication is unaffected by this option.
See the Apache Kafka SSL security guide for background on this check: https://kafka.apache.org/42/security/encryption-and-authentication-using-ssl/#host-name-verification
- schema_configobjectRequired
Schema definitions for the stream, provided either directly on the Stream or resolved from an external schema registry through a UC Connection.
Show child attributesHide child attributes
- direct_schemasobject
Schema definitions provided directly on the Stream.
Show child attributesHide child attributes
- payload_schemaobject
Schema for the message payload. For Kafka, this is the value schema. Unless the platform supports another schema (e.g. keys for Kafka), this must be specified.
Show child attributesHide child attributes
- json_schemastring
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastring
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobject
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringRequired
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringRequired
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- key_schemaobject
Schema for the message key. This is only used for Kafka streams. For Kafka, at least one of payload_schema or key_schema must be specified.
Show child attributesHide child attributes
- json_schemastring
Schema of the JSON object in standard IETF JSON schema format (https://json-schema.org/).
- avro_schemastring
Avro schema in JSON format (https://avro.apache.org/docs/current/specification/).
- proto_schemaobject
Protocol Buffer schema with its payload message name.
Show child attributesHide child attributes
- schema_textstringRequired
The raw .proto file text (proto2 and proto3 syntax supported, see https://protobuf.dev/programming-guides/proto3/ and https://protobuf.dev/programming-guides/proto2/).
- message_namestringRequired
The fully-qualified name of the message within schema_text that describes the Kafka payload (e.g. "Event" or "com.example.Event" if schema_text declares a package). Identifies which message is used to decode each Kafka record — a .proto file may declare multiple messages but only one represents the payload. Must not be empty.
- schema_registry_configobject
Resolve schemas from an external schema registry.
Show child attributesHide child attributes
- uc_connectionstring
A Schema Registry UC Connection object.
- api_secret_refobject
Reference to the schema registry API secret in a <Databricks> secret scope. Set this only if required for authentication for the schema registry.
Show child attributesHide child attributes
- scopestringRequired
The <Databricks> secret scope name.
- keystringRequired
The key within the scope.
- payload_schema_locatorobject
Schema locator for the message payload. For Kafka this is the value. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobject
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringRequired
The Confluent schema registry subject name.
- formatstringRequired
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- key_schema_locatorobject
Schema locator for the message key. Only used for Kafka streams. At least one of payload_schema_locator or key_schema_locator must be set.
Show child attributesHide child attributes
- confluent_schemaobject
Confluent Schema Registry schema locator.
Show child attributesHide child attributes
- subjectstringRequired
The Confluent schema registry subject name.
- formatstringRequired
Serialization format for this schema.
FORMAT_UNSPECIFIEDFORMAT_AVROFORMAT_PROTOBUFFORMAT_JSON
- ingestion_configobjectRequired
Configuration for streaming data ingestion: the managed table storing an offline copy of forward fill data and optional historical backfill.
Show child attributesHide child attributes
- ingestion_destinationobjectRequired
Destination for the <Databricks>-managed Delta table that holds an offline copy of the streaming data for querying and training. This table contains both 1) forward-filled data from the Stream and 2) backfilled data from the BackfillSource (if provided). This table is created and managed by <Databricks> and is deleted when the Stream is deleted.
Show child attributesHide child attributes
- delta_table_namestring
The full three-part name (catalog, schema, name) of the Delta table to be created for ingestion.
- backfill_sourceobject
A user-provided source for backfilling data. Historical data is used when creating a training set from streaming features linked to this Stream. The backfill data stored in this location will be copied into the ingestion table for offline querying and training. The schema for this source must match exactly that of the key and payload schemas specified for this Stream, except that it may omit any columns listed in excluded_columns.
Show child attributesHide child attributes
- delta_table_namestring
The full three-part name (catalog, schema, name) of the Delta table containing the historical data to backfill.
- deduplication_columnsarray of string
Column paths used to identify duplicate rows during ingestion; only one row per distinct combination of these values is kept. Use dot notation for nested fields (e.g.
value.user_id). Empty list means every column is compared.
- record_type_filterstring
Optional SQL predicate to filter which record types from a streaming channel (e.g. a topic for Kafka) belong to this Stream. Events that do not match are not written to the ingestion table and are not used in materialization. Example: "value.event_type = 'transaction'".
- excluded_columnsarray of string
Column paths (dot notation, e.g. "value.email" for Kafka) to drop. A path may reference a struct, in which case all of its nested fields are dropped (e.g. "value.address" drops "value.address.city" and "value.address.zip"). These columns are not written to the ingestion table and cannot be referenced by any feature. They are dropped from ingestion, backfill, and materialization. For direct schemas, each column must exist in the relevant key or payload schema. With a schema registry, a column can be excluded before it exists. A column cannot also be a deduplication column in the ingestion_config.
Response
Returns the Stream object.