Feature
Feature object
- full_namestringBetaImmutable
The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.
- sourceobjectBetaImmutable
The data source of the feature.
Show child attributesHide child attributes
- delta_table_sourceobjectBeta
A Delta table data source.
Show child attributesHide child attributes
- full_namestringBeta
The full three-part (catalog, schema, table) name of the Delta table.
- filter_conditionstringBeta
Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.
- transformation_sqlstringBeta
A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have
transformation_sql", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.
- dataframe_schemastringBeta
Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}
- kafka_sourceobjectBeta
A Kafka stream data source.
Show child attributesHide child attributes
- namestringBeta
Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.
- filter_conditionstringBeta
The filter condition applied to the source data before aggregation.
- request_sourceobjectBeta
A request-time data source.
Show child attributesHide child attributes
- flat_schemaobjectBeta
A flat schema with scalar-typed fields only.
Show child attributesHide child attributes
- fieldsarray of objectBeta
The list of fields in this schema.
Show child attributesHide child attributes
- namestringBeta
The name of the field.
- data_typestringBeta
The scalar data type of the field.
SCALAR_DATA_TYPE_UNSPECIFIEDINTEGERFLOATBOOLEANSTRINGDOUBLELONGTIMESTAMPDATESHORTBINARYDECIMAL
- stream_sourceobjectBeta
A Stream data source.
Show child attributesHide child attributes
- full_namestringBeta
Three-part full name of the Stream (catalog.schema.stream).
- filter_conditionstringBeta
The filter condition applied to the source data before aggregation.
- functionobjectBetaImmutable
The function by which the feature is computed.
Show child attributesHide child attributes
- aggregation_functionobjectBeta
An aggregation function applied over a time window.
Show child attributesHide child attributes
- avgobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- count_functionobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- sumobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- minobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the minimum is computed.
- maxobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the maximum is computed.
- firstobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the first value is returned.
- lastobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the last value is returned.
- approx_count_distinctobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the approximate count of distinct values is computed.
- relative_sddoubleBeta
The maximum relative standard deviation allowed (default defined by Spark).
- approx_percentileobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the approximate percentile is computed.
- percentiledoubleBeta
The percentile value to compute (between 0 and 1).
- accuracyint64Beta
The accuracy parameter (higher is more accurate but slower).
- stddev_popobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- stddev_sampobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the sample standard deviation is computed.
- var_popobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the population variance is computed.
- var_sampobjectBeta
Show child attributesHide child attributes
- inputstringBeta
The input column from which the sample variance is computed.
- time_windowobjectBeta
The time window over which the aggregation is computed.
Show child attributesHide child attributes
- tumblingobjectBeta
Show child attributesHide child attributes
- window_durationstringBeta
The duration of each tumbling window (non-overlapping, fixed-duration windows).
- slidingobjectBeta
Show child attributesHide child attributes
- window_durationstringBeta
The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- slide_durationstringBeta
The slide duration (interval by which windows advance, must be positive and less than duration).
- rollingobjectBeta
Show child attributesHide child attributes
- window_durationstringBeta
The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- delaystringBeta
Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.
- column_selectionobjectBeta
Selects the latest value of a single column in a data source
Show child attributesHide child attributes
- columnstringBeta
Column name from source to select as the feature value.
- descriptionstringBeta
The description of the feature.
- entitiesarray of objectBeta
The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringBeta
The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- timeseries_columnobjectBeta
Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringBeta
The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- catalog_namestringBetaOutput only
Name of parent catalog.
- schema_namestringBetaOutput only
Name of parent schema relative to its parent catalog.
- namestringBetaOutput only
Name of the feature, extracted from the full three-part name (catalog.schema.name).
- created_atstringBetaOutput only
Time at which this feature was created.
- created_bystringBetaOutput only
Username of the feature creator.
Get a feature Beta
GET
Get a Feature.
API scopes: mlflow
Parameters
- full_namestringRequiredpath
Name of the feature to get.
Response
Returns the Feature object.
List features Beta
GET
List Features.
API scopes: mlflow
Parameters
- page_tokenstringquery
Pagination token to go to the next page based on a previous query.
- page_sizeint32query
The maximum number of results to return.
- catalog_namestringRequiredquery
Name of parent catalog for features of interest.
- schema_namestringRequiredquery
Name of parent schema relative to its parent catalog.
Response
Returns a list of Feature objects.
Create a feature Beta
POST
Create a Feature.
API scopes: mlflow
Request body
Feature to create.
- full_namestringRequiredImmutable
The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.
- sourceobjectRequiredImmutable
The data source of the feature.
Show child attributesHide child attributes
- delta_table_sourceobject
A Delta table data source.
Show child attributesHide child attributes
- full_namestringRequired
The full three-part (catalog, schema, table) name of the Delta table.
- filter_conditionstring
Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.
- transformation_sqlstring
A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have
transformation_sql", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.
- dataframe_schemastring
Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}
- kafka_sourceobject
A Kafka stream data source.
Show child attributesHide child attributes
- namestringRequired
Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.
- filter_conditionstring
The filter condition applied to the source data before aggregation.
- request_sourceobject
A request-time data source.
Show child attributesHide child attributes
- flat_schemaobject
A flat schema with scalar-typed fields only.
Show child attributesHide child attributes
- fieldsarray of objectRequired
The list of fields in this schema.
Show child attributesHide child attributes
- namestringRequired
The name of the field.
- data_typestringRequired
The scalar data type of the field.
SCALAR_DATA_TYPE_UNSPECIFIEDINTEGERFLOATBOOLEANSTRINGDOUBLELONGTIMESTAMPDATESHORTBINARYDECIMAL
- stream_sourceobject
A Stream data source.
Show child attributesHide child attributes
- full_namestringRequired
Three-part full name of the Stream (catalog.schema.stream).
- filter_conditionstring
The filter condition applied to the source data before aggregation.
- functionobjectRequiredImmutable
The function by which the feature is computed.
Show child attributesHide child attributes
- aggregation_functionobject
An aggregation function applied over a time window.
Show child attributesHide child attributes
- avgobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- count_functionobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- sumobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- minobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the minimum is computed.
- maxobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the maximum is computed.
- firstobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the first value is returned.
- lastobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the last value is returned.
- approx_count_distinctobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the approximate count of distinct values is computed.
- relative_sddouble
The maximum relative standard deviation allowed (default defined by Spark).
- approx_percentileobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the approximate percentile is computed.
- percentiledoubleRequired
The percentile value to compute (between 0 and 1).
- accuracyint64
The accuracy parameter (higher is more accurate but slower).
- stddev_popobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- stddev_sampobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sample standard deviation is computed.
- var_popobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the population variance is computed.
- var_sampobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sample variance is computed.
- time_windowobject
The time window over which the aggregation is computed.
Show child attributesHide child attributes
- tumblingobject
Show child attributesHide child attributes
- window_durationstringRequired
The duration of each tumbling window (non-overlapping, fixed-duration windows).
- slidingobject
Show child attributesHide child attributes
- window_durationstring
The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- slide_durationstringRequired
The slide duration (interval by which windows advance, must be positive and less than duration).
- rollingobject
Show child attributesHide child attributes
- window_durationstring
The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- delaystring
Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.
- column_selectionobject
Selects the latest value of a single column in a data source
Show child attributesHide child attributes
- columnstringRequired
Column name from source to select as the feature value.
- descriptionstring
The description of the feature.
- entitiesarray of object
The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringRequired
The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- timeseries_columnobject
Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringRequired
The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.
Response
Returns the Feature object.
Update a feature's description (all other fields are immutable) Beta
PATCH
Update a Feature.
API scopes: mlflow
Parameters
- full_namestringRequiredImmutablepath
The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.
- update_maskstringRequiredquery
Fields to update. The only supported path is description.
Request body
Feature whose full_name identifies the target. Only description is mutable.
- sourceobjectRequiredImmutable
The data source of the feature.
Show child attributesHide child attributes
- delta_table_sourceobject
A Delta table data source.
Show child attributesHide child attributes
- full_namestringRequired
The full three-part (catalog, schema, table) name of the Delta table.
- filter_conditionstring
Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.
- transformation_sqlstring
A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have
transformation_sql", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.
- dataframe_schemastring
Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}
- kafka_sourceobject
A Kafka stream data source.
Show child attributesHide child attributes
- namestringRequired
Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.
- filter_conditionstring
The filter condition applied to the source data before aggregation.
- request_sourceobject
A request-time data source.
Show child attributesHide child attributes
- flat_schemaobject
A flat schema with scalar-typed fields only.
Show child attributesHide child attributes
- fieldsarray of objectRequired
The list of fields in this schema.
Show child attributesHide child attributes
- namestringRequired
The name of the field.
- data_typestringRequired
The scalar data type of the field.
SCALAR_DATA_TYPE_UNSPECIFIEDINTEGERFLOATBOOLEANSTRINGDOUBLELONGTIMESTAMPDATESHORTBINARYDECIMAL
- stream_sourceobject
A Stream data source.
Show child attributesHide child attributes
- full_namestringRequired
Three-part full name of the Stream (catalog.schema.stream).
- filter_conditionstring
The filter condition applied to the source data before aggregation.
- functionobjectRequiredImmutable
The function by which the feature is computed.
Show child attributesHide child attributes
- aggregation_functionobject
An aggregation function applied over a time window.
Show child attributesHide child attributes
- avgobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- count_functionobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- sumobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- minobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the minimum is computed.
- maxobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the maximum is computed.
- firstobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the first value is returned.
- lastobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the last value is returned.
- approx_count_distinctobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the approximate count of distinct values is computed.
- relative_sddouble
The maximum relative standard deviation allowed (default defined by Spark).
- approx_percentileobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the approximate percentile is computed.
- percentiledoubleRequired
The percentile value to compute (between 0 and 1).
- accuracyint64
The accuracy parameter (higher is more accurate but slower).
- stddev_popobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- stddev_sampobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sample standard deviation is computed.
- var_popobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the population variance is computed.
- var_sampobject
Show child attributesHide child attributes
- inputstringRequired
The input column from which the sample variance is computed.
- time_windowobject
The time window over which the aggregation is computed.
Show child attributesHide child attributes
- tumblingobject
Show child attributesHide child attributes
- window_durationstringRequired
The duration of each tumbling window (non-overlapping, fixed-duration windows).
- slidingobject
Show child attributesHide child attributes
- window_durationstring
The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- slide_durationstringRequired
The slide duration (interval by which windows advance, must be positive and less than duration).
- rollingobject
Show child attributesHide child attributes
- window_durationstring
The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).
- delaystring
Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.
- column_selectionobject
Selects the latest value of a single column in a data source
Show child attributesHide child attributes
- columnstringRequired
Column name from source to select as the feature value.
- descriptionstring
The description of the feature.
- entitiesarray of object
The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringRequired
The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.
- timeseries_columnobject
Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.
Show child attributesHide child attributes
- namestringRequired
The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.
Response
Returns the Feature object.