Skip to main content

Feature

View as Markdown

Feature object​

full_namestringBetaImmutable

The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.

sourceobjectBetaImmutable

The data source of the feature.

Show child attributesHide child attributes
delta_table_sourceobjectBeta

A Delta table data source.

Show child attributesHide child attributes
full_namestringBeta

The full three-part (catalog, schema, table) name of the Delta table.

filter_conditionstringBeta

Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.

transformation_sqlstringBeta

A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have transformation_sql ", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.

Default: *

dataframe_schemastringBeta

Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}

kafka_sourceobjectBeta

A Kafka stream data source.

Show child attributesHide child attributes
namestringBeta

Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.

filter_conditionstringBeta

The filter condition applied to the source data before aggregation.

request_sourceobjectBeta

A request-time data source.

Show child attributesHide child attributes
flat_schemaobjectBeta

A flat schema with scalar-typed fields only.

Show child attributesHide child attributes
fieldsarray of objectBeta

The list of fields in this schema.

Show child attributesHide child attributes
namestringBeta

The name of the field.

data_typestringBeta

The scalar data type of the field.

Values:

  • SCALAR_DATA_TYPE_UNSPECIFIED
  • INTEGER
  • FLOAT
  • BOOLEAN
  • STRING
  • DOUBLE
  • LONG
  • TIMESTAMP
  • DATE
  • SHORT
  • BINARY
  • DECIMAL
stream_sourceobjectBeta

A Stream data source.

Show child attributesHide child attributes
full_namestringBeta

Three-part full name of the Stream (catalog.schema.stream).

filter_conditionstringBeta

The filter condition applied to the source data before aggregation.

functionobjectBetaImmutable

The function by which the feature is computed.

Show child attributesHide child attributes
aggregation_functionobjectBeta

An aggregation function applied over a time window.

Show child attributesHide child attributes
avgobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

count_functionobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

sumobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

minobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the minimum is computed.

maxobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the maximum is computed.

firstobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the first value is returned.

lastobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the last value is returned.

approx_count_distinctobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the approximate count of distinct values is computed.

relative_sddoubleBeta

The maximum relative standard deviation allowed (default defined by Spark).

approx_percentileobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the approximate percentile is computed.

percentiledoubleBeta

The percentile value to compute (between 0 and 1).

accuracyint64Beta

The accuracy parameter (higher is more accurate but slower).

stddev_popobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

stddev_sampobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the sample standard deviation is computed.

var_popobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the population variance is computed.

var_sampobjectBeta
Show child attributesHide child attributes
inputstringBeta

The input column from which the sample variance is computed.

time_windowobjectBeta

The time window over which the aggregation is computed.

Show child attributesHide child attributes
tumblingobjectBeta
Show child attributesHide child attributes
window_durationstringBeta

The duration of each tumbling window (non-overlapping, fixed-duration windows).

slidingobjectBeta
Show child attributesHide child attributes
window_durationstringBeta

The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

slide_durationstringBeta

The slide duration (interval by which windows advance, must be positive and less than duration).

rollingobjectBeta
Show child attributesHide child attributes
window_durationstringBeta

The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

delaystringBeta

Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.

column_selectionobjectBeta

Selects the latest value of a single column in a data source

Show child attributesHide child attributes
columnstringBeta

Column name from source to select as the feature value.

descriptionstringBeta

The description of the feature.

entitiesarray of objectBeta

The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringBeta

The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.

timeseries_columnobjectBeta

Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringBeta

The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.

catalog_namestringBetaOutput only

Name of parent catalog.

schema_namestringBetaOutput only

Name of parent schema relative to its parent catalog.

namestringBetaOutput only

Name of the feature, extracted from the full three-part name (catalog.schema.name).

created_atstringBetaOutput only

Time at which this feature was created.

created_bystringBetaOutput only

Username of the feature creator.

Get a feature Beta​

GET /api/2.0/feature-engineering/features/{full_name}

Get a Feature.

API scopes: mlflow

Parameters​

full_namestringRequiredpath

Name of the feature to get.

Response​

Returns the Feature object.

List features Beta​

GET /api/2.0/feature-engineering/features

List Features.

API scopes: mlflow

Parameters​

page_tokenstringquery

Pagination token to go to the next page based on a previous query.

page_sizeint32query

The maximum number of results to return.

catalog_namestringRequiredquery

Name of parent catalog for features of interest.

schema_namestringRequiredquery

Name of parent schema relative to its parent catalog.

Response​

Returns a list of Feature objects.

Create a feature Beta​

POST /api/2.0/feature-engineering/features

Create a Feature.

API scopes: mlflow

Request body​

Feature to create.

full_namestringRequiredImmutable

The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.

sourceobjectRequiredImmutable

The data source of the feature.

Show child attributesHide child attributes
delta_table_sourceobject

A Delta table data source.

Show child attributesHide child attributes
full_namestringRequired

The full three-part (catalog, schema, table) name of the Delta table.

filter_conditionstring

Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.

transformation_sqlstring

A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have transformation_sql ", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.

Default: *

dataframe_schemastring

Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}

kafka_sourceobject

A Kafka stream data source.

Show child attributesHide child attributes
namestringRequired

Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.

filter_conditionstring

The filter condition applied to the source data before aggregation.

request_sourceobject

A request-time data source.

Show child attributesHide child attributes
flat_schemaobject

A flat schema with scalar-typed fields only.

Show child attributesHide child attributes
fieldsarray of objectRequired

The list of fields in this schema.

Show child attributesHide child attributes
namestringRequired

The name of the field.

data_typestringRequired

The scalar data type of the field.

Values:

  • SCALAR_DATA_TYPE_UNSPECIFIED
  • INTEGER
  • FLOAT
  • BOOLEAN
  • STRING
  • DOUBLE
  • LONG
  • TIMESTAMP
  • DATE
  • SHORT
  • BINARY
  • DECIMAL
stream_sourceobject

A Stream data source.

Show child attributesHide child attributes
full_namestringRequired

Three-part full name of the Stream (catalog.schema.stream).

filter_conditionstring

The filter condition applied to the source data before aggregation.

functionobjectRequiredImmutable

The function by which the feature is computed.

Show child attributesHide child attributes
aggregation_functionobject

An aggregation function applied over a time window.

Show child attributesHide child attributes
avgobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

count_functionobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

sumobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

minobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the minimum is computed.

maxobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the maximum is computed.

firstobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the first value is returned.

lastobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the last value is returned.

approx_count_distinctobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the approximate count of distinct values is computed.

relative_sddouble

The maximum relative standard deviation allowed (default defined by Spark).

approx_percentileobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the approximate percentile is computed.

percentiledoubleRequired

The percentile value to compute (between 0 and 1).

accuracyint64

The accuracy parameter (higher is more accurate but slower).

stddev_popobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

stddev_sampobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sample standard deviation is computed.

var_popobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the population variance is computed.

var_sampobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sample variance is computed.

time_windowobject

The time window over which the aggregation is computed.

Show child attributesHide child attributes
tumblingobject
Show child attributesHide child attributes
window_durationstringRequired

The duration of each tumbling window (non-overlapping, fixed-duration windows).

slidingobject
Show child attributesHide child attributes
window_durationstring

The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

slide_durationstringRequired

The slide duration (interval by which windows advance, must be positive and less than duration).

rollingobject
Show child attributesHide child attributes
window_durationstring

The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

delaystring

Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.

column_selectionobject

Selects the latest value of a single column in a data source

Show child attributesHide child attributes
columnstringRequired

Column name from source to select as the feature value.

descriptionstring

The description of the feature.

entitiesarray of object

The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringRequired

The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.

timeseries_columnobject

Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringRequired

The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.

Response​

Returns the Feature object.

Update a feature's description (all other fields are immutable) Beta​

PATCH /api/2.0/feature-engineering/features/{full_name}

Update a Feature.

API scopes: mlflow

Parameters​

full_namestringRequiredImmutablepath

The full three-part name (catalog, schema, name) of the feature. This is the feature's resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.

update_maskstringRequiredquery

Fields to update. The only supported path is description.

Request body​

Feature whose full_name identifies the target. Only description is mutable.

sourceobjectRequiredImmutable

The data source of the feature.

Show child attributesHide child attributes
delta_table_sourceobject

A Delta table data source.

Show child attributesHide child attributes
full_namestringRequired

The full three-part (catalog, schema, table) name of the Delta table.

filter_conditionstring

Single WHERE clause to filter delta table before applying transformations. Will be row-wise evaluated, so should only include conditionals and projections.

transformation_sqlstring

A single SQL SELECT expression applied after filter_condition. Should contains all the columns needed (eg. "SELECT , col_a + col_b AS col_c FROM x.y.z WHERE col_a > 0" would have transformation_sql ", col_a + col_b AS col_c") If transformation_sql is not provided, all columns of the delta table are present in the DataSource dataframe.

Default: *

dataframe_schemastring

Schema of the resulting dataframe after transformations, in Spark StructType JSON format (from df.schema.json()). Required if transformation_sql is specified. Example: {"type":"struct","fields":[{"name":"col_a","type":"integer","nullable":true,"metadata":{}},{"name":"col_c","type":"integer","nullable":true,"metadata":{}}]}

kafka_sourceobject

A Kafka stream data source.

Show child attributesHide child attributes
namestringRequired

Name of the Kafka source, used to identify it. This is used to look up the corresponding KafkaConfig object. Can be distinct from topic name.

filter_conditionstring

The filter condition applied to the source data before aggregation.

request_sourceobject

A request-time data source.

Show child attributesHide child attributes
flat_schemaobject

A flat schema with scalar-typed fields only.

Show child attributesHide child attributes
fieldsarray of objectRequired

The list of fields in this schema.

Show child attributesHide child attributes
namestringRequired

The name of the field.

data_typestringRequired

The scalar data type of the field.

Values:

  • SCALAR_DATA_TYPE_UNSPECIFIED
  • INTEGER
  • FLOAT
  • BOOLEAN
  • STRING
  • DOUBLE
  • LONG
  • TIMESTAMP
  • DATE
  • SHORT
  • BINARY
  • DECIMAL
stream_sourceobject

A Stream data source.

Show child attributesHide child attributes
full_namestringRequired

Three-part full name of the Stream (catalog.schema.stream).

filter_conditionstring

The filter condition applied to the source data before aggregation.

functionobjectRequiredImmutable

The function by which the feature is computed.

Show child attributesHide child attributes
aggregation_functionobject

An aggregation function applied over a time window.

Show child attributesHide child attributes
avgobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the average is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

count_functionobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the count is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

sumobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sum is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

minobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the minimum is computed.

maxobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the maximum is computed.

firstobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the first value is returned.

lastobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the last value is returned.

approx_count_distinctobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the approximate count of distinct values is computed.

relative_sddouble

The maximum relative standard deviation allowed (default defined by Spark).

approx_percentileobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the approximate percentile is computed.

percentiledoubleRequired

The percentile value to compute (between 0 and 1).

accuracyint64

The accuracy parameter (higher is more accurate but slower).

stddev_popobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the population standard deviation is computed. For Kafka sources, use dot-prefixed path notation (e.g., "value.amount"). For nested fields, the leaf node name is used. Colon-prefixed notation (e.g., "value:amount") is supported for backwards compatibility but is deprecated; migrate to dot notation.

stddev_sampobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sample standard deviation is computed.

var_popobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the population variance is computed.

var_sampobject
Show child attributesHide child attributes
inputstringRequired

The input column from which the sample variance is computed.

time_windowobject

The time window over which the aggregation is computed.

Show child attributesHide child attributes
tumblingobject
Show child attributesHide child attributes
window_durationstringRequired

The duration of each tumbling window (non-overlapping, fixed-duration windows).

slidingobject
Show child attributesHide child attributes
window_durationstring

The duration of the sliding window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

slide_durationstringRequired

The slide duration (interval by which windows advance, must be positive and less than duration).

rollingobject
Show child attributesHide child attributes
window_durationstring

The duration of the rolling window. Must be positive when set; absent means lifetime (aggregate over the entity's entire history).

delaystring

Non-negative analytic lag that evaluates the window this far in the past. Use this for timing variations unrelated to source lateness, such as a 30-day count as of one week ago. If unset, the analytic lag is zero. It composes with source.lateness when both are set.

column_selectionobject

Selects the latest value of a single column in a data source

Show child attributesHide child attributes
columnstringRequired

Column name from source to select as the feature value.

descriptionstring

The description of the feature.

entitiesarray of object

The entity columns for the feature, used as aggregation keys and for query-time lookup. Optional since entities are not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringRequired

The name of the entity column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.user_id", "key.partition_key"). For nested fields, the leaf node name (e.g., "user_id" from "value.trip_details.user_id") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:user_id") is supported for backwards compatibility but is deprecated; migrate to dot notation.

timeseries_columnobject

Column recording time, used for point-in-time joins, backfills, and aggregations. Optional since a timeseries column is not set for RequestSource features or on-demand calculated features.

Show child attributesHide child attributes
namestringRequired

The name of the timeseries column. For Kafka sources, use dot-prefixed path notation to reference fields within the key or value schema (e.g., "value.event_timestamp"). For nested fields, the leaf node name (e.g., "event_timestamp" from "value.event_details.event_timestamp") is what will be present in materialized tables and expected to match at query time. Colon-prefixed notation (e.g., "value:event_timestamp") is supported for backwards compatibility but is deprecated; migrate to dot notation.

Response​

Returns the Feature object.

Delete a feature Beta​

DELETE /api/2.0/feature-engineering/features/{full_name}

Delete a Feature.

API scopes: mlflow

Parameters​

full_namestringRequiredpath

Name of the feature to delete.