Skip to main content

sampleBy (DataFrameStatFunctions)

Returns a stratified sample without replacement based on the fraction given on each stratum.

Syntax​

sampleBy(col, fractions, seed=None)

Parameters​

Parameter

Type

Description

col

str

The column that defines strata.

fractions

dict

The sampling fraction for each stratum. Strata not specified are treated as having a fraction of zero.

seed

int, optional

Random seed.

Parameter

Type

Description

col

str

The column that defines strata.

fractions

dict

The sampling fraction for each stratum. Strata not specified are treated as having a fraction of zero.

seed

int, optional

Random seed.

Returns​

DataFrame

Examples​

Python
from pyspark.sql import functions as sf
dataset = spark.range(0, 100, 1, 5).select((sf.col("id") % 3).alias("key"))
sampled = dataset.stat.sampleBy("key", fractions={0: 0.1, 1: 0.2}, seed=0)
sampled.groupBy("key").count().orderBy("key").show()
# +---+-----+
# |key|count|
# +---+-----+
# | 0| 4|
# | 1| 9|
# +---+-----+