Practical SQL

Approximate Aggregate Functions in BigQuery

Constantin LunguUpdated 1 min read

Photo by gretta vosper on Unsplash

Sometimes you don't need perfect, but just good enough. Take approximate aggregate functions in BigQuery, for example.

These are a type of aggregate functions that produce approximate results instead of exact ones but have the upside of typically requiring fewer resources for the computation.

When would I use one? This would be suitable where we can live with an uncertainty or small difference, especially for huge tables, during a preliminary check or data exploration.

Let's look at a practical example. Suppose we have the following data:

BigQuery console preview of the sample data for approximate aggregates, with columns id, value and ds_date; the first 10 rows are all dated 2020-12-18 and have single-digit values such as 1, 8, 6, 4 and 0.

APPROX_TOP_COUNT will compute the approx top N elements and their value counts

SELECT 
    APPROX_TOP_COUNT(value, 5) AS top_value_counts 
FROM `learning.data_source`

BigQuery console result of APPROX_TOP_COUNT(value, 5): one row holding an array of value and count pairs, 4 with 40258, 0 with 40057, 5 with 40051, 3 with 39979 and 9 with 39944.

APPROX_COUNT_DISTINCT will compute the approx distinct count (also can be grouped)

SELECT 
    APPROX_COUNT_DISTINCT(value) AS approx_distinct_value_count 
FROM `learning.data_source`

BigQuery console result of APPROX_COUNT_DISTINCT(value): a single row in the approx_distinct_value_count column, header truncated, with the value 11.

You can discover more approximate aggregate functions in the documentation.

Thanks for reading!


Enjoyed this? Here are some related articles you might find useful: