Practical SQL

Calculating the MODE in BigQuery

Constantin LunguUpdated 1 min read

Photo by Chris Liverani on Unsplash

How do you compute the MODE (most frequent value) in BigQuery?

For the other measures of central tendency like MEAN and MEDIAN, there are straightforward ways to compute results - functions AVG and PERCENTILE_CONT/PERCENTILE_DIST respectively, but there's no dedicated function for MODE.

By the way, if you have a huge dataset and can bear some lack of precision, take a look at APPROX_TOP_COUNT.

Say we have the following input data:

Input table for calculating the mode, with columns country, value and value_date: US and UK rows twice a month from 2023-01-01 to 2023-03-15 with values 1 or 2, and empty NULL values on 2023-02-15 and 2023-03-15.

Now here's how we can compute them otherwise:
- filter out NULLS (if we want to ignore them) or do nothing if we want to keep them
- compute value counts for our desired grain
- take the most frequent one per our grain using QUALIFY + RANK

SELECT country, value, COUNT(1) AS times_seen

FROM input_data

-- comment this if you want to include NULLS
WHERE value IS NOT NULL

GROUP BY country, value

QUALIFY RANK() OVER(PARTITION BY country
                    ORDER BY times_seen DESC) = 1

Here's how the output would look with NULLS excluded.

Mode query output with NULLs excluded, columns country, value and times_seen: UK has mode 1 and US has mode 2, each seen 2 times.

And with them included:

Mode query output with NULLs included, columns country, value and times_seen: UK returns 1 and NULL, US returns 2 and NULL, all tied at 2 occurrences, so RANK keeps two rows per country.