The PARTITIONS view in INFORMATION_SCHEMA

Senior Data Engineer • Contractor / Freelancer • GCP & AWS Certified
Search for a command to run...

Senior Data Engineer • Contractor / Freelancer • GCP & AWS Certified
No comments yet. Be the first to comment.
Practical guides to making BigQuery queries faster and cheaper — partitioning, clustering, search indexes, time travel, cost optimization, and query tuning strategies.
Have you ever used ingestion-time partitioning in BigQuery? It's a separate type of partitioning that distributes rows into partitions based on the time they land in BQ. Once such a table is defined, you can query the pseudocolumns PARTITIONDATE and ...
Here's a useful Dataform concept: pre_operations and post_operations. As the name implies, these represent a set of actions that run before and after the main operation (table, view, or SQL operations

BigQuery has always been a SQL engine for tabular data. Object tables add an interesting twist to that. Instead of rows containing values, an object table gives you one row per file — pointing at da

Query your data lake with warehouse-grade security and performance — without moving a single file.

Ever run a heavy BigQuery SQL query, processed gigabytes of data — and then accidentally closed the tab or forgot to save the results? 😬 Don't re-run it. Your results are still there. BigQuery automa

You can use query parameters in BigQuery hashtag#SQL (now in the console as well!) — but how are they different from variables, and when should you use each? Both parameters and variables act as place

Riding on the back of recent news that BigQuery table partition limit has just increased from 4k to 10k partitions, I wanted to talk a bit about the PARTITIONS view in INFORMATION_SCHEMA.
I've previously posted about information schema views, but this particular view allows us to get information about partitions in our partitioned tables.
Here's what we can find there:
- total rows in that partition
- logical & billable bytes for that partition
- storage tier (ACTIVE if modified in the last 90 days, LONG_TERM otherwise which is 50% cheaper)
- last modified time
Now, let's focus this last modified time as it is quite useful when building incremental SQL pipelines. Looking at this field could tell you if data in one of your many partitions was changed since your last run and needs to be reprocessed.
Such a feature should help you in cases when you don't have a reliable watermark column to determine what changed since your last run.
Found it useful? Subscribe to my Analytics newsletter at notjustsql.com.