Extended data type support for BigQuery search indexes

Senior Data Engineer • Contractor / Freelancer • GCP & AWS Certified
Search for a command to run...

Senior Data Engineer • Contractor / Freelancer • GCP & AWS Certified
No comments yet. Be the first to comment.
Practical guides to making BigQuery queries faster and cheaper — partitioning, clustering, search indexes, time travel, cost optimization, and query tuning strategies.
Query your data lake with warehouse-grade security and performance — without moving a single file.
Here's a useful Dataform concept: pre_operations and post_operations. As the name implies, these represent a set of actions that run before and after the main operation (table, view, or SQL operations

BigQuery has always been a SQL engine for tabular data. Object tables add an interesting twist to that. Instead of rows containing values, an object table gives you one row per file — pointing at da

Query your data lake with warehouse-grade security and performance — without moving a single file.

Ever run a heavy BigQuery SQL query, processed gigabytes of data — and then accidentally closed the tab or forgot to save the results? 😬 Don't re-run it. Your results are still there. BigQuery automa

You can use query parameters in BigQuery hashtag#SQL (now in the console as well!) — but how are they different from variables, and when should you use each? Both parameters and variables act as place

So, a couple of months ago, I posted about BigQuery search indexes, interesting to those who work with large volumes of STRING or JSON data.
Recently, I came across a blog post introducing, in preview, the indexing of INT64 and TIMESTAMP columns as well.
Now, this brings interesting applications when working with big and huge tables, logs, or JSON data. You can partition only by one field and cluster by another four. You can’t partition or cluster by fields belonging to a STRUCT. Now this , using a SEARCH INDEX with integers and timestamps allows for a more efficient retrieval in such scenarios.
The referenced blog post (in comments) presents some pretty impressive improvements over not using search indexes.
I’ve played a bit with it but still haven’t managed to get the index to be used (which, last time, with STRING columns, I did). Perhaps something is wrong with the sample data (although not the size; I tried with a ~200 GB table), need to dig some more.
Does anyone use this feature in the real world?
Found it useful? Subscribe to my Analytics newsletter at https://www.notjustsql.com.