Google Cloud BigQuery Vector Search lets you use GoogleSQL to do semantic search, using vector indexes for fast approximate results, or using brute force for exact results.This tutorial illustrates how to work with an end-to-end data and embedding management system in LangChain, and provides a scalable semantic search in BigQuery using the
BigQueryVectorStore
class. This class is part of a set of 2 classes capable of providing a unified data storage and flexible vector search in Google Cloud:
- BigQuery Vector Search: with
BigQueryVectorStore
class, which is ideal for rapid prototyping with no infrastructure setup and batch retrieval. - Feature Store Online Store: with
VertexFSVectorStore
class, enables low-latency retrieval with manual or scheduled data sync. Perfect for production-ready user-facing GenAI applications.
Getting started
Install the library
Before you begin
Set your project ID
If you don’t know your project ID, try the following:- Run
gcloud config list
. - Run
gcloud projects list
. - See the support page: Locate the project ID.
Set the region
You can also change theREGION
variable used by BigQuery. Learn more about BigQuery regions.
Set the dataset and table names
They will be your BigQuery Vector Store.Authenticating your notebook environment
- If you are using Colab to run this notebook, uncomment the cell below and continue.
- If you are using Vertex AI Workbench, check out the setup instructions here.
Demo: BigQueryVectorStore
Create an embedding class instance
You may need to enable Vertex AI API in your project by runninggcloud services enable aiplatform.googleapis.com --project {PROJECT_ID}
(replace {PROJECT_ID}
with the name of your project).
You can use any LangChain embeddings model.
Initialize BigQueryVectorStore
BigQuery Dataset and Table will be automatically created if they do not exist. See class definition here for all optional paremeters.Add texts
Search for documents
Search for documents by vector
Searching Documents with Metadata Filters
The vectorstore supports two methods for applying filters to metadata fields when performing document searches:- Dictionary-based Filters
- You can pass a dictionary (dict) where the keys represent metadata fields and the values specify the filter condition. This method applies an equality filter between the key and the corresponding value. When multiple key-value pairs are provided, they are combined using a logical AND operation.
- SQL-based Filters
- Alternatively, you can provide a string representing an SQL WHERE clause to define more complex filtering conditions. This allows for greater flexibility, supporting SQL expressions such as comparison operators and logical operators. Learn more about BigQuery operators.
Batch search
BigQueryVectorStore offers abatch_search
method for scalable Vector similarity search.
Add text with embeddings
You can also bring your own embeddings with theadd_texts_with_embeddings
method.
This is particularly useful for multimodal data which might require custom preprocessing before the embedding generation.
Low-latency serving with Feature Store
You can simply use the method.to_vertex_fs_vector_store()
to get a VertexFSVectorStore object, which offers low latency for online use cases. All mandatory parameters will be automatically transferred from the existing BigQueryVectorStore class. See the class definition for all the other parameters you can use.
Moving back to BigQueryVectorStore is equivalently easy with the .to_bq_vector_store()
method.