BigQuery AI.AGG
AI.AGG uses a Vertex AI Gemini model to aggregate data based on natural
language instructions and returns a STRING. It enables reasoning over groups of
rows, up to an entire table.
Common use cases:
- Sentiment analysis of many user reviews
- Summarization of content (e.g. log analysis, object table image descriptions)
- Analysis of agent prompts/responses or unstructured feedback
- Categorization of user feedback, reviews, or support tickets into common themes
Syntax Reference
AI.AGG(
[ DISTINCT ]
input,
instruction
[, connection_id => 'CONNECTION_ID']
[, endpoint => 'ENDPOINT']
)Input Arguments
| Argument | Requirement | Type | Description |
|---|---|---|---|
input |
Required | String/Struct | The data to be aggregated, as either a STRING value or a STRUCT object consisting of STRING values, [ObjectRefRuntime values](/bigquery/docs /reference/standard -sql/objectref _functions #objectrefruntime), and arrays of STRING and ObjectRefRuntime values. ObjectRefRuntime values reference text or image data in {{gcs_name}} and are generated by the [OBJ.GET_ACCESS_URL function](/bigquery /docs/reference /standard-sql /objectref_functions #objget_access_url). |
instruction |
Required | String | The user instruction on how to aggregate data (must be a string literal or query parameter). |
connection_id |
Optional | String | Identifies a cloud resource connection. |
endpoint |
Optional | String | The model endpoint (e.g., 'gemini-2.5-flash'). |
Output Schema
| Column Name | Type | Description |
|---|---|---|
| (Scalar Result) | STRING |
The summarized/aggregated result. It |
| : : : returns a STRING value. If you use the : | ||
: : : AI.AGG function with a GROUP_BY : |
||
| : : : statement, then the function returns a : | ||
| : : : STRING value for each input group. Capped : | ||
| : : : at 10000 tokens per group. : |
Execution Limitations
- Final output is capped at 10000 tokens per group. Longer outputs will be truncated.
- Arrays of object refs using OBJ.GET_ACCESS_URL() may cause rows to be skipped.
Examples
User sentiment analysis
SELECT
title,
movie_id,
AI.AGG(
review,
'You will be given user-provided reviews of a movie. Summarize the overall sentiment towards the movie.'
) AS sentiment
FROM `bigquery-public-data.imdb.reviews`
WHERE movie_id IN ('tt0339384', 'tt0084787', 'tt0029850')
GROUP BY movie_id, title;Common categories
SELECT AI.AGG(
TO_JSON_STRING(t),
'These are Wikipedia comments. What are the most used languages?'
)
FROM (SELECT * FROM `bigquery-public-data.samples.wikipedia` LIMIT 30000) AS t;Most frequent items
SELECT
AI.AGG(
TO_JSON_STRING(t),
'Among these tech news articles, what are the top 3 most referenced companies and what are they famous for?'
)
FROM `bigquery-public-data.bbc_news.fulltext` AS t
WHERE t.category = "tech";Image content summarization
SELECT AI.AGG(
STRUCT(OBJ.GET_ACCESS_URL(ref, 'r')),
'You will be provided with a series of images. What are the most common categories these images belong to?'
)
FROM multimodal.images;