Generate text by using the ML.GENERATE_TEXT function

This document shows you how to create a BigQuery ML remote model that references a Vertex AI foundation model. You can then use that model in conjunction with the ML.GENERATE_TEXT function to analyze text or visual content in a BigQuery table.

Required permissions

To create a connection, you need membership in the following Identity and Access Management (IAM) role:
- roles/bigquery.connectionAdmin
To grant permissions to the connection's service account, you need the following permission:
- resourcemanager.projects.setIamPolicy
To create the model using BigQuery ML, you need the following IAM permissions:
- bigquery.jobs.create
- bigquery.models.create
- bigquery.models.getData
- bigquery.models.updateData
- bigquery.models.updateMetadata
To run inference, you need the following permissions:
- bigquery.tables.getData on the table
- bigquery.models.getData on the model
- bigquery.jobs.create

Before you begin

In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

Note: If you don't plan to keep the resources that you create in this procedure, create a project instead of selecting an existing project. After you finish these steps, you can delete the project, removing all resources associated with the project.

Go to project selector
Make sure that billing is enabled for your Google Cloud project.
Enable the BigQuery, BigQuery Connection, and Vertex AI APIs.
Enable the APIs

If you want to use ML.GENERATE_TEXT with a gemini-pro-vision model in order to analyze visual content in an object table, you must have an Enterprise or Enterprise Plus reservation. For more information, see Create reservations.

Create a connection

Create a Cloud resource connection and get the connection's service account.

Select one of the following options:

Console

Go to the BigQuery page.

Go to BigQuery
To create a connection, click Add, and then click Connections to external data sources.
In the Connection type list, select Vertex AI remote models, remote functions and BigLake (Cloud Resource).
In the Connection ID field, enter a name for your connection.
Click Create connection.
Click Go to connection.
In the Connection info pane, copy the service account ID for use in a later step.

bq

In a command-line environment, create a connection:
```
bq mk --connection --location=REGION --project_id=PROJECT_ID \
    --connection_type=CLOUD_RESOURCE CONNECTION_ID
```
The --project_id parameter overrides the default project.

Replace the following:
- REGION: your connection region
- PROJECT_ID: your Google Cloud project ID
- CONNECTION_ID: an ID for your connection
When you create a connection resource, BigQuery creates a unique system service account and associates it with the connection.

Troubleshooting: If you get the following connection error, update the Google Cloud SDK:
```
Flags parsing error: flag --connection_type=CLOUD_RESOURCE: value should be one of...
```

Retrieve and copy the service account ID for use in a later step:

bq show --connection PROJECT_ID.REGION.CONNECTION_ID

The output is similar to the following:

name                          properties
1234.REGION.CONNECTION_ID     {"serviceAccountId": "connection-1234-9u56h9@gcp-sa-bigquery-condel.iam.gserviceaccount.com"}

Terraform

Append the following section into your main.tf file.

 ## This creates a cloud resource connection.
 ## Note: The cloud resource nested object has only one output only field - serviceAccountId.
 resource "google_bigquery_connection" "connection" {
    connection_id = "CONNECTION_ID"
    project = "PROJECT_ID"
    location = "REGION"
    cloud_resource {}
}

Replace the following:

CONNECTION_ID: an ID for your connection
PROJECT_ID: your Google Cloud project ID
REGION: your connection region

Give the service account access

Give your service account permission to use the connection. Failure to give permission results in an error. Select one of the following options:

Console

Go to the IAM & Admin page.

Go to IAM & Admin
Click Add.

The Add principals dialog opens.
In the New principals field, enter the service account ID that you copied earlier.
In the Select a role field, select Vertex AI, and then select Vertex AI User.
Click Save.

gcloud

Use the gcloud projects add-iam-policy-binding command:

gcloud projects add-iam-policy-binding 'PROJECT_NUMBER' --member='serviceAccount:MEMBER' --role='roles/aiplatform.user' --condition=None

Replace the following:

PROJECT_NUMBER: your project number
MEMBER: the service account ID that you copied earlier

Create a model

In the Google Cloud console, go to the BigQuery page.

Go to BigQuery
Using the SQL editor, create a remote model:
```
CREATE OR REPLACE MODEL
`PROJECT_ID.DATASET_ID.MODEL_NAME`
REMOTE WITH CONNECTION `PROJECT_ID.REGION.CONNECTION_ID`
OPTIONS (ENDPOINT = 'ENDPOINT');
```
Replace the following:
- PROJECT_ID: your project ID
- DATASET_ID: the ID of the dataset to contain the model. This dataset must be in the same location as the connection that you are using
- MODEL_NAME: the name of the model
- REGION: the region used by the connection
- CONNECTION_ID: the ID of your BigQuery connection
  When you view the connection details in the Google Cloud console, this is the value in the last section of the fully qualified connection ID that is shown in Connection ID, for example projects/myproject/locations/connection_location/connections/myconnection
- ENDPOINT: the name of the supported Vertex AI model to use. For example, ENDPOINT='gemini-pro'.
  For some types of models, you can specify a particular version of the model by appending @version to the model name. For example, text-bison@001. For information about supported model versions for different model types, see ENDPOINT.

Generate text from text data by using a prompt from a table

Generate text by using the ML.GENERATE_TEXT function with a remote model based on a supported Vertex AI Gemini API or Vertex AI PaLM API text model and a prompt from a table column:

`gemini-pro`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  TABLE PROJECT_ID.DATASET_ID.TABLE_NAME,
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
TABLE_NAME: the name of the table that contains the prompt. This table must have a column that's named prompt, or you can use an alias to use a differently named column.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,8192]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example

The following example shows a request with these characteristics:

Uses the prompt column of the prompts table for the prompt.
Returns a short and moderately probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    TABLE mydataset.prompts,
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      40 AS top_k, TRUE AS flatten_json_output));

`text-bison`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  TABLE PROJECT_ID.DATASET_ID.TABLE_NAME,
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
TABLE_NAME: the name of the table that contains the prompt. This table must have a column that's named prompt, or you can use an alias to use a differently named column.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,1024]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example

The following example shows a request with these characteristics:

Uses the prompt column of the prompts table for the prompt.
Returns a short and moderately probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    TABLE mydataset.prompts,
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      40 AS top_k, TRUE AS flatten_json_output));

`text-bison32`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  TABLE PROJECT_ID.DATASET_ID.TABLE_NAME,
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
TABLE_NAME: the name of the table that contains the prompt. This table must have a column that's named prompt, or you can use an alias to use a differently named column.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,8192]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example

The following example shows a request with these characteristics:

Uses the prompt column of the prompts table for the prompt.
Returns a short and moderately probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    TABLE mydataset.prompts,
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      40 AS top_k, TRUE AS flatten_json_output));

`text-unicorn`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  TABLE PROJECT_ID.DATASET_ID.TABLE_NAME,
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
TABLE_NAME: the name of the table that contains the prompt. This table must have a column that's named prompt, or you can use an alias to use a differently named column.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,1024]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example

The following example shows a request with these characteristics:

Uses the prompt column of the prompts table for the prompt.
Returns a short and moderately probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    TABLE mydataset.prompts,
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      40 AS top_k, TRUE AS flatten_json_output));

Generate text from text data by using a prompt from a query

Generate text by using the ML.GENERATE_TEXT function with a remote model based on a supported Gemini API or PaLM API text model and a query that provides the prompt:

`gemini-pro`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  (PROMPT_QUERY),
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
PROMPT_QUERY: a query that provides the prompt data.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,8192]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example 1

The following example shows a request with these characteristics:

Prompts for a summary of the text in the body column of the articles table.
Returns a moderately long and more probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT('Summarize this text', body) AS prompt
      FROM mydataset.articles
    ),
    STRUCT(
      0.2 AS temperature, 650 AS max_output_tokens, 0.2 AS top_p,
      15 AS top_k, TRUE AS flatten_json_output));

Example 2

The following example shows a request with these characteristics:

Uses a query to create the prompt data by concatenating strings that provide prompt prefixes with table columns.
Returns a short and moderately probable response.
Doesn't return the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT(question, 'Text:', description, 'Category') AS prompt
      FROM mydataset.input_table
    ),
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      30 AS top_k, FALSE AS flatten_json_output));

`text-bison`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  (PROMPT_QUERY),
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
PROMPT_QUERY: a query that provides the prompt data.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,1024]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example 1

The following example shows a request with these characteristics:

Prompts for a summary of the text in the body column of the articles table.
Returns a moderately long and more probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT('Summarize this text', body) AS prompt
      FROM mydataset.articles
    ),
    STRUCT(
      0.2 AS temperature, 650 AS max_output_tokens, 0.2 AS top_p,
      15 AS top_k, TRUE AS flatten_json_output));

Example 2

The following example shows a request with these characteristics:

Uses a query to create the prompt data by concatenating strings that provide prompt prefixes with table columns.
Returns a short and moderately probable response.
Doesn't return the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT(question, 'Text:', description, 'Category') AS prompt
      FROM mydataset.input_table
    ),
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      30 AS top_k, FALSE AS flatten_json_output));

`text-bison32`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  (PROMPT_QUERY),
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
PROMPT_QUERY: a query that provides the prompt data.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,8192]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example 1

The following example shows a request with these characteristics:

Prompts for a summary of the text in the body column of the articles table.
Returns a moderately long and more probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT('Summarize this text', body) AS prompt
      FROM mydataset.articles
    ),
    STRUCT(
      0.2 AS temperature, 650 AS max_output_tokens, 0.2 AS top_p,
      15 AS top_k, TRUE AS flatten_json_output));

Example 2

The following example shows a request with these characteristics:

Uses a query to create the prompt data by concatenating strings that provide prompt prefixes with table columns.
Returns a short and moderately probable response.
Doesn't return the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT(question, 'Text:', description, 'Category') AS prompt
      FROM mydataset.input_table
    ),
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      30 AS top_k, FALSE AS flatten_json_output));

`text-unicorn`

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  (PROMPT_QUERY),
  STRUCT(TOKENS AS max_output_tokens, TEMPERATURE AS temperature,
  TOP_K AS top_k, TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
PROMPT_QUERY: a query that provides the prompt data.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,1024]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 128.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 40.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example 1

The following example shows a request with these characteristics:

Prompts for a summary of the text in the body column of the articles table.
Returns a moderately long and more probable response.
Returns the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT('Summarize this text', body) AS prompt
      FROM mydataset.articles
    ),
    STRUCT(
      0.2 AS temperature, 650 AS max_output_tokens, 0.2 AS top_p,
      15 AS top_k, TRUE AS flatten_json_output));

Example 2

The following example shows a request with these characteristics:

Uses a query to create the prompt data by concatenating strings that provide prompt prefixes with table columns.
Returns a short and moderately probable response.
Doesn't return the generated text and the safety attributes in separate columns.

SELECT *
FROM
  ML.GENERATE_TEXT(
    MODEL `mydataset.llm_model`,
    (
      SELECT CONCAT(question, 'Text:', description, 'Category') AS prompt
      FROM mydataset.input_table
    ),
    STRUCT(
      0.4 AS temperature, 100 AS max_output_tokens, 0.5 AS top_p,
      30 AS top_k, FALSE AS flatten_json_output));

Generate text that describes visual content

Generate text by using the ML.GENERATE_TEXT function with a remote model based on a gemini-pro-vision multimodal model:

SELECT *
FROM ML.GENERATE_TEXT(
  MODEL `PROJECT_ID.DATASET_ID.MODEL_NAME`,
  TABLE PROJECT_ID.DATASET_ID.TABLE_NAME,
  STRUCT(PROMPT AS prompt, TOKENS AS max_output_tokens,
  TEMPERATURE AS temperature, TOP_K AS top_k,
  TOP_P AS top_p, FLATTEN_JSON AS flatten_json_output,
  STOP_SEQUENCES AS stop_sequences)
);

Replace the following:

PROJECT_ID: your project ID.
DATASET_ID: the ID of the dataset that contains the model.
MODEL_NAME: the name of the model.
TABLE_NAME: the name of the object table that contains the visual content to analyze. For more information on what types of visual content you can analyze, see Supported visual content.
The Cloud Storage bucket used by the object table must be in the same project where you have created the model and where you are calling the ML.GENERATE_TEXT function.
PROMPT: the prompt to use to analyze the visual content.
TOKENS: an INT64 value that sets the maximum number of tokens that can be generated in the response. This value must be in the range [1,2048]. Specify a lower value for shorter responses and a higher value for longer responses. The default is 2048.
TEMPERATURE: a FLOAT64 value in the range [0.0,1.0] that controls the degree of randomness in token selection. The default is 0.4.
Lower values for temperature are good for prompts that require a more deterministic and less open-ended or creative response, while higher values for temperature can lead to more diverse or creative results. A value of 0 for temperature is deterministic, meaning that the highest probability response is always selected.
TOP_K: an INT64 value in the range [1,40] that determines the initial pool of tokens the model considers for selection. Specify a lower value for less random responses and a higher value for more random responses. The default is 32.
TOP_P: a FLOAT64 value in the range [0.0,1.0] helps determine which tokens from the pool determined by TOP_K are selected. Specify a lower value for less random responses and a higher value for more random responses. The default is 0.95.
FLATTEN_JSON: a BOOL value that determines whether to return the generated text and the safety attributes in separate columns. The default is FALSE.
STOP_SEQUENCES: an ARRAY<STRING> value that removes the specified strings if they are included in responses from the model. Strings are matched exactly, including capitalization. The default is an empty array.

Example

This example analyzes visual content from an object table that's named videos and describes the content in each video:

SELECT
  uri,
  ml_generate_text_llm_result
FROM
  ML.GENERATE_TEXT(
    MODEL
      `mydataset.gemini_pro_vision_model`
        TABLE `mydataset.videos`
          STRUCT('What is happening in this video?' AS PROMPT,
          TRUE AS FLATTEN_JSON_OUTPUT));