Configure and customize queries for managed knowledge bases
You can configure and customize retrieval, further improving the relevancy of results. For example, you can apply filters to document metadata fields/attributes to use the most recently updated documents or documents with recent modification times.
Note
All of the following configurations are only applicable to unstructured data sources.
To learn more about these configurations in the console or the API, select from the following topics:
When you query a knowledge base, Amazon Bedrock returns up to five results in the response by default. Each result corresponds to a source chunk.
Note
The actual number of results in the response might be less than the specified numberOfResults value, since this parameter sets the maximum number of results to return. If you have configured hierarchical chunking for your chunking strategy, the numberOfResults parameter maps to the number of child chunks that the knowledge base will retrieve. Since child chunks that share the same parent chunk are replaced with the parent chunk in the final response, the number of results returned might be less than the requested amount.
To modify the maximum number of results to return, choose the tab for your preferred method, and then follow the steps:
You can apply filters to document fields/attributes to help you further improve the relevancy of responses. Your data sources can include document metadata attributes/fields to filter on and can specify which fields to include in the embeddings.
Managed knowledge base considerations
When using metadata filtering with a managed knowledge base:
-
The
startsWithandstringContainsmetadata filters are not supported. Useequals,greaterThan,lessThan,in, ornotInoperators instead. -
The range operators (
greaterThan,greaterThanOrEquals,lessThan, andlessThanOrEquals) accept either a number or a date-time value. To filter on a date-time, provide the value as a string in ISO-8601 offset date-time format. Use the full offset form, for example,2026-03-02T19:02:18Z. Range operators support date-time string values only for managed knowledge bases. -
For custom knowledge bases, metadata fields prefixed with
x-amz-bedrockare reserved by the service. For fully managed knowledge bases, reserved metadata fields use an underscore prefix (for example,_source_uri,_data_source_id). You cannot override reserved metadata fields in either knowledge base type.
For example, "epoch_modification_time" represents the time in number of seconds since January 1, 1970 (UTC) when the document was last updated. You can filter on the most recent data, where "epoch_modification_time" is greater than a certain number. These most recent documents can be used for the query.
To use filters when querying a knowledge base, check that your knowledge base fulfills the following requirements:
-
When configuring your data source connector, most connectors crawl the main metadata fields of your documents. If you're using an Amazon S3 bucket as your data source, the bucket must include at least one
fileName.extension.metadata.jsonfor the file or document it's associated with. See Document metadata fields in Connection configuration for more information about configuring the metadata file. -
If your knowledge base's vector index is in an Amazon OpenSearch Serverless vector store, check that the vector index is configured with the
faissengine. If the vector index is configured with thenmslibengine, you'll have to do one of the following:-
Create a new knowledge base in the console and let Amazon Bedrock automatically create a vector index in Amazon OpenSearch Serverless for you.
-
Create another vector index in the vector store and select
faissas the Engine. Then Create a new knowledge base and specify the new vector index.
-
-
If your knowledge base uses a vector index in an S3 vector bucket, you cannot use the
startsWithandstringContainsfilters. -
If you're adding metadata to an existing vector index in an Amazon Aurora database cluster, we recommend that you provide the field name of the custom metadata column to store all your metadata in a single column. During data ingestion, this column will be used to populate all the information in your metadata files from your data sources. If you choose to provide this field, you must create an index on this column.
-
When you create a new knowledge base in the console and let Amazon Bedrock configure your Amazon Aurora database, it will automatically create a single column for you and populate it with the information from your metadata files.
-
When you choose to create another vector index in the vector store, you must provide the custom metadata field name to store information from your metadata files. If you don't provide this field name, you must create a column for each metadata attribute in your files and specify the data type (text, number, or boolean). For example, if the attribute
genreexists in your data source, you would add a column namedgenreand specifytextas the data type. During ingestion, these separate columns will be populated with the corresponding attribute values.
-
If you have PDF documents in your data source and use Amazon OpenSearch Serverless or Amazon Aurora for your vector store: Amazon Bedrock knowledge bases will generate document page numbers and store them in a metadata field/attribute called x-amz-bedrock-kb-document-page-number. Note that page numbers stored in a metadata field is not supported if you choose no chunking for your documents.
You can use the following filtering operators to filter results when you query:
| Operator | Console | API filter name | Supported attribute data types | Filtered results |
|---|---|---|---|---|
| Equals | = | equals | string, number, boolean | Attribute matches the value you provide |
| Not equals | != | notEquals | string, number, boolean | Attribute doesn’t match the value you provide |
| Greater than | > | greaterThan | number | Attribute is greater than the value you provide |
| Greater than or equals | >= | greaterThanOrEquals | number | Attribute is greater than or equal to the value you provide |
| Less than | < | lessThan | number | Attribute is less than the value you provide |
| Less than or equals | <= | lessThanOrEquals | number | Attribute is less than or equal to the value you provide |
| In | : | in | string list | Attribute is in the list you provide (currently best supported with Amazon OpenSearch Serverless and Neptune Analytics GraphRAG vector stores) |
| Not in | !: | notIn | string list | Attribute isn’t in the list you provide (currently best supported with Amazon OpenSearch Serverless and Neptune Analytics GraphRAG vector stores) |
| String contains | Not available | stringContains | string | Attribute must be a string. Attribute name matches the key and whose value is a string that contains the value that you provided as a substring, or a list with a member that contains the value that you provided as a substring (currently best supported with Amazon OpenSearch Serverless vector store. The Neptune Analytics GraphRAG vector store supports the string variant but not the list variant of this filter). |
| List contains | Not available | listContains | string | Attribute must be a string list. Attribute name matches the key and whose value is a list that contains the value that you provided as one of its members (currently best supported with Amazon OpenSearch Serverless vector stores). |
To combine filtering operators, you can use the following logical operators:
To learn how to filter results using metadata, choose the tab for your preferred method, and then follow the steps:
You can implement safeguards for your knowledge base for your use cases and responsible AI policies. You can create multiple guardrails tailored to different use cases and apply them across multiple request and response conditions, providing a consistent user experience and standardizing safety controls across your knowledge base. You can configure denied topics to disallow undesirable topics and content filters to block harmful content in model inputs and responses. For more information, see Detect and filter harmful content by using Amazon Bedrock Guardrails.
Note
Using guardrails with contextual grounding for knowledge bases is currently not supported on Claude 3 Sonnet and Haiku.
For general prompt engineering guidelines, see Prompt engineering concepts.
Choose the tab for your preferred method, and then follow the steps:
You can use a reranker model to rerank results from knowledge base query. Follow the console steps at Query a knowledge base and retrieve data. When you open the Configurations pane, expand the Reranking section. Select a reranker model, update permissions if necessary, and modify any additional options. Enter a prompt and select Run to test the results after reranking.