OpenSearch/docs/reference/search/suggesters/term-suggest.asciidoc

[[term-suggester]]
=== Term suggester

NOTE: In order to understand the format of suggestions, please
read the <<search-suggesters>> page first.

The `term` suggester suggests terms based on edit distance. The provided
suggest text is analyzed before terms are suggested. The suggested terms
are provided per analyzed suggest text token. The `term` suggester
doesn't take the query into account that is part of request.

==== Common suggest options:

[horizontal]
`text`::
    The suggest text. The suggest text is a required option that
    needs to be set globally or per suggestion.

`field`::
    The field to fetch the candidate suggestions from. This is
    a required option that either needs to be set globally or per
    suggestion.

`analyzer`::
    The analyzer to analyse the suggest text with. Defaults
    to the search analyzer of the suggest field.

`size`::
    The maximum corrections to be returned per suggest text
    token.

`sort`::
    Defines how suggestions should be sorted per suggest text
    term. Two possible values:
+
    ** `score`:     Sort by score first, then document frequency and
                    then the term itself.
    ** `frequency`: Sort by document frequency first, then similarity
                    score and then the term itself.
+
`suggest_mode`::
    The suggest mode controls what suggestions are
    included or controls for what suggest text terms, suggestions should be
    suggested. Three possible values can be specified:
+
     ** `missing`:  Only provide suggestions for suggest text terms that are
                    not in the index. This is the default.
     ** `popular`:  Only suggest suggestions that occur in more docs than
                    the original suggest text term.
     ** `always`:   Suggest any matching suggestions based on terms in the
                    suggest text.

==== Other term suggest options:

[horizontal]
`lowercase_terms`::
    Lowercases the suggest text terms after text analysis.

`max_edits`::
    The maximum edit distance candidate suggestions can
    have in order to be considered as a suggestion. Can only be a value
    between 1 and 2. Any other value results in a bad request error being
    thrown. Defaults to 2.

`prefix_length`::
    The number of minimal prefix characters that must
    match in order be a candidate for suggestions. Defaults to 1. Increasing
    this number improves spellcheck performance. Usually misspellings don't
    occur in the beginning of terms. (Old name "prefix_len" is deprecated)

`min_word_length`::
    The minimum length a suggest text term must have in
    order to be included. Defaults to 4. (Old name "min_word_len" is deprecated)

`shard_size`::
    Sets the maximum number of suggestions to be retrieved
    from each individual shard. During the reduce phase only the top N
    suggestions are returned based on the `size` option. Defaults to the
    `size` option. Setting this to a value higher than the `size` can be
    useful in order to get a more accurate document frequency for spelling
    corrections at the cost of performance. Due to the fact that terms are
    partitioned amongst shards, the shard level document frequencies of
    spelling corrections may not be precise. Increasing this will make these
    document frequencies more precise.

`max_inspections`::
    A factor that is used to multiply with the
    `shards_size` in order to inspect more candidate spelling corrections on
    the shard level. Can improve accuracy at the cost of performance.
    Defaults to 5.

`min_doc_freq`::
    The minimal threshold in number of documents a
    suggestion should appear in. This can be specified as an absolute number
    or as a relative percentage of number of documents. This can improve
    quality by only suggesting high frequency terms. Defaults to 0f and is
    not enabled. If a value higher than 1 is specified, then the number
    cannot be fractional. The shard level document frequencies are used for
    this option.

`max_term_freq`::
    The maximum threshold in number of documents in which a
    suggest text token can exist in order to be included. Can be a relative
    percentage number (e.g., 0.4) or an absolute number to represent document
    frequencies. If a value higher than 1 is specified, then fractional can
    not be specified. Defaults to 0.01f. This can be used to exclude high
    frequency terms -- which are usually spelled correctly -- from being spellchecked.
    This also improves the spellcheck performance. The shard level document frequencies
    are used for this option.

`string_distance`::
    Which string distance implementation to use for comparing how similar
    suggested terms are. Five possible values can be specified:
    
    ** `internal`: The default based on damerau_levenshtein but highly optimized
    for comparing string distance for terms inside the index.
    ** `damerau_levenshtein`: String distance algorithm based on
    Damerau-Levenshtein algorithm.
    ** `levenshtein`: String distance algorithm based on Levenshtein edit distance
    algorithm.
    ** `jaro_winkler`: String distance algorithm based on Jaro-Winkler algorithm.
    ** `ngram`: String distance algorithm based on character n-grams.
[DOCS] Move Elasticsearch APIs to REST APIs section. (#44238) (#44372) Moves the following API sections under the REST APIs navigations: - API Conventions - Document APIs - Search APIs - Index APIs (previously named Indices APIs) - cat APIs - Cluster APIs Other supporting changes: - Removes the previous index APIs page under REST APIs. Adds a redirect for the removed page. - Removes several [partintro] macros so the docs build correctly. - Changes anchors for pages that become sections of a parent page. - Adds several redirects for existing pages that become sections of a parent page. This commit re-applies changes from #44238. Changes from that PR were reverted due to broken links in several repos. This commit adds redirects for those broken links. 2019-07-17 08:49:22 -04:00			`[[term-suggester]]`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`=== Term suggester`

			`NOTE: In order to understand the format of suggestions, please`
			`read the <<search-suggesters>> page first.`

			The `term` suggester suggests terms based on edit distance. The provided
			`suggest text is analyzed before terms are suggested. The suggested terms`
			are provided per analyzed suggest text token. The `term` suggester
			`doesn't take the query into account that is part of request.`

[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`==== Common suggest options:`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
			`[horizontal]`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`text`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The suggest text. The suggest text is a required option that`
			`needs to be set globally or per suggestion.`

[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`field`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The field to fetch the candidate suggestions from. This is`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`a required option that either needs to be set globally or per`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`suggestion.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`analyzer`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The analyzer to analyse the suggest text with. Defaults`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`to the search analyzer of the suggest field.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`size`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The maximum corrections to be returned per suggest text`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`token.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`sort`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`Defines how suggestions should be sorted per suggest text`
			`term. Two possible values:`
			`+`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			** `score`: Sort by score first, then document frequency and
			`then the term itself.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			** `frequency`: Sort by document frequency first, then similarity
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`score and then the term itself.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`+`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`suggest_mode`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The suggest mode controls what suggestions are`
			`included or controls for what suggest text terms, suggestions should be`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`suggested. Three possible values can be specified:`
			`+`
Clarify `missing` behavior. 2014-05-05 09:53:37 -04:00			** `missing`: Only provide suggestions for suggest text terms that are
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`not in the index. This is the default.`
then -> than (#21829) 2016-11-28 11:04:24 -05:00			** `popular`: Only suggest suggestions that occur in more docs than
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`the original suggest text term.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			** `always`: Suggest any matching suggestions based on terms in the
			`suggest text.`

[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`==== Other term suggest options:`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
			`[horizontal]`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`lowercase_terms`::
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`Lowercases the suggest text terms after text analysis.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`max_edits`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The maximum edit distance candidate suggestions can`
			`have in order to be considered as a suggestion. Can only be a value`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`between 1 and 2. Any other value results in a bad request error being`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`thrown. Defaults to 2.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`prefix_length`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The number of minimal prefix characters that must`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`match in order be a candidate for suggestions. Defaults to 1. Increasing`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`this number improves spellcheck performance. Usually misspellings don't`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`occur in the beginning of terms. (Old name "prefix_len" is deprecated)`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`min_word_length`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The minimum length a suggest text term must have in`
Standardized use of “_length” for parameter names rather than “_len”. Java Builder apis drop old “len” methods in favour of new “length” Rest APIs support both old “len: and new “length” forms using new ParseField class to a) provide compiler-checked consistency between Builder and Parser classes and b) a common means of handling deprecated syntax in the DSL. Documentation and rest specs only document the new “*length” forms Closes #4083 2014-01-02 11:11:20 -05:00			`order to be included. Defaults to 4. (Old name "min_word_len" is deprecated)`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`shard_size`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`Sets the maximum number of suggestions to be retrieved`
			`from each individual shard. During the reduce phase only the top N`
			suggestions are returned based on the `size` option. Defaults to the
			`size` option. Setting this to a value higher than the `size` can be
			`useful in order to get a more accurate document frequency for spelling`
			`corrections at the cost of performance. Due to the fact that terms are`
			`partitioned amongst shards, the shard level document frequencies of`
			`spelling corrections may not be precise. Increasing this will make these`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`document frequencies more precise.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`max_inspections`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`A factor that is used to multiply with the`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`shards_size` in order to inspect more candidate spelling corrections on
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`the shard level. Can improve accuracy at the cost of performance.`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`Defaults to 5.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`min_doc_freq`::
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`The minimal threshold in number of documents a`
			`suggestion should appear in. This can be specified as an absolute number`
			`or as a relative percentage of number of documents. This can improve`
			`quality by only suggesting high frequency terms. Defaults to 0f and is`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`not enabled. If a value higher than 1 is specified, then the number`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`cannot be fractional. The shard level document frequencies are used for`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`this option.`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`max_term_freq`::
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`The maximum threshold in number of documents in which a`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`suggest text token can exist in order to be included. Can be a relative`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`percentage number (e.g., 0.4) or an absolute number to represent document`
			`frequencies. If a value higher than 1 is specified, then fractional can`
Migrated documentation into the main repo 2013-08-28 19:24:34 -04:00			`not be specified. Defaults to 0.01f. This can be used to exclude high`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			`frequency terms -- which are usually spelled correctly -- from being spellchecked.`
			`This also improves the spellcheck performance. The shard level document frequencies`
			`are used for this option.`
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00
			`string_distance`::
			`Which string distance implementation to use for comparing how similar`
Fix typos in docs. 2016-02-09 05:07:32 -05:00			`suggested terms are. Five possible values can be specified:`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00
			** `internal`: The default based on damerau_levenshtein but highly optimized
Fix typos in docs. 2016-02-09 05:07:32 -05:00			`for comparing string distance for terms inside the index.`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			** `damerau_levenshtein`: String distance algorithm based on
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`Damerau-Levenshtein algorithm.`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			** `levenshtein`: String distance algorithm based on Levenshtein edit distance
[DOCS] Document the `string_distance` parameter for term suggestor 2016-01-21 12:00:46 -05:00			`algorithm.`
Edits to text & formatting in Term Suggester doc (#38963) 2019-02-15 15:51:33 -05:00			** `jaro_winkler`: String distance algorithm based on Jaro-Winkler algorithm.
			** `ngram`: String distance algorithm based on character n-grams.