OpenSearch/docs/plugins/analysis-phonetic.asciidoc

[[analysis-phonetic]]
=== Phonetic Analysis Plugin

The Phonetic Analysis plugin provides token filters which convert tokens to
their phonetic representation using Soundex, Metaphone, and a variety of other
algorithms.

:plugin_name: analysis-phonetic
include::install_remove.asciidoc[]


[[analysis-phonetic-token-filter]]
==== `phonetic` token filter

The `phonetic` token filter takes the following settings:

`encoder`::

    Which phonetic encoder to use.  Accepts `metaphone` (default),
    `double_metaphone`, `soundex`, `refined_soundex`, `caverphone1`,
    `caverphone2`, `cologne`, `nysiis`, `koelnerphonetik`, `haasephonetik`,
    `beider_morse`, `daitch_mokotoff`.

`replace`::

    Whether or not the original token should be replaced by the phonetic
    token. Accepts `true` (default) and `false`.  Not supported by
    `beider_morse` encoding.

[source,console]
--------------------------------------------------
PUT phonetic_sample
{
  "settings": {
    "index": {
      "analysis": {
        "analyzer": {
          "my_analyzer": {
            "tokenizer": "standard",
            "filter": [
              "lowercase",
              "my_metaphone"
            ]
          }
        },
        "filter": {
          "my_metaphone": {
            "type": "phonetic",
            "encoder": "metaphone",
            "replace": false
          }
        }
      }
    }
  }
}

GET phonetic_sample/_analyze
{
  "analyzer": "my_analyzer",
  "text": "Joe Bloggs" <1>
}
--------------------------------------------------

<1> Returns: `J`, `joe`, `BLKS`, `bloggs`

It is important to note that `"replace": false` can lead to unexpected behavior since
the original and the phonetically analyzed version are both kept at the same token position.
Some queries handle these stacked tokens in special ways. For example, the fuzzy `match`
query does not apply {ref}/common-options.html#fuzziness[fuzziness] to stacked synonym tokens.
This can lead to issues that are difficult to diagnose and reason about. For this reason, it
is often beneficial to use separate fields for analysis with and without phonetic filtering.
That way searches can be run against both fields with differing boosts and trade-offs (e.g.
only run a fuzzy `match` query on the original text field, but not on the phonetic version).

[float]
===== Double metaphone settings

If the `double_metaphone` encoder is used, then this additional setting is
supported:

`max_code_len`::

    The maximum length of the emitted metaphone token.  Defaults to `4`.

[float]
===== Beider Morse settings

If the `beider_morse` encoder is used, then these additional settings are
supported:

`rule_type`::

    Whether matching should be `exact` or `approx` (default).

`name_type`::

    Whether names are `ashkenazi`, `sephardic`, or `generic` (default).

`languageset`::

    An array of languages to check. If not specified, then the language will
    be guessed. Accepts: `any`, `common`, `cyrillic`, `english`, `french`,
    `german`, `hebrew`, `hungarian`, `polish`, `romanian`, `russian`,
    `spanish`.
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`[[analysis-phonetic]]`
			`=== Phonetic Analysis Plugin`

			`The Phonetic Analysis plugin provides token filters which convert tokens to`
			`their phonetic representation using Soundex, Metaphone, and a variety of other`
			`algorithms.`

Added "release-state" support to plugin docs 2017-04-20 09:01:37 -04:00			`:plugin_name: analysis-phonetic`
			`include::install_remove.asciidoc[]`
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00

			`[[analysis-phonetic-token-filter]]`
			==== `phonetic` token filter

			The `phonetic` token filter takes the following settings:

			`encoder`::

			Which phonetic encoder to use. Accepts `metaphone` (default),
Consistent encoder names (#29492) This commit updates encoder names to be consistent within documentation and align with snake casing convention. 2018-07-23 19:21:43 -04:00			`double_metaphone`, `soundex`, `refined_soundex`, `caverphone1`,
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`caverphone2`, `cologne`, `nysiis`, `koelnerphonetik`, `haasephonetik`,
Consistent encoder names (#29492) This commit updates encoder names to be consistent within documentation and align with snake casing convention. 2018-07-23 19:21:43 -04:00			`beider_morse`, `daitch_mokotoff`.
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00
			`replace`::

			`Whether or not the original token should be replaced by the phonetic`
			token. Accepts `true` (default) and `false`. Not supported by
Consistent encoder names (#29492) This commit updates encoder names to be consistent within documentation and align with snake casing convention. 2018-07-23 19:21:43 -04:00			`beider_morse` encoding.
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00
[DOCS] [2 of 5] Change // CONSOLE comments to [source,console] (#46353) (#46502) 2019-09-09 13:38:14 -04:00			`[source,console]`
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`--------------------------------------------------`
Remove `include_type_name` in asciidoc where possible (#37568) The "include_type_name" parameter was temporarily introduced in #37285 to facilitate moving the default parameter setting to "false" in many places in the documentation code snippets. Most of the places can simply be reverted without causing errors. In this change I looked for asciidoc files that contained the "include_type_name=true" addition when creating new indices but didn't look likey they made use of the "_doc" type for mappings. This is mostly the case e.g. in the analysis docs where index creating often only contains settings. I manually corrected the use of types in some places where the docs still used an explicit type name and not the dummy "_doc" type. 2019-01-18 03:34:11 -05:00			`PUT phonetic_sample`
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`{`
			`"settings": {`
			`"index": {`
			`"analysis": {`
			`"analyzer": {`
			`"my_analyzer": {`
			`"tokenizer": "standard",`
			`"filter": [`
			`"lowercase",`
			`"my_metaphone"`
			`]`
			`}`
			`},`
			`"filter": {`
			`"my_metaphone": {`
			`"type": "phonetic",`
			`"encoder": "metaphone",`
			`"replace": false`
			`}`
			`}`
			`}`
			`}`
			`}`
			`}`

Removing request parameters in _analyze API Remove unused imports Replace POST method by GET method in docs Add breaking changes explanation Fix small issue in Kuromoji docs Closes #20246 2016-09-30 16:42:45 -04:00			`GET phonetic_sample/_analyze`
Removing request parameters in _analyze API Remove request params in _analyze API without index param Change rest-api-test using JSON Change docs using JSON Closes #20246 2016-09-22 07:54:30 -04:00			`{`
			`"analyzer": "my_analyzer",`
			`"text": "Joe Bloggs" <1>`
			`}`
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`--------------------------------------------------`

			<1> Returns: `J`, `joe`, `BLKS`, `bloggs`

[Docs] Clarify caveats for phonetic filters replace option (#42807) The `replace` option in the phonetic token filter can have suprising side effects, e.g. such as described in #26921. This PR adds a note to be mindful about such scenarios and offers alternatives to using the `replace` option. Closes #26921 2019-06-05 16:02:17 -04:00			It is important to note that `"replace": false` can lead to unexpected behavior since
			`the original and the phonetically analyzed version are both kept at the same token position.`
			Some queries handle these stacked tokens in special ways. For example, the fuzzy `match`
			`query does not apply {ref}/common-options.html#fuzziness[fuzziness] to stacked synonym tokens.`
			`This can lead to issues that are difficult to diagnose and reason about. For this reason, it`
			`is often beneficial to use separate fields for analysis with and without phonetic filtering.`
			`That way searches can be run against both fields with differing boosts and trade-offs (e.g.`
			only run a fuzzy `match` query on the original text field, but not on the phonetic version).
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00
			`[float]`
			`===== Double metaphone settings`

			If the `double_metaphone` encoder is used, then this additional setting is
			`supported:`

			`max_code_len`::

			The maximum length of the emitted metaphone token. Defaults to `4`.

			`[float]`
			`===== Beider Morse settings`

			If the `beider_morse` encoder is used, then these additional settings are
			`supported:`

			`rule_type`::

			Whether matching should be `exact` or `approx` (default).

			`name_type`::

			Whether names are `ashkenazi`, `sephardic`, or `generic` (default).

			`languageset`::

			`An array of languages to check. If not specified, then the language will`
[DOCS] Various spelling corrections (#37046) 2019-01-07 08:44:12 -05:00			be guessed. Accepts: `any`, `common`, `cyrillic`, `english`, `french`,
Docs: Prepare plugin and integration docs for 2.0 * Centralised plugin docs in docs/plugins/ * Moved integrations into same docs * Moved community clients into the clients section of the docs * Removed docs/community Closes #11734 Closes #11724 Closes #11636 Closes #11635 Closes #11632 Closes #11630 Closes #12046 Closes #12438 Closes #12579 2015-08-15 12:00:55 -04:00			`german`, `hebrew`, `hungarian`, `polish`, `romanian`, `russian`,
			`spanish`.