OpenSearch/docs/reference/analysis/analyzers/pattern-analyzer.asciidoc

[[analysis-pattern-analyzer]]
=== Pattern Analyzer

The `pattern` analyzer uses a regular expression to split the text into terms.
The regular expression should match the *token separators*  not the tokens
themselves. The regular expression defaults to `\W+` (or all non-word characters).

[float]
=== Definition

It consists of:

Tokenizer::
* <<analysis-pattern-tokenizer,Pattern Tokenizer>>

Token Filters::
*  <<analysis-lowercase-tokenfilter,Lower Case Token Filter>>
*  <<analysis-stop-tokenfilter,Stop Token Filter>> (disabled by default)

[float]
=== Example output

[source,js]
---------------------------
POST _analyze
{
  "analyzer": "pattern",
  "text": "The 2 QUICK Brown-Foxes jumped over the lazy dog's bone."
}
---------------------------
// CONSOLE

/////////////////////

[source,js]
----------------------------
{
  "tokens": [
    {
      "token": "the",
      "start_offset": 0,
      "end_offset": 3,
      "type": "word",
      "position": 0
    },
    {
      "token": "2",
      "start_offset": 4,
      "end_offset": 5,
      "type": "word",
      "position": 1
    },
    {
      "token": "quick",
      "start_offset": 6,
      "end_offset": 11,
      "type": "word",
      "position": 2
    },
    {
      "token": "brown",
      "start_offset": 12,
      "end_offset": 17,
      "type": "word",
      "position": 3
    },
    {
      "token": "foxes",
      "start_offset": 18,
      "end_offset": 23,
      "type": "word",
      "position": 4
    },
    {
      "token": "jumped",
      "start_offset": 24,
      "end_offset": 30,
      "type": "word",
      "position": 5
    },
    {
      "token": "over",
      "start_offset": 31,
      "end_offset": 35,
      "type": "word",
      "position": 6
    },
    {
      "token": "the",
      "start_offset": 36,
      "end_offset": 39,
      "type": "word",
      "position": 7
    },
    {
      "token": "lazy",
      "start_offset": 40,
      "end_offset": 44,
      "type": "word",
      "position": 8
    },
    {
      "token": "dog",
      "start_offset": 45,
      "end_offset": 48,
      "type": "word",
      "position": 9
    },
    {
      "token": "s",
      "start_offset": 49,
      "end_offset": 50,
      "type": "word",
      "position": 10
    },
    {
      "token": "bone",
      "start_offset": 51,
      "end_offset": 55,
      "type": "word",
      "position": 11
    }
  ]
}
----------------------------
// TESTRESPONSE

/////////////////////


The above sentence would produce the following terms:

[source,text]
---------------------------
[ the, 2, quick, brown, foxes, jumped, over, the, lazy, dog, s, bone ]
---------------------------

[float]
=== Configuration

The `pattern` analyzer accepts the following parameters:

[horizontal]
`pattern`::

    A http://docs.oracle.com/javase/8/docs/api/java/util/regex/Pattern.html[Java regular expression], defaults to `\W+`.

`flags`::

    Java regular expression http://docs.oracle.com/javase/8/docs/api/java/util/regex/Pattern.html#field.summary[flags].
    Flags should be pipe-separated, eg `"CASE_INSENSITIVE|COMMENTS"`.

`lowercase`::

    Should terms be lowercased or not. Defaults to `true`.

`max_token_length`::

    The maximum token length. If a token is seen that exceeds this length then
    it is split at `max_token_length` intervals. Defaults to `255`.

`stopwords`::

    A pre-defined stop words list like `_english_` or an array  containing a
    list of stop words.  Defaults to `_none_`.

`stopwords_path`::

    The path to a file containing stop words.

See the <<analysis-stop-tokenfilter,Stop Token Filter>> for more information
about stop word configuration.


[float]
=== Example configuration

In this example, we configure the `pattern` analyzer to split email addresses
on non-word characters or on underscores (`\W|_`), and to lower-case the result:

[source,js]
----------------------------
PUT my_index
{
  "settings": {
    "analysis": {
      "analyzer": {
        "my_email_analyzer": {
          "type":      "pattern",
          "pattern":   "\\W|_", <1>
          "lowercase": true
        }
      }
    }
  }
}

POST my_index/_analyze
{
  "analyzer": "my_email_analyzer",
  "text": "John_Smith@foo-bar.com"
}
----------------------------
// CONSOLE

<1> The backslashes in the pattern need to be escaped when specifying the
    pattern as a JSON string.

/////////////////////

[source,js]
----------------------------
{
  "tokens": [
    {
      "token": "john",
      "start_offset": 0,
      "end_offset": 4,
      "type": "word",
      "position": 0
    },
    {
      "token": "smith",
      "start_offset": 5,
      "end_offset": 10,
      "type": "word",
      "position": 1
    },
    {
      "token": "foo",
      "start_offset": 11,
      "end_offset": 14,
      "type": "word",
      "position": 2
    },
    {
      "token": "bar",
      "start_offset": 15,
      "end_offset": 18,
      "type": "word",
      "position": 3
    },
    {
      "token": "com",
      "start_offset": 19,
      "end_offset": 22,
      "type": "word",
      "position": 4
    }
  ]
}
----------------------------
// TESTRESPONSE

/////////////////////


The above example produces the following terms:

[source,text]
---------------------------
[ john, smith, foo, bar, com ]
---------------------------

[float]
==== CamelCase tokenizer

The following more complicated example splits CamelCase text into tokens:

[source,js]
--------------------------------------------------
PUT my_index
{
  "settings": {
    "analysis": {
      "analyzer": {
        "camel": {
          "type": "pattern",
          "pattern": "([^\\p{L}\\d]+)|(?<=\\D)(?=\\d)|(?<=\\d)(?=\\D)|(?<=[\\p{L}&&[^\\p{Lu}]])(?=\\p{Lu})|(?<=\\p{Lu})(?=\\p{Lu}[\\p{L}&&[^\\p{Lu}]])"
        }
      }
    }
  }
}

GET my_index/_analyze
{
  "analyzer": "camel",
  "text": "MooseX::FTPClass2_beta"
}
--------------------------------------------------
// CONSOLE

/////////////////////

[source,js]
----------------------------
{
  "tokens": [
    {
      "token": "moose",
      "start_offset": 0,
      "end_offset": 5,
      "type": "word",
      "position": 0
    },
    {
      "token": "x",
      "start_offset": 5,
      "end_offset": 6,
      "type": "word",
      "position": 1
    },
    {
      "token": "ftp",
      "start_offset": 8,
      "end_offset": 11,
      "type": "word",
      "position": 2
    },
    {
      "token": "class",
      "start_offset": 11,
      "end_offset": 16,
      "type": "word",
      "position": 3
    },
    {
      "token": "2",
      "start_offset": 16,
      "end_offset": 17,
      "type": "word",
      "position": 4
    },
    {
      "token": "beta",
      "start_offset": 18,
      "end_offset": 22,
      "type": "word",
      "position": 5
    }
  ]
}
----------------------------
// TESTRESPONSE

/////////////////////


The above example produces the following terms:

[source,text]
---------------------------
[ moose, x, ftp, class, 2, beta ]
---------------------------

The regex above is easier to understand as:

[source,js]
--------------------------------------------------

  ([^\p{L}\d]+)                 # swallow non letters and numbers,
| (?<=\D)(?=\d)                 # or non-number followed by number,
| (?<=\d)(?=\D)                 # or number followed by non-number,
| (?<=[ \p{L} && [^\p{Lu}]])    # or lower case
  (?=\p{Lu})                    #   followed by upper case,
| (?<=\p{Lu})                   # or upper case
  (?=\p{Lu}                     #   followed by upper case
    [\p{L}&&[^\p{Lu}]]          #   then lower case
  )
--------------------------------------------------
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`[[analysis-pattern-analyzer]]`
			`=== Pattern Analyzer`

First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			The `pattern` analyzer uses a regular expression to split the text into terms.
			`The regular expression should match the token separators not the tokens`
			themselves. The regular expression defaults to `\W+` (or all non-word characters).
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`[float]`
			`=== Definition`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`It consists of:`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`Tokenizer::`
			`* <<analysis-pattern-tokenizer,Pattern Tokenizer>>`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`Token Filters::`
			`* <<analysis-lowercase-tokenfilter,Lower Case Token Filter>>`
			`* <<analysis-stop-tokenfilter,Stop Token Filter>> (disabled by default)`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[float]`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`=== Example output`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[source,js]`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`---------------------------`
			`POST _analyze`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`{`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`"analyzer": "pattern",`
			`"text": "The 2 QUICK Brown-Foxes jumped over the lazy dog's bone."`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`}`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`---------------------------`
			`// CONSOLE`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
Docs: Improved tokenizer docs (#18356) * Docs: Improved tokenizer docs Added descriptions and runnable examples * Addressed Nik's comments * Added TESTRESPONSEs for all tokenizer examples * Added TESTRESPONSEs for all analyzer examples too * Added docs, examples, and TESTRESPONSES for character filters * Skipping two tests: One interprets "$1" as a stack variable - same problem exists with the REST tests The other because the "took" value is always different * Fixed tests with "took" * Fixed failing tests and removed preserve_original from fingerprint analyzer 2016-05-19 19:42:23 +02:00			`/////////////////////`

			`[source,js]`
			`----------------------------`
			`{`
			`"tokens": [`
			`{`
			`"token": "the",`
			`"start_offset": 0,`
			`"end_offset": 3,`
			`"type": "word",`
			`"position": 0`
			`},`
			`{`
			`"token": "2",`
			`"start_offset": 4,`
			`"end_offset": 5,`
			`"type": "word",`
			`"position": 1`
			`},`
			`{`
			`"token": "quick",`
			`"start_offset": 6,`
			`"end_offset": 11,`
			`"type": "word",`
			`"position": 2`
			`},`
			`{`
			`"token": "brown",`
			`"start_offset": 12,`
			`"end_offset": 17,`
			`"type": "word",`
			`"position": 3`
			`},`
			`{`
			`"token": "foxes",`
			`"start_offset": 18,`
			`"end_offset": 23,`
			`"type": "word",`
			`"position": 4`
			`},`
			`{`
			`"token": "jumped",`
			`"start_offset": 24,`
			`"end_offset": 30,`
			`"type": "word",`
			`"position": 5`
			`},`
			`{`
			`"token": "over",`
			`"start_offset": 31,`
			`"end_offset": 35,`
			`"type": "word",`
			`"position": 6`
			`},`
			`{`
			`"token": "the",`
			`"start_offset": 36,`
			`"end_offset": 39,`
			`"type": "word",`
			`"position": 7`
			`},`
			`{`
			`"token": "lazy",`
			`"start_offset": 40,`
			`"end_offset": 44,`
			`"type": "word",`
			`"position": 8`
			`},`
			`{`
			`"token": "dog",`
			`"start_offset": 45,`
			`"end_offset": 48,`
			`"type": "word",`
			`"position": 9`
			`},`
			`{`
			`"token": "s",`
			`"start_offset": 49,`
			`"end_offset": 50,`
			`"type": "word",`
			`"position": 10`
			`},`
			`{`
			`"token": "bone",`
			`"start_offset": 51,`
			`"end_offset": 55,`
			`"type": "word",`
			`"position": 11`
			`}`
			`]`
			`}`
			`----------------------------`
			`// TESTRESPONSE`

			`/////////////////////`


First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`The above sentence would produce the following terms:`

			`[source,text]`
			`---------------------------`
			`[ the, 2, quick, brown, foxes, jumped, over, the, lazy, dog, s, bone ]`
			`---------------------------`

			`[float]`
			`=== Configuration`

			The `pattern` analyzer accepts the following parameters:

			`[horizontal]`
			`pattern`::

			A http://docs.oracle.com/javase/8/docs/api/java/util/regex/Pattern.html[Java regular expression], defaults to `\W+`.

			`flags`::

			`Java regular expression http://docs.oracle.com/javase/8/docs/api/java/util/regex/Pattern.html#field.summary[flags].`
[docs] s/lags/Flags/ Copy and paste lots an `F`. 2016-06-09 13:08:53 -04:00			Flags should be pipe-separated, eg `"CASE_INSENSITIVE\|COMMENTS"`.
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00
			`lowercase`::

			Should terms be lowercased or not. Defaults to `true`.

			`max_token_length`::

			`The maximum token length. If a token is seen that exceeds this length then`
			it is split at `max_token_length` intervals. Defaults to `255`.

			`stopwords`::

			A pre-defined stop words list like `_english_` or an array containing a
			list of stop words. Defaults to `_none_`.

			`stopwords_path`::

			`The path to a file containing stop words.`

			`See the <<analysis-stop-tokenfilter,Stop Token Filter>> for more information`
			`about stop word configuration.`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[float]`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`=== Example configuration`

			In this example, we configure the `pattern` analyzer to split email addresses
			on non-word characters or on underscores (`\W\|_`), and to lower-case the result:
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[source,js]`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`----------------------------`
			`PUT my_index`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`{`
			`"settings": {`
			`"analysis": {`
			`"analyzer": {`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`"my_email_analyzer": {`
			`"type": "pattern",`
			`"pattern": "\\W\|_", <1>`
			`"lowercase": true`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`}`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`}`
			`}`
			`}`
			`}`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`POST my_index/_analyze`
			`{`
			`"analyzer": "my_email_analyzer",`
			`"text": "John_Smith@foo-bar.com"`
			`}`
			`----------------------------`
Renamed all AUTOSENSE snippets to CONSOLE (#18210) 2016-05-09 15:42:23 +02:00			`// CONSOLE`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`<1> The backslashes in the pattern need to be escaped when specifying the`
			`pattern as a JSON string.`

Docs: Improved tokenizer docs (#18356) * Docs: Improved tokenizer docs Added descriptions and runnable examples * Addressed Nik's comments * Added TESTRESPONSEs for all tokenizer examples * Added TESTRESPONSEs for all analyzer examples too * Added docs, examples, and TESTRESPONSES for character filters * Skipping two tests: One interprets "$1" as a stack variable - same problem exists with the REST tests The other because the "took" value is always different * Fixed tests with "took" * Fixed failing tests and removed preserve_original from fingerprint analyzer 2016-05-19 19:42:23 +02:00			`/////////////////////`

			`[source,js]`
			`----------------------------`
			`{`
			`"tokens": [`
			`{`
			`"token": "john",`
			`"start_offset": 0,`
			`"end_offset": 4,`
			`"type": "word",`
			`"position": 0`
			`},`
			`{`
			`"token": "smith",`
			`"start_offset": 5,`
			`"end_offset": 10,`
			`"type": "word",`
			`"position": 1`
			`},`
			`{`
			`"token": "foo",`
			`"start_offset": 11,`
			`"end_offset": 14,`
			`"type": "word",`
			`"position": 2`
			`},`
			`{`
			`"token": "bar",`
			`"start_offset": 15,`
			`"end_offset": 18,`
			`"type": "word",`
			`"position": 3`
			`},`
			`{`
			`"token": "com",`
			`"start_offset": 19,`
			`"end_offset": 22,`
			`"type": "word",`
			`"position": 4`
			`}`
			`]`
			`}`
			`----------------------------`
			`// TESTRESPONSE`

			`/////////////////////`


First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`The above example produces the following terms:`

			`[source,text]`
			`---------------------------`
			`[ john, smith, foo, bar, com ]`
			`---------------------------`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[float]`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`==== CamelCase tokenizer`

			`The following more complicated example splits CamelCase text into tokens:`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
			`[source,js]`
			`--------------------------------------------------`
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`PUT my_index`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`{`
			`"settings": {`
			`"analysis": {`
			`"analyzer": {`
			`"camel": {`
			`"type": "pattern",`
			`"pattern": "([^\\p{L}\\d]+)\|(?<=\\D)(?=\\d)\|(?<=\\d)(?=\\D)\|(?<=[\\p{L}&&[^\\p{Lu}]])(?=\\p{Lu})\|(?<=\\p{Lu})(?=\\p{Lu}[\\p{L}&&[^\\p{Lu}]])"`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`}`
Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`}`
			`}`
			`}`
			`}`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`GET my_index/_analyze`
			`{`
			`"analyzer": "camel",`
			`"text": "MooseX::FTPClass2_beta"`
			`}`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`--------------------------------------------------`
Renamed all AUTOSENSE snippets to CONSOLE (#18210) 2016-05-09 15:42:23 +02:00			`// CONSOLE`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00
Docs: Improved tokenizer docs (#18356) * Docs: Improved tokenizer docs Added descriptions and runnable examples * Addressed Nik's comments * Added TESTRESPONSEs for all tokenizer examples * Added TESTRESPONSEs for all analyzer examples too * Added docs, examples, and TESTRESPONSES for character filters * Skipping two tests: One interprets "$1" as a stack variable - same problem exists with the REST tests The other because the "took" value is always different * Fixed tests with "took" * Fixed failing tests and removed preserve_original from fingerprint analyzer 2016-05-19 19:42:23 +02:00			`/////////////////////`

			`[source,js]`
			`----------------------------`
			`{`
			`"tokens": [`
			`{`
			`"token": "moose",`
			`"start_offset": 0,`
			`"end_offset": 5,`
			`"type": "word",`
			`"position": 0`
			`},`
			`{`
			`"token": "x",`
			`"start_offset": 5,`
			`"end_offset": 6,`
			`"type": "word",`
			`"position": 1`
			`},`
			`{`
			`"token": "ftp",`
			`"start_offset": 8,`
			`"end_offset": 11,`
			`"type": "word",`
			`"position": 2`
			`},`
			`{`
			`"token": "class",`
			`"start_offset": 11,`
			`"end_offset": 16,`
			`"type": "word",`
			`"position": 3`
			`},`
			`{`
			`"token": "2",`
			`"start_offset": 16,`
			`"end_offset": 17,`
			`"type": "word",`
			`"position": 4`
			`},`
			`{`
			`"token": "beta",`
			`"start_offset": 18,`
			`"end_offset": 22,`
			`"type": "word",`
			`"position": 5`
			`}`
			`]`
			`}`
			`----------------------------`
			`// TESTRESPONSE`

			`/////////////////////`


First pass at improving analyzer docs (#18269) * Docs: First pass at improving analyzer docs I've rewritten the intro to analyzers plus the docs for all analyzers to provide working examples. I've also removed: * analyzer aliases (see #18244) * analyzer versions (see #18267) * snowball analyzer (see #8690) Next steps will be tokenizers, token filters, char filters * Fixed two typos 2016-05-11 14:17:56 +02:00			`The above example produces the following terms:`

			`[source,text]`
			`---------------------------`
			`[ moose, x, ftp, class, 2, beta ]`
			`---------------------------`

Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`The regex above is easier to understand as:`

			`[source,js]`
			`--------------------------------------------------`

Docs: Fixed the backslash escaping on the pattern analyzer docs Closes #11099 2015-05-15 18:40:16 +02:00			`([^\p{L}\d]+) # swallow non letters and numbers,`
			`\| (?<=\D)(?=\d) # or non-number followed by number,`
			`\| (?<=\d)(?=\D) # or number followed by non-number,`
			`\| (?<=[ \p{L} && [^\p{Lu}]]) # or lower case`
			`(?=\p{Lu}) # followed by upper case,`
			`\| (?<=\p{Lu}) # or upper case`
			`(?=\p{Lu} # followed by upper case`
			`[\p{L}&&[^\p{Lu}]] # then lower case`
			`)`
Migrated documentation into the main repo 2013-08-29 01:24:34 +02:00			`--------------------------------------------------`