Commit Graph

205 Commits

Author SHA1 Message Date
Martijn van Groningen 98a674fc6e Added suggest api.
# Suggest feature
The suggest feature suggests similar looking terms based on a provided text by using a suggester. At the moment there the only supported suggester is `fuzzy`. The suggest feature is available since version `0.21.0`.

# Fuzzy suggester
The `fuzzy` suggester suggests terms based on edit distance. The provided suggest text is analyzed before terms are suggested. The suggested terms are provided per analyzed suggest text token. The `fuzzy` suggester doesn't take the query into account that is part of request.

# Suggest API
The suggest request part is defined along side the query part as top field in the json request.

```
curl -s -XPOST 'localhost:9200/_search' -d '{
    "query" : {
        ...
    },
    "suggest" : {
        ...
    }
}'
```

Several suggestions can be specified per request. Each suggestion is identified with an arbitary name. In the example below two suggestions are requested. The `my-suggest-1` suggestion uses the `body` field and `my-suggest-2` uses the `title` field. The `type` field is a required field and defines what suggester to use for a suggestion.

```
"suggest" : {
    "suggestions" : {
        "my-suggest-1" : {
            "type" : "fuzzy",
            "field" : "body",
            "text" : "the amsterdma meetpu"
        },
        "my-suggest-2" : {
            "type" : "fuzzy",
            "field" : "title",
            "text" : "the rottredam meetpu"
        }
    }
}
```

The below suggest response example includes the suggestions part for `my-suggest-1` and `my-suggest-2`. Each suggestion part contains a terms array, that contains all terms outputted by the analyzed suggest text. Each term object includes the term itself, the original start and end offset in the suggest text and if found an arbitary number of suggestions.

```
{
    ...
    "suggest": {
        "my-suggest-1": {
            "terms" : [
              {
                "term" : "amsterdma",
                "start_offset": 5,
                "end_offset": 14,
                "suggestions": [
                   ...
                ]
              }
              ...
            ]
        },
        "my-suggest-2" : {
          "terms" : [
            ...
          ]
        }
    }
```

Each suggestions array contains a suggestion object that includes the suggested term, its document frequency and score compared to the suggest text term. The meaning of the score depends on the used suggester. The fuzzy suggester's score is based on the edit distance.

```
"suggestions": [
    {
        "term": "amsterdam",
        "frequency": 77,
        "score": 0.8888889
    },
    ...
]
```

# Global suggest text

To avoid repitition of the suggest text, it is possible to define a global text. In the example below the suggest text is a global option and applies to the `my-suggest-1` and `my-suggest-2` suggestions.

```
"suggest" : {
    "suggestions" : {
        "text" : "the amsterdma meetpu",
        "my-suggest-1" : {
            "type" : "fuzzy",
            "field" : "title"
        },
        "my-suggest-2" : {
            "type" : "fuzzy",
            "field" : "body"
        }
    }
}
```

The suggest text can be specied as global option or as suggestion specific option. The suggest text specified on suggestion level override the suggest text on the global level.

# Other suggest example.

In the below example we request suggestions for the following suggest text: `devloping distibutd saerch engies` on the `title` field with a maximum of 3 suggestions per term inside the suggest text. Note that in this example we use the `count` search type. This isn't required, but a nice optimalization. The suggestions are gather in the `query` phase and in the case that we only care about suggestions (so no hits) we don't need to execute the `fetch` phase.

```
curl -s -XPOST 'localhost:9200/_search?search_type=count' -d '{
  "suggest" : {
      "suggestions" : {
        "my-title-suggestions" : {
          "suggester" : "fuzzy",
          "field" : "title",
          "text" : "devloping distibutd saerch engies",
          "size" : 3
        }
      }
  }
}'
```

The above request could yield the response as stated in the code example below. As you can see if we take the first suggested term of each suggest text term we get `developing distributed search engines` as result.

```
{
  ...
  "suggest": {
    "my-title-suggestions": {
      "terms": [
        {
          "term": "devloping",
          "start_offset": 0,
          "end_offset": 9,
          "suggestions": [
            {
              "term": "developing",
              "frequency": 77,
              "score": 0.8888889
            },
            {
              "term": "deloping",
              "frequency": 1,
              "score": 0.875
            },
            {
              "term": "deploying",
              "frequency": 2,
              "score": 0.7777778
            }
          ]
        },
        {
          "term": "distibutd",
          "start_offset": 10,
          "end_offset": 19,
          "suggestions": [
            {
              "term": "distributed",
              "frequency": 217,
              "score": 0.7777778
            },
            {
              "term": "disributed",
              "frequency": 1,
              "score": 0.7777778
            },
            {
              "term": "distribute",
              "frequency": 1,
              "score": 0.7777778
            }
          ]
        },
        {
          "term": "saerch",
          "start_offset": 20,
          "end_offset": 26,
          "suggestions": [
            {
              "term": "search",
              "frequency": 1038,
              "score": 0.8333333
            },
            {
              "term": "smerch",
              "frequency": 3,
              "score": 0.8333333
            },
            {
              "term": "serch",
              "frequency": 2,
              "score": 0.8
            }
          ]
        },
        {
          "term": "engies",
          "start_offset": 27,
          "end_offset": 33,
          "suggestions": [
            {
              "term": "engines",
              "frequency": 568,
              "score": 0.8333333
            },
            {
              "term": "engles",
              "frequency": 3,
              "score": 0.8333333
            },
            {
              "term": "eggies",
              "frequency": 1,
              "score": 0.8333333
            }
          ]
        }
      ]
    }
  }
  ...
}
```

# Common suggest options:
* `suggester` - The suggester implementation type. The only supported value is 'fuzzy'. This is a required option.
* `text` - The suggest text. The suggest text is a required option that needs to be set globally or per suggestion.

# Common fuzzy suggest options
* `field` - The field to fetch the candidate suggestions from. This is an required option that either needs to be set globally or per suggestion.
* `analyzer` - The analyzer to analyse the suggest text with. Defaults to the search analyzer of the suggest field.
* `size` - The maximum corrections to be returned per suggest text token.
* `sort` - Defines how suggestions should be sorted per suggest text term. Two possible value:
** `score` - Sort by sore first, then document frequency and then the term itself.
** `frequency` - Sort by document frequency first, then simlarity score and then the term itself.
* `suggest_mode` - The suggest mode controls what suggestions are included or controls for what suggest text terms, suggestions should be suggested. Three possible values can be specified:
** `missing` - Only suggest terms in the suggest text that aren't in the index. This is the default.
** `popular` - Only suggest suggestions that occur in more docs then the original suggest text term.
** `always` - Suggest any matching suggestions based on terms in the suggest text.

# Other fuzzy suggest options:
* `lowercase_terms` - Lower cases the suggest text terms after text analyzation.
* `max_edits` - The maximum edit distance candidate suggestions can have in order to be considered as a suggestion. Can only be a value between 1 and 2. Any other value result in an bad request error being thrown. Defaults to 2.
* `min_prefix` - The number of minimal prefix characters that must match in order be a candidate suggestions. Defaults to 1. Increasing this number improves spellcheck performance. Usually misspellings don't occur in the beginning of terms.
* `min_query_length` -  The minimum length a suggest text term must have in order to be included. Defaults to 4.
* `shard_size` - Sets the maximum number of suggestions to be retrieved from each individual shard. During the reduce phase only the top N suggestions are returned based on the `size` option. Defaults to the `size` option. Setting this to a value higher than the `size` can be useful in order to get a more accurate document frequency for spelling corrections at the cost of performance. Due to the fact that terms are partitioned amongst shards, the shard level document frequencies of spelling corrections may not be precise. Increasing this will make these document frequencies more precise.
* `max_inspections` - A factor that is used to multiply with the `shards_size` in order to inspect more candidate spell corrections on the shard level. Can improve accuracy at the cost of performance. Defaults to 5.
* `threshold_frequency` - The minimal threshold in number of documents a suggestion should appear in. This can be specified as an absolute number or as a relative percentage of number of documents. This can improve quality by only suggesting high frequency terms. Defaults to 0f and is not enabled. If a value higher than 1 is specified then the number cannot be fractional. The shard level document frequencies are used for this option.
* `max_query_frequency` - The maximum threshold in number of documents a sugges text token can exist in order to be included. Can be a relative percentage number (e.g 0.4) or an absolute number to represent document frequencies. If an value higher than 1 is specified then fractional can not be specified. Defaults to 0.01f. This can be used to exclude high frequency terms from being spellchecked. High frequency terms are usually spelled correctly on top of this this also improves the spellcheck performance.  The shard level document frequencies are used for this option.

 Closes 
2013-01-24 15:41:06 +01:00
Simon Willnauer 2880cd0172 Upgrade to Lucene 4.1
* Removed CustmoMemoryIndex in favor of MemoryIndex which as of 4.1 supports adding the same field twice
* Replaced duplicated logic in X[*]FSDirectory for rate limiting with a RateLimitedFSDirectory wrapper
* Remove hacks to find out merge context in rate limiting in favor of IOContext
* replaced Scorer#freq() return type (from float to int)
* Upgraded FVHighlighter to new 'centered' highlighting
* Fixed RobinEngine to use seperate setCommitData
2013-01-23 11:54:11 +01:00
Shay Banon 468295dc37 upgrade to netty 3.6.2 2013-01-18 12:47:13 +01:00
Shay Banon 7726a4a9dd upgrade to netty 3.6.1 2013-01-04 16:06:15 +01:00
Jilles van Gurp 2649c6c758 add plugin configuration to make m2e ignore maven dependency plugin configuration that it cannot handle 2013-01-04 08:23:30 +01:00
Shay Banon bab91bb9fd upgrade to jackson 2.1.1 2012-12-27 12:03:59 -08:00
Shay Banon d739498498 upgrade to netty 3.6.0 2012-12-27 11:55:37 -08:00
Shay Banon f17ad829ac remove snappy support
relates to 
2012-12-03 12:30:13 +01:00
Shay Banon b10cec1908 Upgrade to Netty 3.5.11
closes 
2012-12-02 22:29:49 +01:00
Simon Willnauer 479f1784e8 lucene 4: converted queryparser to lucene classic query parser 2012-11-12 13:44:32 +01:00
Shay Banon a4d0e3a0e8 lucene 4: add codes dependency 2012-11-12 13:44:31 +01:00
Igor Motov 05138bb2fb lucene 4: upgrade analyzers 2012-11-12 13:44:30 +01:00
Shay Banon 55a31f7ac5 change to lucene 4.0 dependency
upgrade has begun...
2012-11-12 13:44:30 +01:00
Shay Banon ac3501bc72 Upgrade to netty 3.5.10
closes 
2012-11-12 10:39:27 +01:00
Shay Banon ec64c8c907 Upgrade to Netty 3.5.9
closes 
2012-11-01 16:16:06 +01:00
Shay Banon c1489be816 forgot to mark 0.21.0 as snap 2012-10-22 01:55:33 +02:00
Shay Banon 6ec0071fa4 move master to be 0.21.0 Beta1 snap 2012-10-22 01:51:20 +02:00
Shay Banon fe980343be release 0.20.0.RC1 2012-10-22 01:48:09 +02:00
Shay Banon c350dcd9cf upgrade to jackson 2.1.0 2012-10-14 15:37:08 -04:00
Shay Banon c56cc6ecbe Upgrade to netty 3.5.8
closes .
2012-09-29 19:24:43 -04:00
Shay Banon 81775cc763 upgrade to mvel 2.1.3 2012-09-28 10:02:52 +02:00
Chris Male 6fc0b83e07 Upgraded to Spatial4j 0.3 2012-09-20 23:55:51 +02:00
Shay Banon 2275b82549 upgrade to testng 6.8 2012-09-17 16:23:18 +02:00
Shay Banon 19fdd46c87 upgrade to log4j 1.2.17 2012-09-15 10:16:18 +02:00
Shay Banon dd970752e7 upgrade to Netty 3.5.7.Final 2012-09-07 11:12:35 +02:00
Shay Banon 162dfb7011 Upgrade to LZF 0.9.6 2012-09-06 14:55:01 +02:00
Shay Banon 82b36e5bb8 upgrade to maven surefire 2.12.3 2012-09-03 20:48:44 +02:00
Shay Banon b055b5a94c Upgrade to mvel 2.1.1, closes . 2012-09-02 21:29:10 +02:00
Shay Banon 8b499dd4fd upgrade to guava 13.0.1 2012-09-02 17:52:40 +02:00
Shay Banon 6c3847b0a9 move spatial4j and jts to be optional dependencies
allowing data and client nodes to work without them, disabling shapes if needed
2012-09-01 00:05:49 +02:00
Shay Banon 377914bd07 Upgrade to netty 3.5.5, closes . 2012-08-23 01:30:49 +02:00
Shay Banon a80639abf5 exclude xerces from the deps for jts, we don't need it 2012-08-20 17:48:18 +02:00
Shay Banon dbe2f53a00 Support YAML as content type 2012-08-20 17:40:19 +02:00
Shay Banon 341c53b580 Upgrade to Netty 3.5.4, closes . 2012-08-16 12:15:24 -07:00
Shay Banon 980fc6ca34 add new spatial jars to deb package as well 2012-08-13 14:51:53 +02:00
Chris Male bea4346f3a Added GeoShape indexing and querying support 2012-08-13 13:44:29 +02:00
Shay Banon 53d29e5d8d Upgrade to guava 13.0 2012-08-08 20:52:09 +02:00
Shay Banon 82cfe0e8b2 upgrade to latest testng, improve console output when running test, add more options as env vars when using maven 2012-07-31 20:24:39 +02:00
Shay Banon 6e20056619 Upgrade to Netty 3.5.3, closes . 2012-07-27 19:01:21 +02:00
Shay Banon 90e94ebab9 Upgrade to Lucene 3.6.1, closes . 2012-07-22 13:28:39 -07:00
Shay Banon 73a34ee537 upgrade to guava 12.0.1 2012-07-14 13:02:21 +02:00
Shay Banon bfb4a29700 fix new jackson version to be properly shaded 2012-07-11 01:18:59 +02:00
Shay Banon 57e966e9d7 upgrade to jackson 2.0.4 2012-07-10 23:44:02 +02:00
Shay Banon 5f1b1c6f69 Upgrade to Netty 3.5.2, closes . 2012-07-05 23:09:26 +02:00
Shay Banon 57023c8ba9 Compression: Support snappy as a compression option, closes . 2012-07-04 17:14:12 +02:00
Shay Banon 1ffd68f2de Upgrade to netty 3.5.1 2012-06-28 00:51:37 +02:00
Shay Banon 2b893fe1e5 Use bloom filter when flushing (applying deletes), closes . 2012-06-26 16:45:29 +02:00
Shay Banon aebd27afbd abstract compression
abstract the LZF compression into a compress package allowing for different implementation in the future
2012-06-19 04:07:11 +02:00
Shay Banon 2280915d3c upgrade to joda 2.1
with the hack of duplicating BaseDateTime to remove the volatile
2012-06-14 21:57:01 +02:00
Shay Banon 2a80e14284 .DS_Store file in .deb package, closes . 2012-06-11 21:06:52 +02:00
Shay Banon 8f0bc799c6 Upgrade to latest jst166y and jsr166e
Embed the code now in our source, since jsr166e jar generation with 1.6 instead of 1.7 is complicated when doing it on its own as it relies on ThreadLocalRandom, and we have it in jsr166y
2012-06-10 00:42:54 +02:00
Shay Banon 43483b1237 upgrade to trove 3.0.3 2012-06-09 23:01:51 +02:00
Shay Banon 6c67570589 fix wrong change to version 2012-06-08 23:22:38 +02:00
Shay Banon 77e3cc1790 Upgrade to Netty 3.5.0.Final, closes . 2012-06-08 23:20:45 +02:00
Shay Banon 398e8a597f Upgrade to netty 3.4.6
Contains no applicable changes to our usage, but still be on track with latest
2012-05-29 19:33:21 +02:00
Shay Banon 11a1f69257 Upgrade to Netty 3.4.5, closes . 2012-05-16 21:58:34 +03:00
Shay Banon 76e1c3a017 Upgrade to guava 12.0, closes . 2012-05-07 20:08:41 +03:00
Shay Banon 3556020065 Upgrade to netty 3.4.3.Final, closes . 2012-05-06 19:52:00 +03:00
Shay Banon 2d26cbcde3 Upgrade to Netty 3.4.2, closes . 2012-05-03 22:03:34 +03:00
Shay Banon d0ab672982 upgrade to jackson 1.9.7 2012-05-03 21:28:51 +03:00
Shay Banon eb057d4ce6 Upgrade to Netty 3.4.1.Final, closes . 2012-04-21 16:03:15 +03:00
Shay Banon 62954e6d1f Upgrade to Netty 3.4.0, closes . 2012-04-15 18:15:17 +03:00
Shay Banon 16cd159a38 Upgrade to Lucene 3.6, closes . 2012-04-15 17:39:41 +03:00
Shay Banon 5131dc44ff upgrade to guava 11.0.2 2012-03-21 15:35:05 +02:00
Shay Banon 32e1eff05a move to 0.20.0.Beta1 snapshot 2012-03-02 11:23:24 +02:00
Shay Banon b8d583a06e release 0.19.0 2012-03-02 11:17:55 +02:00
Shay Banon 9a51179c22 upgrade to jackson 1.9.5 2012-02-26 23:06:04 +02:00
Shay Banon 7a10f18645 upgrade to latest 2.3 assembly plugin 2012-02-26 21:07:55 +02:00
Shay Banon d8123efac8 upgrade to jdeb 0.9 plugin 2012-02-26 10:06:44 +02:00
Shay Banon 2dfee54de7 upgrade to the latest assembly plugin and use artifact id (though not in distributions) 2012-02-26 10:06:18 +02:00
Shay Banon ca5f6ec0f6 move to 0.19.0.RC4 snap 2012-02-21 14:54:41 +02:00
Shay Banon aeaed0a1b0 release 0.19.0.RC3 2012-02-21 14:52:49 +02:00
Shay Banon ff3ebe4a4b upgrade to jackson 1.7.4 2012-02-17 12:50:43 +02:00
Shay Banon 5836c24b80 move to use sigar jar from maven central, which means that when using in the IDEA or running tests, sigar will not be enabled (since native files are not present), interim fix until we have a better one 2012-02-14 12:49:31 +02:00
Shay Banon bf11c19e49 move to use sigar jar from maven central, which means that when using in the IDEA or running tests, sigar will not be enabled (since native files are not present), interim fix until we have a better one 2012-02-14 12:49:15 +02:00
Shay Banon 676f115a26 update versions to 0.19.3 snap 2012-02-09 00:37:12 +02:00
Shay Banon f5676780ae release 0.19.0.RC2 2012-02-09 00:35:59 +02:00
Peter 4492b14f5f mini layout change to pom.xml 2012-02-08 08:45:40 +01:00
Peter 93134bd89e merged master 2012-02-08 08:31:41 +01:00
Shay Banon 7b3d9efe2e upgrade to netty 3.3.1 2012-02-07 20:15:55 +02:00
Shay Banon a89878ce6c move to 0.19.0.RC2 snap 2012-02-07 15:17:09 +02:00
Shay Banon b160ddbb2c release 0.19.0.RC1 2012-02-07 15:02:42 +02:00
Shay Banon 73266d1e44 upgrade to mvel 2.1.Beta8 2012-02-06 04:08:26 +02:00
Shay Banon f282081361 upgrade to latest compiled jsr166 libs 2012-01-31 21:00:58 +02:00
Shay Banon 62809bb62a Upgrade to netty 3.3.0 2012-01-19 23:32:59 +02:00
Shay Banon 002c8a6599 upgrade to 11.0.1 guava 2012-01-11 15:39:09 +02:00
Shay Banon 3d51553cf2 Move phonetic token filter to a plugin, closes . 2012-01-07 23:18:30 +02:00
Shay Banon a18021c778 Filter cache to have just weighted (node) and none, and index query parser cache to be size based, closes . 2012-01-05 20:44:09 +02:00
Shay Banon a488424404 Analysis: Add phonetic encodder called `bm` or `beider_morse`, closes . 2011-12-21 03:53:44 +02:00
Shay Banon e827c56be4 move to 1.9.3 jackson 2011-12-19 13:51:50 +02:00
Shay Banon 367a608707 remove jline from distribution to simplify it (no longer painting log levels though...) 2011-12-16 15:36:22 +02:00
Shay Banon 41a9dc1e6d make sigar optional 2011-12-13 19:03:32 +02:00
Shay Banon 16c2c6b2c3 fix deb to properly copy over sigar 2011-12-08 13:40:09 +02:00
Peter 6d0b94b81b improved deb packaging and bound assembly:single to package 2011-12-08 13:22:09 +02:00
Peter bccf0b193e moving debian package to maven 2011-12-08 13:22:09 +02:00
Peter 17e8b7f40c improved deb packaging and bound assembly:single to package 2011-12-08 12:05:05 +01:00
Peter 3c574ba398 moving debian package to maven 2011-12-08 10:31:53 +01:00
Shay Banon 6632f8cb56 upgrade to trove 3.0.2 2011-12-07 19:42:07 +02:00
Shay Banon c1b86f5786 upgrade to guava 10.0.1 and fix assembly 2011-12-06 21:09:28 +02:00
Shay Banon f190ec4396 move some deps from provided to compile with optional set to true 2011-12-06 15:43:59 +02:00
Shay Banon 2bcc14a6ea increase mem allocated for tests, add ability to run local tests 2011-12-06 15:15:14 +02:00
Shay Banon c6d5672a17 fix to provided scope from runtime since we need to compile with it 2011-12-06 13:52:40 +02:00
Shay Banon 19d80a7f1b assemblies 2011-12-06 13:41:49 +02:00
Shay Banon 6a71eab51f finalize structure, tests pass 2011-12-06 02:43:17 +02:00
Shay Banon fa782b1210 initial pom 2011-12-06 01:05:44 +02:00