OpenSearch/docs/reference/aggregations/metrics/percentile-rank-aggregation...

[[search-aggregations-metrics-percentile-rank-aggregation]]
=== Percentile Ranks Aggregation

A `multi-value` metrics aggregation that calculates one or more percentile ranks
over numeric values extracted from the aggregated documents.  These values
can be extracted either from specific numeric fields in the documents, or
be generated by a provided script.

[NOTE]
==================================================
Please see <<search-aggregations-metrics-percentile-aggregation-approximation>>
and <<search-aggregations-metrics-percentile-aggregation-compression>> for advice
regarding approximation and memory use of the percentile ranks aggregation
==================================================

Percentile rank show the percentage of observed values which are below certain
value.  For example, if a value is greater than or equal to 95% of the observed values
it is said to be at the 95th percentile rank.

Assume your data consists of website load times.  You may have a service agreement that
95% of page loads completely within 15ms and 99% of page loads complete within 30ms.

Let's look at a range of percentiles representing load time:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "load_time_outlier" : {
            "percentile_ranks" : {
                "field" : "load_time", <1>
                "values" : [15, 30]
            }
        }
    }
}
--------------------------------------------------
<1> The field `load_time` must be a numeric field

The response will look like this:

[source,js]
--------------------------------------------------
{
    ...

   "aggregations": {
      "load_time_outlier": {
         "values" : {
            "15": 92,
            "30": 100
         }
      }
   }
}
--------------------------------------------------

From this information you can determine you are hitting the 99% load time target but not quite
hitting the 95% load time target


==== Script

The percentile rank metric supports scripting.  For example, if our load times
are in milliseconds but we want to specify values in seconds, we could use
a script to convert them on-the-fly:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "load_time_outlier" : {
            "percentile_ranks" : {
                "values" : [3, 5],
                "script" : {
                    "lang": "painless",
                    "inline": "doc['load_time'].value / params.timeUnit", <1>
                    "params" : {
                        "timeUnit" : 1000   <2>
                    }
                }
            }
        }
    }
}
--------------------------------------------------
<1> The `field` parameter is replaced with a `script` parameter, which uses the
script to generate values which percentile ranks are calculated on
<2> Scripting supports parameterized input just like any other script

This will interpret the `script` parameter as an `inline` script with the `painless` script language and no script parameters. To use a file script use the following syntax:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "load_time_outlier" : {
            "percentile_ranks" : {
                "values" : [3, 5],
                "script" : {
                    "file": "my_script",
                    "params" : {
                        "timeUnit" : 1000
                    }
                }
            }
        }
    }
}
--------------------------------------------------

TIP: for indexed scripts replace the `file` parameter with an `id` parameter.

==== HDR Histogram

experimental[]

https://github.com/HdrHistogram/HdrHistogram[HDR Histogram] (High Dynamic Range Histogram) is an alternative implementation
that can be useful when calculating percentile ranks for latency measurements as it can be faster than the t-digest implementation
with the trade-off of a larger memory footprint. This implementation maintains a fixed worse-case percentage error (specified as a
number of significant digits). This means that if data is recorded with values from 1 microsecond up to 1 hour (3,600,000,000
microseconds) in a histogram set to 3 significant digits, it will maintain a value resolution of 1 microsecond for values up to
1 millisecond and 3.6 seconds (or better) for the maximum tracked value (1 hour).

The HDR Histogram can be used by specifying the `method` parameter in the request:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "load_time_outlier" : {
            "percentile_ranks" : {
                "field" : "load_time",
                "values" : [15, 30],
                "hdr": { <1>
                  "number_of_significant_value_digits" : 3 <2>
                }
            }
        }
    }
}
--------------------------------------------------
<1> `hdr` object indicates that HDR Histogram should be used to calculate the percentiles and specific settings for this algorithm can be specified inside the object
<2> `number_of_significant_value_digits` specifies the resolution of values for the histogram in number of significant digits

The HDRHistogram only supports positive values and will error if it is passed a negative value. It is also not a good idea to use
the HDRHistogram if the range of values is unknown as this could lead to high memory usage.

==== Missing value

The `missing` parameter defines how documents that are missing a value should be treated.
By default they will be ignored but it is also possible to treat them as if they
had a value.

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "grade_ranks" : {
            "percentile_ranks" : {
                "field" : "grade",
                "missing": 10 <1>
            }
        }
    }
}
--------------------------------------------------

<1> Documents without a value in the `grade` field will fall into the same bucket as documents that have the value `10`.
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`[[search-aggregations-metrics-percentile-rank-aggregation]]`
			`=== Percentile Ranks Aggregation`

			A `multi-value` metrics aggregation that calculates one or more percentile ranks
			`over numeric values extracted from the aggregated documents. These values`
			`can be extracted either from specific numeric fields in the documents, or`
			`be generated by a provided script.`

			`[NOTE]`
			`==================================================`
Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00			`Please see <<search-aggregations-metrics-percentile-aggregation-approximation>>`
			`and <<search-aggregations-metrics-percentile-aggregation-compression>> for advice`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`regarding approximation and memory use of the percentile ranks aggregation`
			`==================================================`

Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00			`Percentile rank show the percentage of observed values which are below certain`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`value. For example, if a value is greater than or equal to 95% of the observed values`
			`it is said to be at the 95th percentile rank.`

Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00			`Assume your data consists of website load times. You may have a service agreement that`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`95% of page loads completely within 15ms and 99% of page loads complete within 30ms.`

			`Let's look at a range of percentiles representing load time:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"load_time_outlier" : {`
			`"percentile_ranks" : {`
[DOCS] add missing comma in percentile_rank aggregation example 2015-02-21 04:19:11 -05:00			`"field" : "load_time", <1>`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`"values" : [15, 30]`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
			<1> The field `load_time` must be a numeric field

			`The response will look like this:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`...`

			`"aggregations": {`
			`"load_time_outlier": {`
			`"values" : {`
			`"15": 92,`
			`"30": 100`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00			`From this information you can determine you are hitting the 99% load time target but not quite`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`hitting the 95% load time target`


			`==== Script`

			`The percentile rank metric supports scripting. For example, if our load times`
			`are in milliseconds but we want to specify values in seconds, we could use`
			`a script to convert them on-the-fly:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"load_time_outlier" : {`
			`"percentile_ranks" : {`
			`"values" : [3, 5],`
Scripting: Unify script and template requests across codebase This change unifies the way scripts and templates are specified for all instances in the codebase. It builds on the Script class added previously and adds request building and parsing support as well as the ability to transfer script objects between nodes. It also adds a Template class which aims to provide the same functionality for template APIs Closes #11091 2015-05-12 05:37:22 -04:00			`"script" : {`
cutover some docs to painless 2016-06-27 09:55:16 -04:00			`"lang": "painless",`
			`"inline": "doc['load_time'].value / params.timeUnit", <1>`
Scripting: Unify script and template requests across codebase This change unifies the way scripts and templates are specified for all instances in the codebase. It builds on the Script class added previously and adds request building and parsing support as well as the ability to transfer script objects between nodes. It also adds a Template class which aims to provide the same functionality for template APIs Closes #11091 2015-05-12 05:37:22 -04:00			`"params" : {`
			`"timeUnit" : 1000 <2>`
			`}`
Aggregations: Added percentile rank aggregation Percentile Rank Aggregation is the reverse of the Percetiles aggregation. It determines the percentile rank (the proportion of values less than a given value) of the provided array of values. Closes #6386 2014-06-06 10:25:21 -04:00			`}`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
			<1> The `field` parameter is replaced with a `script` parameter, which uses the
			`script to generate values which percentile ranks are calculated on`
			`<2> Scripting supports parameterized input just like any other script`
Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00
cutover some docs to painless 2016-06-27 09:55:16 -04:00			This will interpret the `script` parameter as an `inline` script with the `painless` script language and no script parameters. To use a file script use the following syntax:
Scripting: Unify script and template requests across codebase This change unifies the way scripts and templates are specified for all instances in the codebase. It builds on the Script class added previously and adds request building and parsing support as well as the ability to transfer script objects between nodes. It also adds a Template class which aims to provide the same functionality for template APIs Closes #11091 2015-05-12 05:37:22 -04:00
			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"load_time_outlier" : {`
			`"percentile_ranks" : {`
			`"values" : [3, 5],`
			`"script" : {`
			`"file": "my_script",`
			`"params" : {`
			`"timeUnit" : 1000`
			`}`
			`}`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

			TIP: for indexed scripts replace the `file` parameter with an `id` parameter.
Aggs: Make it possible to configure missing values. Most aggregations (terms, histogram, stats, percentiles, geohash-grid) now support a new `missing` option which defines the value to consider when a field does not have a value. This can be handy if you eg. want a terms aggregation to handle the same way documents that have "N/A" or no value for a `tag` field. This works in a very similar way to the `missing` option on the `sort` element. One known issue is that this option sometimes cannot make the right decision in the unmapped case: it needs to replace all values with the `missing` value but might not know what kind of values source should be produced (numerics, strings, geo points?). For this reason, we might want to add an `unmapped_type` option in the future like we did for sorting. Related to #5324 2015-05-07 10:46:40 -04:00
Aggregations: Add HDRHistogram as an option in percentiles and percentile_ranks aggregations HDRHistogram has been added as an option in the percentiles and percentile_ranks aggregation. It has one option `number_significant_digits` which controls the accuracy and memory size for the algorithm Closes #8324 2015-07-20 07:23:21 -04:00			`==== HDR Histogram`

			`experimental[]`

[DOCS] fix documentation for selecting algorithm for percentiles agg 2016-07-27 03:48:35 -04:00			`https://github.com/HdrHistogram/HdrHistogram[HDR Histogram] (High Dynamic Range Histogram) is an alternative implementation`
			`that can be useful when calculating percentile ranks for latency measurements as it can be faster than the t-digest implementation`
			`with the trade-off of a larger memory footprint. This implementation maintains a fixed worse-case percentage error (specified as a`
			`number of significant digits). This means that if data is recorded with values from 1 microsecond up to 1 hour (3,600,000,000`
			`microseconds) in a histogram set to 3 significant digits, it will maintain a value resolution of 1 microsecond for values up to`
Aggregations: Add HDRHistogram as an option in percentiles and percentile_ranks aggregations HDRHistogram has been added as an option in the percentiles and percentile_ranks aggregation. It has one option `number_significant_digits` which controls the accuracy and memory size for the algorithm Closes #8324 2015-07-20 07:23:21 -04:00			`1 millisecond and 3.6 seconds (or better) for the maximum tracked value (1 hour).`

			The HDR Histogram can be used by specifying the `method` parameter in the request:

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"load_time_outlier" : {`
			`"percentile_ranks" : {`
			`"field" : "load_time",`
			`"values" : [15, 30],`
[DOCS] fix documentation for selecting algorithm for percentiles agg 2016-07-27 03:48:35 -04:00			`"hdr": { <1>`
			`"number_of_significant_value_digits" : 3 <2>`
			`}`
Aggregations: Add HDRHistogram as an option in percentiles and percentile_ranks aggregations HDRHistogram has been added as an option in the percentiles and percentile_ranks aggregation. It has one option `number_significant_digits` which controls the accuracy and memory size for the algorithm Closes #8324 2015-07-20 07:23:21 -04:00			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
[DOCS] fix documentation for selecting algorithm for percentiles agg 2016-07-27 03:48:35 -04:00			<1> `hdr` object indicates that HDR Histogram should be used to calculate the percentiles and specific settings for this algorithm can be specified inside the object
Aggregations: Add HDRHistogram as an option in percentiles and percentile_ranks aggregations HDRHistogram has been added as an option in the percentiles and percentile_ranks aggregation. It has one option `number_significant_digits` which controls the accuracy and memory size for the algorithm Closes #8324 2015-07-20 07:23:21 -04:00			<2> `number_of_significant_value_digits` specifies the resolution of values for the histogram in number of significant digits

[DOCS] fix documentation for selecting algorithm for percentiles agg 2016-07-27 03:48:35 -04:00			`The HDRHistogram only supports positive values and will error if it is passed a negative value. It is also not a good idea to use`
Aggregations: Add HDRHistogram as an option in percentiles and percentile_ranks aggregations HDRHistogram has been added as an option in the percentiles and percentile_ranks aggregation. It has one option `number_significant_digits` which controls the accuracy and memory size for the algorithm Closes #8324 2015-07-20 07:23:21 -04:00			`the HDRHistogram if the range of values is unknown as this could lead to high memory usage.`

Aggs: Make it possible to configure missing values. Most aggregations (terms, histogram, stats, percentiles, geohash-grid) now support a new `missing` option which defines the value to consider when a field does not have a value. This can be handy if you eg. want a terms aggregation to handle the same way documents that have "N/A" or no value for a `tag` field. This works in a very similar way to the `missing` option on the `sort` element. One known issue is that this option sometimes cannot make the right decision in the unmapped case: it needs to replace all values with the `missing` value but might not know what kind of values source should be produced (numerics, strings, geo points?). For this reason, we might want to add an `unmapped_type` option in the future like we did for sorting. Related to #5324 2015-05-07 10:46:40 -04:00			`==== Missing value`

			The `missing` parameter defines how documents that are missing a value should be treated.
			`By default they will be ignored but it is also possible to treat them as if they`
			`had a value.`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"grade_ranks" : {`
			`"percentile_ranks" : {`
			`"field" : "grade",`
			`"missing": 10 <1>`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

			<1> Documents without a value in the `grade` field will fall into the same bucket as documents that have the value `10`.