OpenSearch/docs/reference/aggregations/pipeline.asciidoc

[[search-aggregations-pipeline]]

== Pipeline Aggregations

coming[2.0.0-beta1]

experimental[]

Pipeline aggregations work on the outputs produced from other aggregations rather than from document sets, adding
information to the output tree. There are many different types of pipeline aggregation, each computing different information from
other aggregations, but these types can be broken down into two families:

_Parent_::
                A family of pipeline aggregations that is provided with the output of its parent aggregation and is able
                to compute new buckets or new aggregations to add to existing buckets.

_Sibling_::
                Pipeline aggregations that are provided with the output of a sibling aggregation and are able to compute a
                new aggregation which will be at the same level as the sibling aggregation.

Pipeline aggregations can reference the aggregations they need to perform their computation by using the `buckets_path`
parameter to indicate the paths to the required metrics. The syntax for defining these paths can be found in the
<<buckets-path-syntax, `buckets_path` Syntax>> section below.

Pipeline aggregations cannot have sub-aggregations but depending on the type it can reference another pipeline in the `buckets_path`
allowing pipeline aggregations to be chained.  For example, you can chain together two derivatives to calculate the second derivative
(e.g. a derivative of a derivative).

NOTE: Because pipeline aggregations only add to the output, when chaining pipeline aggregations the output of each pipeline aggregation 
will be included in the final output.

[[buckets-path-syntax]]
[float]
=== `buckets_path` Syntax

Most pipeline aggregations require another aggregation as their input.  The input aggregation is defined via the `buckets_path`
parameter, which follows a specific format:

--------------------------------------------------
AGG_SEPARATOR       :=  '>'
METRIC_SEPARATOR    :=  '.'
AGG_NAME            :=  <the name of the aggregation>
METRIC              :=  <the name of the metric (in case of multi-value metrics aggregation)>
PATH                :=  <AGG_NAME>[<AGG_SEPARATOR><AGG_NAME>]*[<METRIC_SEPARATOR><METRIC>]
--------------------------------------------------

For example, the path `"my_bucket>my_stats.avg"` will path to the `avg` value in the `"my_stats"` metric, which is
contained in the `"my_bucket"` bucket aggregation.

Paths are relative from the position of the pipeline aggregation; they are not absolute paths, and the path cannot go back "up" the
aggregation tree. For example, this moving average is embedded inside a date_histogram and refers to a "sibling"
metric `"the_sum"`:

[source,js]
--------------------------------------------------
{
    "my_date_histo":{
        "date_histogram":{
            "field":"timestamp",
            "interval":"day"
        },
        "aggs":{
            "the_sum":{
                "sum":{ "field": "lemmings" } <1>
            },
            "the_movavg":{
                "moving_avg":{ "buckets_path": "the_sum" } <2>
            }
        }
    }
}
--------------------------------------------------
<1> The metric is called `"the_sum"`
<2> The `buckets_path` refers to the metric via a relative path `"the_sum"`

`buckets_path` is also used for Sibling pipeline aggregations, where the aggregation is "next" to a series of buckets
instead of embedded "inside" them.  For example, the `max_bucket` aggregation uses the `buckets_path` to specify
a metric embedded inside a sibling aggregation:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "sales_per_month" : {
            "date_histogram" : {
                "field" : "date",
                "interval" : "month"
            },
            "aggs": {
                "sales": {
                    "sum": {
                        "field": "price"
                    }
                }
            }
        },
        "max_monthly_sales": {
            "max_bucket": {
                "buckets_path": "sales_per_month>sales" <1>
            }
        }
    }
}
--------------------------------------------------
<1> `buckets_path` instructs this max_bucket aggregation that we want the maximum value of the `sales` aggregation in the
`sales_per_month` date histogram.

[float]
==== Special Paths

Instead of pathing to a metric, `buckets_path` can use a special `"_count"` path.  This instructs
the pipeline aggregation to use the document count as it's input.  For example, a moving average can be calculated on the document
count of each bucket, instead of a specific metric:

[source,js]
--------------------------------------------------
{
    "my_date_histo":{
        "date_histogram":{
            "field":"timestamp",
            "interval":"day"
        },
        "aggs":{
            "the_movavg":{
                "moving_avg":{ "buckets_path": "_count" } <1>
            }
        }
    }
}
--------------------------------------------------
<1> By using `_count` instead of a metric name, we can calculate the moving average of document counts in the histogram

[[gap-policy]]
[float]
=== Dealing with gaps in the data

Data in the real world is often noisy and sometimes contains *gaps* -- places where data simply doesn't exist.  This can
occur for a variety of reasons, the most common being:

* Documents falling into a bucket do not contain a required field
* There are no documents matching the query for one or more buckets
* The metric being calculated is unable to generate a value, likely because another dependent bucket is missing a value.
Some pipeline aggregations have specific requirements that must be met (e.g. a derivative cannot calculate a metric for the
first value because there is no previous value, HoltWinters moving average need "warmup" data to begin calculating, etc)

Gap policies are a mechanism to inform the pipeline aggregation about the desired behavior when "gappy" or missing
data is encountered.  All pipeline aggregations accept the `gap_policy` parameter.  There are currently two gap policies
to choose from:

_skip_::
                This option treats missing data as if the bucket does not exist.  It will skip the bucket and continue
                calculating using the next available value.

_insert_zeros_::
                This option will replace missing values with a zero (`0`) and pipeline aggregation computation will
                proceed as normal.


include::pipeline/avg-bucket-aggregation.asciidoc[]
include::pipeline/derivative-aggregation.asciidoc[]
include::pipeline/max-bucket-aggregation.asciidoc[]
include::pipeline/min-bucket-aggregation.asciidoc[]
include::pipeline/sum-bucket-aggregation.asciidoc[]
include::pipeline/percentiles-bucket-aggregation.asciidoc[]
include::pipeline/movavg-aggregation.asciidoc[]
include::pipeline/cumulative-sum-aggregation.asciidoc[]
include::pipeline/bucket-script-aggregation.asciidoc[]
include::pipeline/bucket-selector-aggregation.asciidoc[]
include::pipeline/serial-diff-aggregation.asciidoc[]
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`[[search-aggregations-pipeline]]`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`== Pipeline Aggregations`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
Docs: Updated annotations for 2.0.0-beta1 2015-08-14 10:51:09 +02:00			`coming[2.0.0-beta1]`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
			`experimental[]`

Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`Pipeline aggregations work on the outputs produced from other aggregations rather than from document sets, adding`
			`information to the output tree. There are many different types of pipeline aggregation, each computing different information from`
Fixing typo 2015-08-08 14:14:59 -07:00			`other aggregations, but these types can be broken down into two families:`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
			`_Parent_::`
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`A family of pipeline aggregations that is provided with the output of its parent aggregation and is able`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`to compute new buckets or new aggregations to add to existing buckets.`

			`_Sibling_::`
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`Pipeline aggregations that are provided with the output of a sibling aggregation and are able to compute a`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`new aggregation which will be at the same level as the sibling aggregation.`

Docs: Fixed variations of spelling of buckets_path Closes #13201 2015-08-31 13:47:40 +02:00			Pipeline aggregations can reference the aggregations they need to perform their computation by using the `buckets_path`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`parameter to indicate the paths to the required metrics. The syntax for defining these paths can be found in the`
Docs: Fixed variations of spelling of buckets_path Closes #13201 2015-08-31 13:47:40 +02:00			<<buckets-path-syntax, `buckets_path` Syntax>> section below.
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			Pipeline aggregations cannot have sub-aggregations but depending on the type it can reference another pipeline in the `buckets_path`
			`allowing pipeline aggregations to be chained. For example, you can chain together two derivatives to calculate the second derivative`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`(e.g. a derivative of a derivative).`

Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`NOTE: Because pipeline aggregations only add to the output, when chaining pipeline aggregations the output of each pipeline aggregation`
			`will be included in the final output.`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
Docs: Fixed variations of spelling of buckets_path Closes #13201 2015-08-31 13:47:40 +02:00			`[[buckets-path-syntax]]`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`[float]`
			=== `buckets_path` Syntax

Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			Most pipeline aggregations require another aggregation as their input. The input aggregation is defined via the `buckets_path`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`parameter, which follows a specific format:`

			`--------------------------------------------------`
			`AGG_SEPARATOR := '>'`
			`METRIC_SEPARATOR := '.'`
			`AGG_NAME := <the name of the aggregation>`
			`METRIC := <the name of the metric (in case of multi-value metrics aggregation)>`
			`PATH := <AGG_NAME>[<AGG_SEPARATOR><AGG_NAME>]*[<METRIC_SEPARATOR><METRIC>]`
			`--------------------------------------------------`

			For example, the path `"my_bucket>my_stats.avg"` will path to the `avg` value in the `"my_stats"` metric, which is
			contained in the `"my_bucket"` bucket aggregation.

Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`Paths are relative from the position of the pipeline aggregation; they are not absolute paths, and the path cannot go back "up" the`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`aggregation tree. For example, this moving average is embedded inside a date_histogram and refers to a "sibling"`
			metric `"the_sum"`:

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"my_date_histo":{`
			`"date_histogram":{`
			`"field":"timestamp",`
			`"interval":"day"`
			`},`
			`"aggs":{`
			`"the_sum":{`
			`"sum":{ "field": "lemmings" } <1>`
			`},`
			`"the_movavg":{`
			`"moving_avg":{ "buckets_path": "the_sum" } <2>`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
			<1> The metric is called `"the_sum"`
			<2> The `buckets_path` refers to the metric via a relative path `"the_sum"`

Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`buckets_path` is also used for Sibling pipeline aggregations, where the aggregation is "next" to a series of buckets
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			instead of embedded "inside" them. For example, the `max_bucket` aggregation uses the `buckets_path` to specify
			`a metric embedded inside a sibling aggregation:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"sales_per_month" : {`
			`"date_histogram" : {`
			`"field" : "date",`
			`"interval" : "month"`
			`},`
			`"aggs": {`
			`"sales": {`
			`"sum": {`
			`"field": "price"`
			`}`
			`}`
			`}`
			`},`
			`"max_monthly_sales": {`
			`"max_bucket": {`
Docs: Fixed variations of spelling of buckets_path Closes #13201 2015-08-31 13:47:40 +02:00			`"buckets_path": "sales_per_month>sales" <1>`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
Docs: Fixed variations of spelling of buckets_path Closes #13201 2015-08-31 13:47:40 +02:00			<1> `buckets_path` instructs this max_bucket aggregation that we want the maximum value of the `sales` aggregation in the
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`sales_per_month` date histogram.

			`[float]`
			`==== Special Paths`

			Instead of pathing to a metric, `buckets_path` can use a special `"_count"` path. This instructs
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`the pipeline aggregation to use the document count as it's input. For example, a moving average can be calculated on the document`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`count of each bucket, instead of a specific metric:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"my_date_histo":{`
			`"date_histogram":{`
			`"field":"timestamp",`
			`"interval":"day"`
			`},`
			`"aggs":{`
			`"the_movavg":{`
			`"moving_avg":{ "buckets_path": "_count" } <1>`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
			<1> By using `_count` instead of a metric name, we can calculate the moving average of document counts in the histogram

Aggregations: Adding Average Bucket Aggregation Also includes changes to the other bucket metric aggregations to share code Closes #11006 2015-05-06 12:54:42 +01:00			`[[gap-policy]]`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00			`[float]`
			`=== Dealing with gaps in the data`

[DOCS] Update section about gap_policy 2015-07-07 15:37:42 -04:00			`Data in the real world is often noisy and sometimes contains gaps -- places where data simply doesn't exist. This can`
			`occur for a variety of reasons, the most common being:`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
[DOCS] Update section about gap_policy 2015-07-07 15:37:42 -04:00			`* Documents falling into a bucket do not contain a required field`
			`* There are no documents matching the query for one or more buckets`
			`* The metric being calculated is unable to generate a value, likely because another dependent bucket is missing a value.`
			`Some pipeline aggregations have specific requirements that must be met (e.g. a derivative cannot calculate a metric for the`
			`first value because there is no previous value, HoltWinters moving average need "warmup" data to begin calculating, etc)`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
[DOCS] Update section about gap_policy 2015-07-07 15:37:42 -04:00			`Gap policies are a mechanism to inform the pipeline aggregation about the desired behavior when "gappy" or missing`
			data is encountered. All pipeline aggregations accept the `gap_policy` parameter. There are currently two gap policies
			`to choose from:`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
Aggregations: Adding Average Bucket Aggregation Also includes changes to the other bucket metric aggregations to share code Closes #11006 2015-05-06 12:54:42 +01:00			`_skip_::`
[DOCS] Update section about gap_policy 2015-07-07 15:37:42 -04:00			`This option treats missing data as if the bucket does not exist. It will skip the bucket and continue`
			`calculating using the next available value.`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00
			`_insert_zeros_::`
[DOCS] Update section about gap_policy 2015-07-07 15:37:42 -04:00			This option will replace missing values with a zero (`0`) and pipeline aggregation computation will
			`proceed as normal.`
[DOCS] Restructure Aggs documentation 2015-05-01 16:04:55 -04:00



Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`include::pipeline/avg-bucket-aggregation.asciidoc[]`
			`include::pipeline/derivative-aggregation.asciidoc[]`
			`include::pipeline/max-bucket-aggregation.asciidoc[]`
			`include::pipeline/min-bucket-aggregation.asciidoc[]`
			`include::pipeline/sum-bucket-aggregation.asciidoc[]`
Aggregations: Add percentiles_bucket pipeline aggregations This pipeline will calculate percentiles over a set of sibling buckets. This is an exact implementation, meaning it needs to cache a copy of the series in memory and sort it to determine the percentiles. This comes with a few limitations: to prevent serializing data around, only the requested percentiles are calculated (unlike the TDigest version, which allows the java API to ask for any percentile). It also needs to store the data in-memory, resulting in some overhead if the requested series is very large. 2015-08-28 12:23:19 -04:00			`include::pipeline/percentiles-bucket-aggregation.asciidoc[]`
Aggregations: Renaming reducers to Pipeline Aggregators 2015-05-21 10:39:38 +01:00			`include::pipeline/movavg-aggregation.asciidoc[]`
Aggregations: Adds cumulative sum aggregation This adds a new pipeline aggregation, the cumulative sum aggregation. This is a parent aggregation which must be specified as a sub-aggregation to a histogram or date_histogram aggregation. It will add a new aggregation to each bucket containing the sum of a specified metrics over this and all previous buckets. 2015-06-22 16:30:42 +01:00			`include::pipeline/cumulative-sum-aggregation.asciidoc[]`
Aggregations: Rename `series_arithmetic` agg to `bucket_script` 2015-06-17 10:48:21 +01:00			`include::pipeline/bucket-script-aggregation.asciidoc[]`
Aggregations: Pipeline Aggregation to filter buckets based on a script This pipeline aggregation runs a script on each bucket in the parent aggregation to determine whether the bucket is kept in the final aggregation tree. If the script returns true the bucket is retained, if it returns false the bucket is dropped 2015-06-25 14:22:15 +01:00			`include::pipeline/bucket-selector-aggregation.asciidoc[]`
[DOCS] Fix link to serial_diff docs 2015-07-10 19:01:18 -04:00			`include::pipeline/serial-diff-aggregation.asciidoc[]`