OpenSearch/docs/reference/aggregations/metrics/cardinality-aggregation.asc...

[[search-aggregations-metrics-cardinality-aggregation]]
=== Cardinality Aggregation

A `single-value` metrics aggregation that calculates an approximate count of
distinct values. Values can be extracted either from specific fields in the
document or generated by a script.

Assume you are indexing books and would like to count the unique authors that
match a query:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "author_count" : {
            "cardinality" : {
                "field" : "author"
            }
        }
    }
}
--------------------------------------------------

==== Precision control

This aggregation also supports the `precision_threshold` and `rehash` options:

experimental[The `precision_threshold` and `rehash` options are specific to the current internal implementation of the `cardinality` agg, which may change in the future]

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "author_count" : {
            "cardinality" : {
                "field" : "author_hash",
                "precision_threshold": 100, <1>
                "rehash": false <2>
            }
        }
    }
}
--------------------------------------------------

<1> The `precision_threshold` options allows to trade memory for accuracy, and
defines a unique count below which counts are expected to be close to
accurate. Above this value, counts might become a bit more fuzzy. The maximum
supported value is 40000, thresholds above this number will have the same
effect as a threshold of 40000.
Default value depends on the number of parent aggregations that multiple
create buckets (such as terms or histograms).
<2> If you computed a hash on client-side, stored it into your documents and want
Elasticsearch to use them to compute counts using this hash function without
rehashing values, it is possible to specify `rehash: false`. Default value is
`true`. Please note that the hash must be indexed as a long when `rehash` is
false.

==== Counts are approximate

Computing exact counts requires loading values into a hash set and returning its
size. This doesn't scale when working on high-cardinality sets and/or large
values as the required memory usage and the need to communicate those
per-shard sets between nodes would utilize too many resources of the cluster.

This `cardinality` aggregation is based on the
http://static.googleusercontent.com/media/research.google.com/fr//pubs/archive/40671.pdf[HyperLogLog++]
algorithm, which counts based on the hashes of the values with some interesting
properties:

 * configurable precision, which decides on how to trade memory for accuracy,
 * excellent accuracy on low-cardinality sets,
 * fixed memory usage: no matter if there are tens or billions of unique values,
   memory usage only depends on the configured precision.

For a precision threshold of `c`, the implementation that we are using requires
about `c * 8` bytes.

The following chart shows how the error varies before and after the threshold:

image:images/cardinality_error.png[]

For all 3 thresholds, counts have been accurate up to the configured threshold
(although not guaranteed, this is likely to be the case). Please also note that
even with a threshold as low as 100, the error remains under 5%, even when
counting millions of items.

==== Pre-computed hashes

If you don't want Elasticsearch to re-compute hashes on every run of this
aggregation, it is possible to use pre-computed hashes, either by computing a
hash on client-side, indexing it and specifying `rehash: false`, or by using
the special `murmur3` field mapper, typically in the context of a `multi-field`
in the mapping:

[source,js]
--------------------------------------------------
{
    "author": {
        "type": "string",
        "fields": {
            "hash": {
                "type": "murmur3"
            }
        }
    }
}
--------------------------------------------------

With such a mapping, Elasticsearch is going to compute hashes of the `author`
field at indexing time and store them in the `author.hash` field. This
way, unique counts can be computed using the cardinality aggregation by only
loading the hashes into memory, not the values of the `author` field, and
without computing hashes on the fly:

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "author_count" : {
            "cardinality" : {
                "field" : "author.hash"
            }
        }
    }
}
--------------------------------------------------

NOTE: `rehash` is automatically set to `false` when computing unique counts on
a `murmur3` field.

NOTE: Pre-computing hashes is usually only useful on very large and/or
high-cardinality fields as it saves CPU and memory. However, on numeric
fields, hashing is very fast and storing the original values requires as much
or less memory than storing the hashes. This is also true on low-cardinality
string fields, especially given that those have an optimization in order to
make sure that hashes are computed at most once per unique value per segment.

==== Script

The `cardinality` metric supports scripting, with a noticeable performance hit
however since hashes need to be computed on the fly.

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "author_count" : {
            "cardinality" : {
                "script": "doc['author.first_name'].value + ' ' + doc['author.last_name'].value"
            }
        }
    }
}
--------------------------------------------------

TIP: The `script` parameter expects an inline script. Use `script_id` for indexed scripts and `script_file` for scripts in the `config/scripts/` directory.

==== Missing value

The `missing` parameter defines how documents that are missing a value should be treated.
By default they will be ignored but it is also possible to treat them as if they
had a value.

[source,js]
--------------------------------------------------
{
    "aggs" : {
        "tag_cardinality" : {
            "cardinality" : {
                "field" : "tag",
                "missing": "N/A" <1>
            }
        }
    }
}
--------------------------------------------------

<1> Documents without a value in the `tag` field will fall into the same bucket as documents that have the value `N/A`.
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`[[search-aggregations-metrics-cardinality-aggregation]]`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00			`=== Cardinality Aggregation`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00
			A `single-value` metrics aggregation that calculates an approximate count of
			`distinct values. Values can be extracted either from specific fields in the`
			`document or generated by a script.`

			`Assume you are indexing books and would like to count the unique authors that`
			`match a query:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00			`"author_count" : {`
			`"cardinality" : {`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`"field" : "author"`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

Docs: Remove the experimental status of the cardinality and percentiles(-ranks) aggregations These aggregations are not experimental anymore but some of their parameters still are: - `precision_threshold` and `rehash` on `cardinality` - `compression` on percentiles(-ranks) Close #9560 2015-02-04 09:05:26 -05:00			`==== Precision control`

Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			This aggregation also supports the `precision_threshold` and `rehash` options:

Docs: Updated the experimental annotations in the docs as follows: * Removed the docs for `index.compound_format` and `index.compound_on_flush` - these are expert settings which should probably be removed (see https://github.com/elastic/elasticsearch/issues/10778) * Removed the docs for `index.index_concurrency` - another expert setting * Labelled the segments verbose output as experimental * Marked the `compression`, `precision_threshold` and `rehash` options as experimental in the cardinality and percentile aggs * Improved the experimental text on `significant_terms`, `execution_hint` in the terms agg, and `terminate_after` param on count and search * Removed the experimental flag on the `geobounds` agg * Marked the settings in the `merge` and `store` modules as experimental, rather than the modules themselves Closes #10782 2015-04-26 12:49:15 -04:00			experimental[The `precision_threshold` and `rehash` options are specific to the current internal implementation of the `cardinality` agg, which may change in the future]

Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00			`"author_count" : {`
			`"cardinality" : {`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`"field" : "author_hash",`
			`"precision_threshold": 100, <1>`
			`"rehash": false <2>`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

Docs: Updated the experimental annotations in the docs as follows: * Removed the docs for `index.compound_format` and `index.compound_on_flush` - these are expert settings which should probably be removed (see https://github.com/elastic/elasticsearch/issues/10778) * Removed the docs for `index.index_concurrency` - another expert setting * Labelled the segments verbose output as experimental * Marked the `compression`, `precision_threshold` and `rehash` options as experimental in the cardinality and percentile aggs * Improved the experimental text on `significant_terms`, `execution_hint` in the terms agg, and `terminate_after` param on count and search * Removed the experimental flag on the `geobounds` agg * Marked the settings in the `merge` and `store` modules as experimental, rather than the modules themselves Closes #10782 2015-04-26 12:49:15 -04:00			<1> The `precision_threshold` options allows to trade memory for accuracy, and
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`defines a unique count below which counts are expected to be close to`
			`accurate. Above this value, counts might become a bit more fuzzy. The maximum`
			`supported value is 40000, thresholds above this number will have the same`
			`effect as a threshold of 40000.`
			`Default value depends on the number of parent aggregations that multiple`
			`create buckets (such as terms or histograms).`
Docs: Updated the experimental annotations in the docs as follows: * Removed the docs for `index.compound_format` and `index.compound_on_flush` - these are expert settings which should probably be removed (see https://github.com/elastic/elasticsearch/issues/10778) * Removed the docs for `index.index_concurrency` - another expert setting * Labelled the segments verbose output as experimental * Marked the `compression`, `precision_threshold` and `rehash` options as experimental in the cardinality and percentile aggs * Improved the experimental text on `significant_terms`, `execution_hint` in the terms agg, and `terminate_after` param on count and search * Removed the experimental flag on the `geobounds` agg * Marked the settings in the `merge` and `store` modules as experimental, rather than the modules themselves Closes #10782 2015-04-26 12:49:15 -04:00			`<2> If you computed a hash on client-side, stored it into your documents and want`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`Elasticsearch to use them to compute counts using this hash function without`
			rehashing values, it is possible to specify `rehash: false`. Default value is
			`true`. Please note that the hash must be indexed as a long when `rehash` is
			`false.`

			`==== Counts are approximate`

			`Computing exact counts requires loading values into a hash set and returning its`
			`size. This doesn't scale when working on high-cardinality sets and/or large`
			`values as the required memory usage and the need to communicate those`
			`per-shard sets between nodes would utilize too many resources of the cluster.`

			This `cardinality` aggregation is based on the
			`http://static.googleusercontent.com/media/research.google.com/fr//pubs/archive/40671.pdf[HyperLogLog++]`
			`algorithm, which counts based on the hashes of the values with some interesting`
			`properties:`

			`* configurable precision, which decides on how to trade memory for accuracy,`
			`* excellent accuracy on low-cardinality sets,`
			`* fixed memory usage: no matter if there are tens or billions of unique values,`
			`memory usage only depends on the configured precision.`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			For a precision threshold of `c`, the implementation that we are using requires
			about `c * 8` bytes.

			`The following chart shows how the error varies before and after the threshold:`

			`image:images/cardinality_error.png[]`

			`For all 3 thresholds, counts have been accurate up to the configured threshold`
			`(although not guaranteed, this is likely to be the case). Please also note that`
Docs: Update cardinality-aggregation.asciidoc Closes #7516 2014-08-29 13:24:24 -04:00			`even with a threshold as low as 100, the error remains under 5%, even when`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`counting millions of items.`

			`==== Pre-computed hashes`

			`If you don't want Elasticsearch to re-compute hashes on every run of this`
			`aggregation, it is possible to use pre-computed hashes, either by computing a`
			hash on client-side, indexing it and specifying `rehash: false`, or by using
			the special `murmur3` field mapper, typically in the context of a `multi-field`
			`in the mapping:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"author": {`
			`"type": "string",`
			`"fields": {`
			`"hash": {`
			`"type": "murmur3"`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

			With such a mapping, Elasticsearch is going to compute hashes of the `author`
			field at indexing time and store them in the `author.hash` field. This
			`way, unique counts can be computed using the cardinality aggregation by only`
			loading the hashes into memory, not the values of the `author` field, and
			`without computing hashes on the fly:`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00			`"author_count" : {`
			`"cardinality" : {`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`"field" : "author.hash"`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

			NOTE: `rehash` is automatically set to `false` when computing unique counts on
			a `murmur3` field.

			`NOTE: Pre-computing hashes is usually only useful on very large and/or`
			`high-cardinality fields as it saves CPU and memory. However, on numeric`
			`fields, hashing is very fast and storing the original values requires as much`
			`or less memory than storing the hashes. This is also true on low-cardinality`
			`string fields, especially given that those have an optimization in order to`
			`make sure that hashes are computed at most once per unique value per segment.`

			`==== Script`

			The `cardinality` metric supports scripting, with a noticeable performance hit
			`however since hashes need to be computed on the fly.`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
[DOCS] Added "Aggregation" to all aggs titles 2014-05-12 19:35:58 -04:00			`"author_count" : {`
			`"cardinality" : {`
Cardinality aggregation. This aggregation computes unique term counts using the hyperloglog++ algorithm which uses linear counting to estimate low cardinalities and hyperloglog on higher cardinalities. Since this algorithm works on hashes, it is useful for high-cardinality fields to store the hash of values directly in the index, which is the purpose of the new `murmur3` field type. This is less necessary on low-cardinality string fields because the aggregator is smart enough to only compute the hash once per unique value per segment thanks to ordinals, or on numeric fields since hashing them is very fast. Close #5426 2014-02-16 13:20:33 -05:00			`"script": "doc['author.first_name'].value + ' ' + doc['author.last_name'].value"`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`
Docs: Mentioned script_id and script_file parameters across all aggs Closes #10760 2015-04-26 11:30:38 -04:00
			TIP: The `script` parameter expects an inline script. Use `script_id` for indexed scripts and `script_file` for scripts in the `config/scripts/` directory.

Aggs: Make it possible to configure missing values. Most aggregations (terms, histogram, stats, percentiles, geohash-grid) now support a new `missing` option which defines the value to consider when a field does not have a value. This can be handy if you eg. want a terms aggregation to handle the same way documents that have "N/A" or no value for a `tag` field. This works in a very similar way to the `missing` option on the `sort` element. One known issue is that this option sometimes cannot make the right decision in the unmapped case: it needs to replace all values with the `missing` value but might not know what kind of values source should be produced (numerics, strings, geo points?). For this reason, we might want to add an `unmapped_type` option in the future like we did for sorting. Related to #5324 2015-05-07 10:46:40 -04:00			`==== Missing value`

			The `missing` parameter defines how documents that are missing a value should be treated.
			`By default they will be ignored but it is also possible to treat them as if they`
			`had a value.`

			`[source,js]`
			`--------------------------------------------------`
			`{`
			`"aggs" : {`
			`"tag_cardinality" : {`
			`"cardinality" : {`
			`"field" : "tag",`
			`"missing": "N/A" <1>`
			`}`
			`}`
			`}`
			`}`
			`--------------------------------------------------`

			<1> Documents without a value in the `tag` field will fall into the same bucket as documents that have the value `N/A`.