OpenSearch/docs/reference/sql/limitations.asciidoc

[role="xpack"]
[testenv="basic"]
[[sql-limitations]]
== SQL Limitations

[float]
=== Nested fields in `SYS COLUMNS` and `DESCRIBE TABLE`

{es} has a special type of relationship fields called `nested` fields. In {es-sql} they can be used by referencing their inner
sub-fields. Even though `SYS COLUMNS` in non-driver mode (in the CLI and in REST calls) and `DESCRIBE TABLE` will still display
them as having the type `NESTED`, they cannot be used in a query. One can only reference its sub-fields in the form:

[source, sql]
--------------------------------------------------
[nested_field_name].[sub_field_name]
--------------------------------------------------

For example:

[source, sql]
--------------------------------------------------
SELECT dep.dep_name.keyword FROM test_emp GROUP BY languages;
--------------------------------------------------

[float]
=== Multi-nested fields

{es-sql} doesn't support multi-nested documents, so a query cannot reference more than one nested field in an index.
This applies to multi-level nested fields, but also multiple nested fields defined on the same level. For example, for this index:

[source, sql]
----------------------------------------------------
       column         |     type      |    mapping
----------------------+---------------+-------------
nested_A              |STRUCT         |NESTED
nested_A.nested_X     |STRUCT         |NESTED
nested_A.nested_X.text|VARCHAR        |KEYWORD
nested_A.text         |VARCHAR        |KEYWORD
nested_B              |STRUCT         |NESTED
nested_B.text         |VARCHAR        |KEYWORD
----------------------------------------------------

`nested_A` and `nested_B` cannot be used at the same time, nor `nested_A`/`nested_B` and `nested_A.nested_X` combination.
For such situations, {es-sql} will display an error message.

[float]
=== Paginating nested inner hits

When SELECTing a nested field, pagination will not work as expected, {es-sql} will return __at least__ the page size records. 
This is because of the way nested queries work in {es}: the root nested field will be returned and it's matching inner nested fields as well,
pagination taking place on the **root nested document and not on its inner hits**.

[float]
=== Normalized `keyword` fields

`keyword` fields in {es} can be normalized by defining a `normalizer`. Such fields are not supported in {es-sql}.

[float]
=== Array type of fields

Array fields are not supported due to the "invisible" way in which {es} handles an array of values: the mapping doesn't indicate whether
a field is an array (has multiple values) or not, so without reading all the data, {es-sql} cannot know whether a field is a single or multi value.
When multiple values are returned for a field, by default, {es-sql} will throw an exception. However, it is possible to change this behavior through `field_multi_value_leniency` parameter in REST (disabled by default) or
`field.multi.value.leniency` in drivers (enabled by default).

[float]
=== Sorting by aggregation

When doing aggregations (`GROUP BY`) {es-sql} relies on {es}'s `composite` aggregation for its support for paginating results.
However this type of aggregation does come with a limitation: sorting can only be applied on the key used for the aggregation's buckets. 
{es-sql} overcomes this limitation by doing client-side sorting however as a safety measure, allows only up to *512* rows.

It is recommended to use `LIMIT` for queries that use sorting by aggregation, essentially indicating the top N results that are desired:

[source, sql]
--------------------------------------------------
SELECT * FROM test GROUP BY age ORDER BY COUNT(*) LIMIT 100;
--------------------------------------------------

It is possible to run the same queries without a `LIMIT` however in that case if the maximum size (*512*) is passed, an exception will be
returned as {es-sql} is unable to track (and sort) all the results returned.

[float]
=== Using aggregation functions on top of scalar functions

Aggregation functions like <<sql-functions-aggs-min,`MIN`>>, <<sql-functions-aggs-max,`MAX`>>, etc. can only be used
directly on fields, and so queries like `SELECT MAX(abs(age)) FROM test` are not possible.

[float]
=== Using a sub-select

Using sub-selects (`SELECT X FROM (SELECT Y)`) is **supported to a small degree**: any sub-select that can be "flattened" into a single
`SELECT` is possible with {es-sql}. For example:

["source","sql",subs="attributes,macros"]
--------------------------------------------------
include-tagged::{sql-specs}/docs/docs.csv-spec[limitationSubSelect]
--------------------------------------------------

The query above is possible because it is equivalent with:

["source","sql",subs="attributes,macros"]
--------------------------------------------------
include-tagged::{sql-specs}/docs/docs.csv-spec[limitationSubSelectRewritten]
--------------------------------------------------

But, if the sub-select would include a `GROUP BY` or `HAVING` or the enclosing `SELECT` would be more complex than `SELECT X
FROM (SELECT ...) WHERE [simple_condition]`, this is currently **un-supported**.

[float]
=== Using <<sql-functions-aggs-first, `FIRST`>>/<<sql-functions-aggs-last,`LAST`>> aggregation functions in `HAVING` clause

Using `FIRST` and `LAST` in the `HAVING` clause is not supported. The same applies to
<<sql-functions-aggs-min,`MIN`>> and <<sql-functions-aggs-max,`MAX`>> when their target column
is of type <<keyword, `keyword`>> as they are internally translated to `FIRST` and `LAST`.

[float]
=== Using TIME data type in GROUP BY or <<sql-functions-grouping-histogram>>

Using `TIME` data type as a grouping key is currently not supported. For example:

[source, sql]
-------------------------------------------------------------
SELECT count(*) FROM test GROUP BY CAST(date_created AS TIME);
-------------------------------------------------------------

On the other hand, it can still be used if it's wrapped with a scalar function that returns another data type,
for example:

[source, sql]
-------------------------------------------------------------
SELECT count(*) FROM test GROUP BY MINUTE((CAST(date_created AS TIME));
-------------------------------------------------------------

`TIME` data type is also currently not supported in histogram grouping function. For example:

[source, sql]
-------------------------------------------------------------
SELECT HISTOGRAM(CAST(birth_date AS TIME), INTERVAL '10' MINUTES) as h, COUNT(*) FROM t GROUP BY h
-------------------------------------------------------------
SQL: documentation improvements and updates (#36918) * Added Limitations page * Made the aggregations page follow the common template for functions * Modified all tables to have the first row's cells content centered * Polishing in other various sections 2018-12-21 23:25:54 +02:00			`[role="xpack"]`
			`[testenv="basic"]`
			`[[sql-limitations]]`
			`== SQL Limitations`

			`[float]`
			=== Nested fields in `SYS COLUMNS` and `DESCRIBE TABLE`

			{es} has a special type of relationship fields called `nested` fields. In {es-sql} they can be used by referencing their inner
SQL: ignore UNSUPPORTED fields for JDBC and ODBC modes in 'SYS COLUMNS' (#39518) * SYS COLUMNS will skip UNSUPPORTED field types in ODBC and JDBC, as well. NESTED and OBJECT types were already skipped in ODBC mode, now they are skipped in JDBC mode, as well. (cherry picked from commit 9e0df64b2d36c9069dfa506570468f0522c86417) 2019-03-01 15:23:15 +02:00			sub-fields. Even though `SYS COLUMNS` in non-driver mode (in the CLI and in REST calls) and `DESCRIBE TABLE` will still display
			them as having the type `NESTED`, they cannot be used in a query. One can only reference its sub-fields in the form:
SQL: documentation improvements and updates (#36918) * Added Limitations page * Made the aggregations page follow the common template for functions * Modified all tables to have the first row's cells content centered * Polishing in other various sections 2018-12-21 23:25:54 +02:00
			`[source, sql]`
			`--------------------------------------------------`
			`[nested_field_name].[sub_field_name]`
			`--------------------------------------------------`

			`For example:`

			`[source, sql]`
			`--------------------------------------------------`
			`SELECT dep.dep_name.keyword FROM test_emp GROUP BY languages;`
			`--------------------------------------------------`

			`[float]`
			`=== Multi-nested fields`

			`{es-sql} doesn't support multi-nested documents, so a query cannot reference more than one nested field in an index.`
			`This applies to multi-level nested fields, but also multiple nested fields defined on the same level. For example, for this index:`

			`[source, sql]`
			`----------------------------------------------------`
			`column \| type \| mapping`
			`----------------------+---------------+-------------`
			`nested_A \|STRUCT \|NESTED`
			`nested_A.nested_X \|STRUCT \|NESTED`
			`nested_A.nested_X.text\|VARCHAR \|KEYWORD`
			`nested_A.text \|VARCHAR \|KEYWORD`
			`nested_B \|STRUCT \|NESTED`
			`nested_B.text \|VARCHAR \|KEYWORD`
			`----------------------------------------------------`

			`nested_A` and `nested_B` cannot be used at the same time, nor `nested_A`/`nested_B` and `nested_A.nested_X` combination.
			`For such situations, {es-sql} will display an error message.`

			`[float]`
			`=== Paginating nested inner hits`

			`When SELECTing a nested field, pagination will not work as expected, {es-sql} will return __at least__ the page size records.`
			`This is because of the way nested queries work in {es}: the root nested field will be returned and it's matching inner nested fields as well,`
			`pagination taking place on the root nested document and not on its inner hits.`

			`[float]`
			=== Normalized `keyword` fields

			`keyword` fields in {es} can be normalized by defining a `normalizer`. Such fields are not supported in {es-sql}.

			`[float]`
			`=== Array type of fields`

			`Array fields are not supported due to the "invisible" way in which {es} handles an array of values: the mapping doesn't indicate whether`
			`a field is an array (has multiple values) or not, so without reading all the data, {es-sql} cannot know whether a field is a single or multi value.`
SQL: Add multi_value_field_leniency inside FieldHitExtractor (#40113) For cases where fields can have multi values, allow the behavior to be customized through a dedicated configuration field. By default this will be enabled on the drivers so that existing datasets work instead of throwing an exception. For regular SQL usage, the behavior is false so that the user is aware of the underlying data. Fix #39700 (cherry picked from commit 2b351571961f172fd59290ee079126bbd081ceaf) 2019-03-18 14:56:00 +02:00			When multiple values are returned for a field, by default, {es-sql} will throw an exception. However, it is possible to change this behavior through `field_multi_value_leniency` parameter in REST (disabled by default) or
			`field.multi.value.leniency` in drivers (enabled by default).
SQL: documentation improvements and updates (#36918) * Added Limitations page * Made the aggregations page follow the common template for functions * Modified all tables to have the first row's cells content centered * Polishing in other various sections 2018-12-21 23:25:54 +02:00
			`[float]`
			`=== Sorting by aggregation`

			When doing aggregations (`GROUP BY`) {es-sql} relies on {es}'s `composite` aggregation for its support for paginating results.
SQL: Allow sorting of groups by aggregates (#38042) Introduce client-side sorting of groups based on aggregate functions. To allow this, the Analyzer has been extended to push down to underlying Aggregate, aggregate function and the Querier has been extended to identify the case and consume the results in order and sort them based on the given columns. The underlying QueryContainer has been slightly modified to allow a view of the underlying values being extracted as the columns used for sorting might not be requested by the user. The PR also adds minor tweaks, mainly related to tree output. Close #35118 2019-02-02 01:38:25 +02:00			`However this type of aggregation does come with a limitation: sorting can only be applied on the key used for the aggregation's buckets.`
			`{es-sql} overcomes this limitation by doing client-side sorting however as a safety measure, allows only up to 512 rows.`

			It is recommended to use `LIMIT` for queries that use sorting by aggregation, essentially indicating the top N results that are desired:

			`[source, sql]`
			`--------------------------------------------------`
			`SELECT * FROM test GROUP BY age ORDER BY COUNT(*) LIMIT 100;`
			`--------------------------------------------------`

			It is possible to run the same queries without a `LIMIT` however in that case if the maximum size (512) is passed, an exception will be
			`returned as {es-sql} is unable to track (and sort) all the results returned.`
SQL: add sub-selects to the Limitations page (#37012) 2019-01-07 10:08:51 +02:00
SQL: [Docs] Add limitation for aggregate functions on scalars (#38186) Currently aggregate functions can operate only directly on fields. They cannot be used on top of scalar functions as painless scripting is currently not supported. 2019-02-01 16:13:51 +02:00			`[float]`
			`=== Using aggregation functions on top of scalar functions`

			Aggregation functions like <<sql-functions-aggs-min,`MIN`>>, <<sql-functions-aggs-max,`MAX`>>, etc. can only be used
			directly on fields, and so queries like `SELECT MAX(abs(age)) FROM test` are not possible.

SQL: add sub-selects to the Limitations page (#37012) 2019-01-07 10:08:51 +02:00			`[float]`
			`=== Using a sub-select`

			Using sub-selects (`SELECT X FROM (SELECT Y)`) is supported to a small degree: any sub-select that can be "flattened" into a single
			`SELECT` is possible with {es-sql}. For example:

			`["source","sql",subs="attributes,macros"]`
			`--------------------------------------------------`
SQL: Spec tests now use classpath discovery (#40388) To avoid having to specify each spec by hand (which can miss specs to be added), the test infrastructure now performs classpath discovery so that each spec added, is automatically considered. Relates #40358 (cherry picked from commit d0f60b4425c731509aa8ca765d55f563f866ef90) 2019-03-25 15:22:59 +02:00			`include-tagged::{sql-specs}/docs/docs.csv-spec[limitationSubSelect]`
SQL: add sub-selects to the Limitations page (#37012) 2019-01-07 10:08:51 +02:00			`--------------------------------------------------`

			`The query above is possible because it is equivalent with:`

			`["source","sql",subs="attributes,macros"]`
			`--------------------------------------------------`
SQL: Spec tests now use classpath discovery (#40388) To avoid having to specify each spec by hand (which can miss specs to be added), the test infrastructure now performs classpath discovery so that each spec added, is automatically considered. Relates #40358 (cherry picked from commit d0f60b4425c731509aa8ca765d55f563f866ef90) 2019-03-25 15:22:59 +02:00			`include-tagged::{sql-specs}/docs/docs.csv-spec[limitationSubSelectRewritten]`
SQL: add sub-selects to the Limitations page (#37012) 2019-01-07 10:08:51 +02:00			`--------------------------------------------------`

			But, if the sub-select would include a `GROUP BY` or `HAVING` or the enclosing `SELECT` would be more complex than `SELECT X
			FROM (SELECT ...) WHERE [simple_condition]`, this is currently un-supported.
SQL: Implement FIRST/LAST aggregate functions (#37936) FIRST and LAST can be used with one argument and work similarly to MIN and MAX but they are implemented using a Top Hits aggregation and therefore can also operate on keyword fields. When a second argument is provided then they return the first/last value of the first arg when its values are ordered ascending/descending (respectively) by the values of the second argument. Currently because of the usage of a Top Hits aggregation FIRST and LAST cannot be used in the HAVING clause of a GROUP BY query to filter on the results of the aggregation. Closes: #35639 2019-01-31 16:33:05 +02:00
			`[float]`
SQL: [Docs] Add limitation for aggregate functions on scalars (#38186) Currently aggregate functions can operate only directly on fields. They cannot be used on top of scalar functions as painless scripting is currently not supported. 2019-02-01 16:13:51 +02:00			=== Using <<sql-functions-aggs-first, `FIRST`>>/<<sql-functions-aggs-last,`LAST`>> aggregation functions in `HAVING` clause
SQL: Implement FIRST/LAST aggregate functions (#37936) FIRST and LAST can be used with one argument and work similarly to MIN and MAX but they are implemented using a Top Hits aggregation and therefore can also operate on keyword fields. When a second argument is provided then they return the first/last value of the first arg when its values are ordered ascending/descending (respectively) by the values of the second argument. Currently because of the usage of a Top Hits aggregation FIRST and LAST cannot be used in the HAVING clause of a GROUP BY query to filter on the results of the aggregation. Closes: #35639 2019-01-31 16:33:05 +02:00
			Using `FIRST` and `LAST` in the `HAVING` clause is not supported. The same applies to
			<<sql-functions-aggs-min,`MIN`>> and <<sql-functions-aggs-max,`MAX`>> when their target column
			is of type <<keyword, `keyword`>> as they are internally translated to `FIRST` and `LAST`.
SQL: Introduce SQL TIME data type (#39802) Support ANSI SQL's TIME type by introductin a runtime-only ES SQL time type. Closes: #38174 (cherry picked from commit 046ccd4cf0a251b2a3ddff6b072ab539a6711900) 2019-04-01 23:30:39 +02:00
			`[float]`
			`=== Using TIME data type in GROUP BY or <<sql-functions-grouping-histogram>>`

			Using `TIME` data type as a grouping key is currently not supported. For example:

			`[source, sql]`
			`-------------------------------------------------------------`
			`SELECT count(*) FROM test GROUP BY CAST(date_created AS TIME);`
			`-------------------------------------------------------------`

			`On the other hand, it can still be used if it's wrapped with a scalar function that returns another data type,`
			`for example:`

			`[source, sql]`
			`-------------------------------------------------------------`
			`SELECT count(*) FROM test GROUP BY MINUTE((CAST(date_created AS TIME));`
			`-------------------------------------------------------------`

			`TIME` data type is also currently not supported in histogram grouping function. For example:

			`[source, sql]`
			`-------------------------------------------------------------`
			`SELECT HISTOGRAM(CAST(birth_date AS TIME), INTERVAL '10' MINUTES) as h, COUNT(*) FROM t GROUP BY h`
			`-------------------------------------------------------------`