druid

Commit Graph

Author	SHA1	Message	Date
Maytas Monsereenusorn	fac6a48a8f	add impl (#12201 )	2022-01-27 11:39:59 -08:00
zachjsh	f906f2f577	Fix HttpRemoteTaskRunner LifecycleStart / LifecycleStop race condition (#12184 ) * * stop workers, remove listener, and call exitStop() on HttpRemoteTaskRunner @LifecycleStop * * fix test failure	2022-01-27 13:15:14 -06:00
Abhishek Agarwal	1b8808cce8	Fix SQL queries for inline datasource with null values (#12092 ) Fixes a bug because of which some SQL queries cannot be parsed using druid convention. Specifically, these queries translate to an inline datasource and have some null values. Calcite internally uses NULL as SQL type for these literals and that is not supported by the druid. I am now allowing null column types to be returned while building RowSignature in org.apache.druid.sql.calcite.table.RowSignatures#fromRelDataType. RowSignature already allows null column type for any column. Doing so should also fix bindable queries such as select (1,2). When such queries are run with headers set to true, we get an exception in org.apache.druid.sql.http.ArrayWriter#writeHeader. This is again a similar exception to the one addressed in this PR. Because SQL type for the result column is RECORD and that doesn't have a corresponding columnType.	2022-01-27 18:04:12 +05:30
TSFenwick	a813816fb1	add module test for QueryableModule to allow for better runtime.properties testing (#12202 ) added a default GetRequestLoggerProviderTest and GetEmitterRequestLoggerProviderTest	2022-01-25 22:26:11 -08:00
Karan Kumar	96b3498a40	Grouping on arrays as arrays (#12078 ) * init multiValue column group by * Changing sorting to Lexicographic as default * Adding initial tests * 1.Fixing test cases adding 2.Optimized inmem structs * Linking SQL layer to native layer * Adding multiDimension support to group by column strategy * 1. Removing array coercion in Calcite layer 2. Removing ResultRowDeserializer * 1. Supporting all primitive array types 2. Removing dimension spec as part of columnSelector * 1. Supporting all primitive array types 2. Removing dimension spec as part of columnSelector * 1. Checkstyle things 2. Removing flag * Minor naming things * CheckStyle Things * Fixing test case * Fixing hashing * 1. Adding the MV function 2. Added few test cases * 1. Adding MV function test cases * Adding Selector strategy function test cases * Fixing ClientQuerySegmentWalkerTest * Adding GroupByQueryRunnerTest test cases * Fixing test cases * Adding few more test cases * Fixing Exception asset statement and intellij inspection * Adding null compatibility tests * Review comments * Fixing few failing tests * Fixing few failing tests * Do no convert to topN Q incase of group by on array * Fixing checkstyle * Fixing differences between jdk's class cast exception message * 1. Fixing ordering if the grouping key is an array * Fixing DefaultLimitSpec * Fixing CalciteArraysQueryTest * Dummy commit for LGTM * changes: * only coerce multi-value string null values when `ExpressionPlan.Trait.NEEDS_APPLIED` is set * correct return type inference for ARRAY_APPEND,ARRAY_PREPEND,ARRAY_SLICE,ARRAY_CONCAT * fix bug with ExprEval.ofType when actual type of object from binding doesn't match its claimed type * Review comments * Fixing test cases * Fixing spot bugs * Fixing strict compile Co-authored-by: Clint Wylie <cwylie@apache.org>	2022-01-25 20:30:56 -08:00
Vadim Ogievetsky	8ae5de5114	Web console: fix multi-value dimension column detection and tidy up (#12160 ) * streamline services view query * better column type detection * fix query view page size bug * fill out MetricSpec interface * fix pagination in status dialog * update tests * adjust pagination * better type guessing * better test fixtures * add Avg. row size to segments view	2022-01-25 15:46:29 -08:00
Clint Wylie	fce62b2643	fix StringAnyAggregatorFactory to use single value selector for non-existent columns (#12194 )	2022-01-25 12:52:30 -08:00
Jihoon Son	20347e0c86	Wait for datasource to be ready for SQL in integration tests (#12189 ) * Wait for datasource to be ready for SQL in integration tests * add limit to the check query	2022-01-25 10:14:26 -08:00
Suneet Saldanha	2b32d86f3b	Enable automatic metdata cleanup by default (#12188 )	2022-01-24 20:04:17 -08:00
somu-imply	cc8b9c0b6e	Handling OOM error in ExpressionVector setup by reducing number of rows (#12186 ) * Handling OOM error in ExpressionVector setup by reducing number of rows * Removing row size to 10K in sanity tests	2022-01-24 08:37:13 -08:00
Laksh Singla	dc1703d5f9	Change value of `druid.sql.planner.useGroupingSetForExactDistinct` in common.runtime.properties (#12182 ) This PR changes the value of the property `druid.sql.planner.useGroupingSetForExactDistinct` from `false` to `true` in the runtime.properties files, so that newer installations have this property as `true`, while the default still remains as `false`. The flag determines how queries which contain an aggregation over `DISTINCT` like `SELECT COUNT(DISTINCT foo.dim1) FILTER(WHERE foo.cnt = 1), SUM(foo.cnt) FROM druid.foo` get planned by Calcite. With the flag being set to false, it plans it via joins, whereas with it being set to true, the query is set using grouping sets. There is a known issue with Calcite (https://github.com/apache/druid/issues/7953), where an NPE is thrown while planning the above query with joins. There is no such issue while planning the query using grouping sets.	2022-01-24 14:00:03 +05:30
JoyKing	ac87bdd736	fix typo in materialized view (#12174 )	2022-01-22 11:32:22 +08:00
AmatyaAvadhanula	1f63b447c4	Mitigate Kinesis stream LimitExceededException by using listShards API (#12161 ) Makes kinesis ingestion resilient to `LimitExceededException` caused by resharding. Replace `describeStream` with `listShards` (recommended) to get shard related info. `describeStream` has a limit (100) to the number of shards returned per call and a low default TPS limit of 10. `listShards` returns the info for at most 1000 shards and has a higher TPS limit of 100 as well. Key changed/added classes in this PR * `KinesisRecordSupplier` * `KinesisAdminClient`	2022-01-21 10:15:51 +05:30
zachjsh	376d7c069d	Close provisioner during HttpRemotetaskRunner LifecycleStop (#12176 ) Fixed an issue where the provisionerService which can be used to spawn resources as needed is left running on a non-leader coordinator/overlord, after it is removed from leadership. Provisioning should only be done by the leader. To fix the issue, a call to stop the provisionerService was added to the stop() method of HttpRemoteTaskRunner class. The provisionerService was properly closed on other TaskRunner types.	2022-01-20 13:32:08 -05:00
Vadim Ogievetsky	6ce14e6b17	allow W in durations (#12175 )	2022-01-20 09:17:43 +05:30
Uwe Schindler	1f7dd6d86c	Forbiddenapis: Split the guava16-only signatures file from main signatures file (#12170 )	2022-01-19 17:50:28 -08:00
Jihoon Son	cacfcfcdab	ignore hadoop-gcs directory already exists error for integration tests (#12169 )	2022-01-19 09:35:50 -08:00
Jihoon Son	cc2ffc6c0f	Fix node discovery to ignore unknown DruidServices (#12157 ) * Fix node discovery to ignore unknown DruidServices * ignore all runtime exceptions * fix test * add custom deserializer * custom serializer * log host for unparseable druidService	2022-01-18 22:08:59 -08:00
Abhishek Agarwal	53c0e489c2	Fix infinite retrying during task pausing (#12167 ) This fixes a bug that causes TaskClient in overlord to continuously retry to pause tasks. This can happen when a task is not responding to the pause command. Ideally, in such a case when the task is unresponsive, the overlord would have given up after a few retries and would have killed the task. However, due to this bug, retries go on forever.	2022-01-19 09:03:36 +05:30
Gian Merlino	cf7191d2bc	Validate target dataSource for INSERT. (#12129 )	2022-01-18 09:34:23 -08:00
Abhishek Agarwal	7bdb9ebdf1	Suppress Avro CVEs (#12166 )	2022-01-18 21:09:48 +05:30
Maytas Monsereenusorn	bd7fe45da0	Support adding metrics in Auto Compaction (#12125 ) * add impl * add impl * add unit tests * add unit tests * add unit tests * add unit tests * add unit tests * add integration tests * add integration tests * fix LGTM * fix test * remove doc	2022-01-17 20:19:31 -08:00
Benedict Jin	b55f7a25fe	Fix forbiddenapis causing travis failing (#12158 ) * Fix forbiddenapis causing travis failing * Use failOnUnresolvableSignatures instead	2022-01-15 16:13:37 -08:00
Victoria Lim	d2ac146365	Docs for cluster tiering to improve query concurrency (#12128 ) * add new doc * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> * reorder query laning properties * rename doc * new name in doc header * organize material into "service tiering" section * text edits and update sidebars.json * update query laning * how queries get assigned to lanes * add more details to intro; use more consistent terminology * more content * Apply suggestions from code review Co-authored-by: Jihoon Son <jihoonson@apache.org> * Update docs/operations/mixed-workloads.md * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> * typo Co-authored-by: Charles Smith <techdocsmith@gmail.com> Co-authored-by: Jihoon Son <jihoonson@apache.org>	2022-01-15 12:22:08 +08:00
Ivan Vankovich	6a93872586	OpenTelemetry emitter extension (#12015 ) * Add OpenTelemetry emitter extension * Fix build * Fix checkstyle * Add used undeclared dependencies * Ignore unused declared dependencies	2022-01-15 12:18:04 +08:00
Clint Wylie	e0c4c568cb	fix incorrect ColumnInspector in IncrementalIndex.makeColumnSelectorFactory (#12155 )	2022-01-13 18:09:06 -08:00
aggarwalakshay	eb4fafe08f	Upgrading follow-redirects to 1.14.7 (#12153 ) * Upgrading follow-redirects to 1.14.7 * removed the existing follow-redirects i.e. 1.14.4 from package-lock.json	2022-01-13 14:01:36 -08:00
Jonathan Wei	74c876e578	Throw parse exceptions on schema get errors for SchemaRegistryBasedAvroBytesDecoder (#12080 ) * Add option to throw parse exceptions on schema get errors for SchemaRegistryBasedAvroBytesDecoder * Remove option	2022-01-13 12:36:51 -06:00
Clint Wylie	1dba089a62	fix array type strategy write size tracking (#12150 ) * fix array type strategy write size tracking * fix checkstyle	2022-01-13 10:22:40 -08:00
Xavier Léauté	e56ea31697	follow-up to fix formatting broken in #12147 (#12148 ) follow-up to #12147 to fix the build	2022-01-12 20:59:32 -08:00
Jihoon Son	58378aa967	Move gcs-connector from lib to hadoop-dependencies for integration test (#12144 )	2022-01-12 16:47:34 -08:00
Xavier Léauté	168187e6df	avoid unnecessary String.format calls in IdUtils.validateId (#12147 ) Based on profiling data, about 25% of the time de-serializing DataSchema is spent on formatting strings in validateId. This can add up quickly, especially when de-serializing task information in the overlord, where in can consume almost 2% of CPU if there are many tasks. Since the formatting is unnecessary unless the checks fail, we can leverage the built-in formatting of Preconditions.checkArgument instead to avoid the cost.	2022-01-12 16:34:40 -08:00
Vadim Ogievetsky	9cd52ed914	Web console: make range partitioning a first class citizen of the console (#12146 ) * first class support for range partitioning * update e2e tests	2022-01-12 03:50:10 -08:00
Clint Wylie	f2ce76966c	add EARLIEST_BY/LATEST_BY to make EARLIEST/LATEST function signatures less ambiguous (#12145 ) * add EARLIEST_BY/LATEST_BY to make EARLIEST/LATEST function signatures unambiguous * switcheroo * EARLIEST_BY/LATEST_BY use timestamp instead of numeric types, update docs * revert unintended change * fix docs * fix docs better	2022-01-12 03:48:53 -08:00
Laksh Singla	fae73800a7	Set plannerContext error when cannot query external datasources and when insert is not supported. (#12136 ) This PR aims to add plannerContext.setPlanningError whenever external table scan rule is invoked, without the queryMaker having the ability to do so.	2022-01-12 15:11:17 +05:30
Rohan Garg	81f0aba6cb	Use ListFilteredVirtualColumn for left/fact table expression in join condition (#12127 ) * Pass VirtualColumnRegistry in PlannerContext for join expression planning * Allow for including VCs from join fact table expression * Optmize MV_FILTER functions to use a VC when in join fact table expression * fixup! Allow for including VCs from join fact table expression * Address review comments	2022-01-11 14:47:13 -08:00
imply-cheddar	b153cb2342	Add a small LRU cache and use utf8 bytes in ArrayOfDoubles (#12130 ) * Add a small LRU cache and use utf8 bytes in ArrayOfDoubles * Add tests for extra branches * Even more tests for branch coverage * Fix Style	2022-01-11 13:04:11 -08:00
somu-imply	08fea7a46a	input type validation for datasketches hll "build" aggregator factory (#12131 ) * Ingestion will fail for HLLSketchBuild instead of creating with incorrect values * Addressing review comments for HLL< updated error message introduced test case	2022-01-11 12:00:14 -08:00
imply-cheddar	eb0bae49ec	Update PostAggregator to be backwards compat (#12138 ) This change mimics what was done in PR #11917 to fix the incompatibilities produced by #11713. #11917 fixed it with AggregatorFactory by creating default methods to allow for extensions built against old jars to still work. This does the same for PostAggregator	2022-01-11 02:18:14 -08:00
Clint Wylie	7cf9192765	fix delegated smoosh writer and some new facilities for segment writeout medium (#12132 ) * fix delegated smoosh writer and some new facilities for segment writeout medium changes: * fixed issue with delegated `SmooshedWriter` when writing files that look like paths, causing `NoSuchFileException` exceptions when attempting to open a channel to the file * `FileSmoosher.addWithSmooshedWriter` when _not_ delegating now checks that it is still open when closing, making it a no-op if already closed (allowing column serializers to add additional files and avoid delegated mode if they are finished writing out their own content and ned to add additional files) * add `makeChildWriteOutMedium` to `SegmentWriteOutMedium` interface, which allows users of a shared medium to clean up `WriteOutBytes` if they fully control the lifecycle. there are no callers of this yet, adding for future functionality * `OnHeapByteBufferWriteOutBytes` now can be marked as not open so it `OnHeapMemorySegmentWriteOutMedium` can now behave identically to other medium implementations * fix to address nit - use AtomicLong	2022-01-10 22:25:19 -08:00
Frank Chen	c8ddf60851	Upgrade RSA Key from 1024 bit to 4096 to eliminate warnings (#11743 ) * eliminate warnings * Change the keyStore type to PKCS12	2022-01-11 13:24:09 +08:00
Clint Wylie	e583033231	add 'TypeStrategy' to types (#11888 ) * add TypeStrategy - value comparators and binary serialization for any TypeSignature	2022-01-10 17:12:14 -08:00
Vadim Ogievetsky	2a41b7bffa	Web console: correctly cancel JSON shaped SQL queries (#12134 ) * misc fixes * type typo	2022-01-10 14:24:05 -08:00
Laksh Singla	7c17341caa	Return empty result when a group by gets optimized to a timeseries query (#12065 ) Related to #11188 The above mentioned PR allowed timeseries queries to return a default result, when queries of type: select count() from table where dim1="_not_present_dim_" were executed. Before the PR, it returned no row, after the PR, it would return a row with value of count() as 0 (as expected by SQL standards of different dbs). In Grouping#applyProject, we can sometimes perform optimization of a groupBy query to a timeseries query if possible (when the keys of the groupBy are constants, as generated by automated tools). For example, in select count() from table where dim1="_present_dim_" group by "dummy_key", the groupBy clause can be removed. However, in the case when the filter doesn't return anything, i.e. select count() from table where dim1="_not_present_dim_" group by "dummy_key", the behavior of general databases would be to return nothing, while druid (due to above change) returns an empty row. This PR aims to fix this divergence of behavior. Example cases: select count() from table where dim1="_not_present_dim_" group by "dummy_key". CURRENT: Returns a row with count() = 0 EXPECTED: Return no row select 'A', dim1 from foo where m1 = 123123 and dim1 = '_not_present_again_' group by dim1 CURRENT: Returns a row with ('A', 'wat') EXPECTED: Return no row To do this, a boolean droppedDimensionsWhileApplyingProject has been added to Grouping which is true whenever we make changes to the original shape with optimization. Hence if a timeseries query has a grouping with this set to true, we set skipEmptyBuckets=true in the query context (i.e. donot return any row).	2022-01-07 21:53:48 +05:30
Vadim Ogievetsky	2299eb321e	Standardizing SQL function docs (#12091 ) * fix typos in SQL function docs * more code * Update docs/querying/sql.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update docs/querying/sql.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update docs/querying/sql.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update docs/querying/sql.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update docs/querying/sql.md Co-authored-by: Frank Chen <frankchen@apache.org> * a few more expr, fixes * more fixes * quote TIME_SHIFT * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * Update docs/querying/sql.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> * undo header change Co-authored-by: Frank Chen <frankchen@apache.org> Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2022-01-06 23:57:03 -08:00
Marcelo R Costa	c28b2834a1	Add http response status code to org.eclipse.jetty.server.RequestLog (#12116 ) * Add http response status code to org.eclipse.jetty.server.RequestLog * http response code is expressed as an int. Set log msg interpolation based on digit * trying to add an unit test to verify if the logger.debug method is called * trying to add an unit test to verify if the logger.debug method is called * fix compilation issues * remove test	2022-01-06 20:10:01 +08:00
Jihoon Son	4a74c5adcc	Use Druid's extension loading for integration test instead of maven (#12095 ) * Use Druid's extension loading for integration test instead of maven * fix maven command * override config path * load input format extensions and kafka by default; add prepopulated-data group * all docker-composes are overridable * fix s3 configs * override config for all * fix docker_compose_args * fix security tests * turn off debug logs for overlord api calls * clean up stuff * revert docker-compose.yml * fix override config for query error test; fix circular dependency in docker compose * add back some dependencies in docker compose * new maven profile for integration test * example file filter	2022-01-05 23:33:04 -08:00
Victoria Lim	6846622080	Docs: add FILTER to sql query syntax (#12093 ) * docs: add FILTER to sql query syntax * Update docs/querying/sql.md * Update docs/querying/sql.md * Update docs/querying/sql.md * Update docs/querying/sql.md * move and update FILTER section	2022-01-05 12:59:41 -08:00
Maytas Monsereenusorn	b53e7f4d12	Support overlapping segment intervals in auto compaction (#12062 ) * add impl * add impl * fix more bugs * add tests * fix checkstyle * address comments * address comments * fix test	2022-01-04 11:47:38 -08:00
Frank Chen	fe71fc414f	Update log4j2 to 2.17.1 (#12106 ) Signed-off-by: frank chen <frank.chen021@outlook.com>	2021-12-30 19:18:16 -06:00

1 2 3 4 5 ...

11513 Commits All Branches Search

11513 Commits

All Branches