druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	a3affe1471	make EncodedKeyComponent constructor public, remove nullable from DimensionIndexer.processRowValsToUnsortedEncodedKeyComponent (#12229 )	2022-02-03 15:02:32 -08:00
Clint Wylie	8fd587b28c	remove duplicate Broker ServerInventoryView, improve HttpServerInventoryView logging (#12209 ) * changes: * remove SystemSchema duplicate ServerInventoryView in broker * suppress duplicate segment added/removed warnings in HttpServerInventoryView when doing a full sync * fixes	2022-02-03 12:57:34 -08:00
Maytas Monsereenusorn	3717693633	Fix java.lang.ClassCastException error when using useApproximateCountDistinct false for aggregation query (#12216 ) * add imply * add test * add unit test * add test	2022-02-03 12:01:13 -08:00
Vadim Ogievetsky	fc76b014d1	Web console: fix supervisor stats table pagination (#12227 ) * fixes #11627 supervisor stats table pagination * use spread instead of assign	2022-02-03 00:09:21 -08:00
Kashif Faraz	e648b01afb	Improve memory estimates in Aggregator and DimensionIndexer (#12073 ) Fixes #12022 ### Description The current implementations of memory estimation in `OnHeapIncrementalIndex` and `StringDimensionIndexer` tend to over-estimate which leads to more persistence cycles than necessary. This PR replaces the max estimation mechanism with getting the incremental memory used by the aggregator or indexer at each invocation of `aggregate` or `encode` respectively. ### Changes - Add new flag `useMaxMemoryEstimates` in the task context. This overrides the same flag in DefaultTaskConfig i.e. `druid.indexer.task.default.context` map - Add method `AggregatorFactory.factorizeWithSize()` that returns an `AggregatorAndSize` which contains the aggregator instance and the estimated initial size of the aggregator - Add method `Aggregator.aggregateWithSize()` which returns the incremental memory used by this aggregation step - Update the method `DimensionIndexer.processRowValsToKeyComponent()` to return the encoded key component as well as its effective size in bytes - Update `OnHeapIncrementalIndex` to use the new estimations only if `useMaxMemoryEstimates = false`	2022-02-03 10:34:02 +05:30
Vadim Ogievetsky	bc408bacc8	Web console: Adding a shard detail column to the segments view (#12212 ) * shard spec details * improve pattern match * refactor spec cleanup * better format detection * update JSONbig * add multiline option to autoform	2022-02-02 18:46:17 -08:00
AshishKapoor	801d9e7f1b	[Web Console] fix deprecated keyboard event method "keyCode" with "key" (#11947 ) * 11946 fix keyboard event keyCode method * fix with key and respective cases * e.which method required since it's anyway deprecated too. * updated as per feedback	2022-02-02 18:26:21 -08:00
zachjsh	f47e1e0dcc	Reduce RemoteTaskRunnerTest flakiness (#12211 ) * * add more logging to start / stop of RemoteTaskRunner * * add more logging * Increase timeout on RemoteTaskRunnerTest * Apply suggestions from code review Co-authored-by: Suneet Saldanha <suneet@apache.org> Co-authored-by: Suneet Saldanha <suneet@apache.org>	2022-02-01 15:35:18 -08:00
Clint Wylie	f9b406c8f2	add backwards compatibility mode for multi-value string array null value coercion (#12210 )	2022-01-31 22:38:15 -08:00
Clint Wylie	978b8f7dde	do not explode if mysql transient exception class does not exist (#12213 ) Follow up to #12205 to allow druid-mysql-extensions to work with mysql connector/j 8.x again, which does not contain MySQLTransientException, and while would have had the same problem as mariadb if a transient exception was checked, the new check eagerly loads the class when starting up, causing immediate failure.	2022-02-01 09:06:24 +05:30
Rohan Garg	c4fa3ccfc4	Fix load-drop-load sequence for same segment and historical in http loadqueue peon (#11717 ) Fixes an issue where a load-drop-load sequence for a segment and historical doesn't work correctly for http based load queue peon. The first cycle of load-drop works fine - the problem comes when there is an attempt to reload the segment. The historical caches load success for some recent segments and makes the reload as a no-op. But it doesn't consider that fact that the segment was also dropped in between the load requests. This change invalidates the cache after a client tries to fetch a success result.	2022-01-31 13:16:58 +05:30
Vadim Ogievetsky	fe8530dac4	Change link to Apache Druid Slack (#12206 )	2022-01-27 21:10:40 -08:00
Jihoon Son	eeed156dc0	Fix compile error in VirtualizedColumnSelectorFactoryTest (#12208 )	2022-01-27 17:35:50 -08:00
Gian Merlino	99a5c2f3d3	Harmonize behavior when virtual columns reference each other. (#11955 ) * VirtualizedColumnSelectorFactory: Allow virtual columns to reference each other. This matches the behavior of QueryableIndex and IncrementalIndex based cursors. * Fixes to getColumnCapabilities.	2022-01-27 14:31:48 -08:00
Clint Wylie	5d2291991e	use reflection to check for mysql transient exception type (#12205 ) * use reflection to check for mysql transient exception type * better * oops	2022-01-27 13:13:16 -08:00
Victoria Lim	24716bfedc	Doc updates for metadata cleanup and storage (#12190 ) * doc updates for metadata storage/cleanup * Add comments for disabling cleanup * Apply suggestions from code review * updated for https://github.com/apache/druid/pull/12201 * Apply suggestions from code review Co-authored-by: Maytas Monsereenusorn <maytasm@apache.org> * move retention period line earlier; more concise text * fix typo Co-authored-by: Maytas Monsereenusorn <maytasm@apache.org>	2022-01-27 11:40:54 -08:00
Maytas Monsereenusorn	fac6a48a8f	add impl (#12201 )	2022-01-27 11:39:59 -08:00
zachjsh	f906f2f577	Fix HttpRemoteTaskRunner LifecycleStart / LifecycleStop race condition (#12184 ) * * stop workers, remove listener, and call exitStop() on HttpRemoteTaskRunner @LifecycleStop * * fix test failure	2022-01-27 13:15:14 -06:00
Abhishek Agarwal	1b8808cce8	Fix SQL queries for inline datasource with null values (#12092 ) Fixes a bug because of which some SQL queries cannot be parsed using druid convention. Specifically, these queries translate to an inline datasource and have some null values. Calcite internally uses NULL as SQL type for these literals and that is not supported by the druid. I am now allowing null column types to be returned while building RowSignature in org.apache.druid.sql.calcite.table.RowSignatures#fromRelDataType. RowSignature already allows null column type for any column. Doing so should also fix bindable queries such as select (1,2). When such queries are run with headers set to true, we get an exception in org.apache.druid.sql.http.ArrayWriter#writeHeader. This is again a similar exception to the one addressed in this PR. Because SQL type for the result column is RECORD and that doesn't have a corresponding columnType.	2022-01-27 18:04:12 +05:30
TSFenwick	a813816fb1	add module test for QueryableModule to allow for better runtime.properties testing (#12202 ) added a default GetRequestLoggerProviderTest and GetEmitterRequestLoggerProviderTest	2022-01-25 22:26:11 -08:00
Karan Kumar	96b3498a40	Grouping on arrays as arrays (#12078 ) * init multiValue column group by * Changing sorting to Lexicographic as default * Adding initial tests * 1.Fixing test cases adding 2.Optimized inmem structs * Linking SQL layer to native layer * Adding multiDimension support to group by column strategy * 1. Removing array coercion in Calcite layer 2. Removing ResultRowDeserializer * 1. Supporting all primitive array types 2. Removing dimension spec as part of columnSelector * 1. Supporting all primitive array types 2. Removing dimension spec as part of columnSelector * 1. Checkstyle things 2. Removing flag * Minor naming things * CheckStyle Things * Fixing test case * Fixing hashing * 1. Adding the MV function 2. Added few test cases * 1. Adding MV function test cases * Adding Selector strategy function test cases * Fixing ClientQuerySegmentWalkerTest * Adding GroupByQueryRunnerTest test cases * Fixing test cases * Adding few more test cases * Fixing Exception asset statement and intellij inspection * Adding null compatibility tests * Review comments * Fixing few failing tests * Fixing few failing tests * Do no convert to topN Q incase of group by on array * Fixing checkstyle * Fixing differences between jdk's class cast exception message * 1. Fixing ordering if the grouping key is an array * Fixing DefaultLimitSpec * Fixing CalciteArraysQueryTest * Dummy commit for LGTM * changes: * only coerce multi-value string null values when `ExpressionPlan.Trait.NEEDS_APPLIED` is set * correct return type inference for ARRAY_APPEND,ARRAY_PREPEND,ARRAY_SLICE,ARRAY_CONCAT * fix bug with ExprEval.ofType when actual type of object from binding doesn't match its claimed type * Review comments * Fixing test cases * Fixing spot bugs * Fixing strict compile Co-authored-by: Clint Wylie <cwylie@apache.org>	2022-01-25 20:30:56 -08:00
Vadim Ogievetsky	8ae5de5114	Web console: fix multi-value dimension column detection and tidy up (#12160 ) * streamline services view query * better column type detection * fix query view page size bug * fill out MetricSpec interface * fix pagination in status dialog * update tests * adjust pagination * better type guessing * better test fixtures * add Avg. row size to segments view	2022-01-25 15:46:29 -08:00
Clint Wylie	fce62b2643	fix StringAnyAggregatorFactory to use single value selector for non-existent columns (#12194 )	2022-01-25 12:52:30 -08:00
Jihoon Son	20347e0c86	Wait for datasource to be ready for SQL in integration tests (#12189 ) * Wait for datasource to be ready for SQL in integration tests * add limit to the check query	2022-01-25 10:14:26 -08:00
Suneet Saldanha	2b32d86f3b	Enable automatic metdata cleanup by default (#12188 )	2022-01-24 20:04:17 -08:00
somu-imply	cc8b9c0b6e	Handling OOM error in ExpressionVector setup by reducing number of rows (#12186 ) * Handling OOM error in ExpressionVector setup by reducing number of rows * Removing row size to 10K in sanity tests	2022-01-24 08:37:13 -08:00
Laksh Singla	dc1703d5f9	Change value of `druid.sql.planner.useGroupingSetForExactDistinct` in common.runtime.properties (#12182 ) This PR changes the value of the property `druid.sql.planner.useGroupingSetForExactDistinct` from `false` to `true` in the runtime.properties files, so that newer installations have this property as `true`, while the default still remains as `false`. The flag determines how queries which contain an aggregation over `DISTINCT` like `SELECT COUNT(DISTINCT foo.dim1) FILTER(WHERE foo.cnt = 1), SUM(foo.cnt) FROM druid.foo` get planned by Calcite. With the flag being set to false, it plans it via joins, whereas with it being set to true, the query is set using grouping sets. There is a known issue with Calcite (https://github.com/apache/druid/issues/7953), where an NPE is thrown while planning the above query with joins. There is no such issue while planning the query using grouping sets.	2022-01-24 14:00:03 +05:30
JoyKing	ac87bdd736	fix typo in materialized view (#12174 )	2022-01-22 11:32:22 +08:00
AmatyaAvadhanula	1f63b447c4	Mitigate Kinesis stream LimitExceededException by using listShards API (#12161 ) Makes kinesis ingestion resilient to `LimitExceededException` caused by resharding. Replace `describeStream` with `listShards` (recommended) to get shard related info. `describeStream` has a limit (100) to the number of shards returned per call and a low default TPS limit of 10. `listShards` returns the info for at most 1000 shards and has a higher TPS limit of 100 as well. Key changed/added classes in this PR * `KinesisRecordSupplier` * `KinesisAdminClient`	2022-01-21 10:15:51 +05:30
zachjsh	376d7c069d	Close provisioner during HttpRemotetaskRunner LifecycleStop (#12176 ) Fixed an issue where the provisionerService which can be used to spawn resources as needed is left running on a non-leader coordinator/overlord, after it is removed from leadership. Provisioning should only be done by the leader. To fix the issue, a call to stop the provisionerService was added to the stop() method of HttpRemoteTaskRunner class. The provisionerService was properly closed on other TaskRunner types.	2022-01-20 13:32:08 -05:00
Vadim Ogievetsky	6ce14e6b17	allow W in durations (#12175 )	2022-01-20 09:17:43 +05:30
Uwe Schindler	1f7dd6d86c	Forbiddenapis: Split the guava16-only signatures file from main signatures file (#12170 )	2022-01-19 17:50:28 -08:00
Jihoon Son	cacfcfcdab	ignore hadoop-gcs directory already exists error for integration tests (#12169 )	2022-01-19 09:35:50 -08:00
Jihoon Son	cc2ffc6c0f	Fix node discovery to ignore unknown DruidServices (#12157 ) * Fix node discovery to ignore unknown DruidServices * ignore all runtime exceptions * fix test * add custom deserializer * custom serializer * log host for unparseable druidService	2022-01-18 22:08:59 -08:00
Abhishek Agarwal	53c0e489c2	Fix infinite retrying during task pausing (#12167 ) This fixes a bug that causes TaskClient in overlord to continuously retry to pause tasks. This can happen when a task is not responding to the pause command. Ideally, in such a case when the task is unresponsive, the overlord would have given up after a few retries and would have killed the task. However, due to this bug, retries go on forever.	2022-01-19 09:03:36 +05:30
Gian Merlino	cf7191d2bc	Validate target dataSource for INSERT. (#12129 )	2022-01-18 09:34:23 -08:00
Abhishek Agarwal	7bdb9ebdf1	Suppress Avro CVEs (#12166 )	2022-01-18 21:09:48 +05:30
Maytas Monsereenusorn	bd7fe45da0	Support adding metrics in Auto Compaction (#12125 ) * add impl * add impl * add unit tests * add unit tests * add unit tests * add unit tests * add unit tests * add integration tests * add integration tests * fix LGTM * fix test * remove doc	2022-01-17 20:19:31 -08:00
Benedict Jin	b55f7a25fe	Fix forbiddenapis causing travis failing (#12158 ) * Fix forbiddenapis causing travis failing * Use failOnUnresolvableSignatures instead	2022-01-15 16:13:37 -08:00
Victoria Lim	d2ac146365	Docs for cluster tiering to improve query concurrency (#12128 ) * add new doc * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> * reorder query laning properties * rename doc * new name in doc header * organize material into "service tiering" section * text edits and update sidebars.json * update query laning * how queries get assigned to lanes * add more details to intro; use more consistent terminology * more content * Apply suggestions from code review Co-authored-by: Jihoon Son <jihoonson@apache.org> * Update docs/operations/mixed-workloads.md * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> * typo Co-authored-by: Charles Smith <techdocsmith@gmail.com> Co-authored-by: Jihoon Son <jihoonson@apache.org>	2022-01-15 12:22:08 +08:00
Ivan Vankovich	6a93872586	OpenTelemetry emitter extension (#12015 ) * Add OpenTelemetry emitter extension * Fix build * Fix checkstyle * Add used undeclared dependencies * Ignore unused declared dependencies	2022-01-15 12:18:04 +08:00
Clint Wylie	e0c4c568cb	fix incorrect ColumnInspector in IncrementalIndex.makeColumnSelectorFactory (#12155 )	2022-01-13 18:09:06 -08:00
aggarwalakshay	eb4fafe08f	Upgrading follow-redirects to 1.14.7 (#12153 ) * Upgrading follow-redirects to 1.14.7 * removed the existing follow-redirects i.e. 1.14.4 from package-lock.json	2022-01-13 14:01:36 -08:00
Jonathan Wei	74c876e578	Throw parse exceptions on schema get errors for SchemaRegistryBasedAvroBytesDecoder (#12080 ) * Add option to throw parse exceptions on schema get errors for SchemaRegistryBasedAvroBytesDecoder * Remove option	2022-01-13 12:36:51 -06:00
Clint Wylie	1dba089a62	fix array type strategy write size tracking (#12150 ) * fix array type strategy write size tracking * fix checkstyle	2022-01-13 10:22:40 -08:00
Xavier Léauté	e56ea31697	follow-up to fix formatting broken in #12147 (#12148 ) follow-up to #12147 to fix the build	2022-01-12 20:59:32 -08:00
Jihoon Son	58378aa967	Move gcs-connector from lib to hadoop-dependencies for integration test (#12144 )	2022-01-12 16:47:34 -08:00
Xavier Léauté	168187e6df	avoid unnecessary String.format calls in IdUtils.validateId (#12147 ) Based on profiling data, about 25% of the time de-serializing DataSchema is spent on formatting strings in validateId. This can add up quickly, especially when de-serializing task information in the overlord, where in can consume almost 2% of CPU if there are many tasks. Since the formatting is unnecessary unless the checks fail, we can leverage the built-in formatting of Preconditions.checkArgument instead to avoid the cost.	2022-01-12 16:34:40 -08:00
Vadim Ogievetsky	9cd52ed914	Web console: make range partitioning a first class citizen of the console (#12146 ) * first class support for range partitioning * update e2e tests	2022-01-12 03:50:10 -08:00
Clint Wylie	f2ce76966c	add EARLIEST_BY/LATEST_BY to make EARLIEST/LATEST function signatures less ambiguous (#12145 ) * add EARLIEST_BY/LATEST_BY to make EARLIEST/LATEST function signatures unambiguous * switcheroo * EARLIEST_BY/LATEST_BY use timestamp instead of numeric types, update docs * revert unintended change * fix docs * fix docs better	2022-01-12 03:48:53 -08:00

... 8 9 10 11 12 ...

11979 Commits All Branches Search

11979 Commits

All Branches