druid

Commit Graph

Author	SHA1	Message	Date
Vadim Ogievetsky	e8635df9e7	clean up some bp3 classes (#12403 )	2022-04-06 15:27:44 -07:00
Victoria Lim	e6229b76a6	Document data format and example for featureSpec (#12394 ) * add data format and example for featureSpec * add second feature in example * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2022-04-06 15:17:15 -07:00
317brian	ac6c24793e	docs(fix): add clarity around granularitySpec (#12362 ) * fix: add clarify around granularitySpec * fix spacing * Update docs/ingestion/compaction.md Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2022-04-06 09:24:37 -07:00
Victoria Lim	d326c681c1	Document config for ingesting null columns (#12389 ) * config for ingesting null columns * add link * edit .spelling * what happens if storeEmptyColumns is disabled	2022-04-05 09:15:42 -07:00
aggarwalakshay	7d5666109c	upgrade surefire 3.0.0-M6 (#12395 ) * upgrade surefire 3.0.0-M6 * increasing memory	2022-04-04 23:56:15 -07:00
Paul Rogers	2cc2088720	Method to specify eternity in the scan query builder (#12223 ) * Method to specify eternity in the scan query builder * Fix checkstyle issue * Renamed eterity() to eternityInterval() * Minor fixes	2022-04-04 15:11:32 -07:00
John Gozde	90680543d0	Blueprint 4 (#12391 ) * Update blueprint dependencies & LICENSES * Switch to bp4 namespace; use bp-ns variable in overrides * Add webpack alias for colors.scss * Snapshots * Update selectors in e2e tests	2022-04-04 10:34:22 -07:00
AmatyaAvadhanula	067254b778	Package kinesis client jar within the extension (#12370 ) amazon-kinesis-client was not covered undered the apache license and required separate insertion in the kinesis extension. This can now be avoided since it is covered, and including it within druid helps prevent incompatibilities. Allows enabling of deaggregation out of the box by packaging amazon-kinesis-client (1.14.4) with druid for kinesis ingestion.	2022-04-04 21:31:18 +05:30
Tejaswini Bandlamudi	984904779b	Increase default DatasourceCompactionConfig.inputSegmentSizeBytes to Long.MAX_VALUE (#12381 ) The current default value of inputSegmentSizeBytes is 400MB, which is pretty low for most compaction use cases. Thus most users are forced to override the default. The default value is now increased to Long.MAX_VALUE.	2022-04-04 16:28:53 +05:30
AmatyaAvadhanula	c5531be553	Add feature flag for Kinesis listShards API usage (#12383 ) listShards API was used to get all the shards for kinesis ingestion to improve its resiliency as part of #12161. However, this may require additional permissions in the IAM policy where the stream is present. (Please refer to: https://docs.aws.amazon.com/kinesis/latest/APIReference/API_ListShards.html). A dynamic configuration useListShards has been added to KinesisSupervisorTuningConfig to control the usage of this API and prevent issues upon upgrade. It can be safely turned on (and is recommended when using kinesis ingestion) by setting this configuration to true.	2022-04-04 14:58:10 +05:30
somu-imply	a1ea658115	Introducing a new config to ignore nulls while computing String Cardinality (#12345 ) * Counting nulls in String cardinality with a config * Adding tests for the new config * Wrapping the vectorize part to allow backward compatibility * Adding different tests, cleaning the code and putting the check at the proper position, handling hasRow() and hasValue() changes * Updating testcase and code * Adding null handling test to improve coverage * Checkstyle fix * Adding 1 more change in docs * Making docs clearer	2022-03-29 14:31:36 -07:00
Peter Marshall	f1841c6444	Docs - S3 masking and nav update to S3 page (#11490 ) * Docs: Masking S3 creds and some rewording Knowledge transfer from https://groups.google.com/g/druid-user/c/FydcpFrA688 * Removed bold in one of the quote sections * Update s3.md * Update s3.md Quick grammar change * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update s3.md Typo * Update docs/development/extensions-core/s3.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update s3.md Active lang * Update s3.md LAng nit * Update native-batch.md LAng nit * Update docs/ingestion/native-batch.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Grammar tidy-up and link fix Corrected 2 x links to old page H2s, resolved the question around precedence, and some other grammatical changes. * Update docs/development/extensions-core/s3.md * Update s3.md Removed an Erroneous E Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2022-03-29 09:13:05 -07:00
Peter Marshall	b9a968e7ff	Docs – expressions link back and timestamp hint (#11674 ) * Update math-expr.md Link back to transformSpec * Update ingestion-spec.md Moved info about using the timestamp inside transforms into the actual timestamp section. * Update ingestion-spec.md Active language.	2022-03-29 09:12:30 -07:00
mark-imply	3c55565398	Update ingestion-spec.md (#12371 ) * Update ingestion-spec.md Added best practice point to dimensions description. * Update docs/ingestion/ingestion-spec.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2022-03-29 09:12:02 -07:00
Jihoon Son	49a3f4291a	Add an integration test for null-only columns (#12365 ) * integration test for null-only-columns * metadata query * fix test	2022-03-28 16:40:45 -07:00
Victoria Lim	9ed7aa33ec	Docs for request logging (#12363 ) * add docs for request logging * remove stray character * Update docs/operations/request-logging.md Co-authored-by: TSFenwick <tsfenwick@gmail.com> * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> Co-authored-by: TSFenwick <tsfenwick@gmail.com> Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2022-03-28 14:09:41 -07:00
Yuanli Han	f2495a67d2	fix messageGap metric (#12337 )	2022-03-28 09:21:06 -07:00
AmatyaAvadhanula	9c6b9abcde	Use javaOptsArray provided in task context (#12326 ) The `javaOpts` property is being read from task context but not `javaOptsArray`. Changes: - Read `javaOptsArray` from task context in `ForkingTaskRunner`. - Add test to verify that `javaOptsArray` in task context takes precedence over `javaOpts`	2022-03-28 16:33:40 +05:30
dependabot[bot]	ee44fe45c6	Bump java-dogstatsd-client from 2.13.0 to 4.0.0 (#12353 ) * Bump java-dogstatsd-client from 2.13.0 to 4.0.0 Bumps [java-dogstatsd-client](https://github.com/DataDog/java-dogstatsd-client) from 2.13.0 to 4.0.0. - [Release notes](https://github.com/DataDog/java-dogstatsd-client/releases) - [Changelog](https://github.com/DataDog/java-dogstatsd-client/blob/master/CHANGELOG.md) - [Commits](https://github.com/DataDog/java-dogstatsd-client/compare/v2.13.0...v4.0.0) * migrate statsd-emitter tests from easymock to mockito * add simple init test to make diff coverage happy Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Xavier Léauté <xvrl@apache.org>	2022-03-26 16:25:13 -07:00
Maytas Monsereenusorn	ea51d8a16c	Duties in Indexing group (such as Auto Compaction) does not report metrics (#12352 ) * add impl * add unit tests * fix checkstyle * address comments * fix checkstyle	2022-03-23 18:18:28 -07:00
Jihoon Son	b6eeef31e5	Store null columns in the segments (#12279 ) * Store null columns in the segments * fix test * remove NullNumericColumn and unused dependency * fix compile failure * use guava instead of apache commons * split new tests * unused imports * address comments	2022-03-23 16:54:04 -07:00
syacobovitz	d7308e9290	Added support in urls, and grouped metrics (#12296 )	2022-03-22 11:22:05 -07:00
Kashif Faraz	0867ca75e1	Fix OOM failures in dimension distribution phase of parallel indexing (#12331 ) Parallel indexing with range partitioning can often cause OOM in the `ParallelIndexSupervisorTask` during the dimension distribution phase. This typically happens because of too many `StringSketch` objects obtained from the different `partial_dimension_distribution` sub-tasks. We need not keep any of the sketches in memory until we need to compute the PartitionBoundaries for the respective interval. Changes - Extract `StringDistribution` from `DimensionDistributionReport`s when they are received and write to disk inside the task/temp/distributions - After all the subtasks have finished, iterate over all the intervals one by one - For each interval, read the distributions from disk, merge them and create `PartitionBoundaries`. - Cleanup task/temp/distributions directory when all `PartitionBoundaries` have been determined	2022-03-22 19:28:15 +05:30
Adarsh Sanjeev	ef45a1551e	Convert inQueryThreshold into query context parameter. (#12357 ) Added Calcites InQueryThreshold as a query context parameter. Setting this parameter appropriately reduces the time taken for queries with large number of values in their IN conditions.	2022-03-22 18:33:57 +05:30
Xavier Léauté	1f0447e613	fix use of deprecated initMocks method (#12351 ) follow-up to #12341 - fix use of deprecated initMocks methods and properly close mocks on teardown	2022-03-19 10:19:02 -07:00
Xavier Léauté	c3377bf744	upgrade maven-pmd-plugin to fix warning (#12349 ) we sometimes see warnings similar to the one mentioned https://issues.apache.org/jira/browse/MPMD-325 Upgrading the plugin should hopefully reduce occurrence of those.	2022-03-19 10:18:26 -07:00
dependabot[bot]	4ed1abca94	Bump slf4j.version from 1.7.12 to 1.7.36 (#11594 ) Bump slf4j.version from 1.7.12 to 1.7.36 - [Release notes](Release notes: https://www.slf4j.org/news.html) Updates `jcl-over-slf4j` from 1.7.12 to 1.7.36 - [Commits](https://github.com/qos-ch/slf4j/compare/v_1.7.12...v_1.7.36) Updates `slf4j-simple` from 1.7.12 to 1.7.36 - [Commits](https://github.com/qos-ch/slf4j/compare/v_1.7.12...v_1.7.36) Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Suneet Saldanha <suneet@apache.org> Co-authored-by: Xavier Léauté <xvrl@apache.org>	2022-03-18 13:45:44 -07:00
Maytas Monsereenusorn	dbb9518f50	Fix auto compaction by adjusting compaction task's interval to align with segmentGranularity when segmentGranularity is set (#12334 ) * add impl * add ITs * address comments * address comments * address comments * fix failure * fix checkstyle * fix checkstyle	2022-03-18 12:46:16 -07:00
Xavier Léauté	6f0e5f25fa	update surefire plugin to 3.0.0-M4 (#12342 ) stay on surefire 3.0.0-M4 until we can upgrade to 3.0.0-M6 with a fix for https://issues.apache.org/jira/browse/SUREFIRE-1815 causing issues in RetryUtilsTest.	2022-03-18 08:20:28 -07:00
Xavier Léauté	c33fa11669	improve test compatibility with Java 17 and remove deprecated methods (#12341 ) * remove use of reflection in EnvironmentVariableDynamicConfigProvider for Java 17 compatibility * fix mocks mock objects not getting closed properly, causing issues with Java 17 * remove use of deprecated methods and rules in tests	2022-03-18 08:19:28 -07:00
Aurélien Dunand	8f3a631cbf	Fix missing conversionFactor in prometheus emitter (#12338 ) query/node/ttfb metrics are in milliseconds.	2022-03-17 21:46:06 -07:00
Xavier Léauté	192e411249	fix build due to com.nimbusds:lang-tag update (#12348 ) the version of com.nimbusds:oauth2-oidc-sdk we depend on does not specific an exact version dependency for com.nimbusds:lang-tag, and instead uses a version range (see https://search.maven.org/artifact/com.nimbusds/oauth2-oidc-sdk/6.5/jar) Recently a new version of lang-tag was released requiring us to update the license file accordingly.	2022-03-17 17:44:08 -07:00
dependabot[bot]	a5dfb911de	Bump maven-site-plugin from 3.1 to 3.11.0 (#12310 ) Bumps [maven-site-plugin](https://github.com/apache/maven-site-plugin) from 3.1 to 3.11.0. - [Release notes](https://github.com/apache/maven-site-plugin/releases) - [Commits](https://github.com/apache/maven-site-plugin/compare/maven-site-plugin-3.1...maven-site-plugin-3.11.0) --- updated-dependencies: - dependency-name: org.apache.maven.plugins:maven-site-plugin dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2022-03-17 15:17:29 +08:00
Jihoon Son	5e23674fe5	Fix a race condition in the '/tasks' Overlord API (#12330 ) * finds complete and active tasks from the same snapshot * overlord resource * unit test * integration test * javadoc and cleanup * more cleanup * fix test and add more	2022-03-17 10:47:45 +09:00
Frank Chen	d745d0b338	Add JDK 11 (#12333 )	2022-03-16 15:03:04 -07:00
Dr. Sizzles	69f928f50e	Adding k8s support for human readable parsing (#12316 ) * Adding k8s support for human readable parsing * Update docs/configuration/human-readable-byte.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update docs/configuration/human-readable-byte.md Co-authored-by: Frank Chen <frankchen@apache.org> * Update core/src/main/java/org/apache/druid/java/util/common/HumanReadableBytes.java Co-authored-by: Frank Chen <frankchen@apache.org> * Changes per review Co-authored-by: Rahul Gidwani <r_gidwani@apple.com> Co-authored-by: Frank Chen <frankchen@apache.org>	2022-03-16 11:18:47 +08:00
Xavier Léauté	5d02a91faa	upgrade Error Prone to 2.11 (requires Java 11) (#12306 ) The latest version of Error Prone now requires Java 11. Upgrading means we can remove a lot of the maven profile complexity required to run checks with Java 8. This also requires switching our strict build to use Java 11. * update error-prone to 2.11 * remove need for specific maven profiles for Java 8 and Java 15 * fix additional Error Prone warnings with Java 11 * update strict build to use Java 11	2022-03-14 19:40:48 -07:00
somu-imply	b5195c5095	Graceful null handling and correctness in DoubleMean Aggregator (#12320 ) * Adding null handling for double mean aggregator * Updating code to handle nulls in DoubleMean aggregator * oops last one should have checkstyle issues. fixed * Updating some code and test cases * Checking on object is null in case of numeric aggregator * Adding one more test to improve coverage * Changing one test as asked in the review * Changing one test as asked in the review for nulls	2022-03-14 16:52:47 -07:00
mchades	3de1272926	bug fix: merge results of group by limit push down (#11969 )	2022-03-11 09:04:34 -08:00
Kyle Larose	db91961af7	kubernetes: restart watch on null response (#12233 ) * kubernetes: restart watch on null response Kubernetes watches allow a client to efficiently processes changes to resources. However, they have some idiosyncrasies. In particular, they can error out for various reasons leading to what would normally be seen as an invalid result. The Druid kubernetes node discovery subsystem does not handle a certain case properly. The watch can return an item with a null object. These leads to a null pointer exception. When this happens, the provider needs to restart the watch, because rerunning the watch from the same resource version leads to the same result: yet another null pointer exception. This commit changes the provider to handle null objects by restarting the watch. * review: add more coverage This adds a bit more coverage to the K8sDruidNodeDiscoveryProvider watch loop, and removes an unnecessay return. * kubernetes: reduce logging verbosity The log messages about items being NULL don't really deserve to be at a level other than DEBUG since they are not actionable, particularly since we automatically recover now. Move them to the DEBUG level.	2022-03-10 12:56:40 -08:00
Gian Merlino	cb2b2b696d	Fix error message for groupByEnableMultiValueUnnesting. (#12325 ) * Fix error message for groupByEnableMultiValueUnnesting. It referred to the incorrect context parameter. Also, create a dedicated exception class, to allow easier detection of this specific error. * Fix other test. * More better error messages. * Test getDimensionName method.	2022-03-10 11:37:24 -08:00
Parag Jain	2efb74ff1e	fix supervisor auto scaler config serde bug (#12317 )	2022-03-09 16:17:12 -08:00
Jihoon Son	d89d4ff588	Git hooks should fail on errors; pass args to git hooks (#12322 ) * Git hooks should fail on errors * don't set shell to pass args	2022-03-10 09:07:50 +09:00
Abhishek Agarwal	6346b9561d	Reuse the InputEntityReader in SettableByteEntityReader (#12269 ) * Reuse the InputEntityReader in SettableByteEntityReader * Fix logic * Fix kafka streaming ingestion * Add Tests for kafka input format change * Address review comments	2022-03-09 14:38:31 -08:00
Clint Wylie	9cfb23935f	push value range and set index get operations into BitmapIndex (#12315 ) * push value range and set index get operations into BitmapIndex * fix bug * oops, fix better * better like, fix test, javadocs * fix checkstyle * simplify and fixes * cache * fix tests * move indexOf into GenericIndexed * oops * fix tests	2022-03-09 13:30:58 -08:00
Rohan Garg	9f6a930462	Fix join query incase of filter explosion during CNF conversion (#12324 )	2022-03-09 12:43:09 -08:00
Clint Wylie	dc0372a28e	improve FileWriteOutBytes.readFully (#12323 ) * improve FileWriteOutBytes.readFully * no need to flush if out of bounds	2022-03-09 11:45:45 -08:00
AmatyaAvadhanula	7bf1d8c5c0	Facilitate lazy initialization of connections to mitigate overwhelming of Coordinator (#12298 ) Add config for eager / lazy connection initialization in ResourcePool Description Currently, when multiple tasks are launched, each of them eagerly initializes a full pool's worth of connections to the coordinator. While this is acceptable when the parameter for number of eagerConnections (== maxSize) is small, this can be problematic in environments where it's a large value (say 1000) and multiple tasks are launched simultaneously, which can cause a large number of connections to be created to the coordinator, thereby overwhelming it. Patch Nodes like the broker may require eager initialization of resources and do not create connections with the Coordinator. It is unnecessary to do this with other types of nodes. A config parameter eagerInitialization is added, which when set to true, initializes the max permissible connections when ResourcePool is initialized. If set to false, lazy initialization of connection resources takes place. NOTE: All nodes except the broker have this new parameter set to false in the quickstart as part of this PR Algorithm The current implementation relies on the creation of maxSize resources eagerly. The new implementation's behaviour is as follows: If a resource has been previously created and is available, lend it. Else if the number of created resources is less than the allowed parameter, create and lend it. Else, wait for one of the lent resources to be returned.	2022-03-09 23:17:43 +05:30
Rohan Garg	56fbd2af6f	Guard against exponential increase of filters during CNF conversion (#12314 ) Currently, the CNF conversion of a filter is unbounded, which means that it can create as many filters as possible thereby also leading to OOMs in historical heap. We should throw an error or disable CNF conversion if the filter count starts getting out of hand. There are ways to do CNF conversion with linear increase in filters as well but that has been left out of the scope of this change since those algorithms add new variables in the predicate - which can be contentious.	2022-03-09 13:19:52 +05:30
Clint Wylie	0600772cce	use a non-concurrent map for lookups-cached-global unless incremental updates are actually required (#12293 ) * use a non-concurrent map for lookups-cached-global unless incremental updates are actually required * adjustments * fix test	2022-03-08 21:54:25 -08:00

1 2 3 4 5 ...

11735 Commits All Branches Search

11735 Commits

All Branches