druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	c2e9ab8100	benchmark schema with numeric dimensions and null column values (#9036 ) * benchmark schema with null column values * oops * adjustments * rename again, different null percentage so rows more variety * more schema	2019-12-19 17:45:19 -08:00
Jihoon Son	3c31493772	Add missing docs for http client configurations (#9054 ) * Add missing docs for http client configurations * fix typo * backticks	2019-12-19 17:41:04 -08:00
Suneet Saldanha	3c13444167	Fix flaky ITBasicAuthConfigurationTest (#9072 ) This test was failing to authenticate using the admin credentials. These should be available by default in the metadata store. This indicates that the credentials are not successfully being syncd before the test is run. This change increases the number of retries to 20 so that the services are syncd before the test runs	2019-12-19 17:38:55 -08:00
Suneet Saldanha	176bc8fd97	Remove resolve-ip dependency for integration-tests (#9065 ) * Remove resolve-ip dependency for integration-tests * use host hostname and fallback to dscacheutil * better shell script comparisons	2019-12-19 14:53:36 -08:00
Fangjin Yang	256b8f69b6	Update README.md (#9078 )	2019-12-19 13:00:27 -08:00
Fangjin Yang	d20d2ff71d	Update README.md (#9077 )	2019-12-19 11:54:14 -08:00
Fangjin Yang	de18f76c8b	Update README.md (#9074 ) Updates to readme	2019-12-19 11:39:27 -08:00
Clint Wylie	84ef8b819e	fix druid-sql issue with filtering numeric columns by null values (#9061 ) * fix druid-sql issue with filtering numeric columns by null values * fix tests * fix tests for reals	2019-12-18 13:30:34 -08:00
Jihoon Son	94a23fb17e	Fix flaky realtime index task tests (#8999 ) * Fix flaky realtime index task tests * fix ITAppenderatorDriverRealtimeIndexTaskTest * fix comment * address comments	2019-12-18 13:25:00 -08:00
Jonathan Wei	15884f6d10	Fix hadoop ingestion property handling when using indexers (#9059 )	2019-12-18 12:13:19 -08:00
Jonathan Wei	b1547a76b1	Update GPG key instructions for ASF release guide (#9006 )	2019-12-18 12:12:48 -08:00
Suneet Saldanha	1fb93d56c3	Add instructions to backport a PR (#9052 ) * Add instructions to backport a PR * Clearer image * Add period in backport instructions	2019-12-18 11:57:01 -08:00
Chi Cao Minh	6178f05da6	Fail superbatch range partition multi dim values (#9058 ) * Fail superbatch range partition multi dim values Change the behavior of parallel indexing range partitioning to fail ingestion if any row had multiple values for the partition dimension. After this change, the behavior matches that of hadoop indexing. (Previously, rows with multiple dimension values would be skipped.) * Improve err msg, rename method, rename test class	2019-12-18 10:14:03 -08:00
Jonathan Wei	131b3f13be	Skip non-Apache repo PRs in milestone tagging script (#9064 )	2019-12-17 18:28:11 -08:00
Vadim Ogievetsky	e7b1653d88	add button to reapply retention rules (#9055 )	2019-12-17 18:08:57 -08:00
Benedict Jin	24be558347	Fix NPE for subquery with limit (#8775 ) * Fix NPE for subquery with limit * Mark it as unplannable by returning null * Migrate testcases from SqlResourceTest to CalciteQueryTest * Throw CannotBuildQueryException * Fix typo * Patch comments	2019-12-17 10:21:12 -08:00
Suneet Saldanha	301c0649a7	Fix equalsAndHashCode in ClientCompactQueryTuningConfig (#9035 ) * Fix equalsAndHashCode in ClientCompactQueryTuningConfig This change introduces a dependency to EqualsVerifier for the test scope. The dependency is licensed under Apache 2. The library makes it trivial to add equals and hashCode checks to prevent bugs like this from happening in the future * fix checkstyle * fix test name	2019-12-16 14:33:00 -08:00
Jihoon Son	298425a33a	Fix handling interruptedException in resource pool (#9044 )	2019-12-16 09:41:13 -08:00
Clint Wylie	bc16ff5e7c	sql auto limit wrapping fix (#9043 ) * sql auto limit wrapping fix * fix tests and style * remove setImportance	2019-12-16 01:38:24 -08:00
Clint Wylie	6881535b48	docs - clarify cache parameters (#9020 )	2019-12-13 16:53:45 -08:00
Gian Merlino	d452cbbb82	GenericIndexedWriter: Fix issue when writing large values to large columns. (#9029 )	2019-12-13 15:33:14 -08:00
Suneet Saldanha	3325da1718	Allow startup scripts to specify java home (#9021 ) * Allow startup scripts to specify java home The startup scripts now look for java in 3 locations. The order is from most related to druid to least, ie ${DRUID_JAVA_HOME} ${JAVA_HOME} ${PATH} * Update fn names and clean up code * final round of fixes * fix spellcheck	2019-12-12 21:36:00 -08:00
Fangyuan Deng	41f30e53a6	[bugfix]fix getAvgSizePerGranularity logic in DerivativeDataSourceManager(materializedview) (#8929 ) * fix getAvgSizePerGranularity in DerivativeDataSourceManager * revert * redo	2019-12-12 17:27:02 -08:00
Himanshu	9236dd9467	optionally enable Jetty ForwardedRequestCustomizer (#9010 ) * optionally enable Jetty ForwardedRequestCustomizer * fix doc build	2019-12-12 17:00:08 -08:00
Himanshu	45101183bc	HRTR: make pending task execution handling to go through all tasks on not finding worker slots (#8697 ) * HRTR: make pending task execution handling to go through all tasks on not finding worker slots * make HRTR methods package private that are meant to be used only in HttpRemoteTaskRunnerResource * mark HttpRemoteTaskRunnerWorkItem.State global variables final * hrtr: move immutableWorker NULL check outside of try-catch or finally block could have NPE * add some explanatory comments * add comment on explaining mechanics around hand off of pending tasks from submission to it getting picked up by a task execution thread * fix spelling	2019-12-12 14:58:52 -08:00
Xavier Léauté	810b85a352	allow druid.host to be undefined to use canonical hostname (#9019 ) It is currently not possible to unset the druid.host property in the docker image to let Druid default to the canonical hostname. It always gets set to the container's IP address. Passing the override environment variable druid_host= unfortunately does not solve the problem, as this gets interpreted as empty string and does not let the default kick in. This change adds the option to pass DRUID_SET_HOST=0 as environment variable to disable the default behavior, and allows passing a common runtime.properties file without druid.host.	2019-12-12 13:51:57 -08:00
Benjamin Hopp	13c33c1766	Update architecture.md (#9015 )	2019-12-11 19:05:50 -08:00
Jihoon Son	66056b2826	Using annotation to distinguish Hadoop Configuration in each module (#9013 ) * Multibinding for NodeRole * Fix endpoints * fix doc * fix test * Using annotation to distinguish Hadoop Configuration in each module	2019-12-11 17:30:44 -08:00
Jihoon Son	e5e1e9c4ee	Fix broken master (#9005 ) * Multibinding for NodeRole * Fix endpoints * fix doc * fix test	2019-12-11 15:56:36 -08:00
Jonathan Wei	8af41d7cd0	Update version to 0.18.0-incubating-SNAPSHOT (#9009 )	2019-12-11 14:04:03 -08:00
Parag Jain	24fe824055	add readiness endpoints to processes having initialization delays (#8841 )	2019-12-10 17:26:13 -08:00
Chi Cao Minh	3de7ab8523	DataSketches jars in core (#9003 ) Having DataSketches jars in core will allow potential improvements, for example: - Provide an alternative implementation of HLL: https://datasketches.github.io/docs/HLL/HllSketchVsDruidHyperLogLogCollector.html - Range partitioning for native parallel batch indexing without having the user load extensions on the classpath Dev mailing list discussion: https://lists.apache.org/thread.html/301410d71ff799cf616bf17c4ebcf9999fc30829f5fa62909f403e6c%40%3Cdev.druid.apache.org%3E	2019-12-10 14:02:34 -08:00
Chi Cao Minh	bab78fc80e	Parallel indexing single dim partitions (#8925 ) * Parallel indexing single dim partitions Implements single dimension range partitioning for native parallel batch indexing as described in #8769. This initial version requires the druid-datasketches extension to be loaded. The algorithm has 5 phases that are orchestrated by the supervisor in `ParallelIndexSupervisorTask#runRangePartitionMultiPhaseParallel()`. These phases and the main classes involved are described below: 1) In parallel, determine the distribution of dimension values for each input source split. `PartialDimensionDistributionTask` uses `StringSketch` to generate the approximate distribution of dimension values for each input source split. If the rows are ungrouped, `PartialDimensionDistributionTask.UngroupedRowDimensionValueFilter` uses a Bloom filter to skip rows that would be grouped. The final distribution is sent back to the supervisor via `DimensionDistributionReport`. 2) The range partitions are determined. In `ParallelIndexSupervisorTask#determineAllRangePartitions()`, the supervisor uses `StringSketchMerger` to merge the individual `StringSketch`es created in the preceding phase. The merged sketch is then used to create the range partitions. 3) In parallel, generate partial range-partitioned segments. `PartialRangeSegmentGenerateTask` uses the range partitions determined in the preceding phase and `RangePartitionCachingLocalSegmentAllocator` to generate `SingleDimensionShardSpec`s. The partition information is sent back to the supervisor via `GeneratedGenericPartitionsReport`. 4) The partial range segments are grouped. In `ParallelIndexSupervisorTask#groupGenericPartitionLocationsPerPartition()`, the supervisor creates the `PartialGenericSegmentMergeIOConfig`s necessary for the next phase. 5) In parallel, merge partial range-partitioned segments. `PartialGenericSegmentMergeTask` uses `GenericPartitionLocation` to retrieve the partial range-partitioned segments generated earlier and then merges and publishes them. * Fix dependencies & forbidden apis * Fixes for integration test * Address review comments * Fix docs, strict compile, sketch check, rollup check * Fix first shard spec, partition serde, single subtask * Fix first partition check in test * Misc rewording/refactoring to address code review * Fix doc link * Split batch index integration test * Do not run parallel-batch-index twice * Adjust last partition * Split ITParallelIndexTest to reduce runtime * Rename test class * Allow null values in range partitions * Indicate which phase failed * Improve asserts in tests	2019-12-09 23:05:49 -08:00
Vadim Ogievetsky	a6dcc99962	better input format detection (#9007 )	2019-12-09 22:31:28 -08:00
Clint Wylie	4327892b84	modify multi-value expression transformation behavior to not treat re-use of the same input as a candidate for cartesian mapping (#8957 )	2019-12-09 20:38:15 -08:00
Vadim Ogievetsky	0330744793	Docs: bold Java 8 requirement (#8996 ) * bold Java 8 req * add warning box	2019-12-09 20:23:07 -08:00
Parag Jain	9640f9649a	fix npe while logging sql/query request (#9001 ) * fix npe while logging sql/query request * forbid forbidden DateTime API	2019-12-09 12:02:11 -08:00
Rye	ca77d576c6	add customize separator for TSV inputFormat (#8993 ) * add customize separator for TSV inputFormat * fix spotbug * code refactor * code refactor * add argument check for delimiter * refine null check * add check for delimiter and listdelimiter can not be same * add unit tests	2019-12-09 11:24:09 -08:00
Roman Leventov	1c62987783	Add SelfDiscoveryResource; rename org.apache.druid.discovery.No… (#6702 ) * Add SelfDiscoveryResource * Rename org.apache.druid.discovery.NodeType to NodeRole. Refactor CuratorDruidNodeDiscoveryProvider. Make SelfDiscoveryResource to listen to updates only about a single node (itself). * Extended docs * Fix brace * Remove redundant throws in Lifecycle.Handler.stop() * Import order * Remove unresolvable link * Address comments * tmp * tmp * Rollback docker changes * Remove extra .sh files * Move filter * Fix SecurityResourceFilterTest	2019-12-08 18:47:58 +03:00
Clint Wylie	441515cb50	update dump-segment docs so example command works (#8998 ) * update dump-segment docs so example command works * not everyone uses bash	2019-12-07 06:36:46 -08:00
Clint Wylie	06cd30460e	add query metrics for broker parallel merges, off by default (#8981 ) * add a bunch of metrics for broker parallel merges, off by default, and tests * fix tests * review stuffs * propogateIfPossible	2019-12-06 13:42:53 -08:00
Clint Wylie	cefcfe26dc	update web-console data loader to support unified s3 and google input sources (#8994 ) * update web-console data loader to support unified s3 and google input source * fixes * add placeholder for objects * only show objects if it already exists	2019-12-06 07:25:26 -08:00
Clint Wylie	ca2a7a1f08	more flush timeout for emitter tests (#8991 ) * more flush timeout for emitter tests * share constant	2019-12-05 16:52:35 -08:00
Jonathan Wei	c949a25210	Add DruidInputSource (replacement for IngestSegmentFirehose) (#8982 ) * Add Druid input source and format * Inherit dims/metrics from segment * Add ingest segment firehose reindexing test * Remove unnecessary module * Fix unit tests, checkstyle * Add doc entry * Fix dimensionExclusions handling, add parallel index integration test * Add spelling exclusion * Address some PR comments * Checkstyle * wip * Address rest of PR comments * Address PR comments	2019-12-05 16:50:00 -08:00
Chi Cao Minh	af74acaa85	Address security vulnerabilities CVSS >= 7 (#8980 ) * Address security vulnerabilities CVSS >= 7 Update dependencies to address security vulnerabilities with CVSS scores of 7 or higher. A new Travis CI job is added to prevent new high/critical security vulnerabilities from being added. Updated dependencies: - api-util 1.0.0 -> 1.0.3 - jackson 2.9.10 -> 2.10.1 - kafka 2.1.0 -> 2.1.1 - libthrift 0.10.0 -> 0.13.0 - protobuf 3.2.0 -> 3.11.0 The following high/critical security vulnerabilities are currently suppressed (so that the new Travis CI job can be added now) and are left as future work to fix: - hibernate-validator:5.2.5 - jackson-mapper-asl:1.9.13 - libthrift:0.6.1 - netty:3.10.6 - nimbus-jose-jwt:4.41.1 * Rename EDL1 license file * Fix inspection errors	2019-12-05 14:34:35 -08:00
Clint Wylie	5ecdf94d83	add 'prefixes' support to google input source (#8930 ) * add prefixes support to google input source, making it symmetrical-ish with s3 * docs * more better, and tests * unused * formatting * javadoc * dependencies * oops * review comments * better javadoc	2019-12-04 21:01:10 -08:00
Vadim Ogievetsky	1cff73f3e0	Web console: support new ingest spec format (#8828 ) * converter v1 * working v1 * update tests * update tests * upgrades * adjust to new API * remove hack * fwd * step * neo cache * fix time selection * smart reset * parquest autodetection * add binaryAsString option * partitionsSpec * add ORC support * ingestSegment -> druid * remove index tasks * better min * load data works * remove downgrade * filter on group_id * fix group_id in test * update auto form for new props * add dropBeforeByPeriod rule * simplify * prettify json	2019-12-04 20:21:07 -08:00
Lucas Capistrant	8dd9a8cb15	Small doc fix for baseTaskDir conf (#8978 )	2019-12-04 14:07:03 -08:00
Clint Wylie	a48784a1fd	dropwizard-emitter doc fixes (#8988 )	2019-12-04 12:52:58 -08:00
Q	391646123e	Fix double-checked locking in predicate suppliers in BoundDimFi… (#8974 ) * Fix double-checked locking in predicate suppliers in BoundDimFilter * Fix double-checked locking in predicate suppliers in BoundDimFilter * 1. Use Suppliers.memoize() to initialize and publish singleton. 2. Fix coding style. * Fix coding style * Fix double-checked locking bug for predicate suppliers in InDimFilter	2019-12-04 20:01:52 +03:00

1 2 3 4 5 ...

10091 Commits All Branches Search

10091 Commits

All Branches