druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	c5d6163c76	add a GeneratorInputSource to fill up a cluster with generated data for testing (#9946 ) * move benchmark data generator into druid-processing, add a GeneratorInputSource to fill up a cluster with data * newlines * make test coverage not fail maybe * remove useless test * Update pom.xml * Update GeneratorInputSourceTest.java * less passive aggressive test names	2020-06-09 19:31:04 -07:00
Atul Mohan	17cf8ea8f2	Add Sql InputSource (#9449 ) * Add Sql InputSource * Add spelling * Use separate DruidModule * Change module name * Fix docs * Use sqltestutils for tests * Add additional tests * Fix inspection * Add module test * Fix md in docs * Remove annotation Co-authored-by: Atul Mohan <atulmohan@yahoo-inc.com>	2020-06-09 12:55:20 -07:00
Lucas Capistrant	2ae1d26aa8	small fixes to configuration documentation (#9975 )	2020-06-09 10:31:08 -07:00
Jonathan Wei	771870ae2d	Load broadcast datasources on broker and tasks (#9971 ) * Load broadcast datasources on broker and tasks * Add javadocs * Support HTTP segment management * Fix indexer maxSize * inspection fix * Make segment cache optional on non-historicals * Fix build * Fix inspections, some coverage, failed tests * More tests * Add CliIndexer to MainTest * Fix inspection * Rename UnprunedDataSegment to LoadableDataSegment * Address PR comments * Fix	2020-06-08 20:15:59 -07:00
Clint Wylie	7f51e44b00	fix NilVectorSelector filter optimization (#9989 )	2020-06-08 17:40:29 -07:00
Yuanli Han	ee7bda5d8a	Fix compact partially overlapping segments (#9905 ) * fix compact overlapping segments * fix comment * fix CI failure	2020-06-08 09:54:39 -07:00
Maytas Monsereenusorn	45b699fa4a	Add git pre-commit hook to source control (#9554 ) * Add git pre-commit hook to source control * Changed hook to pre-push and simply hook to run all checkstyle * Clean up setup-hooks * Add apache header * Add apache header * add documentation to intellij-setup.md * retrigger tests * update Co-authored-by: Maytas Monsereenusorn <52679095+maytasm3@users.noreply.github.com>	2020-06-05 11:19:42 -10:00
Mainak Ghosh	bcc066a27f	Empty partitionDimension has less rollup compared to when explicitly specified (#9861 ) * Empty partitionDimension has less rollup compared to the case when it is explicitly specified * Adding a unit test for the empty partitionDimension scenario. Fixing another test which was failing * Fixing CI Build Inspection Issue * Addressing all review comments * Updating the javadocs for the hash method in HashBasedNumberedShardSpec	2020-06-05 12:42:42 -07:00
Clint Wylie	77dd5b06ae	ColumnCapabilities.hasMultipleValues refactor (#9731 ) * transition ColumnCapabilities.hasMultipleValues to Capable enum, remove ColumnCapabilities.isComplete * remove artifical, always multi-value capabilities from IncrementalIndexStorageAdapter and fix up fallout from that, fix ColumnCapabilities merge in index merger * fix typo * remove unused method * review stuffs, revert IncrementalIndexStorageAdapater capabilities change, plumb lame workaround to SegmentAnalyzer * more comment * use volatile booleans * fix line length * correctly handle missing columns for vector processors * return ColumnCapabilities.Capable for BitmapIndexSelector.hasMultipleValues, fix vector processor selection for complex * false on non-existent	2020-06-04 23:52:37 -07:00
agricenko	e72f490be0	Integration Tests. Small fixes for CI. (#9988 ) Co-authored-by: agritsenko <agritsenko@provectus.com>	2020-06-04 17:10:56 -07:00
Maytas Monsereenusorn	9738a03c83	Fix groupBy with literal in subquery grouping (#9986 ) * fix groupBy with literal in subquery grouping * fix groupBy with literal in subquery grouping * fix groupBy with literal in subquery grouping * address comments * update javadocs	2020-06-04 13:28:05 -10:00
Maytas Monsereenusorn	790e9482ea	Fix Subquery could not be converted to groupBy query (#9959 ) * Fix join * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * Fix Subquery could not be converted to groupBy query * add tests * address comments * fix failing tests	2020-06-03 16:46:28 -07:00
Jihoon Son	474f6fc99b	Fix shutdown reason for unknown tasks in taskQueue (#9954 ) * Fix shutdown reason for unknown tasks in taskQueue * unused imports	2020-06-03 15:40:28 -07:00
Gian Merlino	3dfd7c30c0	Add REGEXP_LIKE, fix bugs in REGEXP_EXTRACT. (#9893 ) * Add REGEXP_LIKE, fix empty-pattern bug in REGEXP_EXTRACT. - Add REGEXP_LIKE function that returns a boolean, and is useful in WHERE clauses. - Fix REGEXP_EXTRACT return type (should be nullable; causes incorrect filter elision). - Fix REGEXP_EXTRACT behavior for empty patterns: should always match (previously, they threw errors). - Improve error behavior when REGEXP_EXTRACT and REGEXP_LIKE are passed non-literal patterns. - Improve documentation of REGEXP_EXTRACT. * Changes based on PR review. * Fix arg check. * Important fixes! * Add speller. * wip * Additional tests. * Fix up tests. * Add validation error tests. * Additional tests. * Remove useless call.	2020-06-03 14:31:37 -07:00
Maytas Monsereenusorn	0d22462e07	Document unsupported Join on multi-value column (#9948 ) * Document Unsupported Join on multi-value column * Document Unsupported Join on multi-value column * address comments * Add unit tests * address comments * add tests	2020-06-03 09:55:52 -10:00
Xavier Léauté	a934b2664c	remove ListenableFutures and revert to using the Guava implementation (#9944 ) This change removes ListenableFutures.transformAsync in favor of the existing Guava Futures.transform implementation. Our own implementation had a bug which did not fail the future if the applied function threw an exception, resulting in the future never completing. An attempt was made to fix this bug, however when running againts Guava's own tests, our version failed another half dozen tests, so it was decided to not continue down that path and scrap our own implementation. Explanation for how was this bug manifested itself: An exception thrown in BaseAppenderatorDriver.publishInBackground when invoked via transformAsync in StreamAppenderatorDriver.publish will cause the resulting future to never complete. This explains why when encountering https://github.com/apache/druid/issues/9845 the task will never complete, forever waiting for the publishFuture to register the handoff. As a result, the corresponding "Error while publishing segments ..." message only gets logged once the index task times out and is forcefully shutdown when the future is force-cancelled by the executor.	2020-06-03 10:46:03 -07:00
Gian Merlino	3d81564a14	Fix various processing buffer leaks and simplify BlockingPool. (#9928 ) * - GroupByQueryEngineV2: Fix leak of intermediate processing buffer when exceptions are thrown before result sequence is created. - PooledTopNAlgorithm: Fix leak of intermediate processing buffer when exceptions are thrown before the PooledTopNParams object is created. - BlockingPool: Remove unused "take" methods. * Add tests to verify that buffers have been returned.	2020-06-02 18:26:18 -07:00
Gian Merlino	309fc04d54	Fix various Yielder leaks. (#9934 ) * Fix various Yielder leaks. - CombiningSequence leaked the input yielder from "toYielder" if it ran into an exception while accumulating the last value from the input yielder. - MergeSequence leaked input yielders from "toYielder" if it ran into an exception while building the initial priority queue. - ScanQueryRunnerFactory leaked the input yielder in its "priorityQueueSortAndLimit" strategy if it ran into an exception while scanning and sorting. - YieldingSequenceBase.accumulate chomped IOExceptions thrown in "accumulate" during yielder closing. * Add tests. * Fix braces.	2020-06-02 18:26:06 -07:00
Chi Cao Minh	3d513b0bec	Adjust code coverage check (#9969 ) Since there is not currently a good way to have fine-grain code coverage check exclusions, lower the coverage thresholds to make the check more lenient for now. Also, display the code coverage report in the Travis CI logs to make it easier to understand how to improve coverage.	2020-06-02 15:34:58 -07:00
Xavier Léauté	4ecf1900c3	fix nullhandling exceptions related to test ordering (#9964 ) follow-up to https://github.com/apache/druid/pull/9570	2020-06-02 10:13:54 -07:00
agricenko	56a9cad532	Integration Tests. (#9854 ) * Integration Tests. Added docker-compose with druid-cluster configuration. Refactored shell scripts. split code in a few files * Integration Tests. Added environment variable: DRUID_INTEGRATION_TEST_GROUP * Integration Tests. Removed nit * Integration Tests. Updated if block in docker_run_cluster.sh. * Integration Tests. Readme. Added Docker-compose section. * Integration Tests. removed yml files for s3, gcs, azure. Renamed variables for skip start/stop/build docker. Updated readme. Rollback maven profile: int-tests-config-file * Integration Tests. Removed docker-compose.test-env.yml file. Added DRUID_INTEGRATION_TEST_GROUP variable to docker-compose.yml * Integration Tests. Readme. Added details about docker-compose * Integration Tests. cleanup shell scripts Co-authored-by: agritsenko <agritsenko@provectus.com>	2020-06-02 09:38:53 -07:00
Clint Wylie	c690d10a7d	support customized factory.json via IndexSpec for segment persist (#9957 ) * support customized factory.json via IndexSpec for segment persist * equals verifier	2020-06-01 16:36:32 -07:00
Maytas Monsereenusorn	821c5d5a5c	Prevent JOIN reducing to a JOIN with constant in the ON condition (#9941 ) * Prevent Join reducing to on constant condition * Prevent Join reducing to on constant condition * addreess comments * set queryContext in tests	2020-06-01 09:39:06 -07:00
Xavier Léauté	acfcfd35b1	fix unsafe concurrent access in StreamAppenderatorDriver (#9943 ) during segment publishing we do streaming operations on a collection not safe for concurrent modification. To guarantee correct results we must also guard any operations on the stream itself. This may explain the issue seen in https://github.com/apache/druid/issues/9845	2020-05-31 09:12:25 -07:00
Clint Wylie	c2c38f6ac2	only close exec if it exists (#9952 )	2020-05-29 20:09:34 -07:00
Suneet Saldanha	e03d38b6c8	Optimize join queries where filter matches nothing (#9931 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests * Refactor JoinFilterAnalyzer - part 2 This change introduces the following components: * RhsRewriteCandidates - a wrapper for a list of candidates and associated functions to operate on the set of candidates. * JoinableClauses - a wrapper for the list of JoinableClause that represent a join condition and the associated functions to operate on the clauses. * Equiconditions - a wrapper representing the equiconditions that are used in the join condition. And associated test changes. This refactoring surfaced 2 bugs: - Missing equals and hashcode implementation for RhsRewriteCandidate, thus allowing potential duplicates in the rhs rewrite candidates - Missing Filter#supportsRequiredColumnRewrite check in analyzeJoinFilterClause, which could result in UnsupportedOperationException being thrown by the filter * fix compile error * remove unused class * Refactor JoinFilterAnalyzer - Correlations Move the correlation related code out into it's own class so it's easier to maintain. Another patch should follow this one so that the query path uses the correlation object instead of it's underlying maps. * Optimize join queries where filter matches nothing Fixes #9787 This PR changes the Joinable interface to return an Optional set of correlated values for a column. This allows the JoinFilterAnalyzer to differentiate between the case where the column has no matching values and when the column could not find matching values. This PR chose not to distinguish between cases where correlated values could not be computed because of a config that has this behavior disabled or because of user error - like a column that could not be found. The reasoning was that the latter is likely an error and the non filter pushdown path will surface the error if it is.	2020-05-29 16:53:03 -07:00
Suneet Saldanha	9c40bebc02	Refactor JoinFilterAnalyzer - part 2 (#9929 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests * Refactor JoinFilterAnalyzer - part 2 This change introduces the following components: * RhsRewriteCandidates - a wrapper for a list of candidates and associated functions to operate on the set of candidates. * JoinableClauses - a wrapper for the list of JoinableClause that represent a join condition and the associated functions to operate on the clauses. * Equiconditions - a wrapper representing the equiconditions that are used in the join condition. And associated test changes. This refactoring surfaced 2 bugs: - Missing equals and hashcode implementation for RhsRewriteCandidate, thus allowing potential duplicates in the rhs rewrite candidates - Missing Filter#supportsRequiredColumnRewrite check in analyzeJoinFilterClause, which could result in UnsupportedOperationException being thrown by the filter * fix compile error * remove unused class	2020-05-29 15:03:35 -07:00
sthetland	a33705f0e3	Querying doc refresh tutorial (#9879 ) * Update tutorial-query.md * First full pass complete * Smoothing over, a bit * link and spell checking * Update querying.md * Review comments; screenshot fixes * Making ports consistent, pending confirmation Switching to the Router port, to make this be consistent with the tutorial ports, but can switch back here and there if it should be 8082 instead. * Resizing screenshot * Update querying.md * Review feedback incorporated.	2020-05-29 14:32:21 -07:00
Surekha	ff551ae412	Modify information schema doc to specify correct value of TABLE_CATALOG (#9950 )	2020-05-29 10:10:28 -07:00
Suneet Saldanha	faef31a0af	Refactor JoinFilterAnalyzer (#9921 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests	2020-05-28 22:32:09 -07:00
Jonathan Wei	880a7943b3	Fix type restriction for Pattern hashcode inspection (#9947 )	2020-05-28 22:26:32 -07:00
Suneet Saldanha	cbd587dbd6	Add parameterized Calcite tests for join queries (#9923 ) * Add parameterized Calcite tests for join queries * new tests * review comments	2020-05-28 19:10:26 -07:00
Xavier Léauté	65280a6953	update kafka client version to 2.5.0 (#9902 ) - remove dependency on deprecated internal Kafka classes - keep LZ4 version in line with the version shipped with Kafka	2020-05-27 13:20:32 -07:00
Suneet Saldanha	b5dfa5322b	Make it easier for devs to add CalciteQueryTests (#9922 ) * Add ingestion specs for CalciteQueryTests This PR introduces ingestion specs that can be used for local testing so that CalciteQueryTests can be built on a druid cluster. * Add README * Update sql/src/test/resources/calcite/tests/README.md	2020-05-27 13:03:53 -07:00
Chi Cao Minh	b8c2266aa0	Disable function code coverage check (#9933 ) As observed in https://github.com/apache/druid/pull/9905 and https://github.com/apache/druid/pull/9915, the function code coverage check flags false positive issues, so it should be disabled for now.	2020-05-26 20:13:08 -07:00
Maytas Monsereenusorn	6130a834c2	Update doc on tmp dir (java.io.tmpdir) best practice (#9910 ) * Update doc on tmp dir best practice * remove local recommendation	2020-05-26 09:37:01 -07:00
Chi Cao Minh	fd6fffc4b8	Suppress CVEs for openstack-keystone (#9903 ) CVE-2020-12689, CVE-2020-12691, and CVE-2020-12690 can be ignored for openstack-keystone as they are for the python SDK and druid uses the java SDK.	2020-05-22 10:32:17 -07:00
Joseph Glanville	132a1c9fe7	Re-order and document format detection in web console (#9887 ) Motivation for this change is to not inadvertently identify binary formats that contain uncompressed string data as TSV or CSV. Moving detection of magic byte headers before heuristics should be more robust in general.	2020-05-21 16:29:39 -07:00
Vadim Ogievetsky	63baa29ad1	Fix web console query view crashing on simple query (#9897 ) * only parse full queries * upgraded sql parser	2020-05-21 12:57:07 -07:00
frank chen	b91d50044e	add some details to the build doc (#9885 ) * update initial build command * add some details for building * fix spelling check errors * fix spelling check warnings Signed-off-by: frank chen <frank.chen021@outlook.com>	2020-05-21 12:35:54 -07:00
Furkan KAMACI	ca5e66dfcb	Fixed the Copyright year of Druid (#9859 )	2020-05-20 21:13:34 -07:00
Maytas Monsereenusorn	9db29b93bf	Fix Hadoop IT Legacy test query json was not parameterized (#9901 )	2020-05-20 21:09:17 -07:00
Maytas Monsereenusorn	f470bcd11f	Fix deleting a data node tier causes load rules to display incorrectly (#9891 ) * Fix Deleting a data node tier causes load rules to malfunction & display incorrectly * add tests * fix style	2020-05-20 16:49:28 -07:00
Jianhuan Liu	2050f2b00a	fix docs error: google to azure and hdfs to http (#9881 )	2020-05-20 10:17:39 -07:00
Chi Cao Minh	427239f451	Enforce code coverage (#9863 ) * Enforce code coverage Add an automated way of checking if new code has adequate unit tests, since merging code coverage reports and check coverage thresholds via coveralls or codecov is unreliable. The following minimum unit test code coverage is now enforced: - 80% functions - 65% branch - 65% line Branch and line coverage thresholds are slightly lower for now as they are harder to achieve. After the code coverage check looks reliable, the thresholds can be increased later if needed. * Add comments	2020-05-20 09:31:37 -07:00
Maytas Monsereenusorn	5b4b5d77a8	Fails creation of TaskResource if availabilityGroup is null (#9892 ) * Fails creation of TaskResource if availabilityGroup is null * add check for requiredCapacity	2020-05-19 22:19:22 -07:00
Samarth Jain	82e5b0573e	Number based columns representing time in custom format cannot be used as timestamp column in Druid. (#9877 ) * Number based columns representing time in custom format cannot be used as timestamp column in Druid. Prior to this fix, if an integer column in parquet is storing dateint in format yyyyMMdd, it cannot be used as timestamp column in Druid as the timestamp parser interprets it as a number storing UTC time instead of treating it as a number representing time in yyyyMMdd format. Data formats like TSV or CSV don't suffer from this problem as the timestamp is passed in an as string which the timestamp parser is able to parse correctly.	2020-05-18 11:17:28 -07:00
Clint Wylie	2e9548d93d	refactor SeekableStreamSupervisor usage of RecordSupplier (#9819 ) * refactor SeekableStreamSupervisor usage of RecordSupplier to reduce contention between background threads and main thread, refactor KinesisRecordSupplier, refactor Kinesis lag metric collection and emitting * fix style and test * cleanup, refactor, javadocs, test * fixes * keep collecting current offsets and lag if unhealthy in background reporting thread * review stuffs * add comment	2020-05-16 14:09:39 -07:00
Alexander Saydakov	522df300c2	Datasketches 1 3 0 (#9880 ) * use the latest datasketches release * new sketch debug print Co-authored-by: AlexanderSaydakov <AlexanderSaydakov@users.noreply.github.com>	2020-05-16 14:09:23 -07:00
Joseph Glanville	793f386d6a	Add support for Avro OCF using InputFormat (#9671 ) * Add AvroOCFInputFormat * Support supplying a reader schema in AvroOCFInputFormat * Add docs for Avro OCF input format * Address review comments * Address second round of review	2020-05-16 14:09:12 -07:00

1 2 3 4 5 ...

10476 Commits All Branches Search

10476 Commits

All Branches