druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	c690d10a7d	support customized factory.json via IndexSpec for segment persist (#9957 ) * support customized factory.json via IndexSpec for segment persist * equals verifier	2020-06-01 16:36:32 -07:00
Maytas Monsereenusorn	821c5d5a5c	Prevent JOIN reducing to a JOIN with constant in the ON condition (#9941 ) * Prevent Join reducing to on constant condition * Prevent Join reducing to on constant condition * addreess comments * set queryContext in tests	2020-06-01 09:39:06 -07:00
Xavier Léauté	acfcfd35b1	fix unsafe concurrent access in StreamAppenderatorDriver (#9943 ) during segment publishing we do streaming operations on a collection not safe for concurrent modification. To guarantee correct results we must also guard any operations on the stream itself. This may explain the issue seen in https://github.com/apache/druid/issues/9845	2020-05-31 09:12:25 -07:00
Clint Wylie	c2c38f6ac2	only close exec if it exists (#9952 )	2020-05-29 20:09:34 -07:00
Suneet Saldanha	e03d38b6c8	Optimize join queries where filter matches nothing (#9931 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests * Refactor JoinFilterAnalyzer - part 2 This change introduces the following components: * RhsRewriteCandidates - a wrapper for a list of candidates and associated functions to operate on the set of candidates. * JoinableClauses - a wrapper for the list of JoinableClause that represent a join condition and the associated functions to operate on the clauses. * Equiconditions - a wrapper representing the equiconditions that are used in the join condition. And associated test changes. This refactoring surfaced 2 bugs: - Missing equals and hashcode implementation for RhsRewriteCandidate, thus allowing potential duplicates in the rhs rewrite candidates - Missing Filter#supportsRequiredColumnRewrite check in analyzeJoinFilterClause, which could result in UnsupportedOperationException being thrown by the filter * fix compile error * remove unused class * Refactor JoinFilterAnalyzer - Correlations Move the correlation related code out into it's own class so it's easier to maintain. Another patch should follow this one so that the query path uses the correlation object instead of it's underlying maps. * Optimize join queries where filter matches nothing Fixes #9787 This PR changes the Joinable interface to return an Optional set of correlated values for a column. This allows the JoinFilterAnalyzer to differentiate between the case where the column has no matching values and when the column could not find matching values. This PR chose not to distinguish between cases where correlated values could not be computed because of a config that has this behavior disabled or because of user error - like a column that could not be found. The reasoning was that the latter is likely an error and the non filter pushdown path will surface the error if it is.	2020-05-29 16:53:03 -07:00
Suneet Saldanha	9c40bebc02	Refactor JoinFilterAnalyzer - part 2 (#9929 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests * Refactor JoinFilterAnalyzer - part 2 This change introduces the following components: * RhsRewriteCandidates - a wrapper for a list of candidates and associated functions to operate on the set of candidates. * JoinableClauses - a wrapper for the list of JoinableClause that represent a join condition and the associated functions to operate on the clauses. * Equiconditions - a wrapper representing the equiconditions that are used in the join condition. And associated test changes. This refactoring surfaced 2 bugs: - Missing equals and hashcode implementation for RhsRewriteCandidate, thus allowing potential duplicates in the rhs rewrite candidates - Missing Filter#supportsRequiredColumnRewrite check in analyzeJoinFilterClause, which could result in UnsupportedOperationException being thrown by the filter * fix compile error * remove unused class	2020-05-29 15:03:35 -07:00
sthetland	a33705f0e3	Querying doc refresh tutorial (#9879 ) * Update tutorial-query.md * First full pass complete * Smoothing over, a bit * link and spell checking * Update querying.md * Review comments; screenshot fixes * Making ports consistent, pending confirmation Switching to the Router port, to make this be consistent with the tutorial ports, but can switch back here and there if it should be 8082 instead. * Resizing screenshot * Update querying.md * Review feedback incorporated.	2020-05-29 14:32:21 -07:00
Surekha	ff551ae412	Modify information schema doc to specify correct value of TABLE_CATALOG (#9950 )	2020-05-29 10:10:28 -07:00
Suneet Saldanha	faef31a0af	Refactor JoinFilterAnalyzer (#9921 ) * Refactor JoinFilterAnalyzer This patch attempts to make it easier to follow the join filter analysis code with the hope of making it easier to add rewrite optimizations in the future. To keep the patch small and easy to review, this is the first of at least 2 patches that are planned. This patch adds a builder to the Pre-Analysis, so that it is easier to instantiate the preAnalysis. It also moves some of the filter normalization code out to Fitlers with associated tests. * fix tests	2020-05-28 22:32:09 -07:00
Jonathan Wei	880a7943b3	Fix type restriction for Pattern hashcode inspection (#9947 )	2020-05-28 22:26:32 -07:00
Suneet Saldanha	cbd587dbd6	Add parameterized Calcite tests for join queries (#9923 ) * Add parameterized Calcite tests for join queries * new tests * review comments	2020-05-28 19:10:26 -07:00
Xavier Léauté	65280a6953	update kafka client version to 2.5.0 (#9902 ) - remove dependency on deprecated internal Kafka classes - keep LZ4 version in line with the version shipped with Kafka	2020-05-27 13:20:32 -07:00
Suneet Saldanha	b5dfa5322b	Make it easier for devs to add CalciteQueryTests (#9922 ) * Add ingestion specs for CalciteQueryTests This PR introduces ingestion specs that can be used for local testing so that CalciteQueryTests can be built on a druid cluster. * Add README * Update sql/src/test/resources/calcite/tests/README.md	2020-05-27 13:03:53 -07:00
Chi Cao Minh	b8c2266aa0	Disable function code coverage check (#9933 ) As observed in https://github.com/apache/druid/pull/9905 and https://github.com/apache/druid/pull/9915, the function code coverage check flags false positive issues, so it should be disabled for now.	2020-05-26 20:13:08 -07:00
Maytas Monsereenusorn	6130a834c2	Update doc on tmp dir (java.io.tmpdir) best practice (#9910 ) * Update doc on tmp dir best practice * remove local recommendation	2020-05-26 09:37:01 -07:00
Chi Cao Minh	fd6fffc4b8	Suppress CVEs for openstack-keystone (#9903 ) CVE-2020-12689, CVE-2020-12691, and CVE-2020-12690 can be ignored for openstack-keystone as they are for the python SDK and druid uses the java SDK.	2020-05-22 10:32:17 -07:00
Joseph Glanville	132a1c9fe7	Re-order and document format detection in web console (#9887 ) Motivation for this change is to not inadvertently identify binary formats that contain uncompressed string data as TSV or CSV. Moving detection of magic byte headers before heuristics should be more robust in general.	2020-05-21 16:29:39 -07:00
Vadim Ogievetsky	63baa29ad1	Fix web console query view crashing on simple query (#9897 ) * only parse full queries * upgraded sql parser	2020-05-21 12:57:07 -07:00
frank chen	b91d50044e	add some details to the build doc (#9885 ) * update initial build command * add some details for building * fix spelling check errors * fix spelling check warnings Signed-off-by: frank chen <frank.chen021@outlook.com>	2020-05-21 12:35:54 -07:00
Furkan KAMACI	ca5e66dfcb	Fixed the Copyright year of Druid (#9859 )	2020-05-20 21:13:34 -07:00
Maytas Monsereenusorn	9db29b93bf	Fix Hadoop IT Legacy test query json was not parameterized (#9901 )	2020-05-20 21:09:17 -07:00
Maytas Monsereenusorn	f470bcd11f	Fix deleting a data node tier causes load rules to display incorrectly (#9891 ) * Fix Deleting a data node tier causes load rules to malfunction & display incorrectly * add tests * fix style	2020-05-20 16:49:28 -07:00
Jianhuan Liu	2050f2b00a	fix docs error: google to azure and hdfs to http (#9881 )	2020-05-20 10:17:39 -07:00
Chi Cao Minh	427239f451	Enforce code coverage (#9863 ) * Enforce code coverage Add an automated way of checking if new code has adequate unit tests, since merging code coverage reports and check coverage thresholds via coveralls or codecov is unreliable. The following minimum unit test code coverage is now enforced: - 80% functions - 65% branch - 65% line Branch and line coverage thresholds are slightly lower for now as they are harder to achieve. After the code coverage check looks reliable, the thresholds can be increased later if needed. * Add comments	2020-05-20 09:31:37 -07:00
Maytas Monsereenusorn	5b4b5d77a8	Fails creation of TaskResource if availabilityGroup is null (#9892 ) * Fails creation of TaskResource if availabilityGroup is null * add check for requiredCapacity	2020-05-19 22:19:22 -07:00
Samarth Jain	82e5b0573e	Number based columns representing time in custom format cannot be used as timestamp column in Druid. (#9877 ) * Number based columns representing time in custom format cannot be used as timestamp column in Druid. Prior to this fix, if an integer column in parquet is storing dateint in format yyyyMMdd, it cannot be used as timestamp column in Druid as the timestamp parser interprets it as a number storing UTC time instead of treating it as a number representing time in yyyyMMdd format. Data formats like TSV or CSV don't suffer from this problem as the timestamp is passed in an as string which the timestamp parser is able to parse correctly.	2020-05-18 11:17:28 -07:00
Clint Wylie	2e9548d93d	refactor SeekableStreamSupervisor usage of RecordSupplier (#9819 ) * refactor SeekableStreamSupervisor usage of RecordSupplier to reduce contention between background threads and main thread, refactor KinesisRecordSupplier, refactor Kinesis lag metric collection and emitting * fix style and test * cleanup, refactor, javadocs, test * fixes * keep collecting current offsets and lag if unhealthy in background reporting thread * review stuffs * add comment	2020-05-16 14:09:39 -07:00
Alexander Saydakov	522df300c2	Datasketches 1 3 0 (#9880 ) * use the latest datasketches release * new sketch debug print Co-authored-by: AlexanderSaydakov <AlexanderSaydakov@users.noreply.github.com>	2020-05-16 14:09:23 -07:00
Joseph Glanville	793f386d6a	Add support for Avro OCF using InputFormat (#9671 ) * Add AvroOCFInputFormat * Support supplying a reader schema in AvroOCFInputFormat * Add docs for Avro OCF input format * Address review comments * Address second round of review	2020-05-16 14:09:12 -07:00
Jihoon Son	46beaa0640	Fix potential resource leak in ParquetReader (#9852 ) * Fix potential resource leak in ParquetReader * add test * never thrown exception * catch potential exceptions	2020-05-16 09:57:12 -07:00
Maytas Monsereenusorn	0a8bf83bc5	Bad plan for table-lookup-lookup join with filter on first lookup and outer limit (#9773 ) * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * Bad plan for table-lookup-lookup join with filter on first lookup and outer limit * address comments * address comments * fix checkstyle * address comments * address comments	2020-05-14 16:56:40 -07:00
zachjsh	80b212fe43	druid.storage.maxListingLength should default to 1000 for s3 (#9858 ) * druid.storage.maxListingLength should default to 1000 for s3 * * Address review comments * * Address review comments * * Address comments	2020-05-14 07:00:51 -07:00
Chi Cao Minh	41cf826928	Console E2E test docs (#9864 )	2020-05-13 16:41:04 -07:00
Suneet Saldanha	d38d77cb3a	Add back FieldMayBeFinal inspection (#9865 )	2020-05-13 16:32:35 -07:00
Suneet Saldanha	b0167295d7	Fail incorrectly constructed join queries (#9830 ) * Fail incorrectly constructed join queries * wip annotation for equals implementations * Add equals tests * fix tests * Actually fix the tests * Address review comments * prohibit Pattern.hashCode()	2020-05-13 14:23:04 -07:00
Clint Wylie	6bc1d1b33f	fix license registry for com.nimbusds lang-tag (#9860 )	2020-05-13 09:18:18 -07:00
Jihoon Son	c06d3f14b1	Add javadoc for stream ingestion integration tests (#9795 )	2020-05-12 08:56:43 -07:00
awelsh93	6f25a84d2e	Add TaskCountStatsMonitor to config docs (#9447 )	2020-05-11 14:08:46 -07:00
Jonathan Wei	16d293d6e0	Directly rewrite filters on RHS join columns into LHS equivalents (#9818 ) * Directly rewrite filters on RHS join columns into LHS equivalents * PR comments * Fix inspection * Revert unnecessary ExprMacroTable change * Fix build after merge * Address PR comments	2020-05-08 23:45:35 -07:00
mcbrewster	28be107a1c	add flag to flattenSpec to keep null columns (#9814 ) * add flag to flattenSpec to keep null columns * remove changes to inputFormat interface * add comment * change comment message * update web console e2e test * move keepNullColmns to JSONParseSpec * fix merge conflicts * fix tests * set keepNullColumns to false by default * fix lgtm * change Boolean to boolean, add keepNullColumns to hash, add tests for keepKeepNullColumns false + true with no nuulul columns * Add equals verifier tests	2020-05-08 21:53:39 -07:00
Clint Wylie	339876b69d	fill out missing test coverage for druid-stats, druid-momentsketch, druid-tdigestsketch postaggs (#9740 ) * postagg test coverage for druid-stats, druid-momentsketch, druid-tdigestsketch and fixes * style fixes * fix comparator for TDigestQuantilePostAggregator	2020-05-07 13:48:33 -07:00
Clint Wylie	267a6cc175	low hanging fruit - presize hash map for DruidSegmentReader (#9836 )	2020-05-07 12:39:14 -07:00
Clint Wylie	2c0746cfab	increase druid-histogram postagg test coverage (#9732 )	2020-05-07 00:10:29 -07:00
Maytas Monsereenusorn	accd710115	Add equivalent test coverage for all RHS join impls (#9831 ) * Add equivalent test coverage for all RHS join impls * address comments	2020-05-06 16:10:41 -07:00
Jihoon Son	6674d721bc	Avoid sorting values in InDimFilter if possible (#9800 ) * Avoid sorting values in InDimFilter if possible * tests * more tests * fix and and or filters * fix build * false and true vector matchers * fix vector matchers * checkstyle * in filter null handling * remove wrong test * address comments * remove unnecessary null check * redundant separator * address comments * typo * tests	2020-05-06 15:26:36 -07:00
sthetland	ce03f31a73	Clarifying workerThreads and a few other nits (#9804 ) * Update data-formats.md Per Suneet, "Since you're editing this file can you also fix the json on line 177 please - it's missing a comma after the }" * Light text cleanup * Removing discussion of sample data, since it's repeated in the data loading tutorial, and not immediately relevant here. * Clarifying accepted values for URI lookup * Update index.md * original quickstart full first pass * original quickstart full first pass * first pass all the way through * straggler * image touchups and finished old tutorial * a bit of finishing up * druid-caffeine-cache ext previously removed * Sample MaxDirectMemorySize value unrealistic * Review comments * fixing links * spell checking gymnastics * workerThreads desc slightly expanded * typo * Typo * Reversing Kafka config order * Changing order of configs for Kinesis * Trying this again: ioConfig then tuningConfig	2020-05-06 09:05:18 -07:00
Suneet Saldanha	1e857c5303	Ignore druid-processing benchmarks in tests (#9821 )	2020-05-06 08:59:48 -07:00
Jihoon Son	964a1fc9df	Remove ParseSpec.toInputFormat() (#9815 ) * Remove toInputFormat() from ParseSpec * fix test	2020-05-05 11:17:57 -07:00
Jihoon Son	c6caae9a24	Fix filtering on boolean values in transformation (#9812 ) * Fix filter on boolean value in Transform * assert * more descriptive test * remove assert * add assert for cached string; disable tests * typo	2020-05-04 18:47:10 -07:00
Alexander Saydakov	844d626738	added number of bins parameter (#9436 ) * added number of bins parameter * addressed review points * test equals Co-authored-by: AlexanderSaydakov <AlexanderSaydakov@users.noreply.github.com>	2020-05-04 16:53:09 -07:00

1 2 3 4 5 ...

10355 Commits All Branches Search

10355 Commits

All Branches