druid

Commit Graph

Author	SHA1	Message	Date
Abhishek Radhakrishnan	beccc401e1	Segments created in the same batch have the same `created_date` entry & rename metric (#15977 ) * All segments stored in the same batch have the same created_date entry. In the absence of a group_id column, this metadata would allow us to easily reason about and troubleshoot ingestion-related issues. * Rename metric name and code references to eligibleUnusedSegments. Address review comment from https://github.com/apache/druid/pull/15941#discussion_r1503631992	2024-02-27 17:28:43 +05:30
Karan Kumar	5bb5b41b18	Adding task pending time in MSQ reports (#15966 ) Added a new field pendingMs in MSQ task reports. This helps in figuring out the exact run time of the MSQ worker tasks. Fixed data races.	2024-02-27 14:41:28 +05:30
Abhishek Radhakrishnan	38ecf980d0	Refactor and add tests and metric to KillUnusedSegments duty (auto-kill) (#15941 ) * Kill duty and test improvements. Initial commit with: - Bug fixes - auto-kill can throw NPE when there are no datasources present and defaults mismatch. - Add new stat for candidate segment intervals killed. - Move a couple of debug logs to info logs for improved visibility (should only log once per kill period). - Remove redundant checks for code readability. - Updated tests from using mocks (also the mocks weren't using last updated timestamp) and add more test coverage for different config parameters. - Add a couple of unit tests that are ignored for the eternity case to prove that the kill duty doesn't clean up segments with ALL grain or that end in DateTimes.MAX. - Migrate Druid exception from user to operator persona. * Address review comments. * Remove unused methods. * fix up format specifier and validate bad config tests. * Consolidate the helpers a bit more and add another test. * Update test names. Add javadoc placeholders for slightly involved tests. * Add docs for metric kill/candidateUnusedSegments/count. Also, rename to disambiguate. * Comments. * Apply logging suggestions from code review Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> * Review comments - Clarify docs on eligibility. - Add test for multiple segments in the same interval. Clarify comment. - Remove log line from test. - Remove lastUpdatedDate = now.plus(10) from test. * minor cleanup. * Clarify javadocs for getUnusedSegmentIntervals(). --------- Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com>	2024-02-27 12:14:41 +05:30
Laksh Singla	17e4f3ac60	Refactor GroupBy and TopN code to relax the constraint of dimensions being comparable (#15559 ) The code in the groupBy engine and the topN engine assume that the dimensions are comparable and can call dimA.compareTo(dimB) to sort the dimensions and group them together. This works well for the primitive dimensions, because they are Comparable, however falls apart when the dimensions can be arrays (or in future scenarios complex columns). In cases when the dimensions are not comparable, Druid resorts to having a wrapper type ComparableStringArray and ComparableList, which is a Comparable, based on the list comparator.	2024-02-27 11:39:29 +05:30
Soumyava	51cc729fd1	Enforcing type checking for flatten concat (#15903 )	2024-02-26 21:53:49 -08:00
Vadim Ogievetsky	a81429746d	Web console: fix typos in Kinesis suggestions, add regions and groups (#15900 ) * fix typo * update regions * add China * Update web-console/src/druid-models/ingestion-spec/ingestion-spec.tsx Co-authored-by: Benedict Jin <asdf2014@apache.org> * add , --------- Co-authored-by: Benedict Jin <asdf2014@apache.org>	2024-02-27 10:00:02 +08:00
Vadim Ogievetsky	bf3139562c	Web console: support for the export execution state (#15969 ) * init * add CSV keyword	2024-02-26 11:28:25 -08:00
Vadim Ogievetsky	28b3e117cf	Web console: Add input format props (#15950 ) * fix typo * add Protobuf * better padding	2024-02-26 11:28:09 -08:00
Abhishek Radhakrishnan	67a6224d91	Fix up incorrect `PARTITIONED BY` error messages (#15961 ) * Fix up typos, inaccuracies and clean up code related to PARTITIONED BY. * Remove wrapper function and update tests to use DruidExceptionMatcher. * Checkstyle and Intellij inspection fixes.	2024-02-26 14:17:53 -05:00
Abhishek Agarwal	ddfc31d7ed	Reduce the size of distribution docker image (#15968 ) This PR creates symlinks when there are duplicate jars present in the extension. Docker image includes contrib extensions, too, and the size of the image has bloated up quite a lot of late. This change also fixes "ITNestedQueryPushDownTest integration test"	2024-02-26 21:18:55 +05:30
AmatyaAvadhanula	e2b7289dea	Try to fetch the task status for an active from memory (#15724 ) * Reduce metadata calls to fetch the status for an active task	2024-02-26 13:53:05 +05:30
Benjamin Hopp	ebb7190545	Docs: Change single-dim to hashed in example for index task (#15529 )	2024-02-26 09:16:10 +05:30
Zoltan Haindrich	06deda9415	ScanAndSort query fails with NPE for simple queries (#15914 ) * some stuff * add dummy fields * draft-fix * rename test * cleanup * add null * cleanup * cleanup * add test * updates * move check tp constructore * cleanup * updates/etc * fix some more * add rowSignatureMode * checkstyle/etc * override * missing msqIncompat * fix test * fixes * undo * updates * remove param	2024-02-24 15:33:50 -08:00
Clint Wylie	6145c8dd01	fix bug with expression virtual column indexes on missing columns for expressions that turn null values into not null values (#15959 )	2024-02-23 15:07:32 -08:00
Gian Merlino	b69f89d9f8	Clarify where to set druid.monitoring.monitors. (#15729 )	2024-02-23 18:49:37 +05:30
Adithya Chakilam	1f443d218c	Enable partition stats on streaming task completion report (#15930 ) Changes: - Add visibility into number of records processed by each streaming task per partition - Add field `recordsProcessed` to `IngestionStatsAndErrorsTaskReportData` - Populate number of records processed per partition in `SeekableStreamIndexTaskRunner`	2024-02-23 16:29:03 +05:30
dependabot[bot]	3011829419	Bump log4j.version from 2.18.0 to 2.22.1 (#15934 ) * Bump log4j.version from 2.18.0 to 2.22.1 Bumps `log4j.version` from 2.18.0 to 2.22.1. Updates `org.apache.logging.log4j:log4j-api` from 2.18.0 to 2.22.1 Updates `org.apache.logging.log4j:log4j-core` from 2.18.0 to 2.22.1 Updates `org.apache.logging.log4j:log4j-slf4j-impl` from 2.18.0 to 2.22.1 Updates `org.apache.logging.log4j:log4j-1.2-api` from 2.18.0 to 2.22.1 Updates `org.apache.logging.log4j:log4j-jul` from 2.18.0 to 2.22.1 --- updated-dependencies: - dependency-name: org.apache.logging.log4j:log4j-api dependency-type: direct:production update-type: version-update:semver-minor - dependency-name: org.apache.logging.log4j:log4j-core dependency-type: direct:production update-type: version-update:semver-minor - dependency-name: org.apache.logging.log4j:log4j-slf4j-impl dependency-type: direct:production update-type: version-update:semver-minor - dependency-name: org.apache.logging.log4j:log4j-1.2-api dependency-type: direct:production update-type: version-update:semver-minor - dependency-name: org.apache.logging.log4j:log4j-jul dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Update License --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: frank chen <frank.chen021@outlook.com>	2024-02-23 16:19:35 +08:00
dependabot[bot]	936ba25e85	Bump org.postgresql:postgresql from 42.6.0 to 42.7.2 (#15931 ) * Bump org.postgresql:postgresql from 42.6.0 to 42.7.2 Bumps [org.postgresql:postgresql](https://github.com/pgjdbc/pgjdbc) from 42.6.0 to 42.7.2. - [Release notes](https://github.com/pgjdbc/pgjdbc/releases) - [Changelog](https://github.com/pgjdbc/pgjdbc/blob/master/CHANGELOG.md) - [Commits](https://github.com/pgjdbc/pgjdbc/commits) --- updated-dependencies: - dependency-name: org.postgresql:postgresql dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> * Update License --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: frank chen <frank.chen021@outlook.com>	2024-02-23 16:19:26 +08:00
Vadim Ogievetsky	c52ddd0b86	make flattenSpec location adaptive (#15946 )	2024-02-22 14:07:04 -08:00
zachjsh	8ebf237576	Move INSERT & REPLACE validation to the Calcite validator (#15908 ) This PR contains a portion of the changes from the inactive draft PR for integrating the catalog with the Calcite planner https://github.com/apache/druid/pull/13686 from @paul-rogers, Refactoring the IngestHandler and subclasses to produce a validated SqlInsert instance node instead of the previous Insert source node. The SqlInsert node is then validated in the calcite validator. The validation that is implemented as part of this pr, is only that for the source node, and some of the validation that was previously done in the ingest handlers. As part of this change, the partitionedBy clause can be supplied by the table catalog metadata if it exists, and can be omitted from the ingest time query in this case.	2024-02-22 14:01:59 -05:00
Katya Macedo	f37d019fe6	Fix redirects for streaming ingestion (#15943 )	2024-02-22 22:34:19 +05:30
Clint Wylie	cc5964fbcb	fix NestedCommonFormatColumnHandler to use nullable comparator when castToType is set (#15921 ) Fixes a bug when the undocumented castToType parameter is set on 'auto' column schema, which should have been using the 'nullable' comparator to allow null values to be present when merging columns, but wasn't which would lead to null pointer exceptions. Also fixes an issue I noticed while adding tests that if 'FLOAT' type was specified for the castToType parameter it would be an exception because that type is not expected to be present, since 'auto' uses the native expressions to determine the input types and expressions don't have direct support for floats, only doubles. In the future I should probably split this functionality out of the 'auto' schema (maybe even have a simpler version of the auto indexer dedicated to handling non-nested data) but still have the same results of writing out the newer 'nested common format' columns used by 'auto', but I haven't taken that on in this PR.	2024-02-22 21:35:50 +05:30
Jamie	80942d5754	Feature: add support for ingesting from rabbitmq super streams (#14137 ) * Add support for ingesting from Rabbit MQ Super Streams	2024-02-22 10:50:37 +05:30
George Shiqi Wu	59bb72a926	Fix parsing of env variables when properties have underscores (#15919 ) * Fix parsing of env variables when properties have underscores * Add documentation * Use a % sign instead	2024-02-21 13:18:21 -05:00
Zoltan Haindrich	bcce0806d7	Support Union in decoupled mode (#15870 )	2024-02-21 10:54:50 -05:00
Zoltan Haindrich	170d37f188	add check to build docker image (#15894 )	2024-02-21 10:53:35 -05:00
Benedict Jin	0f38a98368	Update the link of Helm Chart to avoid 404 error (#15905 )	2024-02-21 16:57:41 +08:00
Suneet Saldanha	cbc53d53b4	Update k8sTaskRunner log message (#15871 )	2024-02-21 14:34:00 +08:00
Gian Merlino	e20004c7df	Remove helm chart. (#15904 ) The helm chart was originally moved here in #11163 from https://github.com/helm/charts/tree/master/incubator/druid after the helm/charts repository was deprecated. However, it has been excluded from releases since then, due to uncertainty around whether we need IP clearance. We have not had volunteers willing to sort this out, so this patch removes the code. It can be re-added if a volunteer is available to sort out the IP clearance process. See thread at: https://lists.apache.org/thread/ygyzt23m06vc775nq5dsm349rf0j47dg	2024-02-21 14:21:37 +08:00
Laksh Singla	a1b2c7326e	Numeric array support for columnar frames (#15917 ) Columnar frames used in subquery materialization and window functions now support numeric arrays.	2024-02-21 11:32:33 +05:30
George Shiqi Wu	2c0d1128f8	Fix pod template reading logic (#15915 ) * Fix pod template reading * PR changes * Fix unit tests	2024-02-20 11:13:51 -05:00
Adarsh Sanjeev	9eaaeb5c16	Add security ITs to the revised integration tests (#15885 ) * Add IT for security * Add admin client * Clean up code * Clean up code * Address review comments	2024-02-20 11:32:08 +05:30
Gian Merlino	9c41827dba	Globally disable AUTO_CLOSE_JSON_CONTENT. (#15880 ) * Globally disable AUTO_CLOSE_JSON_CONTENT. This JsonGenerator feature is on by default. It causes problems with code like this: try (JsonGenerator jg = ...) { jg.writeStartArray(); for (x : xs) { jg.writeObject(x); } jg.writeEndArray(); } If a jg.writeObject call fails due to some problem with the data it's reading, the JsonGenerator will write the end array marker automatically when closed as part of the try-with-resources. If the generator is writing to a stream where the reader does not have some other mechanism to realize that an exception was thrown, this leads the reader to believe that the array is complete when it actually isn't. Prior to this patch, we disabled AUTO_CLOSE_JSON_CONTENT for JSON-wrapped SQL result formats in #11685, which fixed an issue where such results could be erroneously interpreted as complete. This patch fixes a similar issue with task reports, and all similar issues that may exist elsewhere, by disabling the feature globally. * Update test.	2024-02-16 08:52:48 -08:00
Clint Wylie	fe2ba8cc28	fix return type inference of parse_long, which can also be null if string is not parseable into a long (#15909 ) * fix return type inference of parse_long, which can also be null if string is not parseable into a long * fix msq test	2024-02-15 08:45:34 -08:00
Vadim Ogievetsky	66f54f2066	allow compaction config slots to drop to 0 (#15877 )	2024-02-15 15:27:15 +08:00
Parth Agrawal	495e66f2e7	CVE Fix: Update json-path version (#15772 ) Apache Druid brings the dependency json-path which is affected by CVE-2023-51074. Its latest version 2.9.0 fixes the above CVE. Append function has been added to json-path and so the unit test to check for the append function not present has been updated. --------- Co-authored-by: Xavier Léauté <xvrl@apache.org>	2024-02-14 20:58:27 -08:00
Tom	f224035c7e	Fix Flakiness in KafkaEmitterTest (#15907 ) * thrust of the fix to allow for the json values to be out of order The existing problem is that toMap doesn't turn some values into json primitive values, for example segmentMetadata just has DateTime objects for it's time in the EventMap, but Alert event converts those into strings when calling toMap. This creates an issue because when we check the emitted events the mapper deserializing the string value for dateTime leaves it as a string in the EventMap. So the question is do we alter the events toMap() to return string/map version of objects or to make the expected events do a round trip of eventMap -> string -> eventMap to turn everything into json primitives * fix issue by making toMap events convert Objects into strings, or maps * fix linting errors * use method of using mapper to round trip expected data to make it have same type as those of the events emitted * remove unnecessary comment	2024-02-15 10:01:55 +05:30
317brian	c98d54f3c4	docs: delete unused file that causes confusion (#15910 )	2024-02-14 16:42:02 -08:00
YongGang	19ed5c863f	Enhance rolling Supervisor restarts at taskDuration (#15859 )	2024-02-14 15:44:34 -08:00
Abhishek Radhakrishnan	c324e37751	Add javadocs to `KafkaEmitterTest` & fix flaky test (#15898 ) * Address review comment: add test javadocs * Fix flaky assertion failure. Use ConcurrentHashMap instead of HashMap because the producer callback can trigger concurrently and override the map initialization. * fixup intellij inspection	2024-02-14 11:52:06 -08:00
Sam Rash	be0ee2ee33	update version check for profiling to >= 17 (#15686 )	2024-02-14 21:44:20 +05:30
Peter Marshall	cae9cbd7d7	Update tasks.md (#15887 ) Remove erroneous white space causing render issues on this page.	2024-02-13 05:20:09 -08:00
Clint Wylie	dad8398a4d	start process of deprecating non-sql compatible legacy configurations (#15713 ) Starting the process to officially deprecate non SQL compatible modes by updating docs to aggressively call out that Druids non SQL compliant modes are deprecated and will go away someday. There are no code or behavior changes at this PR.	2024-02-13 15:31:45 +05:30
Tom	c225c19f81	fix copy paste issue in earlier PR (#15890 )	2024-02-12 19:49:19 -05:00
Gian Merlino	0f6a895372	Rework ExprMacro base classes to simplify implementations. (#15622 ) * Rework ExprMacro base classes to simplify implementations. This patch removes BaseScalarUnivariateMacroFunctionExpr, adds BaseMacroFunctionExpr at the top of the hierarchy (a suitable base class for ExprMacros that take either arrays or scalars), and adds an implementation for "visit" to BaseMacroFunctionExpr. The effect on implementations is generally cleaner code: - Exprs no longer need to implement "visit". - Exprs no longer need to implement "stringify", even if they don't use all of their args at runtime, because BaseMacroFunctionExpr has access to even unused args. - Exprs that accept arrays can extend BaseMacroFunctionExpr and inherit a bunch of useful methods. The only one they need to implement themselves that scalar exprs don't is "supplyAnalyzeInputs". * Make StringDecodeBase64UTFExpression a static class. * Remove unused import. * Formatting, annotation changes.	2024-02-12 15:50:45 -08:00
Katya Macedo	0f29ece6a9	[Docs] Refactor streaming ingestion section (#15591 ) Merging the work so far. @ektravel , @vogievetsky if there are additional improvements, let's track them & make another pr. * Refactor streaming ingestion docs * Update property definition * Update after review * Update known issues * Move kinesis and kafka topics to ingestion, add redirects * Saving changes * Saving * Add input format text * Update after review * Minor text edit * Update example syntax * Revert back to colon * Fix merge conflicts * Fix broken links * Fix spelling error	2024-02-12 13:52:42 -08:00
Charles Smith	2a42b11660	remove legacy Jupyter tutorial files (#15834 ) * remove legacy files * redirection for the jupyter tutorial page * remove tutorial from sidebar * remove redirection	2024-02-12 13:45:47 -08:00
Abhishek Radhakrishnan	51fd79ee58	Clean up kafka emitter tests, add more validations and code coverage. (#15878 ) * Clean up kafka emitter tests a bit and add more validations. The test wasn't validating what events were sent, but simply the dropped counters, which aren't that useful. Additionally, this module has fewer tests, so folks often run into code coverage issue in this extension. Hopefully this change helps with that too. * Change things to feed-based rather than topic-based. * Another test for shared topic * Switch to DruidException, add test dependencies and sad path config tests. * missing test dependency * minor renames. * Add more tests - to test unknown events and drop when queue is full	2024-02-12 16:22:19 -05:00
Gian Merlino	7fea34abdd	LOOKUP docs: clarify behavior of replaceMissingValueWith. (#15879 ) Clarify behavior when expr is null.	2024-02-11 13:11:00 -08:00
zachjsh	f9ee2c353b	Extend the PARTITION BY clause to accept string literals for the time partitioning (#15836 ) This PR contains a portion of the changes from the inactive draft PR for integrating the catalog with the Calcite planner https://github.com/apache/druid/pull/13686 from @paul-rogers, extending the PARTITION BY clause to accept string literals for the time partitioning	2024-02-09 11:45:38 -05:00

... 2 3 4 5 6 ...

13875 Commits All Branches Search

13875 Commits

All Branches