druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	d6c07270a5	fix issues with join filter pushdown and virtual column resolution (#16702 )	2024-07-11 04:26:07 -07:00
YongGang	4b293fc2a9	Docs: Fix k8s dynamic config URL (#16720 )	2024-07-11 10:05:47 +05:30
Kashif Faraz	616ae631c6	Fix NPE in CompactSegments (#16713 )	2024-07-10 14:51:52 +08:00
Adarsh Sanjeev	7c625356c5	Add logging for sketches on workers (#16697 ) Improve the logging of sketches on workers.	2024-07-09 14:37:43 +05:30
Adarsh Sanjeev	af5399cd9d	Fixes a bug when running queries with a limit clause (#16643 ) Add a shuffling based on the resultShuffleSpecFactory after a limit processor depending on the query destination. LimitFrameProcessors currently do not update the partition boosting column, so we also add the boost column to the previous stage, if one is required.	2024-07-09 14:29:12 +05:30
Zoltan Haindrich	a9bd0eea2a	Fix queries filtering for the same condition with both an IN and EQUALS to not return empty results (#16597 ) temp fix until CALCITE-6435 gets fixed (released&upgraded to) added a custom rule (FixIncorrectInExpansionTypes) to fix-up types of the affected literals added a testcase which will alert on upgrade	2024-07-09 12:28:21 +05:30
Clint Wylie	09e0eefdc3	modify equality and typed in filter behavior for numeric match values on string columns (#16593 ) * fix equality and typed in filter behavior for numeric match values on string columns changes: * EqualityFilter and TypedInfilter numeric match values against string columns will now cast strings to numeric values instead of converting the numeric values directly to string for pure string equality, which is consistent with the casts which are eaten in the SQL layer, as well as classic druid behavior * added tests to cover numeric equality matching. Double match values in particular would fail to match the string values since `1.0` would become `'1.0'` which does not match `'1'`.	2024-07-08 10:58:05 -07:00
Kashif Faraz	7c6f2b1e20	Minor log cleanup in K8sDruidNodeDiscoveryProvider (#16701 )	2024-07-08 18:32:39 +05:30
Abhishek Radhakrishnan	bf2be938a9	Refactor `SegmentLoadDropHandler` code (#16685 ) Motivation: - Improve code hygeiene - Make `SegmentLoadDropHandler` easily extensible Changes: - Add `SegmentBootstrapper` - Move code for bootstrapping segments already cached on disk and fetched from coordinator to `SegmentBootstrapper`. - No functional change - Use separate executor service in `SegmentBootstrapper` - Bind `SegmentBootstrapper` to `ManageLifecycle` explicitly in `CliBroker`, `CliHistorical` etc.	2024-07-08 09:29:55 +05:30
Alberic Liu	c6c2652c89	unified the code format in NestedDataOperatorConversions (#16695 )	2024-07-08 10:06:24 +08:00
Lars Francke	586c713d12	Updates build documentation to not mention explicit Java version as it was out of sync with the dedicated Java page. (#16674 ) This means there is one less place to keep information in sync.	2024-07-03 20:53:15 +05:30
Virushade	f290cf083a	Update examples/bin/dsql scripts to accept Python 3 (#16677 ) * Update examples/bin/dsql scripts to accept Python 3 Remove redundant urllib import Translating to Python3: Changing xrange to range Translating to Python3: Changing long to int Translating to Python3: Change urllib2 methods, and fix encoding/decoding issues Remove unnecessary import Add option for Python2 Rename files * Update examples/bin/dsql Co-authored-by: Benedict Jin <asdf2014@apache.org> * Resolve PR comments Add comment in files indicating updates need to be made in both places Update examples/bin/dsql Co-authored-by: Benedict Jin <asdf2014@apache.org> * Update error output when using Python 2. Co-authored-by: Abhishek Radhakrishnan <abhishek.rb19@gmail.com> --------- Co-authored-by: Benedict Jin <asdf2014@apache.org> Co-authored-by: Abhishek Radhakrishnan <abhishek.rb19@gmail.com>	2024-07-03 15:52:57 +08:00
Kashif Faraz	6c87b1637b	Revert "Downgrade the version of Apache Curator from 5.5.0 to 5.3.0 to avoid a bug in the new version (#16425 )" (#16688 ) This reverts commit `cb7c2c1e37`.	2024-07-03 11:18:50 +05:30
Abhishek Radhakrishnan	35b970935f	Better error handling when retrieving Avro schemas from registry (#16684 ) * Handle RestClientException separately, instead of returning a generic error. - Add tests - Clean up the tests; remove the legacy expected exception pattern - Better test assertions * Rename tests * checkstyle fixes	2024-07-02 16:48:34 -07:00
317brian	d65e015c94	docs: nit for link format (#16687 )	2024-07-02 16:45:09 -07:00
Victoria Lim	adde024e11	docs: Subtitle updates in migration guide overview (#16683 )	2024-07-02 12:56:05 -07:00
zachjsh	5e05858ff7	Catalog granularity accepts query format (#16680 ) Previously, the segment granularity for tables in the catalog had to be defined in period format, ie `'PT1H'` , `'P1D'`, etc. This disallows a user from defining segment granularity of `'ALL'` for a table in the catalog, which may be a valid use case. This change makes it so that a user may define the segment granularity of a table in the catalog, as any string that results in a valid granularity using either the `Granularity.fromString(str)` method, or `new PeriodGranularity(new Period(value), null, null)`, and that granularity maps to a standard supported granularity, where `GranularityType.isStandard(granularity)` returns true. As a result a user may who wants to assign a catalog table's segment granularity to be hourly, may assign the segment granularity property of the table to be either `PT1H`, or `HOUR`. These are the same formats accepted at query time.	2024-07-02 12:14:28 -04:00
Jill Osborne	bd49ecfd29	Addition to subquery limit migration guide (#16671 ) Co-authored-by: Laksh Singla <lakshsingla@gmail.com> Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2024-07-01 14:22:47 -07:00
Akshat Jain	34c80ee3de	Add MSQ engine support for window function drill tests (#16665 ) * Add MSQ engine support for window function drill tests * Address review comments * Revert formatting changes in TestDataBuilder	2024-06-28 11:14:17 +05:30
Rishabh Singh	c96e783750	Fix schema backfill count metric (#16536 ) * Fix build * Fix backfill metric * Address review comment	2024-06-28 11:07:28 +05:30
Rishabh Singh	b9c7664ac3	Fix empty datasource schema on the Broker when metadata query is disabled (#16645 ) * Fix build * Fix empty datasource schema on the broker * review comment * Remove unused import	2024-06-28 11:06:56 +05:30
Clint Wylie	45c020060c	better javadoc for ColumnIndexSupplier (#16663 ) Updated javadoc for `ColumnIndexSupplier.as` to elaborate on the types of indexes callers might want to ask for from the method, as well as help implementors know what kinds of indexes they should implement to participate in filtering	2024-06-27 17:53:20 -07:00
Clint Wylie	d86f25c74a	fix vector grouping expression deferred evaluation to only consider dictionary encoded strings as fixed width (#16666 )	2024-06-27 16:19:16 -07:00
317brian	4401c9d138	docs: add redirect for kafka lookups (#16668 )	2024-06-27 10:56:51 -07:00
Rishabh Singh	f51c7b346f	Add druid parquet extensions to example quickstarts (#16664 ) This change adds druid-parquet-extensions to all example quickstarts	2024-06-27 14:41:58 +05:30
Hugh Evans	920d9020c0	Docs: Fix default value for globalIngestionHeapLimitBytes (#16654 ) Use the new default value added in #8255	2024-06-27 07:01:56 +05:30
Gian Merlino	dbed1b0f50	Defer more expressions in vectorized groupBy. (#16338 ) * Defer more expressions in vectorized groupBy. This patch adds a way for columns to provide GroupByVectorColumnSelectors, which controls how the groupBy engine operates on them. This mechanism is used by ExpressionVirtualColumn to provide an ExpressionDeferredGroupByVectorColumnSelector that uses the inputs of an expression as the grouping key. The actual expression evaluation is deferred until the grouped ResultRow is created. A new context parameter "deferExpressionDimensions" allows users to control when this deferred selector is used. The default is "fixedWidthNonNumeric", which is a behavioral change from the prior behavior. Users can get the prior behavior by setting this to "singleString". * Fix style. * Add deferExpressionDimensions to SqlExpressionBenchmark. * Fix style. * Fix inspections. * Add more testing. * Use valueOrDefault. * Compute exprKeyBytes a bit lighter-weight.	2024-06-26 17:28:36 -07:00
Clint Wylie	d4f2636325	fix greatest/least function non-vectorized processing to ignore null argument types (#16649 )	2024-06-26 12:59:42 -07:00
Andreas Maechler	ab76d851ad	Update docs contribution with correct script (#16581 ) * Spacing * Fix ordering * npm run start	2024-06-26 10:30:52 -07:00
Abhishek Radhakrishnan	82117e8101	Add MSQ query context `maxNumSegments` (#16637 ) * Add MSQ query context maxNumSegments. - Default is MAX_INT (unbounded). - When set and if a time chunk contains more number of segments than set in the query context, the MSQ task will fail with TooManySegments fault. * Fixup hashCode(). * Rename and checkpoint. * Add some insert and replace happy and sad path tests. * Update error msg. * Commentary * Adjust the default to be null (meaning no max bound on number of segments). Also fix formatter. * Fix CodeQL warnings and minor cleanup. * Assert on maxNumSegments tuning config. * Minor test cleanup. * Use null default for the MultiStageQueryContext as well * Review feedback * Review feedback * Move logic to common function getPartitionsByBucket shared by INSERT and REPLACE. * Rename to validateNumSegmentsPerBucketOrThrow() for consistency. * Add segmentGranularity to error message.	2024-06-26 09:29:51 -07:00
Rahul Bansal	b772277d3b	Update intellij-setup.md (#16655 ) updating typing mistakes	2024-06-26 17:38:37 +05:30
Laksh Singla	71b3b5ab5d	Add query context parameter to remove null bytes when writing frames (#16579 ) MSQ cannot process null bytes in string fields, and the current workaround is to remove them using the REPLACE function. 'removeNullBytes' context parameter has been added which sanitizes the input string fields by removing these null bytes.	2024-06-26 15:00:30 +05:30
Kashif Faraz	d9bd02256a	Refactor: Rename UsedSegmentChecker and cleanup task actions (#16644 ) Changes: - Rename `UsedSegmentChecker` to `PublishedSegmentsRetriever` - Remove deprecated single `Interval` argument from `RetrieveUsedSegmentsAction` as it is now unused and has been deprecated since #1988 - Return `Set` of segments instead of a `Collection` from `IndexerMetadataStorageCoordinator.retrieveUsedSegments()`	2024-06-26 10:48:59 +05:30
Tom	52c9929019	Column name in parse exceptions (#16529 ) * first pass * more changes * fix tests and formatting * fix kinesis failing tests * fix kafka tests * add dimension name to float parse errors * double and convertToType handling of dimensionName can report parse errors with dimension name * fix checkstyle issue * fix tests * more cases to have better parse exception messages * fix test * fix tests * partially address comments * annotate method parameter with nullable * address comments * fix tests * let float, double, long dimensionIndexer pass dimensionName down to dimensionHandlerUtils * fix compilation error and clean up formatting * clean up whitespace * address feedback. undo change, pass down report parse exception for convertToType * fix test	2024-06-25 13:42:52 -07:00
Abhishek Radhakrishnan	e01f155209	Add missing `delta-storage` dependency and class loader workaround to Delta table ingestion (#16648 ) * Workaround to ingesting from Delta table in 3.2.0. With the upgrade to Kernel 3.2.0, the Druid Delta connector extension isn't able to ingest from Delta tables successfully. Please see https://github.com/delta-io/delta/issues/3299 The underlying problem seems to be coming from https://github.com/delta-io/delta/blob/master/kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/logstore/LogStoreProvider.java#L99 This patch is a workaround to setting the thread class loader explictly. The Kernel community may consider a fix in the next release as it's affected another connector as well. * Address review comment: clear the CL after the Thread CL is set.	2024-06-25 09:16:13 -07:00
Edgar Melendrez	b43f4063c5	Docs: update link and title of quickstart (#16638 ) * update link and title * Discard changes to website/package.json * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> --------- Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2024-06-25 09:07:00 -07:00
Abhishek Radhakrishnan	2979f73e89	Fix Intellij inspection (#16651 )	2024-06-25 04:32:43 -07:00
Kashif Faraz	f1043d20bc	Support csv input format in Kafka ingestion with header (#16630 ) * Support ListBasedInputRow in Kafka ingestion with header * Fix up buildBlendedEventMap * Add new test for KafkaInputFormat with csv value and headers * Do not use forbidden APIs * Move utility method to TestUtils	2024-06-25 11:50:01 +05:30
Clint Wylie	37a50e6803	Remove index_realtime and index_realtime_appenderator tasks (#16602 ) index_realtime tasks were removed from the documentation in #13107. Even at that time, they weren't really documented per se— just mentioned. They existed solely to support Tranquility, which is an obsolete ingestion method that predates migration of Druid to ASF and is no longer being maintained. Tranquility docs were also de-linked from the sidebars and the other doc pages in #11134. Only a stub remains, so people with links to the page can see that it's no longer recommended. index_realtime_appenderator tasks existed in the code base, but were never documented, nor as far as I am aware were they used for any purpose. This patch removes both task types completely, as well as removes all supporting code that was otherwise unused. It also updates the stub doc for Tranquility to be firmer that it is not compatible. (Previously, the stub doc said it wasn't recommended, and pointed out that it is built against an ancient 0.9.2 version of Druid.) ITUnionQueryTest has been migrated to the new integration tests framework and updated to use Kafka ingestion. Co-authored-by: Gian Merlino <gianmerlino@gmail.com>	2024-06-24 20:13:33 -07:00
317brian	2131917f16	docs: added front-coded dictionaries to upgrade notes (#16647 ) * docs: add front-coded dictionareis to upgrade notes * add it to release notes template	2024-06-24 10:52:26 -07:00
Abhishek Radhakrishnan	7463589b07	Support for bootstrap segments (#16609 ) * Initial support for bootstrap segments. - Adds a new API in the coordinator. - All processes that have storage locations configured (including tasks) talk to the coordinator if they can, and fetch bootstrap segments from it. - Then load the segments onto the segment cache as part of startup. - This addresses the segment bootstrapping logic required by processes before they can start serving queries or ingesting. This patch also lays the foundation to speed up upgrades. * Fail open by default if there are any errors talking to the coordinator. * Add test for failure scenario and cleanup logs. * Cleanup and add debug log * Assert the events so we know the list exactly. * Revert RunRules test. The rules aren't evaluated if there are no clusters. * Revert RunRulesTest too. * Remove debug info. * Make the API POST and update log. * Fix up UTs. * Throw 503 from MetadataResource; clean up exception handling and DruidException. * Remove unused logger, add verification of metrics and docs. * Update error message * Update server/src/main/java/org/apache/druid/server/coordination/SegmentLoadDropHandler.java Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> * Apply suggestions from code review Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> * Adjust test metric expectations with the rename. * Add BootstrapSegmentResponse container in the response for future extensibility. * Rename to BootstrapSegmentsInfo for internal consistency. * Remove unused log. * Use a member variable for broadcast segments instead of segmentAssigner. * Minor cleanup * Add test for loadable bootstrap segments and clarify comment. * Review suggestions. --------- Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com>	2024-06-24 09:27:17 -07:00
Misha	354a3bea0b	The default `WHERE' filter for automatically generated SQL queries is returned (#16608 ) * Returned the default `WHERE` filter for auto-generated SQL queries * Checkstyle fix --------- Co-authored-by: sviatahorau <mikhail.sviatahorau@deep.bi>	2024-06-24 08:52:35 -07:00
Sree Charan Manamala	990fd5f5fb	Make use group iterator for all window frames & support for same bound kinds (#16603 ) Fixes apache/druid#15739	2024-06-24 15:52:41 +02:00
Kashif Faraz	0fe6a2af68	Fix replica task failures with metadata inconsistency while running concurrent append replace (#16614 ) Changes: - Add new task action `RetrieveSegmentsByIdAction` - Use new task action to retrieve segments irrespective of their visibility - During rolling upgrades, this task action would fail as Overlord would be on old version - If new action fails, fall back to just fetching used segments as before	2024-06-24 09:56:04 +05:30
Adarsh Sanjeev	1a883ba1f7	Fix complex columns with export (#16572 ) This PR fixes a few bugs with MSQ export. The main change is calling SqlResults#coerce before writing the column. This allows sketches and json to be correctly deserialized. The format of the exported complex columns are similar to those produced by Async MSQ queries with CSV format. Notes: Fix printing of complex columns during export. Sketches and JSON are now correctly formatted during export. Fix an NPE if the writer has not been initialized. Empty export queries will create an empty file at the location. Fix a bug with counters for MSQ export, where rows were reported for only the first partition.	2024-06-24 09:03:30 +05:30
Akshat Jain	641f739a47	Fix flaky test in RetryableS3OutputStreamTest (#16639 ) As part of #16481, we have started uploading the chunks in parallel. That means that it's not necessary for the part that finished uploading last to be less than or equal to the chunkSize (as the final part could've been uploaded earlier). This made a test in RetryableS3OutputStreamTest flaky where we were asserting that the final part should be smaller than chunk size. This commit fixes the test, and also adds another test where the file size is such that all chunk sizes would be of equal size.	2024-06-24 08:13:47 +05:30
Laksh Singla	00c96432af	Materialize scan results correctly when columns are not present in the segments (#16619 ) Fixes a bug causing maxSubqueryBytes not to work when segments have missing columns.	2024-06-23 23:15:45 +05:30
Rishabh Singh	a63c12bf34	Upload tasklogs along with service logs on Standard IT failure (#16631 ) * Fix build * Push tasklogs alongwith service logs * temp changes to run standard its without unit test results * test * minor change * test * test * Update datasource name for ITSystemTableBatchIndexTaskTest * Publish task logs * Revert other changes * update standard-it yaml	2024-06-22 11:45:54 +05:30
Vadim Ogievetsky	51c73b5a4e	Web console: show formatted JSON value (#16632 ) * show formatted json value * update snapshot * window functions * count star can also have a window * better edit query context	2024-06-21 18:33:15 -07:00
Rishabh Singh	4eced9b3c9	Fix CentralizedDatasourceSchema group IT failure (#16636 ) * Fix build * Update datasource name in ITSystemTableBatchIndexTaskTest	2024-06-21 15:40:12 -07:00

1 2 3 4 5 ...

14165 Commits All Branches Search

14165 Commits

All Branches