druid

Commit Graph

Author	SHA1	Message	Date
Kashif Faraz	5f203725dd	Clean up SqlSegmentsMetadataManager and corresponding tests (#16044 ) Changes: Improve `SqlSegmentsMetadataManager` - Break the loop in `populateUsedStatusLastUpdated` before going to sleep if there are no more segments to update - Add comments and clean up logs Refactor `SqlSegmentsMetadataManagerTest` - Merge `SqlSegmentsMetadataManagerEmptyTest` into this test - Add method `testPollEmpty` - Shave a few seconds off of the tests by reducing poll duration - Simplify creation of test segments - Some renames here and there - Remove unused methods - Move `TestDerbyConnector.allowLastUsedFlagToBeNull` to this class Other minor changes - Add javadoc to `NoneShardSpec` - Use lambda in `SqlSegmentMetadataPublisher`	2024-03-08 07:34:51 +05:30
Charles Smith	3caacba8c5	update window functions doc (#15902 ) Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2024-03-07 15:16:52 -08:00
AmatyaAvadhanula	5871b81a78	Fix race in BaseNodeRoleWatcher tests (#16064 ) * Fix race in BaseNodeRoleWatcher tests * Make non static	2024-03-07 13:41:16 -08:00
Zoltan Haindrich	aaa64832fd	Disable DecoupledPlanningCalciteJoinQueryTest until it gets fixed (#16070 ) Recently this test started other tests from executing by triggering a bug somewhere in surefire. This patch disables the testcases in case of non-sql compat mode.	2024-03-07 12:55:48 -08:00
George Shiqi Wu	80cab51d50	Fix bug with cancelling pending tasks when running kubernetes ingestion. (#16036 ) * Fix bug * Add new test	2024-03-07 15:48:14 -05:00
Jill Osborne	67ae0ff450	Update docs for rabbit community extension (#16069 ) * Updated docs for rabbit community extension * Updated after review	2024-03-07 11:29:53 -08:00
Vishesh Garg	bed5d9c3b2	Remove exception on failure response from GCS delete API (#16047 ) * Throw 404 Exception on failure response from GCS delete API * Replace String.format * Apply suggestions from code review Co-authored-by: Abhishek Radhakrishnan <abhishek.rb19@gmail.com> * Remove exception for file not found and fix tests * Add warn log and fix intellij inspection errors * More intellij inspection fixes * * Change to debug log * change runtime exception class for code coverage * Add file paths for batch delete failures * Move failedPaths computation to inside isDebugEnabled flag * Correct handling of StorageException * Address review comments * Remove unused exceptions * Address code coverage and review comments * Minor corrections --------- Co-authored-by: Abhishek Radhakrishnan <abhishek.rb19@gmail.com>	2024-03-07 17:57:17 +05:30
Laksh Singla	5f588fa45c	Fix bug while materializing scan's result to frames (#15987 ) While converting Sequence<ScanResultValue> to Sequence<Frames>, when maxSubqueryBytes is enabled, we batch the results to prevent creating a single frame per ScanResultValue. Batching requires peeking into the actual value, and checking if the row signature of the scan result’s value matches that of the previous value. Since we can do this indefinitely (in the worst case all of them have the same signature), we keep fetching them and accumulating them in a list (on the heap). We don’t really know how much to batch before we actually write the value as frames. The PR modifies the batching logic to not accumulate the results in an intermediary list	2024-03-07 17:11:44 +05:30
Parth Agrawal	bf39c71d2a	Update protocol for MemcachedCache (#16035 )	2024-03-06 22:28:11 -08:00
Adithya Chakilam	564c44ed85	Add stats segmentsRead and segmentsPublished to compaction task reports (#15947 ) Changes: - Add visibility into number of segments read/published by each parallel compaction - Add new fields `segmentsRead`, `segmentsPublished` to `IngestionStatsAndErrorsTaskReportData` - Update `ParallelIndexSupervisorTask` to populate the new stats	2024-03-07 09:37:23 +05:30
Charles Smith	ebf3bdd909	restore information about truncated responses to sql api (#16001 ) Co-authored-by: 317brian <53799971+317brian@users.noreply.github.com> Co-authored-by: Victoria Lim <lim.t.victoria@gmail.com>	2024-03-06 14:03:58 -08:00
Adithya Chakilam	ae022cc0c9	fixup!: #15981 Missing completion reports on index_parallel tasks (#16042 ) * initial commit * comments * typo * comments * comments * remove var * initialize global var early * remove new line * small test fix * same fix another test	2024-03-06 13:58:34 -05:00
AmatyaAvadhanula	c2841425f4	Handle uninitialized cache in Node role watchers (#15726 ) BaseNodeRoleWatcher counts down cacheInitialized after a timeout, but also sets some flag that it was a timed-out initialization. and call nodeViewInitializationTimedOut (new method on listeners) instead of nodeViewInitialized. Then listeners can do what is most appropriate with this information.	2024-03-06 16:00:24 +05:30
Vishesh Garg	cf9bc507f6	Fix compilation failure due to missing constant MISSING_JOIN_CONVERSION (#16050 ) * Reintroduce variable MISSING_JOIN_CONVERSION * Remove redundant constant MISSING_JOIN_CONVERSION2 * Correct fix to address failing tests	2024-03-06 15:34:39 +08:00
Adarsh Sanjeev	ddd9da2e09	Make servedSegments nullable to maintain compatibility (#16034 ) * Make servedSegments nullable to maintain compatibility	2024-03-06 11:39:24 +05:30
Zoltan Haindrich	65c3b4d31a	Support join in decoupled mode (#15957 ) * plan join(s) in decoupled mode * configure DecoupledPlanningCalciteJoinQueryTest the test has 593 cases; however there are quite a few parameterized from the 107 methods annotated with @Test - 42 is not yet working * replace the isRoot hack in DruidQueryGenerator with a logic that instead looks ahead for the next node; and doesn't let the previous node do the Project - this makes it plan more likely than the existing planner	2024-03-05 19:10:13 -06:00
Sergio Ferragut	d38703281c	updated description of rowsPerPage in export operations (#16048 ) * updated description of rowsPerPage in export operations * Update docs/multi-stage-query/reference.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> --------- Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2024-03-05 15:42:12 -08:00
Zoltan Haindrich	27d7c30c38	Only use ExprEval in ConstantExpr if its known that it will be safe (#15694 ) * `Expr#singleThreaded` which creates a singleThreaded version of the actual expression (caching ExprEval is allowed) * `Expr#makeSingleThreaded` to make a whole subtree of expressions 'singleThreaded' - uses `Shuttle` to create the new expression tree * `ConstantExpr#singleThreaded` creates a specialized `ConstantExpr` which does cache the `ExprEval` * some `@Immutable` annotations were added to make it more likely to notice that there might be something off if a similar change will be made around here for some reason	2024-03-05 12:53:09 -08:00
Gian Merlino	e13ed7b878	Improve javadoc for ReadableFieldPointer. (#16043 ) Since #15175, the javadoc for ReadableFieldPointer is somewhat out of date. It says that the pointer only points to the beginning of the field, but this is no longer true. This patch updates the javadoc to be more accurate.	2024-03-05 11:52:49 -08:00
Zoltan Haindrich	bb882727c0	Fix Windowing/scanAndSort query issues on top of Joins. (#15996 ) allow a hashjoin result to be converted to RowsAndColumns added StorageAdapterRowsAndColumns fix incorrect isConcrete() return values during early phase of planning	2024-03-05 15:05:31 +05:30
317brian	00f80417d0	docs: update package and regenerate package-lock (#16038 ) Running npm run build, you should no longer get the message: [ERROR] Invalid docusaurus-theme-mermaid version 2.4.3. All official @docusaurus/* packages should have the exact same version as @docusaurus/core (2.4.1). Maybe you want to check, or regenerate your yarn.lock or package-lock.json file? yarn build didn't have the same issue, and yarn install did not generate a new yarn.lock	2024-03-05 14:56:43 +05:30
Clint Wylie	6e3b33d8c2	fix bug with like filter on missing columns not correctly considering extractionFn (#16037 )	2024-03-04 23:27:50 -08:00
Gian Merlino	566013a5f5	Docs: Fix spelling of 5 GB. (#16040 ) The spellchecker does not consider "5GB" to be spelled correctly.	2024-03-04 22:37:38 -08:00
Zoltan Haindrich	e469b7ed34	Make setting QUERY_CONTEXT_DEFAULT explicit in tests (#16010 )	2024-03-05 10:54:16 +05:30
zachjsh	720f1e834a	Add support for AzureDNSZone enabled storage accounts used for deep storage (#16016 ) * Add support for AzureDNSZone enabled storage accounts used for deep storage Added a new config to AzureAccountConfig `storageAccountEndpointSuffix` which allows the user to specify a storage account endpoint suffix where the underlying storage account is enabled for AzureDNSZone. The previous config `endpointSuffix`, did not allow support for such accounts. The previous config has been deprecated in favor of this new config. Also fixed an issue where `managedIdentityClientId` was not being set properly * * address review comments * * add back azure government link and docs	2024-03-04 16:13:28 -05:00
Gian Merlino	930655ff18	Move retries into DataSegmentPusher implementations. (#15938 ) * Move retries into DataSegmentPusher implementations. The individual implementations know better when they should and should not retry. They can also generate better error messages. The inspiration for this patch was a situation where EntityTooLarge was generated by the S3DataSegmentPusher, and retried uselessly by the retry harness in PartialSegmentMergeTask. * Fix missing var. * Adjust imports. * Tests, comments, style. * Remove unused import.	2024-03-04 10:36:21 -08:00
Katya Macedo	ced8be3044	docs: Add upgrade notes for Druid 29.0.0 (#16022 )	2024-03-04 08:58:52 -08:00
Gian Merlino	376a41f1e9	Rows.objectToNumber: Accept decimals with output type LONG. (#15999 ) * Rows.objectToNumber: Accept decimals with output type LONG. PR #15615 added an optimization to avoid parsing numbers twice in cases where we know that they should definitely be longs or definitely be doubles. Rather than try parsing as long first, and then try parsing as double, it would use only the parsing routine specific to the requested outputType. This caused a bug: previously, we would accept decimals like "1.0" or "1.23" as longs, by truncating them to "1". After that patch, we would treat such decimals as nulls when the outputType is set to LONG. This patch retains the short-circuit for doubles: if outputType is DOUBLE, we only parse the string as a double. But for outputType LONG, this patch restores the old behavior: try to parse as long first, then double.	2024-03-04 22:00:27 +05:30
Sensor	4e9b758661	Support CPU resource configurable for Kubernates job under MoK Mode (#16008 ) * support CPU resource configurable for Kubernates job * update property doc * fix test name * refine doc format	2024-03-04 10:12:09 -05:00
Adithya Chakilam	ec52f686c0	Fix compaction tasks reports getting overwritten (#15981 ) * Fix compaction tasks reports geting overwrittened * only skip for compactiont task * address comments * fix boolean * move boolean flag to task rather than spec * rename variable * add docs, fix missing case * Update docs/ingestion/tasks.md * rename var * add task report decode check in IT * change assert	2024-03-04 10:10:17 -05:00
Adarsh Sanjeev	93eeb05eaf	Revert explain attributes change to old behaviour. (#16004 ) * Revert explain attributes change * Fix tests * Fix tests * Rename function	2024-03-04 15:56:02 +05:30
Sree Charan Manamala	820febf38c	Improved Connection Count server select strategy (#15975 ) Updated the Direct Druid Client so as to make Connection Count Server Selector Strategy work more efficiently. If creating connection to a node is slow, then that slowness wouldn't be accounted for if we count the open connections after sending the request. So we increment the counter and then send the request.	2024-03-04 15:02:32 +05:30
317brian	b3015bd7ce	docs: mention acid-compliance for meta store (#16014 ) * docs: add mermaid diagram support * fix crash when parsing data in data loader that can not be parsed (#15983) * update jetty to address CVE (#16000) * Concurrent replace should work with supervisors using concurrent locks (#15995) * Concurrent replace should work with supervisors using concurrent locks * Ignore supervisors with useConcurrentLocks set to false * Apply feedback * Add pre-check for heavy debug logs (#15706) Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> Co-authored-by: Benedict Jin <asdf2014@apache.org> * Remove helm paths from CodeQL config (#16006) * docs: mention acid-compliance for metadb --------- Co-authored-by: Vadim Ogievetsky <vadim@ogievetsky.com> Co-authored-by: Jan Werner <105367074+janjwerner-confluent@users.noreply.github.com> Co-authored-by: AmatyaAvadhanula <amatya.avadhanula@imply.io> Co-authored-by: Sensor <fectrain@outlook.com> Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> Co-authored-by: Benedict Jin <asdf2014@apache.org>	2024-03-04 11:00:38 +08:00
Gian Merlino	8d3ed31015	MSQ: Nicer error when sortMerge join falls back to broadcast. (#16002 ) * MSQ: Nicer error when sortMerge join falls back to broadcast. In certain cases, joins run as broadcast even when the user hinted that they wanted sortMerge. This happens when the sortMerge algorithm is unable to process the join, because it isn't a direct comparison between two fields on the LHS and RHS. When this happens, the error message from BroadcastTablesTooLargeFault is quite confusing, since it mentions that you should try sortMerge to fix it. But the user may have already configured sortMerge. This patch fixes it by having two error messages, based on whether broadcast join was used as a primary selection or as a fallback selection. * Style. * Better message.	2024-03-01 13:16:39 -08:00
George Shiqi Wu	ef48aceff8	Fix segment/unavailable/count (#16020 )	2024-03-01 15:38:27 -05:00
Zoltan Haindrich	bf0995f846	Introduce dynamic table append (#15897 )	2024-03-01 04:31:57 -05:00
Vadim Ogievetsky	acb5124679	make double detection better (#15998 )	2024-02-29 15:45:58 -08:00
Vadim Ogievetsky	c5b032799c	Web console: add table and column search (#15990 ) * Make a search * fix snapshot * added message when not found	2024-02-29 15:45:50 -08:00
Clint Wylie	101176590c	adaptive filter partitioning (#15838 ) * cooler cursor filter processing allowing much smart utilization of indexes by feeding selectivity forward, with implementations for range and predicate based filters * added new method Filter.makeFilterBundle which cursors use to get indexes and matchers for building offsets * AND filter partitioning is now pushed all the way down, even to nested AND filters * vector engine now uses same indexed base value matcher strategy for OR filters which partially support indexes	2024-02-29 15:38:12 -08:00
Jan Werner	baaa4a6808	update common-compress to address CVE-2024-25710 CVE-2024-26308 (#16009 ) * Update common-compress to 1.26.0 to address CVEs CVE-2024-25710 CVE-2024-26308 * Add commons-codec as a runtime dependency required by common-compress 1.26.0 --------- Co-authored-by: Xavier Léauté <xl+github@xvrl.net>	2024-02-29 14:05:31 -08:00
Sensor	3acfc95453	Remove helm paths from CodeQL config (#16006 )	2024-02-29 20:02:27 +05:30
Sensor	e0bce0ef90	Add pre-check for heavy debug logs (#15706 ) Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> Co-authored-by: Benedict Jin <asdf2014@apache.org>	2024-02-29 12:58:14 +05:30
AmatyaAvadhanula	7c42e87db9	Concurrent replace should work with supervisors using concurrent locks (#15995 ) * Concurrent replace should work with supervisors using concurrent locks * Ignore supervisors with useConcurrentLocks set to false * Apply feedback	2024-02-29 12:06:47 +05:30
Jan Werner	d6f59d1999	update jetty to address CVE (#16000 )	2024-02-29 09:27:31 +08:00
Vadim Ogievetsky	6e222d47c8	fix crash when parsing data in data loader that can not be parsed (#15983 )	2024-02-28 14:37:24 -08:00
Kashif Faraz	f757231420	Use cache for password hash while validating LDAP password (#15993 )	2024-02-28 18:33:33 +05:30
Adarsh Sanjeev	d2c2036ea2	Optimize MSQ realtime queries (#15399 ) Currently, while reading results from realtime tasks, requests are sent on a segment level. This is slightly wasteful, as when contacting a data servers, it is possible to transfer results for all segments which it is hosting, instead of only one segment at a time. One change this PR makes is to group the segments on the basis of servers. This reduces the number of queries to data servers made. Since we don't have access to the number of rows for realtime segments, the grouping is done with a fixed estimated number of rows for each realtime segment.	2024-02-28 11:32:14 +05:30
317brian	3df161f73c	docs: update security doc for hashing (#15970 ) * docs: add mermaid diagram support * docs: update druid-basic-security doc to mention caching * Update docs/development/extensions-core/druid-basic-security.md Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com> --------- Co-authored-by: Kashif Faraz <kashif.faraz@gmail.com>	2024-02-28 09:48:37 +08:00
benkrug	0c601bf430	Update basic-cluster-tuning.md (#14964 ) * Update basic-cluster-tuning.md The sentence "When free system memory is greater than or equal to druid.segmentCache.locations, the more segment data the Historical can be held in the memory-mapped segment cache" didn't read well. Updated to clarify it. * Update docs/operations/basic-cluster-tuning.md * Update docs/operations/basic-cluster-tuning.md --------- Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2024-02-28 09:48:20 +08:00
AlbericByte	f07d402f48	pin Testng dependencies to 7.3.0 (#15924 )	2024-02-28 09:48:13 +08:00

... 3 4 5 6 7 ...

13976 Commits All Branches Search

13976 Commits

All Branches