druid

Commit Graph

Author	SHA1	Message	Date
Clint Wylie	e8fcf2cac8	minor doc adjustments (#15531 )	2023-12-11 18:22:44 -08:00
Xavier Léauté	6f78049760	remove references to non-existant website maven module (#15540 ) The website pom was removed as part of https://github.com/apache/druid/pull/14411 so we no longer need to reference it as a module and the profile can be removed. Dependabot is currently failing trying to look for this module, so removing it should also fix that.	2023-12-11 16:58:35 -08:00
zachjsh	ab7d9bc6ec	Add api for Retrieving unused segments (#15415 ) ### Description This pr adds an api for retrieving unused segments for a particular datasource. The api supports pagination by the addition of `limit` and `lastSegmentId` parameters. The resulting unused segments are returned with optional `sortOrder`, `ASC` or `DESC` with respect to the matching segments `id`, `start time`, and `end time`, or not returned in any guarenteed order if `sortOrder` is not specified `GET /druid/coordinator/v1/datasources/{dataSourceName}/unusedSegments?interval={interval}&limit={limit}&lastSegmentId={lastSegmentId}&sortOrder={sortOrder}` Returns a list of unused segments for a datasource in the cluster contained within an optionally specified interval. Optional parameters for limit and lastSegmentId can be given as well, to limit results and enable paginated results. The results may be sorted in either ASC, or DESC order depending on specifying the sortOrder parameter. `dataSourceName`: The name of the datasource `interval`: the specific interval to search for unused segments for. `limit`: the maximum number of unused segments to return information about. This property helps to support pagination `lastSegmentId`: the last segment id from which to search for results. All segments returned are > this segment lexigraphically if sortOrder is null or ASC, or < this segment lexigraphically if sortOrder is DESC. `sortOrder`: Specifies the order with which to return the matching segments by start time, end time. A null value indicates that order does not matter. This PR has: - [x] been self-reviewed. - [ ] using the [concurrency checklist](https://github.com/apache/druid/blob/master/dev/code-review/concurrency.md) (Remove this item if the PR doesn't have any relation to concurrency.) - [x] added documentation for new or modified features or behaviors. - [ ] a release note entry in the PR description. - [x] added Javadocs for most classes and all non-trivial methods. Linked related entities via Javadoc links. - [ ] added or updated version, license, or notice information in [licenses.yaml](https://github.com/apache/druid/blob/master/dev/license.md) - [x] added comments explaining the "why" and the intent of the code wherever would not be obvious for an unfamiliar reader. - [x] added unit tests or modified existing tests to cover new code paths, ensuring the threshold for [code coverage](https://github.com/apache/druid/blob/master/dev/code-review/code-coverage.md) is met. - [ ] added integration tests. - [x] been tested in a test Druid cluster.	2023-12-11 16:32:18 -05:00
George Shiqi Wu	4152f1d147	Fix empty logs and status messages for mmless ingestion (#15527 ) * Fix empty logs and status messages for mmless ingestion * Add tests	2023-12-11 13:20:45 -05:00
Katya Macedo	fc222377ae	[Docs] Document decode_base64_complex and decode_base64_utf8 functions (#15444 )	2023-12-11 09:12:06 -08:00
Abhishek Radhakrishnan	96be82a3e6	Clean up duty for non-overlapping eternity tombstones (#15281 ) * Add initial draft of MarkDanglingTombstonesAsUnused duty. * Use overshadowed segments instead of all used segments. * Add unit test for MarkDanglingSegmentsAsUnused duty. * Add mock call * Simplify code. * Docs * shorter lines formatting * metric doc * More tests, refactor and fix up some logic. * update javadocs; other review comments. * Make numCorePartitions as 0 in the TombstoneShardSpec. * fix up test * Add tombstone core partition tests * Update docs/design/coordinator.md Co-authored-by: 317brian <53799971+317brian@users.noreply.github.com> * review comment * Minor cleanup * Only consider tombstones with 0 core partitions * Need to register the test shard type to make jackson happy * test comments * checkstyle * fixup misc typos in comments * Update logic to use overshadowed segments * minor cleanup * Rename duty to eternity tombstone instead of dangling. Add test for full eternity tombstone. * Address review feedback. --------- Co-authored-by: 317brian <53799971+317brian@users.noreply.github.com>	2023-12-11 08:57:15 -08:00
Katya Macedo	099a9825d1	[Docs] Add a release notes template (#15333 ) * Add release notes template * Update spellcheck	2023-12-11 11:35:16 +05:30
Clint Wylie	42f2496b7d	fix bug with nested empty array fields (#15532 )	2023-12-09 12:20:21 -08:00
Rishabh Singh	54df235026	Lazily build Filter in FilteredAggregatorFactory to avoid parsing exceptions in Router (#15526 ) Query with lookups in FilteredAggregator fails with this exception in router, Cannot construct instance of `org.apache.druid.query.aggregation.FilteredAggregatorFactory`, problem: Lookup [campaigns_lookup[campaignId][is_sold][autodsp]] not found at [Source: (org.eclipse.jetty.server.HttpInputOverHTTP); line: 1, column: 913] (through reference chain: org.apache.druid.query.groupby.GroupByQuery["aggregations"]->java.util.ArrayList[1]) T he problem is that constructor of FilteredAggregatorFactory is actually validating if the lookup exists in this statement dimFilter.toFilter(). This is failing on the router, which is to be expected, because, the router isn’t assigned any lookups. The fix is to move to a lazy initialisation of the filter object in the constructor.	2023-12-09 12:18:37 +05:30
Clint Wylie	e7c8f2e208	lift restriction of array_to_mv to only support direct column access (#15528 )	2023-12-08 16:27:17 -08:00
Victoria Lim	e68979e03b	Docs: update SQL API reference (#15515 ) Co-authored-by: 317brian <53799971+317brian@users.noreply.github.com>	2023-12-08 11:53:19 -08:00
Katya Macedo	355c800108	Revamp design page (#15486 ) Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2023-12-08 11:40:24 -08:00
Soumyava	ca4ecdf7d0	Fixing NPE with virtual expression with unnest (#15513 ) * Fixing NPE with virtual expression with unnest * Fixing a comment	2023-12-08 10:51:56 -08:00
Clint Wylie	e64b92eb35	add JSON_QUERY_ARRAY function to pluck ARRAY<COMPLEX<json>> out of COMPLEX<json> (#15521 )	2023-12-08 05:28:46 -08:00
Adarsh Sanjeev	2e45eadc08	Add better error messages for using OVERWRITE with INSERT statments (#15517 ) * Add better error messages for using OVERWRITE with INSERT statments	2023-12-08 15:33:46 +05:30
Zoltan Haindrich	c353ccfdef	Windowed min aggregates null-s as 0 (#15371 )	2023-12-08 01:41:16 -08:00
Clint Wylie	1eafe983ec	fix array presenting columns to not match single element arrays to scalars for equality (#15503 ) * fix array presenting columns to not match single element arrays to scalars for equality * update docs to clarify usage model of mixed type columns	2023-12-08 01:22:07 -08:00
sb89594	5fda8613ad	Feature: Add IPv6 Match Function (#15212 )	2023-12-07 23:09:06 -08:00
Adarsh Sanjeev	254a8eb7e0	Add null checks for HllSketchHolder (#15502 ) Fixes a potential NPE which could occur while folding the HllSketchAggregator. If the sketch is null, druid could return a null HllSketchHolder object. Adding a null check here could help here Resolves a null pointer exception in HllSketchAggregatorFactory	2023-12-08 11:43:04 +05:30
AlbericByte	935aa187a0	add Assert function to verify in the DataGeneratorTest (#15504 ) * add Assert function to verify in the DataGeneratorTest * remove unused log in DataGeneratorTest * add comment for DataGeneratorTest	2023-12-08 09:12:17 +08:00
Charles Smith	db3a633250	update timeseries to reflect NULL filling (#15512 ) Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com>	2023-12-07 14:41:27 -08:00
Clint Wylie	c241c6980c	store auto columns with only empty or null containing arrays as ARRAY<LONG> instead of COMPLEX<json> (#15505 )	2023-12-07 03:31:43 -08:00
Vishesh Garg	801967b75f	Add test logs zipping and archival steps for failures in Static Checks Github Actions (#15506 ) Add test logs zipping and archival steps for failures in Static Checks Github Actions	2023-12-07 15:34:23 +05:30
Abhishek Radhakrishnan	b541000d43	Bump up max heap memory for unit tests from 1.5 GB to 2 GB. (#15507 )	2023-12-07 15:34:04 +05:30
Clint Wylie	82ac48786b	document arrayContainsElement filter (#15455 )	2023-12-07 00:14:00 -08:00
Clint Wylie	557f3f6f57	add array column type support to EXTEND operator (#15458 )	2023-12-06 23:21:35 -08:00
Benjamin Hopp	fea53c7084	Re-arranging sections for append and replace docs. (#15497 )	2023-12-06 13:13:05 -08:00
Rishabh Singh	6a64f72c67	Lookup on incomplete partition set in SegmentMetadataQuerySegmentWalker (#15496 ) Description With CentralizedDatasourceSchema (#14989) feature enabled, metadata for appended segments was not being refreshed. This caused numRows to be 0 for the new segments and would probably cause the datasource schema to not include columns from the new segments. Analysis The problem turned out in the new QuerySegmentWalker implementation in the Coordinator. It first finds the segment to be queried in the Coordinator timeline. Then it creates a new timeline of the segments present in the timeline. The problem was that it is looking up complete partition set in the new timeline. Since the appended segments by themselves do not make a complete partition set, no SegmentMetadataQuery were executed.	2023-12-06 15:25:28 +05:30
Gian Merlino	6f51155ccb	Fix NullFilter getDimensionRangeSet. (#15500 ) It wasn't checking the column name, so it would return a domain regardless of the input column. This means that null filters on data sources with range partitioning would lead to excessive pruning of segments, and therefore missing results.	2023-12-06 15:09:59 +05:30
Abhishek Radhakrishnan	f4949afdd7	clarify and fixup typos related to unused segments in docs and javadocs. (#15498 )	2023-12-05 22:30:32 -08:00
Xavier Léauté	ae6893edc3	unpin guava related dependabot dependencies (#15494 ) Several dependabot ignore directives are no longer relevant. Unpin them to ensure we get again get timely updates via dependabot. * support for Hadoop 2 was dropped as part of #14763 * Guava was upgraded to 31 as part of #14767 * Calcite was upgraded to 1.35 as part of #14510	2023-12-05 16:04:39 -08:00
Vadim Ogievetsky	0b41b05aa0	Web console: Update and prune dependancies (#15487 ) * update the basics * remove babel	2023-12-05 14:25:07 -08:00
Vadim Ogievetsky	aa696b0310	Web console: Log out any request errors in e2e tests for better CI debugging (#15483 )	2023-12-05 14:23:47 -08:00
Pranav	82e3c61514	Update lookup model in console (#15472 ) * Update lookup model in console * ran prettify * move Defaults to info * setting defaultValue and removing placeholder	2023-12-05 13:22:22 -08:00
Jan Werner	ff0e838d30	add gson to dependencyManagement (#15488 ) This change completes the change introduced in #15461 and unifies the version of gson dependency used between all the modules. gson is used by kubernetes-extension, avro-extensions, ranger-security, and as a test dependency in several core modules. --------- Co-authored-by: Xavier Léauté <xl+github@xvrl.net>	2023-12-05 11:50:32 -08:00
Jill Osborne	0e14a2c77f	Update retention rules doc (#15439 )	2023-12-05 09:53:17 -08:00
Jan Werner	f4856bc1c1	ranger-security: exclude jackson-jaxrs from + fix outdated documentation (#15481 ) * Excluding jackson-jaxrs dependency from ranger-plugin-common to address CVE regression introduced by ranger-upgrade: CVE-2019-10202, CVE-2019-10172 * remove the reference to outdated ranger 2.0 from the docs --------- Co-authored-by: Xavier Léauté <xl+github@xvrl.net>	2023-12-05 08:24:37 -08:00
Rishabh Singh	77b929f494	Fix CentralizedDatasourceSchema IT (#15493 )	2023-12-05 20:05:13 +05:30
Vishesh Garg	326b7b731d	Upgrade zookeeper from 3.5.10 to 3.8.3 (#15477 ) Upgrade zookeeper from 3.5.10 to 3.8.3	2023-12-05 18:57:56 +05:30
Rishabh Singh	d968bb3f43	Rename config for enabling CentralizedDatasourceSchema feature (#15476 ) * Rename property to druid.centralizedDatasourceSchema.enabled * Update config name in docker-compose	2023-12-05 16:57:25 +05:30
Jan Werner	a469c53c0c	cleanup already resolved CVEs (#15447 ) Remove the crud from the dependency-check suppression file	2023-12-05 10:30:35 +05:30
Jan Werner	b66d995e6f	remove licenses of removed libraries, update the license checker (#15446 ) - Licenses file contains several licenses for outdated libraries. In this PR we remove licenses for no longer used components. This change is purely cosmetic / cleans up the license database. The candidates were designated by reviewing the output of the license check script and comparing it against the depdency tree. - Minor fix to license check tool to fail more gracefully when the license of used dependency is not listed as known, as well as fix not to fail on multi licensed components when at least one of the licenses is accepted. --------- Co-authored-by: Xavier Léauté <xl+github@xvrl.net>	2023-12-04 13:20:40 -08:00
Jan Werner	8cc256b079	update guava to 32.0.1-jre to address CVEs (#15482 ) Update guava to 32.0.1-jre to address two CVEs: CVE-2020-8908, CVE-2023-2976 This change requires a minor test change to remove assumptions about ordering. --------- Co-authored-by: Xavier Léauté <xl+github@xvrl.net>	2023-12-04 13:18:42 -08:00
Jan Werner	3d3d23c53f	run npm audit fix to update JS packages (#15466 )	2023-12-04 13:17:24 -08:00
Adarsh Sanjeev	ddd2299272	Add null check for VarianceAggregatorCollector	2023-12-04 22:26:44 +05:30
Jan Werner	ddeb55fac1	update few minor dependencies to resolve CVEs (#15464 ) Update multiple dependencies to clear CVEs Update dropwizard-metrics to 4.2.22 to address GHSA-mm8h-8587-p46h in com.rabbitmq:amqp-client Update ant to 1.10.14 to resolve GHSA-f62v-xpxf-3v68 GHSA-4p6w-m9wc-c9c9 GHSA-q5r4-cfpx-h6fh GHSA-5v34-g2px-j4fw Update comomons-compress to resolve GHSA-cgwf-w82q-5jrr Update jose4j to 0.9.3 to resolve GHSA-7g24-qg88-p43q GHSA-jgvc-jfgh-rjvv Update kotlin-stdlib to 1.6.0 to resolve GHSA-cqj8-47ch-rvvq and CVE-2022-24329	2023-12-04 08:49:51 +05:30
Zoltan Haindrich	a1aa4340d0	Changing the queryFrameWork in Calcite*Tests may have sideeffects (#15428 ) changes how its configured a bit to use an annotation instead of methods	2023-12-04 00:38:01 +05:30
Jan Werner	b854058491	remove unnecessary elasticsearch dependencies to fix CVE regressions (#15443 ) Recent upgrade of ranger introduced CVE regressions due to outdated elasticsearch components. Druid-ranger-plugin does not elasticsearch components , and they have been explicitly removed. Update woodstox-core to 6.4.0 to address GHSA-3f7h-mf4q-vrm4	2023-12-03 20:56:40 +05:30
AmatyaAvadhanula	4a594bb9f6	Use task actions to fetch used segments in MSQ (#15284 ) * Use task actions to fetch used segments in MSQ * Fix tests * Fixing tests. * Revert "Fix tests" This reverts commit `95ab6494` * Removing conditional check in tests. * Pulling in latest changes. --------- Co-authored-by: cryptoe <karankumar1100@gmail.com>	2023-12-01 15:29:33 +05:30
Pranav	9f3b26676d	Log full stack when exception message is null (#15467 )	2023-11-30 16:47:37 -08:00

... 11 12 13 14 15 ...

14089 Commits All Branches Search

14089 Commits

All Branches