druid

Commit Graph

Author	SHA1	Message	Date
Zoltan Haindrich	5d16d0edf0	Count distinct returned incorrect results without useApproximateCountDistinct (#14748 ) * fix grouping engine handling of summaries when result set is empty	2023-09-12 13:57:54 -07:00
Abhishek Radhakrishnan	0f38a37b9d	Tweak GHA runner label. (#14963 ) - processing/** can be ingestion, querying or neither. Removing it for now. - Also, add msq extension for the querying label.	2023-09-11 20:09:26 -07:00
Clint Wylie	5cecf6ce8f	fix issue with segment metadata cache and complex types when doing out of order upgrades from 0.22 (#14948 )	2023-09-12 10:54:35 +08:00
Suneet Saldanha	757603a773	Set task location as k8sPodName for mm-less ingestion (#14959 ) * Set task location as k8sPodName for mm-less ingestion * tests	2023-09-11 19:44:26 -07:00
George Shiqi Wu	f773d83914	Mixed task runner for migration to mm-less ingestion (#14918 ) * save work * Working * Fix runner constructor * Working runner * extra log lines * try using lifecycle for everything * clean up configs * cleanup /workers call * Use a single config * Allow selecting runner * debug changes * Work on composite task runner * Unit tests running * Add documentation * Add some javadocs * Fix spelling * Use standard libraries * code review * fix * fix * use taskRunner as string * checkstyl --------- Co-authored-by: Suneet Saldanha <suneet@apache.org>	2023-09-11 18:09:46 -07:00
317brian	3a453f7a3c	docs: add note about transparent_reconnection (#14953 ) * add note about transparent_reconnection * Update docs/api-reference/sql-jdbc.md	2023-09-11 11:58:39 -07:00
Kashif Faraz	7871e633c6	Fix bug in KillStalePendingSegments (#14961 )	2023-09-11 15:18:15 +05:30
Tejaswini Bandlamudi	dec6a0aa14	Update google client apis to latest version (#14414 ) Currently Druid is using google apis client 1.26.0 version and google-oauth-client-1.26.0.jar in particular is bringing following CVEs CVE-2020-7692, CVE-2021-22573. Despite the CVEs being false positives, they're causing red security scans on Druid distribution. Hence updating the version to latest version with these CVE fixes.	2023-09-11 12:27:23 +05:30
Clint Wylie	2b7f2c5119	use VectorValueSelector instead of BaseLongVectorValueSelector for StringFirstAggregatorFactory.factorizeVector (#14957 )	2023-09-09 04:03:05 -07:00
317brian	09f7dfe327	docs: update docusaurus 2 stuff (#14864 )	2023-09-08 14:19:15 -07:00
Zoltan Haindrich	699893bcff	Fix StringLastAggregatorFactory equals/toString (#14907 ) * update test * update test * format * test * fix0 * Revert "fix0" This reverts commit `44992cb393`. * ok resultset * add plan * update test * before rewind * test * fix toString/compare/test * move test * add timeColumn to hashCode	2023-09-08 09:20:54 -07:00
Kashif Faraz	647686aee2	Add test and metrics for KillStalePendingSegments duty (#14951 ) Changes: - Add new metric `kill/pendingSegments/count` with dimension `dataSource` - Add tests for `KillStalePendingSegments` - Reduce no-op logs that spit out for each datasource even when no pending segments have been deleted. This can get particularly noisy at low values of `indexingPeriod`. - Refactor the code in `KillStalePendingSegments` for readability and add javadocs	2023-09-08 10:33:47 +05:30
Abhishek Radhakrishnan	f9cf500a69	Extend GHA autolabeler to other areas (#14903 ) * Automate adding labels. * Add metrics/event emitting label * ingestion and segment format	2023-09-07 20:25:37 -07:00
Hardik Bajaj	e100b18e86	Updated documentation for OshiSysMonitor (#14912 )	2023-09-07 16:54:33 +05:30
Kashif Faraz	88f3c9baed	Fix bug in computed value of balancerComputeThreads (#14947 ) In smartSegmentLoading mode, use computed value of balancerComputeThreads rather than configured value.	2023-09-07 01:14:05 +05:30
Soumyava	a8fa979115	Unnest dont push down not (#14942 ) * Not pushing down not filters * New test case * Updating tests * Removing a stale comment	2023-09-06 08:57:03 -07:00
Zoltan Haindrich	23308c050d	Remove DruidAggregateCaseToFilterRule (#14940 ) The issue due to which the custom rule was added has been fixed as a part of https://issues.apache.org/jira/browse/CALCITE-3763 and accommodated during Calcite upgrade	2023-09-06 19:11:58 +05:30
Laksh Singla	6ee0b06e38	Auto configuration for maxSubqueryBytes (#14808 ) A new monitor SubqueryCountStatsMonitor which emits the metrics corresponding to the subqueries and their execution is now introduced. Moreover, the user can now also use the auto mode to automatically set the number of bytes available per query for the inlining of its subquery's results.	2023-09-06 05:47:19 +00:00
Adarsh Sanjeev	959148ad37	Add code to wait for segments generated to be loaded on historicals (#14322 ) Currently, after an MSQ query, the web console is responsible for waiting for the segments to load. It does so by checking if there are any segments loading into the datasource ingested into, which can cause some issues, like in cases where the segments would never be loaded, or would end up waiting for other ingests as well. This PR shifts this responsibility to the controller, which would have the list of segments created.	2023-09-06 10:35:57 +05:30
Clint Wylie	706b57c0b2	fixup array and mvd sql docs (#14928 )	2023-09-05 16:17:00 -07:00
Jill Osborne	425ebaa387	Query tips doc (#14922 ) Co-authored-by: Katya Macedo <38017980+ektravel@users.noreply.github.com> Co-authored-by: Victoria Lim <vtlim@users.noreply.github.com> Co-authored-by: Katya Macedo <38017980+ektravel@users.noreply.github.com>	2023-09-05 14:16:01 -07:00
Soumyava	8088a763a6	Vectorize earliest aggregator for both numeric and string types (#14408 ) * Vectorizing earliest for numeric * Vectorizing earliest string aggregator * checkstyle fix * Removing unnecessary exceptions * Ignoring tests in MSQ as earliest is not supported for numeric there * Fixing benchmarks * Updating tests as MSQ does not support earliest for some cases * Addressing review comments by adding the following: 1. Checking capabilities first before creating selectors 2. Removing mockito in tests for numeric first aggs 3. Removing unnecessary tests * Addressing issues for dictionary encoded single string columns where we can use the dictionary ids instead of the entire string * Adding a flag for multi value dimension selector * Addressing comments * 1 more change * Handling review comments part 1 * Handling review comments and correctness fix for latest_by when the time expression need not be in sorted order * Updating numeric first vector agg * Revert "Updating numeric first vector agg" This reverts commit `4291709901`. * Updating code for correctness issues * fixing an issue with latest agg * Adding more comments and removing an unnecessary check * Addressing null checks for tie selector and only vectorize false for quantile sketches	2023-09-05 08:41:42 -07:00
Abhishek Radhakrishnan	9d6ca61ac1	Verify statsd mock client interaction in unit test (#14939 )	2023-09-05 07:34:22 -07:00
Kashif Faraz	289ee1e011	Refactor: Cleanup NoopTask (#14938 ) Changes: - Simplify static `create` methods for `NoopTask` - Remove `FirehoseFactory`, `IsReadyResult`, `readyTime` from `NoopTask` as these fields were not being used anywhere - Update tests	2023-09-05 09:15:41 +05:30
panhongan	d4e972e1e4	Add checking for new checkpoint (#14353 ) Check that a checkpoint is non-empty before adding it to the checkpoint sequence in a SeekableStreamSupervisor	2023-09-04 13:18:55 +05:30
Kashif Faraz	ec630e3671	Remove deprecated coordinator dynamic configs (#14923 ) Changes: [A] Remove config `decommissioningMaxPercentOfMaxSegmentsToMove` - It is a complicated config 😅 , - It is always desirable to prioritize move from decommissioning servers so that they can be terminated quickly, so this should always be 100% - It is already handled by `smartSegmentLoading` (enabled by default) [B] Remove config `maxNonPrimaryReplicantsToLoad` This was added in #11135 to address two requirements: - Prevent coordinator runs from getting stuck assigning too many segments to historicals - Prevent load of replicas from competing with load of unavailable segments Both of these requirements are now already met thanks to: - Round-robin segment assignment - Prioritization in the new coordinator - Modifications to `replicationThrottleLimit` - `smartSegmentLoading` (enabled by default)	2023-09-04 11:54:36 +05:30
Kashif Faraz	7f26b80e21	Simplify ServiceMetricEvent.Builder (#14933 ) Changes: - Make ServiceMetricEvent.Builder extend ServiceEventBuilder<ServiceMetricEvent> and thus convert it to a plain builder rather than a builder of builder. - Add methods setCreatedTime , setMetricAndValue to the builder	2023-09-01 11:30:45 +05:30
Clint Wylie	dea9d4f1a7	cleaning DruidProcessingConfig bindings (#14927 )	2023-08-30 22:35:08 -07:00
Vadim Ogievetsky	680669fd3a	show execution dialog in task view (#14930 )	2023-08-30 15:59:34 -07:00
Vadim Ogievetsky	04a1153d0f	line chart fix others not mapping correctly (#14931 )	2023-08-30 15:59:26 -07:00
Sébastien	42cfb999cd	Added brush to time-chart (#14929 )	2023-08-30 10:36:50 -07:00
Vadim Ogievetsky	d295b9158f	Web console: dynamic query parameters UI (#14921 ) * fix nvl in table * add query parameter dialog * pre-wrap in the tables * fix typo	2023-08-29 23:14:25 -07:00
Kashif Faraz	8263f0d1e9	Reduce coordinator logs when operating normally (#14926 ) Changes: - Reduce log level of some coordinator stats, which only denote normal coordinator operation. These stats are still emitted and can be logged by setting debugDimensions in the coordinator dynamic config. - Initialize SegmentLoadingConfig only for historical management duties. This config is not needed in other duties and initializing it creates logs which are misleading.	2023-08-30 11:30:38 +05:30
John Gerassimou	d201ea0ece	prometheus-emitter: add extraLabels parameter (#14728 ) * prometheus-emitter: add extraLabels parameter * prometheus-emitter: update readme to include the extraLabels parameter * prometheus-emitter: remove nullable and surface label name issues * remove import to make linter happy	2023-08-29 12:02:22 -07:00
Gian Merlino	004cd012e1	HttpClient: Include error handler on all connection attempts. (#14915 ) Currently we have an error handler for https connection attempts, but not for plaintext connection attempts. This leads to warnings like the following for plaintext connection errors: EXCEPTION, please implement org.jboss.netty.handler.codec.http.HttpContentDecompressor.exceptionCaught() for proper handling. This happens because if we don't add our own error handler, the last handler in the chain during a connection attempt is HttpContentDecompressor, which doesn't handle errors. The new error handler for plaintext doesn't do much: it just closes the channel.	2023-08-29 14:28:04 +05:30
benkrug	8885805bb3	Update filters.md (#14917 )	2023-08-28 15:29:00 -07:00
Kashif Faraz	d6565f46b0	Increase the computed value of replicationThrottleLimit (#14913 ) Changes - Increase value of `replicationThrottleLimit` computed by `smartSegmentLoading` from 2% to 5% of total number of used segments. - Assign replicas to a tier even when some replicas are already being loaded in that tier - Limit the total number of replicas in load queue at start of run + replica assignments in the run to the `replicationThrottleLimit`. i.e. for every tier, num loading replicas at start of run + num replicas assigned in run <= replicationThrottleLimit	2023-08-28 18:20:22 +05:30
Karan Kumar	9fcbf05c5d	Adjusting `SqlStatementResource` and `SqlTaskResource` to set request attribute via a new method. (#14878 )	2023-08-26 10:59:47 +00:00
Vadim Ogievetsky	30c49c4cfc	Web console: misc fixes and SQL query re-formatting (#14906 ) * better dialog formatting * use CSS to render triangle * can flatten in kafka also * better formatting * better format * fill in empty values in line chart * more fp * add show others	2023-08-25 15:18:37 -07:00
Victoria Lim	9142f4b8d7	docs: update note in automatic compaction doc (#14908 )	2023-08-25 14:14:29 -07:00
George Shiqi Wu	95b0de61d1	Move some lifecycle management from doTask -> shutdown for the mm-less task runner (#14895 ) * save work * Add syncronized * Don't shutdown in run * Adding unit tests * Cleanup lifecycle * Fix tests * remove newline	2023-08-25 10:50:38 -06:00
George Shiqi Wu	ad32f84586	Fix capacity response in mm-less ingestion (#14888 ) Changes: - Fix capacity response in mm-less ingestion. - Add field usedClusterCapacity to the GET /totalWorkerCapacity response. This API should be used to get the total ingestion capacity on the overlord. - Remove method `isK8sTaskRunner` from interface `TaskRunner`	2023-08-25 08:17:38 +05:30
Kashif Faraz	e51181957c	Use num cores to determine balancerComputeThreads (#14902 ) Changes: - Determine the default value of balancerComputeThreads based on number of coordinator cpus rather than number of segments. Even if the number of segments is low and we create more balancer threads, it doesn't hurt the system as threads would mostly be idle. - Remove unused field from SegmentLoadQueueManager Expected values: - Clusters with ~1M segments typically work with Coordinators having 16 cores or more. This would give us 8 balancer threads, which is the same as the current maximum. - On small clusters, even a single thread is enough to do the required balancing work.	2023-08-25 08:15:27 +05:30
Tejaswini Bandlamudi	388d5ecf78	Fix reported CVEs (#14882 ) Suppress CVEs from dependencies with no available fix or false positives hadoop-annotations: CVE-2022-25168, CVE-2021-33036 hadoop-client-runtime: CVE-2023-1370, CVE-2023-37475 okio: CVE-2023-3635 Upgrade grpc version to fix CVE-2023-33953	2023-08-24 19:28:55 +05:30
Abhishek Agarwal	3c7b237c22	Add docs for ingesting Kafka topic name (#14894 ) Add documentation on how to extract the Kafka topic name and ingest it into the data.	2023-08-24 19:19:59 +05:30
Zoltan Haindrich	54336e2a3e	Imporve on incremental compilation (#14860 ) This patch fixes a few issues toward #14858 1. some phony classes were added to enable maven to track the compilation of those classes 2. cyclonedx 2.7.9 seem to handle incremental compilation better; it had a PR relating to that 3. needed to update root pom to 25 4. update antlr to 4.5.3 older one didn't really worked incrementally; 4.5.3 works much better	2023-08-24 16:06:16 +05:30
Laksh Singla	f9f734cde5	Display the output column name in InvalidNullByteException (#14780 ) This PR maps the query column to the output column name while surfacing the fault since that is readily visible to the user while executing the query.	2023-08-24 04:24:41 +00:00
Clint Wylie	36e659a501	remove group-by v1 (#14866 ) * remove group-by v1 * docs * remove unused configs, fix test * fix test * adjustments * why not * adjust * review stuff	2023-08-23 12:44:06 -07:00
zachjsh	0c76df1c7d	Enable Continuous auto kill (#14831 ) ### Description This change enables the `KillUnusedSegments` coordinator duty to be scheduled continuously. Things that prevented this, or made this difficult before were the following: 1. If scheduled at fast enough rate, the duty would find the same intervals to kill for the same datasources, while kill tasks submitted for those same datasources and intervals were already underway, thus wasting task slots on duplicated work. 2. The task resources used by auto kill were previously unbounded. Each duty run period, if unused segments were found for any datasource, a kill task would be submitted to kill them. This pr solves for both of these issues: 1. The duty keeps track of the end time of the last interval found when killing unused segments for each datasource, in a in memory map. The end time for each datasource, if found, is used as the start time lower bound, when searching for unused intervals for that same datasource. Each duty run, we remove any datasource keys from this map that are no longer found to match datasources in the system, or in whitelist, and also remove a datasource entry, if there is found to be no unused segments for the datasource, which happens when we fail to find an interval which includes unused segments. Removing the datasource entry from the map, allows for searching for unusedSegments in the datasource from the beginning of time once again 2. The unbounded task resource usage can be mitigated with coordinator dynamic config added as part of `ba957a9b97` Operators can configure continous auto kill by providing coordinator runtime properties similar to the following: ``` druid.coordinator.period.indexingPeriod=PT60S druid.coordinator.kill.period=PT60S ``` And providing sensible limits to the killTask usage via coordinator dynamic properties.	2023-08-23 09:23:08 -04:00
Adarsh Sanjeev	dfb5a98888	Add coordinator API for unused segments (#14846 ) There is a current issue due to inconsistent metadata between worker and controller in MSQ. A controller can receive one set of segments, which are then marked as unused by, say, a compaction job. The worker would be unable to get the segment information as MetadataResource.	2023-08-23 14:51:25 +05:30

1 2 3 4 5 ...

13280 Commits All Branches Search

13280 Commits

All Branches