druid

Commit Graph

Author	SHA1	Message	Date
Katya Macedo	58d9720b00	docs: notebook only for SQL tutorial (#13465 ) CI Failures seem unrelated to docs * docs: notebook only for SQL tutorial * Update logical operators section * Fix typo * Adopt review suggestions * Update examples/quickstart/jupyter-notebooks/sql-tutorial.ipynb * Update examples, add link to keywords * Update after review * Update per review comments * Add links	2023-02-08 20:04:53 -08:00
Suneet Saldanha	714ac07b52	Allow users to add additional metadata to ingestion metrics (#13760 ) * Allow users to add additional metadata to ingestion metrics When submitting an ingestion spec, users may pass a map of metadata in the ingestion spec config that will be added to ingestion metrics. This will make it possible for operators to tag metrics with other metadata that doesn't necessarily line up with the existing tags like taskId. Druid clusters that ingest these metrics can take advantage of the nested data columns feature to process this additional metadata. * rename to tags * docs * tests * fix test * make code cov happy * checkstyle	2023-02-08 18:07:23 -08:00
Adarsh Sanjeev	d7a15be9bc	Add assertions for counters from reports (#13726 ) Adds assertions for counters to MSQ unit tests	2023-02-08 16:33:37 +05:30
AmatyaAvadhanula	34c04daa9f	Fix infinite iteration in http sync monitoring (#13731 ) * Fix infinite iteration in http task runner * Fix infinite iteration in http server view * Add tests	2023-02-08 15:14:11 +05:30
jamon	d925ebdc9e	Bump app version in Helm Chart from 0.23.0 to 24.0.0 (#13341 ) Co-authored-by: zemin <zemin.piao@adyen.com>	2023-02-08 11:48:47 +05:30
AmatyaAvadhanula	0cf1fc3d55	Indexing on multiple disks (#13476 ) * Initial commit * Simple UTs * Parameterize tests * Parameterized tests for k8s task runner * Fix restore bug * Refactor TaskStorageDirTracker * Change CliPeon args	2023-02-08 11:31:34 +05:30
Jason Witkowski	5934d5fffe	helm: Stop helm chart from failing if zkHosts is not set (#13746 )	2023-02-08 10:43:35 +05:30
imply-cheddar	f684df4c22	Use an HllSketchHolder object to enable optimized merge (#13737 ) * Use an HllSketchHolder object to enable optimized merge HllSketchAggregatorFactory.combine had been implemented using a pure pair-wise, "make a union -> add 2 things to union -> get sketch" algorithm. This algorithm does 2 things that was CPU 1) The Union object always builds an HLL_8 sketch regardless of the target type. This means that when the target type is not HLL_8, we spent CPU cycles converting to HLL_8 and back over and over again 2) By throwing away the Union object and converting back to the HllSketch only to build another Union object, we do lots and lots of copy+conversions of the HllSketch This change introduces an HllSketchHolder object which can hold onto a Union object and delay conversion back into an HllSketch until it is actually needed. This follows the same pattern as the SketchHolder object for theta sketches.	2023-02-07 13:57:48 -08:00
AmatyaAvadhanula	dcdae84888	Add server view initialization metrics (#13716 ) * Add server view init metrics * Test coverage * Rename metrics	2023-02-07 20:02:00 +05:30
Rohan Garg	a0f8889f23	Robust handling and management of S3 streams for MSQ shuffle storage (#13741 )	2023-02-07 14:17:37 +05:30
John Gozde	b33962cab7	Upgrade typescript and other dependencies (#13762 ) * Bump zustand, licenses * Bump TypeScript, Eslint, use type imports * Switch to react-shallow-renderer from enzyme * Update ts-loader	2023-02-06 23:12:54 -08:00
Clint Wylie	2d3bee8545	various nested column (and other) fixes (#13732 ) changes: * modified druid schema column type compution to special case COMPLEX<json> handling to choose COMPLEX<json> if any column in any segment is COMPLEX<json> * NestedFieldVirtualColumn can now work correctly on any type of column, returning either a column selector if a root path, or nil selector if not * fixed a random bug with NilVectorSelector when using a vector size larger than the default and druid.generic.useDefaultValueForNull=false would have the nulls vector set to all false instead of true * fixed an overly aggressive check in ExprEval.ofType when handling complex types which would try to treat any string as base64 without gracefully falling back if it was not in fact base64 encoded, along with special handling for complex<json> * added ExpressionVectorSelectors.castValueSelectorToObject and ExpressionVectorSelectors.castObjectSelectorToNumeric as convience methods to cast vector selectors using cast expressions without the trouble of constructing an expression. the polymorphic nature of the non-vectorized engine (and significantly larger overhead of non-vectorized expression processing) made adding similar methods for non-vectorized selectors less attractive and so have not been added at this time * fix inconsistency between nested column indexer and serializer in handling values (coerce non primitive and non arrays of primitives using asString) * ExprEval best effort mode now handles byte[] as string * added test for ExprEval.bestEffortOf, and add missing conversion cases that tests uncovered * more tests more better	2023-02-06 19:48:02 -08:00
imply-cheddar	9c5b61e114	Fallback virtual column (#13739 ) * Fallback virtual column This virtual columns enables falling back to another column if the original column doesn't exist. This is useful when doing column migrations and you have some old data with column X, new data with column Y and you want to use Y if it exists, X otherwise so that you can run a consistent query against all of the data.	2023-02-06 19:36:50 -08:00
Paul Rogers	f28c06515b	Auto-detect docker-compose (#13754 )	2023-02-06 21:29:45 +05:30
Laksh Singla	9100a61bf6	Fix NPE in postCleanupStage if stage doesn't exist (#13742 ) With fault tolerance enabled in MSQ, not all the work orders might be populated if the worker is restarted. In case it gets the request for cleaning up the stage which is not present in the worker's map, it can throw an NPE. Added a check to ensure that the stage is present in the map before cleaning it up, or else logging it as a warning.	2023-02-06 19:13:39 +05:30
Rohan Garg	c5835c29a1	Use durable super sorter intermediate storage only with composable storage (#13748 ) * This enables usage of durable storage connector only in case the composable storage feature is enabled.	2023-02-06 18:59:18 +05:30
Elliott Freis	e16639121f	Local pathing for tests (#13753 ) Co-authored-by: Elliott Freis <elliottfreis@Elliott-Freis.earth.dynamic.blacklight.net>	2023-02-03 20:11:17 -08:00
Elliott Freis	c06631037d	Moving to SHA based cache key (#13751 ) Co-authored-by: Elliott Freis <elliottfreis@Elliott-Freis.earth.dynamic.blacklight.net>	2023-02-03 15:49:17 -08:00
Suneet Saldanha	bea18dc9e4	Update basic auth examples (#13750 )	2023-02-03 14:45:48 -08:00
drudi-at-coffee	7580248770	Update api.md (#13727 ) Added missing '/status' in HTTP status request	2023-02-02 10:43:22 -08:00
Churro	f022a9f246	When a task fails and doesn't throw an exception, report it correctly… (#13668 ) * When a task fails and doesn't throw an exception, report it correctly in mm-less druid * Removing unthrown exception from test	2023-02-02 09:04:18 -08:00
Tejaswini Bandlamudi	6cb842e76e	update snapshots if cache restore failed otherwise run test normally (not all test mvn dependencies are downloaded during build phase due to skipTests) (#13740 )	2023-02-02 08:09:56 -08:00
Tejaswini Bandlamudi	440212c5f9	Fix GHA workflow maven build erroring incase of version updates (#13735 ) * build maven sequentially * run mvn tests in offline mode after retrieving cache	2023-02-01 20:32:57 -08:00
Kashif Faraz	7c188d80b8	Make batch segment allocation logs less noisy (#13725 )	2023-02-02 09:54:53 +05:30
Victoria Lim	33efd5ab1d	docs: Refresh the update data tutorial (#13641 ) Merging regardless of nit since topic is in better shape. * refresh the update data tutorial * Apply suggestions from code review Co-authored-by: Jill Osborne <jill.osborne@imply.io> --------- Co-authored-by: Jill Osborne <jill.osborne@imply.io>	2023-02-01 18:18:16 -08:00
Kashif Faraz	f629643c50	Fix value of lookup sync period in docs (#13695 ) * Fix lookup docs * Fix spelling * Apply suggestions from code review Co-authored-by: Charles Smith <techdocsmith@gmail.com> --------- Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2023-02-01 18:12:00 -08:00
Sergio Ferragut	7f830b20d7	fixed init commands for both mysql and postgresql (#13713 )	2023-02-01 18:07:31 -08:00
Suneet Saldanha	cfc3115a59	Compaction history returns empty list instead of 404 when not found (#13730 ) * Compaction history returns empty list instead of 404 when not found * checkstyle	2023-02-01 17:44:07 -08:00
AmatyaAvadhanula	76e79c7db7	Suppress CVEs (#13733 )	2023-02-01 04:18:41 -08:00
somu-imply	74ff848ce5	Fixing incorrect filtering of nulls in an array when ingesting for JSON and Avro (#13712 )	2023-02-01 04:15:08 -08:00
Tejaswini Bandlamudi	c95a26cae3	Migrate ITs from Travis to GHA (#13681 )	2023-02-01 03:31:29 -08:00
Jason Koch	7a3bd89a85	Dimension dictionary reduce locking (#13710 ) * perf: introduce benchmark for StringDimensionIndexer jdk11 -- Benchmark Mode Cnt Score Error Units StringDimensionIndexerProcessBenchmark.parallelReadWrite avgt 10 30471.552 ± 456.716 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelReader avgt 10 18069.863 ± 327.923 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelWriter avgt 10 67676.617 ± 2351.311 us/op StringDimensionIndexerProcessBenchmark.soloReader avgt 10 1048.079 ± 1.120 us/op StringDimensionIndexerProcessBenchmark.soloWriter avgt 10 4629.769 ± 29.353 us/op * perf: switch DimensionDictionary to StampedLock jdk11 - Benchmark Mode Cnt Score Error Units StringDimensionIndexerProcessBenchmark.parallelReadWrite avgt 10 37958.372 ± 1685.206 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelReader avgt 10 31192.232 ± 2755.365 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelWriter avgt 10 58256.791 ± 1998.220 us/op StringDimensionIndexerProcessBenchmark.soloReader avgt 10 1079.440 ± 1.753 us/op StringDimensionIndexerProcessBenchmark.soloWriter avgt 10 4585.690 ± 13.225 us/op * perf: use optimistic locking in DimensionDictionary jdk11 - Benchmark Mode Cnt Score Error Units StringDimensionIndexerProcessBenchmark.parallelReadWrite avgt 10 6212.366 ± 162.684 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelReader avgt 10 1807.235 ± 109.339 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelWriter avgt 10 19427.759 ± 611.692 us/op StringDimensionIndexerProcessBenchmark.soloReader avgt 10 194.370 ± 1.050 us/op StringDimensionIndexerProcessBenchmark.soloWriter avgt 10 2871.423 ± 14.426 us/op * perf: refactor DimensionDictionary null handling to need less locks jdk11 - Benchmark Mode Cnt Score Error Units StringDimensionIndexerProcessBenchmark.parallelReadWrite avgt 10 6591.619 ± 470.497 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelReader avgt 10 1387.338 ± 144.587 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelWriter avgt 10 22204.462 ± 1620.806 us/op StringDimensionIndexerProcessBenchmark.soloReader avgt 10 204.911 ± 0.459 us/op StringDimensionIndexerProcessBenchmark.soloWriter avgt 10 2935.376 ± 12.639 us/op * perf: refactor DimensionDictionary add handling to do a little less work jdk11 - Benchmark Mode Cnt Score Error Units StringDimensionIndexerProcessBenchmark.parallelReadWrite avgt 10 2914.859 ± 22.519 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelReader avgt 10 508.010 ± 14.675 us/op StringDimensionIndexerProcessBenchmark.parallelReadWrite:parallelWriter avgt 10 10135.408 ± 82.745 us/op StringDimensionIndexerProcessBenchmark.soloReader avgt 10 205.415 ± 0.158 us/op StringDimensionIndexerProcessBenchmark.soloWriter avgt 10 3098.743 ± 23.603 us/op	2023-02-01 02:59:12 -08:00
Clint Wylie	ec1e6ac840	fix nested column handling of null and "null" (#13714 ) * fix nested column handling of null and "null" * fix issue merging nested column value dictionaries that could incorrect lose dictionary values	2023-01-31 20:59:19 -08:00
Tijo Thomas	1beef30bb2	Support postaggregation function as in Math.pow() (#13703 ) (#13704 ) Support postaggregation function as in Math.pow()	2023-01-31 22:55:04 +05:30
Adarsh Sanjeev	51dfde0284	Add maxInputBytesPerWorker as query context parameter (#13707 ) * Add maxInputBytesPerWorker as query context parameter * Move documenation to msq specific docs * Update tests * Spacing * Address review comments * Fix test * Update docs/multi-stage-query/reference.md * Correct spelling mistake --------- Co-authored-by: Karan Kumar <karankumar1100@gmail.com>	2023-01-31 20:55:28 +05:30
Xavier Léauté	698670c88e	update core Apache Kafka dependencies to 3.3.2 (#13717 ) Release notes: - https://downloads.apache.org/kafka/3.3.2/RELEASE_NOTES.html	2023-01-27 21:00:01 -08:00
Jill Osborne	356b0e37cf	Tutorial: Query view (#13565 ) * Tutorial: Query view * Removed duplicate file * Update tutorial-sql-query-view.md * Update tutorial-sql-query-view.md * Update tutorial-sql-query-view.md * Updated after review * Update docs/tutorials/tutorial-sql-query-view.md Co-authored-by: Charles Smith <techdocsmith@gmail.com> * Update tutorial-sql-query-view.md Update title * Update sidebars.json fix merge conflict w/ sidebar * address spelling ci --------- Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2023-01-27 14:29:43 -08:00
Vadim Ogievetsky	3b62d7929c	Web console: Data loader should allow for multiline JSON messages in kafka (#13709 ) * stricter * data loader should allow for mulit-line json * add await * kinesis also	2023-01-25 21:23:18 -08:00
sairam devarashetty	6164c420a1	Create update.md (#13451 ) * Create update.md Important Line highlighted * Update docs/data-management/update.md Co-authored-by: Charles Smith <techdocsmith@gmail.com>	2023-01-25 16:23:40 -08:00
317brian	9021161c8c	doc: fix markdown spacing (#13683 ) * doc: fix markdown spacing * fix spacing	2023-01-25 16:22:49 -08:00
somu-imply	17c0167248	Additional native query tests for unnest datasource (#13554 ) Native tests for the unnest datasource.	2023-01-25 15:57:52 -08:00
Victoria Lim	00cee329bd	pitfall when using combining input source (#13639 )	2023-01-25 12:50:19 -08:00
imply-cheddar	706b8a0227	Adjust Operators to be Pausable (#13694 ) * Adjust Operators to be Pausable This enables "merge" style operations that combine multiple streams. This change includes a naive implementation of one such merge operator just to provide concrete evidence that the refactoring is effective.	2023-01-23 20:52:06 -08:00
Suneet Saldanha	016c881795	Add API to return automatic compaction config history (#13699 ) Add a new API to return the history of changes to automatic compaction config history to make it easy for users to see what changes have been made to their auto-compaction config. The API is scoped per dataSource to allow users to triage issues with an individual dataSource. The API responds with a list of configs when there is a change to either the settings that impact all auto-compaction configs on a cluster or the dataSource in question.	2023-01-23 13:23:45 -08:00
somu-imply	90d445536d	SQL version of unnest native druid function (#13576 ) * adds the SQL component of the native unnest functionality in Druid to unnest SQL queries on a table dimension, virtual column or a constant array and convert them into native Druid queries * unnest in SQL is implemented as a combination of Correlate (the comma join part) and Uncollect (the unnest part)	2023-01-23 12:53:31 -08:00
Rohan Garg	f76acccff2	Allow using composed storage for SuperSorter intermediate data (#13368 )	2023-01-24 01:02:03 +05:30
Laksh Singla	a516eb1a41	Port Calcite's tests to run with MSQ (#13625 ) * SQL test framework extensions * Capture planner artifacts: logical plan, etc. * Planner test builder validates the logical plan * Validation for the SQL resut schema (we already have validation for the Druid row signature) * Better Guice integration: properties, reuse Guice modules * Avoid need for hand-coded expr, macro tables * Retire some of the test-specific query component creation * Fix query log hook race condition Co-authored-by: Paul Rogers <progers@apache.org>	2023-01-19 08:51:11 -08:00
Clint Wylie	fb26a1093d	discover nested columns when using nested column indexer for schemaless ingestion (#13672 ) * discover nested columns when using nested column indexer for schemaless * move useNestedColumnIndexerForSchemaDiscovery from AppendableIndexSpec to DimensionsSpec	2023-01-18 12:57:28 -08:00
Maytas Monsereenusorn	1582d74f37	Fix Parquet Reader for schema-less ingestion need to read all columns (#13689 ) * fix stuff * address comments	2023-01-18 12:52:12 -08:00
Paul Rogers	fa493f1ebc	Convert from DRUID_INTEGRATION_TEST_INDEXER to USE_INDEXER (#13684 ) The old ITs use DRUID_INTEGRATION_TEST_INDEXER. The new ones use the USE_INDEXER env var passed in from the build environment.	2023-01-18 08:51:42 -08:00

... 8 9 10 11 12 ...

12860 Commits All Branches Search

12860 Commits

All Branches