OpenSearch

Commit Graph

Author	SHA1	Message	Date
Przemko Robakowski	bb357f6aae	[7.x] Move internal index templates to composable templates (#61457 ) (#61661 ) This change moves watcher, ILM history and SLM history templates to composable templates. Versions are updated to reflect the switch. Only change to the templates themselves is added `_meta` to mark them as managed	2020-09-08 11:26:06 +02:00
Andrei Stefan	7d5791b6bd	EQL: create the search request with a list of indices (#62005 ) (#62076 ) * The query client uses an array of indices instead of the comma separated version of the indices names (cherry picked from commit 8ec4a768f4892a4a2faed25836cb333a9deb2ace)	2020-09-08 10:26:59 +03:00
David Kyle	a5b24bf44c	Mute ClassificationIT (#62063 ) testWithOnlyTrainingRowsAndTrainingPercentIsFifty_DependentVariableIsBoolean For #60759	2020-09-07 16:10:48 +01:00
Luca Cavanna	168b448a0f	Rename runtime_script field type to runtime (#62034 ) We've had some discussions around the user experience when using runtime fields. Although we do plan on having multiple runtime fields implementation (e.g. grok, lookup etc.) which could be exposed as different field types, we decided to expose all runtime fields under the same `runtime` type. At the moment, the only implementation will be through scripts, hence a `script` must be specified. In the future, there will be other ways to generate values for runtime fields besides scripts. This translates also to renaming the RuntimeScriptFieldMapper class to RuntimeFieldMapper . Relates to #59332	2020-09-07 15:07:23 +02:00
Jim Ferenczi	fa8e76abb1	Improve reduction of terms aggregations (#61779 ) (#62028 ) Today, the terms aggregation reduces multiple aggregations at once using a map to group same buckets together. This operation can be costly since it requires to lookup every bucket in a global map with no particular order. This commit changes how term buckets are sorted by shards and partial reduces in order to be able to reduce results using a merge-sort strategy. For bwc, results are merged with the legacy code if any of the aggregations use a different sort (if it was returned by a node in prior versions). Relates #51857	2020-09-07 13:13:20 +02:00
Martijn van Groningen	7e566ddd06	Move data stream yaml tests to xpack plugin module. (#62032 ) Backport of #61998 to 7.x branch. Moving the data stream yaml tests to xpack plugin module has the following benefits: * The tests are ran both with security enabled (as part of xpack/plugin integTest) and disabled (as part of xpack/plugin/data-stream/qa/rest integTest). * and running the tests in mixed cluster qa environment.	2020-09-07 11:03:32 +02:00
Tanguy Leroux	ebbf4df9fd	Adapt SearchableSnapshotsBlobStoreCacheIntegTests to Lucene 8.7.0 (#61989 ) (#62030 ) Elasticsearch now uses #61957 which includes https://issues.apache.org/jira/browse/LUCENE-9456. We can remove the corresponding //TODO in SearchableSnapshotsBlobStoreCacheIntegTests.	2020-09-07 10:25:44 +02:00
Luca Cavanna	0c8b438577	Add support for runtime fields (#61776 ) This commit includes the work that has been done on the runtime fields feature branch until now. The high level tasks are listed in #59332. The tasks that have not yet been completed can be worked on after merging the feature branch. We are adding a new x-pack plugin called runtime-fields that plugs in a custom mapper which allows to define runtime fields based on a script. The changes included in this commit that were made outside of the x-pack/plugin/runtime-fields directory are minimal and revolve around 1) making the ScriptService available while parsing index mappings so that the scripts associated to runtime fields can be compiled 2) sharing code to manipulate ranges etc. as it can be reused in runtime fields. Co-authored-by: Nik Everett <nik9000@gmail.com>	2020-09-07 09:14:53 +02:00
Dimitris Athanasiou	d37f197efd	[7.x][ML] Allow training_percent to be any positive double up to hundred (#61977 ) (#61990 ) This changes the valid range of `training_percent` for regression and classification from [1, 100] to (0, 100]. Backport of #61977	2020-09-04 17:34:14 +03:00
Yannick Welsch	6d08b55d4e	Simplify searchable snapshot shard allocation (#61911 ) Simplifies allocation for snapshot-backed shards by always making the recovery source "from snapshot" for those snapshot-backed shards (instead of "recover from local or from empty store"). Also let's the balancer pick a node which to allocate the snapshot-backed shard to (which takes number of shards on each node into account unlike the current implementation which just picks whatever node we are allowed to allocate to, with no notion of "balancing" at all).	2020-09-04 15:45:00 +02:00
Tanguy Leroux	289b1f4ae7	Reduce locking in prewarming (#61837 ) (#61967 ) During prewarming of a Lucene file a CacheFile is acquired and then locked for the duration of the prewarming, ie locked until all the part of the file has been downloaded and written to cache on disk. The locking (executed with CacheFile#fileLock()) is here to prevent the cache file to be evicted while it is prewarming. But holding the lock may take a while for large files, specially since restoring snapshot files now respects the indices.recovery.max_bytes_per_sec setting of 40mb (#58658), and this can have bad consequences like preventing the CacheFile to be evicted, opened or closed. In manual tests this bug slow downs various requests like mounting a new searchable snapshot index or deleting an existing one that is still prewarming. This commit reduces the time the lock is held during prewarming so that the read lock is only required when actively writing to the CacheFile.	2020-09-04 15:06:50 +02:00
Martijn van Groningen	84af9abd76	Fix skip versions fix xpack data stream yaml tests. (#61981 ) Backport of #61926 to 7.x branch. Relates to #61904	2020-09-04 14:53:38 +02:00
Benjamin Trent	cec102a391	[7.x] [ML] adds new n_gram_encoding custom processor (#61578 ) (#61935 ) * [ML] adds new n_gram_encoding custom processor (#61578) This adds a new `n_gram_encoding` feature processor for analytics and inference. The focus of this processor is simple ngram encodings that allow: - multiple ngrams [1..5] - Prefix, infix, suffix	2020-09-04 08:36:50 -04:00
Ignacio Vera	31c026f25c	upgrade to Lucene-8.7.0-snapshot-61ea26a (#61957 ) (#61974 )	2020-09-04 13:46:20 +02:00
Dimitris Athanasiou	bdccab7c7a	[7.x][ML] Add incremental id during data frame analytics reindexing (#61943 ) (#61971 ) Previously, we added a copy of the `_id` during reindexing and sorted the destination index on that. This allowed us to traverse the docs in the destination index in a stable order multiple times and with efficiency. However, the destination index being sorted means we cannot have `nested` typed fields. This is a problem as it does not allow us to provide a good experience with our evaluate API when it comes to computing metrics for specific classes, features, etc. This commit changes the approach in order to result to a destination index that allows nested fields. Instead of adding a copy of the `_id` field, we now add an incremental id that we can use to traverse the docs in a stable order. We also ensure we always assign the same incremental id to the same doc from the source indices by sorting on `_seq_no` during reindexing. That in combination with the reindexing API using scroll gives us a stable order as scroll uses the (`_index`, `_doc`, shard_id) tuple to resolve ties. The extractor now does not need to scroll. Instead we sort on the incremental id and we do ranged searches to avoid the sort-all-docs overhead. Finally, the `TestDocsIterator` is simply changed to search_after the incremental id. With these changes data frame analytics jobs do not use scroll at any part. Having all these in place, the commit adds the `nested` types to the necessary fields of `classification` and `regression` analyses results. Backport of #61943	2020-09-04 13:24:42 +03:00
Tanguy Leroux	10d14ce101	Enable searchable snapshot feature for all test clusters (#61888 ) (#61965 ) This commit reenables the searchable snapshot feature for integration tests after #61802 which changed some build plugins.	2020-09-04 11:20:24 +02:00
Tim Vernum	cdfb163c7c	Add explicit test for DLS with OIDC metadata (#61955 ) When a user authenticates via OpenID Connect we copy information from the OIDC claims into the user's metadata in a particular format. This commit adds a test that metadata in that format can be used in a mustache template for Document Level Security. Backport of: #60030	2020-09-04 16:21:20 +10:00
Tim Vernum	57efda2865	Add DEBUG logging for undefined role mapping field (#61887 ) A role mapping with the following content: "rules": { "field": { "userid" : "admin" } } will never match because `userid` is not a valid field. The correct field is `username`. This change adds DEBUG logging when an undefined field is referenced. The choice to use DEBUG rather than INFO/WARN is that the set of fields is partially dynamic (e.g. the `metadata.*` fields), so it may be perfectly reasonable to check a field that is not defined for that user. For example this rule: "rules": { "field": { "metadata.ranking" : "A" } } would generate a log message for an unranked user, which would erroneously suggest that such a rule is an error. This DEBUG logging will assist in diagnosing problems, without introducing that confusion. Backport of: #61246 Co-authored-by: Elastic Machine <elasticmachine@users.noreply.github.com>	2020-09-04 14:19:05 +10:00
Ryan Ernst	d6e17170c3	Simplify adding plugins and modules to testclusters (#61886 ) There are currently half a dozen ways to add plugins and modules for test clusters to use. All of them require the calling project to peek into the plugin or module they want to use to grab its bundlePlugin task, and then both depend on that task, as well as extract the archive path the task will produce. This creates cross project dependencies that are difficult to detect, and if the dependent plugin/module has not yet been configured, the build will fail because the task does not yet exist. This commit makes the plugin and module methods for testclusters symmetetric, and simply adding a file provider directly, or a project path that will produce the plugin/module zip. Internally this new variant uses normal configuration/dependencies across projects to get the zip artifact. It also has the added benefit of no longer needing the caller to add to the test task a dependsOn for bundlePlugin task.	2020-09-03 19:37:46 -07:00
Jake Landis	ea1e8ad6ea	[7.x] Fix passing params to template or script failed in watcher (#58559 ) (#61885 ) The main changes are: * Fix custom params are missing when using template or script in watcher's logging action or jira action. * Add yaml tests to test passing params to template or script successfully. Relates to #57625 Co-authored-by: bellengao <gbl_long@163.com>	2020-09-03 15:47:51 -05:00
Costin Leau	99ee87e332	EQL: Revert filter pipe (#61907 ) The current implementation of the filter pipe is incomplete hence why it got reverted. Note this is not a complete revert as some of the improvements of said commit (such as the PostAnalyzer) are useful in general. Relates #61805 (cherry picked from commit 7a7eb66f7d39586c3a3bc00dce49e6c47a23b46a)	2020-09-03 22:31:08 +03:00
Martijn van Groningen	3d9c12e2d3	Fix data stream wildcard resolution bug in eql search api.(#61910 ) Backport of #61904 to 7.x branch. The eql search api redirects to the search api. For this reason the eql search api could work with concrete data stream names. However if security is enabled and a data stream name snippet with a wildcard was used then it could not resolve this expressions. This is because the EqlSearchRequest class didn't overwrite the `includeDataStreams()` method. This pr fixes this, so that the security layer can properly expand data stream name wildcard expressions for the eql search api. This commit also moves the eql data stream test to xpack rest tests, so that the test runs with security enabled. This is required to reproduce the bug. Closes #60828	2020-09-03 16:03:57 +02:00
Tanguy Leroux	c90ee32cdc	Mute ClassificationIT.testTooLowConfiguredMemoryStillStarts (#61915 ) Relates #61913	2020-09-03 15:52:01 +02:00
Jake Landis	dbb78e1c45	[7.x] Correct the query dsl for watching elasticsearch version (#58321 ) (#61882 ) The term query should be looking at the cluster_uuid field in elasticsearch_version_mismatch.json. Co-authored-by: bellengao <gbl_long@163.com>	2020-09-02 16:58:21 -05:00
Nik Everett	c19f67ce30	Support longs in BitArray (backport of #61867 ) (#61871 ) We frequently use `long`s with `BitArray` in aggs and right now we have to assert that the `long` fits in an `int`. This adds support for `long` to `BitArray` so we don't need those assertions.	2020-09-02 17:24:31 -04:00
Dimitris Athanasiou	ec405978fc	[7.x][ML] Update reindexing task progress before persisting job progress (#61868 ) (#61875 ) This fixes a bug introduced by #61782. In that PR I thought I could simplify the persistence of progress by using the progress straight from the stats holder in the task instead of calling the get stats action. However, I overlooked that it is then possible to have stale progress for the reindexing task as that is only updated when the get stats API is called. In this commit this is fixed by updating reindexing task progress before persisting the job progress. This seems to be much more lightweight than calling the get stats request. Closes #61852 Backport of #61868	2020-09-02 21:44:18 +03:00
Benjamin Trent	c22415c241	[7.x] [ML] unmute testTooLowConfiguredMemoryStillStarts (#61846 ) (#61869 ) * [ML] unmute testTooLowConfiguredMemoryStillStarts (#61846) Native PR addresses this test failure: https://github.com/elastic/ml-cpp/pull/1465 closes https://github.com/elastic/elasticsearch/issues/61704 closes https://github.com/elastic/elasticsearch/issues/61561	2020-09-02 13:23:23 -04:00
Jake Landis	f6b3148e5e	[7.x] Convert second 1/2 x-pack plugins from integTest to [yaml \| java]RestTest or internalClusterTest (#61802 ) (#61856 ) For 1/2 the plugins in x-pack, the integTest task is now a no-op and all of the tests are now executed via a test, yamlRestTest, javaRestTest, or internalClusterTest. This includes the following projects: security, spatial, stack, transform, vecotrs, voting-only-node, and watcher. A few of the more specialized qa projects within these plugins have not been changed with this PR due to additional complexity which should be addressed separately. related: #60630 related: #56841 related: #59939 related: #55896	2020-09-02 11:20:55 -05:00
Jake Landis	794aac717d	[7.x] Convert first 1/2 x-pack plugins from integTest to [yaml \| java]RestTest or internalClusterTest (#60630 ) (#61855 ) For 1/2 the plugins in x-pack, the integTest task is now a no-op and all of the tests are now executed via a test, yamlRestTest, javaRestTest, or internalClusterTest. This includes the following projects: async-search, autoscaling, ccr, enrich, eql, frozen-indicies, data-streams, graph, ilm, mapper-constant-keyword, mapper-flattened, ml A few of the more specialized qa projects within these plugins have not been changed with this PR due to additional complexity which should be addressed separately. A follow up PR will address the remaining x-pack plugins (this PR is big enough as-is). related: #61802 related: #56841 related: #59939 related: #55896	2020-09-02 11:19:24 -05:00
Dimitris Athanasiou	07ab0beea0	[7.x][ML] Improve handling of exception while starting DFA process (#61838 ) (#61847 ) While starting the data frame analytics process it is possible to get an exception before the process crash handler is in place. In addition, right after starting the process, we check the process is alive to ensure we capture a failed process. However, those exceptions are unhandled. This commit catches any exception thrown while starting the process and sets the task to failed with the root cause error message. I have also taken the chance to remove some unused parameters in `NativeAnalyticsProcessFactory`. Relates #61704 Backport of #61838	2020-09-02 16:32:45 +03:00
Costin Leau	e6dc8054a5	EQL: Introduce filter pipe (#61805 ) Allow filtering through a pipe, across events and sequences. Filter pipes are pushed down to base queries. For now filtering after limit (head/tail) is forbidden as the semantics are still up for debate. Fix #59763 (cherry picked from commit 80569a388b76cecb5f55037fe989c8b6f140761b)	2020-09-02 15:48:51 +03:00
David Roberts	89599ba0a3	[ML] Update ML mappings upgrade test and extend to config index (#61830 ) The ML mappings upgrade test had become useless as it was checking a field that has been the same since 6.5. This commit switches to a field that was changed in 7.9. Additionally, the test only used to check the results index mappings. This commit also adds checking for the config index. Backport of #61340	2020-09-02 12:23:59 +01:00
David Kyle	d268540f20	[ML] Check and install the latest template in the DFA executor (#61589 ) (#61842 ) During a rolling upgrade it is possible that a worker node will be upgraded before the master in which case the DFA templates will not have been installed. Before a DFA task starts check that the latest template is installed and install it if necessary.	2020-09-02 12:16:29 +01:00
Nik Everett	f8158bdb2d	Skip failing test Tracked by https://github.com/elastic/elasticsearch/issues/61561	2020-09-01 13:44:31 -04:00
Dimitris Athanasiou	2547cfbe54	[7.x][ML] Persist progress when setting DFA task to failed (#61782 ) (#61792 ) When an error occurs and we set the task to failed via the `DataFrameAnalyticsTask.setFailed` method we do not persist progress. If the job is later restarted, this means we do not correctly restore from where we can but instead we start the job from scratch and have to redo the reindexing phase. This commit solves this bug by persisting the progress before setting the task to failed. Backport of #61782	2020-09-01 18:33:07 +03:00
Tanguy Leroux	d94d6b5b70	Also account for state not recovered in BlobStoreCacheService Following #61726 after a test failure	2020-09-01 12:10:04 +02:00
Ioannis Kakavas	ced2c140fe	Unmute TokenAuthIntegTests test (#61715 ) @ywangd made an awesome analysis on why this test is failing, over at https://github.com/elastic/elasticsearch/issues/55816#issuecomment-620913282 This change makes it so that we use the same client to perform a refresh of a token, as we use to subsequently attempt to authenticate with the refreshed token. This ensures the tests are failing and is a good approximation of how we expect the same client doing the refresh, to also perform the subsequent authentication in real life uses. The errors we were seeing from users have disappeared after #55114 so we deem our behavior safe.	2020-09-01 13:06:11 +03:00
Tanguy Leroux	787dfda4c1	Prevent snapshots to be mounted as system indices (#61517 ) (#61727 ) System indices can be snapshotted and are therefore potential candidates to be mounted as searchable snapshot indices. As of today nothing prevents a snapshot to be mounted under an index name starting with . and this can lead to conflicting situations because searchable snapshot indices are read-only and Elasticsearch expects some system indices to be writable; because searchable snapshot indices will soon use an internal system index (#60522) to speed up recoveries and we should prevent the system index to be itself a searchable snapshot index (leading to some deadlock situation for recovery). This commit introduces a changes to prevent snapshots to be mounted as a system index.	2020-09-01 11:13:28 +02:00
Tanguy Leroux	92eb6e7844	Remove cluster state listener in BlobStoreCacheService (#61726 ) (#61769 ) BlobStoreCacheService implements ClusterStateListener in order to maintain a ready flag that can be used to know when the snapshot blob cache should be queries or not. Now the getAsync() method correctly handles the various exceptions that can be thrown when the .snapshot-blob-cache index is not available(in isExpectedCacheGetException()) and logs as DEBUG we can safely remove the ready flag.	2020-09-01 11:12:52 +02:00
Benjamin Trent	7dabaad7d9	[ML] refactor ml job node selection into its own class (#61521 ) (#61747 ) This is a minor refactor where the job node load logic (node availability, etc.) is refactored into its own class. This will allow future things (i.e. autoscaling decisions) to use the same node load detection class. backport of #61521	2020-08-31 14:00:23 -04:00
Benjamin Trent	8b33d8813a	[ML] binary classification per-class feature importance for model inference (#61597 ) (#61746 ) This commit addresses two issues: - per class feature importance is now written out for binary classification (logistic regression) - The `class_name` in per class feature importance now matches what is written in the `top_classes` array. backport of https://github.com/elastic/elasticsearch/pull/61597	2020-08-31 13:57:00 -04:00
Mayya Sharipova	fe9c66096c	Small refactoring of AsyncExecutionId (#61640 ) - don't do encoding of asynchExecutionId if it is already provided in the encoded form - create a new instance of AsyncExecutionId after checks for correctness are done	2020-08-31 10:24:36 -04:00
Nhat Nguyen	e37ce561c7	Set timeout of auto put-follow request to unbounded (#61679 ) If the master node of the follower cluster is busy, then the auto-follower will fail to initialize the following process. This also occurs when an auto-follow pattern matches multiple indices. We should set the timeout of put-follow requests issued by the auto-follower to unbounded to avoid this problem. Closes #56891	2020-08-31 09:58:19 -04:00
Jason Tedor	64cd229b35	Upgrade to Lucene 8.6.2 (#61688 ) This commit upgrades the Lucene dependencies to 8.6.2.	2020-08-31 09:54:07 -04:00
Rory Hunter	ff6c071275	Implement deprecation logging using log4j (#61629 ) Backport of #61474. Part of #46106. Simplify the implementation of deprecation logging by relying of log4j more completely, and implementing additional behaviour through custom appenders and filters.	2020-08-31 12:42:04 +01:00
Henning Andersen	4c9fe31da8	Mute testTooLowConfiguredMemoryStillStarts (#61705 ) Related to #61704	2020-08-31 11:19:53 +02:00
Ioannis Kakavas	c621d291d2	Call ActionListener.onResponse exactly once (#61584 ) (#61682 ) Under specific circumstances we would call onResponse twice, which led to unexpected behavior.	2020-08-30 16:47:09 +03:00
Lee Hinman	1bfebd54ea	[7.x] Allocate newly created indices on data_hot tier nodes (#61342 ) (#61650 ) This commit adds the functionality to allocate newly created indices on nodes in the "hot" tier by default when they are created. This does not break existing behavior, as nodes with the `data` role are considered to be part of the hot tier. Users that separate their deployments by using the `data_hot` (and `data_warm`, `data_cold`, `data_frozen`) roles will have their data allocated on the hot tier nodes now by default. This change is a little more complicated than changing the default value for `index.routing.allocation.include._tier` from null to "data_hot". Instead, this adds the ability to have a plugin inject a setting into the builder for a newly created index. This has the benefit of allowing this setting to be visible as part of the settings when retrieving the index, for example: ``` // Create an index PUT /eggplant // Get an index GET /eggplant?flat_settings ``` Returns the default settings now of: ```json { "eggplant" : { "aliases" : { }, "mappings" : { }, "settings" : { "index.creation_date" : "1597855465598", "index.number_of_replicas" : "1", "index.number_of_shards" : "1", "index.provided_name" : "eggplant", "index.routing.allocation.include._tier" : "data_hot", "index.uuid" : "6ySG78s9RWGystRipoBFCA", "index.version.created" : "8000099" } } } ``` After the initial setting of this setting, it can be treated like any other index level setting. This new setting is not set on a new index if any of the following is true: - The index is created with an `index.routing.allocation.include.<anything>` setting - The index is created with an `index.routing.allocation.exclude.<anything>` setting - The index is created with an `index.routing.allocation.require.<anything>` setting - The index is created with a null `index.routing.allocation.include._tier` value - The index was created from an existing source metadata (shrink, clone, split, etc) Relates to #60848	2020-08-27 13:41:12 -06:00
Albert Zaharovits	1cb97a2c4f	Relax the index access control check for scroll searches (#61446 ) The check introduced by #60640 for scroll searches, in which we log if the index access control before the query and fetch phases differs from when the scroll context is created, is too strict, leading to spurious warning log messages. The check verifies instance equality but this assumes that the fetch phase is executed in the same thread context as the scroll context validation. However, this is not true if the scroll search is executed cross-cluster, and even for local scroll searches it is an unfounded assumption. The check is hence reduced to a null check for the index access. The fact that the access control is suitable given the indices that are actually accessed (by the scroll) will be done in a follow-up, after we better regulate the creation of index access controls in general.	2020-08-27 21:16:01 +03:00
Luca Cavanna	f769821bc8	Pass SearchLookup supplier through to fielddataBuilder (#61430 ) (#61638 ) Runtime fields need to have a SearchLookup available, when building their fielddata implementations, so that they can look up other fields, runtime or not. To achieve that, we add a Supplier<SearchLookup> argument to the existing MappedFieldType#fielddataBuilder method. As we introduce the ability to look up other fields while building fielddata for mapped fields, we implicitly add the ability for a field to require other fields. This requires some protection mechanism that detects dependency cycles to prevent stack overflow errors. With this commit we also introduce detection for cycles, as well as a limit on the depth of the references for a runtime field. Note that we also plan on introducing cycles detection at compile time, so the runtime cycles detection is a last resort to prevent stack overflow errors but we hope that we can reject runtime fields from being registered in the mappings when they create a cycle in their definition. Note that this commit does not introduce any production implementation of runtime fields, but is rather a pre-requisite to merge the runtime fields feature branch. This is a breaking change for MapperPlugins that plug in a mapper, as the signature of MappedFieldType#fielddataBuilder changes from taking a single argument (the index name), to also accept a Supplier<SearchLookup>. Relates to #59332 Co-authored-by: Nik Everett <nik9000@gmail.com>	2020-08-27 18:09:56 +02:00
Nik Everett	5a83e89a2b	Migrate histogram field test (#61602 ) (#61632 ) Replaces the superclass of the test for `HistogramFieldMapperTests` with one that doesn't extend `ESSingleNodeTestCase` so we don't depend on the entire world to test the field mapper. Continues #61301.	2020-08-27 11:08:19 -04:00
David Turner	c89fb8b9fa	Avoid listener call under SparseFileTracker#mutex (#61626 ) Today we sometimes notify a listener of completion while holding `SparseFileTracker#mutex`. This commit move all such calls out from under the mutex and adds assertions that the mutex is not held in the listener. Closes #61520	2020-08-27 15:39:38 +01:00
David Kyle	49a5afc6c1	[ML] Increase wait for templates timeout in tests (#61623 ) (#61628 )	2020-08-27 12:57:12 +01:00
David Kyle	25e811ced7	Rewrite Inference yml tests for better clean up (#61180 ) (#61555 ) Inference processors asynchronously usage write stats to the .ml-stats index after they used. In tests the write can leak into the next test causing failures depending on which test follows. This change waits for the usage stats docs to be written at the end of the test	2020-08-27 11:16:26 +01:00
David Turner	f6055dc9b2	Suppress noisy SSL exceptions (#61359 ) If a TLS-protected connection closes unexpectedly then today we often emit a `WARN` log, typically one of the following: io.netty.handler.codec.DecoderException: javax.net.ssl.SSLHandshakeException: Insufficient buffer remaining for AEAD cipher fragment (2). Needs to be more than tag size (16) io.netty.handler.codec.DecoderException: javax.net.ssl.SSLException: Received close_notify during handshake We typically only report unexpectedly-closed connections at `DEBUG` level, but these two messages don't follow that rule and generate a lot of noise as a result. This commit adjusts the logging to report these two exceptions at `DEBUG` level only.	2020-08-27 10:59:39 +01:00
David Turner	b866aaf81c	Use int for number of parts in blob store (#61618 ) Today we use `long` to represent the number of parts of a blob. There's no need for this extra range, it forces us to do some casting elsewhere, and indeed when snapshotting we iterate over the parts using an `int` which would be an infinite loop in case of overflow anyway: for (int i = 0; i < fileInfo.numberOfParts(); i++) { This commit changes the representation of the number of parts of a blob to an `int`.	2020-08-27 10:54:03 +01:00
Ioannis Kakavas	aac9eb6b64	Kerberos doc kibana link (#61466 ) (#61619 ) Add a note in Kerberos documentation that Kibana requires a configuration change too, and link to that documentation page.	2020-08-27 12:42:52 +03:00
Ioannis Kakavas	3640ff1ff2	Add SAML AuthN request signing tests (#61582 ) - Add a unit test for our signing code - Change SAML IT to use signed authentication requests for Shibboleth to consume Backport of #48444	2020-08-27 10:41:56 +03:00
David Turner	5df74cc888	Replace Math.toIntExact with toIntBytes (#61604 ) We convert longs to ints using `Math.toIntExact` in places where we're sure there will be no overflow, but this doesn't explain the intent of these conversions very well. This commit introduces a dedicated method for these conversions, and adds an assertion that we never overflow.	2020-08-27 08:28:54 +01:00
David Turner	e14d9c9514	Introduce cache index for searchable snapshots (#61595 ) If a searchable snapshot shard fails (e.g. its node leaves the cluster) we want to be able to start it up again on a different node as quickly as possible to avoid unnecessarily blocking or failing searches. It isn't feasible to fully restore such shards in an acceptably short time. In particular we would like to be able to deal with the `can_match` phase of a search ASAP so that we can skip unnecessary waiting on shards that may still be warming up but which are not required for the search. This commit solves this problem by introducing a system index that holds much of the data required to start a shard. Today() this means it holds the contents of every file with size <8kB, and the first 4kB of every other file in the shard. This system index acts as a second-level cache, behind the first-level node-local disk cache but in front of the blob store itself. Reading chunks from the index is slower than reading them directly from disk, but faster than reading them from the blob store, and is also replicated and accessible to all nodes in the cluster. () the exact heuristics for what we should put into the system index are still under investigation and may change in future. This second-level cache is populated when we attempt to read a chunk which is missing from both levels of cache and must therefore be read from the blob store. We also introduce `SearchableSnapshotsBlobStoreCacheIntegTests` which verify that we do not hit the blob store more than necessary when starting up a shard that we've seen before, whether due to a node restart or because a snapshot was mounted multiple times. Backport of #60522 Co-authored-by: Tanguy Leroux <tlrx.dev@gmail.com>	2020-08-27 06:38:32 +01:00
Dimitris Athanasiou	3ed65eb418	[7.x][ML] Recover data frame extraction search from latest sort key (#61544 ) (#61572 ) If a search failure occurs during data frame extraction we catch the error and retry once. However, we retry another search that is identical to the first one. This means we will re-fetch any docs that were already processed. This may result either to training a model using duplicate data or in the case of outlier detection to an error message that the process received more records than it expected. This commit fixes this issue by tracking the latest doc's sort key and then using that in a range query in case we restart the search due to a failure. Backport of #61544 Co-authored-by: Elastic Machine <elasticmachine@users.noreply.github.com>	2020-08-26 17:54:00 +03:00
Benjamin Trent	a6e7a3d65f	[7.x] [ML] write warning if configured memory limit is too low for analytics job (#61505 ) (#61528 ) Backports the following commits to 7.x: [ML] write warning if configured memory limit is too low for analytics job (#61505) Having `_start` fail when the configured memory limit is too low can be frustrating. We should instead warn the user that their job might not run properly if their configured limit is too low. It might be that our estimate is too high, and their configured limit works just fine.	2020-08-26 10:35:38 -04:00
Przemyslaw Gomulka	9f566644af	Do not create two loggers for DeprecationLogger backport(#58435 ) (#61530 ) DeprecationLogger's constructor should not create two loggers. It was taking parent logger instance, changing its name with a .deprecation prefix and creating a new logger. Most of the time parent logger was not needed. It was causing Log4j to unnecessarily cache the unused parent logger instance. depends on #61515 backports #58435	2020-08-26 16:04:02 +02:00
Ioannis Kakavas	283eaabc71	[7.x] Refactor SamlAuthenticationIT (#57162 ) (#61568 ) Refactor the tests to not require a mock HTTP Server. This has been the cause of flakiness and removing it doesn't affect the logical coverage of this suite. The "fake UI" is now simulated by an http client that makes the necessary requests to Elasticsearch APIs.	2020-08-26 15:34:56 +03:00
Przemysław Witek	11c2710e7f	[7.x] [ML] Do not mark the DFA job as FAILED when a failure occurs after the node is shutdown (#61331 ) (#61526 )	2020-08-26 09:53:13 +02:00
Igor Motov	f70a59971a	[7.x] Add rate aggregation (#61369 ) (#61554 ) Adds a new rate aggregation that can calculate a document rate for buckets of a date_histogram. Closes #60674	2020-08-25 17:39:00 -04:00
markharwood	8b56441d2b	Search - add case insensitive support for regex queries. (#59441 ) (#61532 ) Backport to add case insensitive support for regex queries. Forks a copy of Lucene’s RegexpQuery and RegExp from Lucene master. This can be removed when 8.7 Lucene is released. Closes #59235	2020-08-25 17:18:59 +01:00
Przemyslaw Gomulka	f3f7d25316	Header warning logging refactoring backport(#55941 ) (#61515 ) Splitting DeprecationLogger into two. HeaderWarningLogger - responsible for adding a response warning headers and ThrottlingLogger - responsible for limiting the duplicated log entries for the same key (previously deprecateAndMaybeLog). Introducing A ThrottlingAndHeaderWarningLogger which is a base for other common logging usages where both response warning header and logging throttling was needed. relates #55699 relates #52369 backports #55941	2020-08-25 16:35:54 +02:00
Costin Leau	bff3c7470e	EQL: Replace SearchHit in response with Event (#61428 ) (#61522 ) The building block of the eql response is currently the SearchHit. This is a problem since it is tied to an actual search, and thus has scoring, highlighting, shard information and a lot of other things that are not relevant for EQL. This becomes a problem when doing sequence queries since the response is not generated from one search query and thus there are no SearchHits to speak of. Emulating one is not just conceptually incorrect but also problematic since most of the data is missed or made-up. As such this PR introduces a simple class, Event, that maps nicely to the terminology while hiding the ES internals (the use of SearchHit or GetResult/GetResponse depending on the API used). Fix #59764 Fix #59779 Co-authored-by: Igor Motov <igor@motovs.org> (cherry picked from commit 997376fbe6ef2894038968842f5e0635731ede65)	2020-08-25 17:32:42 +03:00
Armin Braun	f22ddf822e	Some Optimizations around BytesArray (#61183 ) (#61511 ) * Faster `equals` for `BytesArray` which is nice since with this change we use it for the search cache * Lighter `StreamInput` for `BytesArray` that should save memory and some indirection relative to the one on the abstract bytes reference * Lighter `writeTo` implementation * Build a `BytesArray` instead of a PagedBytesReference whenever possible to save indirection and memory	2020-08-25 07:13:39 +02:00
Armin Braun	806dfcfcf7	Speed up Compression Logic by Pooling Resources (#61358 ) (#61495 ) This is mostly motivated by the performance issues we are seeing around the GET mappings REST API which (in case of a large number of indices) will create decompressing streams in a hot loop which takes a significant amount of time for the system calls involved in instantiating deflaters and inflaters. Also, this fixes a leaked deflater when deserializing cached repository data.	2020-08-25 04:01:55 +02:00
David Kyle	539cf914bc	[ML] handle new model metadata stream from native process (#59725 ) (#61251 ) This adds the serialization handling for the new model_metadata object from the native process. Co-authored-by: Benjamin Trent <ben.w.trent@gmail.com>	2020-08-24 15:52:13 -04:00
James Rodewig	2b852388c5	[DOCS] Fix hyphenation for "time series" (#61472 ) (#61481 )	2020-08-24 11:18:07 -04:00
Dimitris Athanasiou	618dd65d5f	[7.x][ML] Add debug logging for field caps request during DF Analytics (#61459 ) (#61478 ) Adds debug logging for the request and the response that is getting field capabilities during a data frame analytics job. Backport of #61459	2020-08-24 18:01:30 +03:00
Dimitris Athanasiou	18ca8a6be3	[7.x][ML] Remove redundant logging for creation of annotations index (#61461 ) (#61475 ) This commit removes the log info message "Created ML annotations index and aliases". The message comes in addition to elasticsearch's index creation logging and it does not add to it. In addition, since #61107 that message may be logged multiple times. Backport of #61461	2020-08-24 17:46:29 +03:00
Yang Wang	f0615113b6	Report anonymous roles in authenticate response (#61355 ) (#61454 ) Report anonymous roles in response to "GET _security/_authenticate" API call when: * Anonymous role is enabled * User is not the anonymous user * Credentials is not an API Key	2020-08-24 14:51:44 +10:00
Yang Wang	0509465a9e	Warn about unlicensed realms if no auth token can be extracted (#61402 ) (#61419 ) There are warnings about unlicense realms when user lookup fails. This PR adds similar warnings for when no authentication token can be extracted from the request.	2020-08-22 00:04:45 +10:00
Yang Wang	cd52233b94	Include authentication type for the authenticate response (#61247 ) (#61411 ) Add a new "authentication_type" field to the response of "GET _security/_authenticate".	2020-08-21 22:59:43 +10:00
Lloyd	cb83e7011c	[Backport][API keys] Add full_name and email to API key doc and use them to populate authing User (#61354 ) (#61403 ) The API key document currently doesn't include the user's full_name or email attributes, and as a result, when those attributes return `null` when hitting `GET`ing `/_security/_authenticate`, and in the SAML response from the [IdP Plugin](https://github.com/elastic/elasticsearch/pull/54046). This changeset adds those fields to the document and extracts them to fill in the User when authenticating. They're effectively going to be a snapshot of the User from when the key was created, but this is in line with roles and metadata as well. Signed-off-by: lloydmeta <lloydmeta@gmail.com>	2020-08-21 18:32:19 +09:00
Julie Tibshirani	997c73ec17	Correct how field retrieval handles multifields and copy_to. (#61391 ) Before when a value was copied to a field through a parent field or `copy_to`, we parsed it using the `FieldMapper` from the source field. Instead we should parse it using the target `FieldMapper`. This ensures that we apply the appropriate mapping type and options to the copied value. To implement the fix cleanly, this PR refactors the value parsing strategy. Now instead of looking up values directly, field mappers produce a helper object `ValueFetcher`. The value fetchers are responsible for almost all aspects of fetching, including looking up the right paths in the _source. The PR is fairly big but each commit can be reviewed individually. Fixes #61033.	2020-08-20 15:53:35 -07:00
Alan Woodward	a3a0c63ccf	Convert NumberFieldMapper to parametrized form (#61092 ) (#61376 ) In addition, this commit converts ScaledFloatFieldMapper as it was relying on a number of static values taken from NumberFieldMapper that had changed or been removed.	2020-08-20 16:43:26 +01:00
Nik Everett	9789e6d154	Migrate some field mapper tests to ESTestCase (#61301 ) (#61346 ) This switches a few tests for field mappers from `ESSingleNodeTestCase` to `ESTestCase` because, in general, we prefer to avoid `ESSingleNodeTestCase` when we can because it is slow and "big". "Big" here means that it pulls in an entire node, making it difficult to reason about what you are testing.	2020-08-19 15:43:49 -04:00
Francisco Fernández Castaño	89a7f32100	Fix SearchableSnapshotDirectoryTests#testRecoveryStateIsKeptOpenAfterPreWarmFailure (#61343 ) The test didn't take into account the case where 0 documents are indexed into the shard, meaning that files aren't loaded during the pre-warm phase. The test injects FileSystem failures, if the snapshot doesn't contain any files, pre-warm doesn't read any files and the recovery completes normally. Closes #61295 Backport of #61317	2020-08-19 19:28:47 +02:00
Andrei Stefan	a214d7902a	EQL: make endsWith function use a wildcard ES query wherever possible (#61160 ) (#61320 ) (cherry picked from commit 55fdb7e2c74d4fae86ec40686091ecba831caeaf)	2020-08-19 14:17:55 +03:00
Andrei Stefan	a6c0670a14	EQL: make stringContains function use a wildcard ES query (#61189 ) (#61313 ) (cherry picked from commit 039a7d1c68f6f1ed0e7e6cfb86be6b04eec8051c)	2020-08-19 12:40:48 +03:00
Martijn van Groningen	d4a8172f8e	Disable ilm history in data streams rest qa module. (#61312 ) Backport of #61291 to 7.x branch. Closes #61273	2020-08-19 10:34:26 +02:00
Andrei Stefan	93abbb9057	Add data streams wildcard pattern yml test (#61269 ) (#61280 ) (cherry picked from commit e13a365eeb6d8c6a7c9a91f94f0e8e78e3fe4773)	2020-08-18 19:38:07 +03:00
Andrei Stefan	5de0f19cc3	EQL: Return sequence join keys in the original type (#61268 ) (#61282 ) (cherry picked from commit d54957d61faa0d502387656e3cace594017b6ea0)	2020-08-18 19:37:15 +03:00
Martijn van Groningen	cbf60f6c5e	Add tests that simulate new indexing strategy upgrade procedure. (#61263 ) Backport of #61082 to 7.x branch. Closes #58251	2020-08-18 17:02:29 +02:00
Andrei Stefan	ad627c7eab	Introduce ordering in the constant_keyword test for better predictibility. (#61248 ) (#61252 ) (cherry picked from commit 69193f9de8178dbaa1d8467f1686b100dd2b161c)	2020-08-18 12:17:15 +03:00
Mark Tozzi	db1df6cc30	[7.x] Remove a bunch of type boilerplate from Aggs (#60852 ) (#61031 )	2020-08-17 12:13:05 -04:00
James Rodewig	60876a0e32	[DOCS] Replace Wikipedia links with attribute (#61171 ) (#61209 )	2020-08-17 11:27:04 -04:00
Andrei Stefan	db8788e5a2	QL: wildcard field type support (#58062 ) (#61205 ) (cherry picked from commit c874e6cdd3e051ce599b50c18642de038b84105f)	2020-08-17 18:24:32 +03:00
Andrei Stefan	90e116738e	QL: add filtering query dsl support to IndexResolver (#60514 ) (#61200 ) (cherry picked from commit 7b3635d796be26af9f87d19963a8ed4ab4bbf13f)	2020-08-17 17:59:58 +03:00
Nik Everett	1b7bbafd81	Add method to make random DateFormatter pattern (backport of #60613 ) (#61213 ) Adds a method to make a random date `DateFormatter` pattern. We expect this'll be useful for runtime fields to compate their formatting with the standard date field.	2020-08-17 10:57:52 -04:00
David Kyle	ba89af544f	[7.x] Respect ML upgrade mode in TrainedModelStatsService (#61143 ) (#61187 ) When in upgrade mode the ml stats service should not write to the stats index.	2020-08-17 11:09:25 +01:00
Benjamin Trent	43fc6c34bc	Muting analytics integration tests for change new native output model_metadata (#61158 ) relates to elastic/ml-cpp#1456	2020-08-14 11:45:35 -04:00
Benjamin Trent	8f302282f4	[ML] adds new feature_processors field for data frame analytics (#60528 ) (#61148 ) feature_processors allow users to create custom features from individual document fields. These `feature_processors` are the same object as the trained model's pre_processors. They are passed to the native process and the native process then appends them to the pre_processor array in the inference model. closes https://github.com/elastic/elasticsearch/issues/59327	2020-08-14 10:32:20 -04:00
David Roberts	d1b60269f4	[ML] Ensure annotations index mappings are up to date (#61142 ) When the ML annotations index was first added, only the ML UI wrote to it, so the code to create it was designed with this in mind. Now the ML backend also creates annotations, and those mappings can change between versions. In this change: 1. The code that runs on the master node to create the annotations index if it doesn't exist but another ML index does also now ensures the mappings are up-to-date. This is good enough for the ML UI's use of the annotations index, because the upgrade order rules say that the whole Elasticsearch cluster must be upgraded prior to Kibana, so the master node should be on the newer version before Kibana tries to write an annotation with the new fields. 2. We now also check whether the annotations index exists with the correct mappings before starting an autodetect process on a node. This is necessary because ML nodes can be upgraded before the master node, so could write an annotation with the new fields before the master node knows about the new fields. Backport of #61107	2020-08-14 13:51:04 +01:00
Benjamin Trent	7c3bfb9437	[ML] updating feature_importance results mapping (#61104 ) (#61144 ) This updates the feature_importance mapping change from elastic/ml-cpp#1387	2020-08-14 08:43:10 -04:00

1 2 3 4 5 ...

6191 Commits