OpenSearch

Commit Graph

Author	SHA1	Message	Date
Dan Hermann	c53731a0cd	[7.x] Fix wrong pipeline name in debug log (#58817 ) (#61233 )	2020-08-21 11:14:01 -05:00
David Turner	078e8717ee	Stop opening PING conns to remote clusters (#61408 ) Today a remote cluster connection comprises a `PING` and a `REG` channel. The `PING` channel is only used for health checks between the elected master and the members of its own cluster, so is unused in a remote cluster connection. This commit removes this unused connection.	2020-08-21 12:21:57 +01:00
Armin Braun	e09058df1a	Serialize Get Mappings Response on Generic ThreadPool (#57937 ) (#61401 ) For large responses to the get mappings request, the serialization to XContent can be extremely slow (serializing mappings is expensive since we have to decompress and deserialize the mapping source). To not introduce instability on the IO thread handling the get mappings response we should move the serialization to the management pool. The trade-off of introducing one or two new context switches for responses that are small enough to not cause trouble on the transport thread to prevent instability in case of a large number of mappings in the cluster seems worth it.	2020-08-21 08:06:30 +02:00
Armin Braun	22509c95f8	Fix Blackholed Connection Behavior in DisruptableMockTransport (#61310 ) (#61381 ) It is not realistic to drop messages without eventually failing. To retain the coverage of long pauses this PR adjusts the blackholed behavior to fail a send after 24h (which is assumed to be longer than any timeout in the system) instead of never. Closes #61034	2020-08-21 07:54:56 +02:00
Julie Tibshirani	997c73ec17	Correct how field retrieval handles multifields and copy_to. (#61391 ) Before when a value was copied to a field through a parent field or `copy_to`, we parsed it using the `FieldMapper` from the source field. Instead we should parse it using the target `FieldMapper`. This ensures that we apply the appropriate mapping type and options to the copied value. To implement the fix cleanly, this PR refactors the value parsing strategy. Now instead of looking up values directly, field mappers produce a helper object `ValueFetcher`. The value fetchers are responsible for almost all aspects of fetching, including looking up the right paths in the _source. The PR is fairly big but each commit can be reviewed individually. Fixes #61033.	2020-08-20 15:53:35 -07:00
Julie Tibshirani	85ad328df7	Ensure fetch fields aren't dropped when rewriting search. (#61390 ) Previously we didn't retain the requested fields when performing a shallow copy of the search source. This meant that when a search was rewritten, we could drop the requested fields and fail to return them in the response.	2020-08-20 14:58:58 -07:00
Armin Braun	08dbd6d989	Optimize a few Spots on IO Loop (#60865 ) (#61380 ) Saving some cycles here and there on the IO loop: * Don't instantiate new `Runnable` to execute on `SAME` in a few spots * Don't instantiate complicated wrapped stream for empty messages * Stop instantiating almost never used `ClusterStateObserver` in two spots * Some minor cleanup and preventing pointless `Predicate<>` instantiation in transport master node action	2020-08-20 20:22:49 +02:00
Alan Woodward	a3a0c63ccf	Convert NumberFieldMapper to parametrized form (#61092 ) (#61376 ) In addition, this commit converts ScaledFloatFieldMapper as it was relying on a number of static values taken from NumberFieldMapper that had changed or been removed.	2020-08-20 16:43:26 +01:00
Nhat Nguyen	a3906dcef3	Enable cancellation for msearch requests (#61337 ) Today multi-search requests are not cancellable because we create regular tasks instead of cancellable ones for them.	2020-08-19 16:59:17 -04:00
Nik Everett	9789e6d154	Migrate some field mapper tests to ESTestCase (#61301 ) (#61346 ) This switches a few tests for field mappers from `ESSingleNodeTestCase` to `ESTestCase` because, in general, we prefer to avoid `ESSingleNodeTestCase` when we can because it is slow and "big". "Big" here means that it pulls in an entire node, making it difficult to reason about what you are testing.	2020-08-19 15:43:49 -04:00
Armin Braun	4a53ae203e	Fix SharedClusterSnapshotRestoreIT.testThrottling (#61323 ) (#61328 ) We have to set the recovery setting to `0` if we don't want throttling from recoveries. Otherwise the randomized value used for this setting in tests can lead to throttling unexpectedly. Closes #61311	2020-08-19 15:26:32 +02:00
Nik Everett	70128e022b	fix stats aggregator tests With #60683 we stopped forcing aggregating all docs using a single Aggregator which made some of our accuracy assumptions about the stats aggregator incorrect. This adds a test that does the forcing and asserts the old accuracy and adds a test without the forcing with much looser accuracy guarantees. Closes #61132	2020-08-19 08:55:43 -04:00
Alan Woodward	b1aa0d8731	Fix fieldnames field type for pre-6.1 indexes (#61322 ) The FieldNamesFieldMapper field has different behaviour for indexes created in clusters earlier than v6.1, and the code to deal with this was still using the vestigial FieldType field of FieldMapper in its indexing path. This meant that documents added after an upgrade were not correctly indexing their field names field. This commit corrects the parseCreateField method to use the default field type. Fixes #61305	2020-08-19 12:59:09 +01:00
David Turner	389f7779e7	Report more details of unobtainable ShardLock (#61255 ) Today a common reason for a `ShardLockObtainFailedException` is when a shard is removed from a node and then assigned straight back to it again before the node has had a chance to shut the previous shard instance down. For instance, this can happen if a node briefly leaves the cluster holding a primary with no in-sync replicas. The message in this case is typically as follows: obtaining shard lock timed out after 5000ms, previous lock details: [shard creation] trying to lock for [shard creation] This is pretty hard to interpret, and doesn't raise the important question: "why didn't the shard shut down sooner?" With this change we reword the message a bit, report the age of the shard lock, and adjust the details to report that the lock is held by a closing shard: obtaining shard lock for [starting shard] timed out after [5000ms], lock already held for [closing shard] with age [12345ms] Relates #38807	2020-08-19 06:36:28 +01:00
Nhat Nguyen	08b0e78ef4	Log more info when search ops higher than expected (#61108 ) We have seen a situation where the total search operations are higher than expected. Unfortunately, we did not have enough info to figure it out. This commit adds the failures to the error to provide more context and adjusts the log level in case of failure to debug.	2020-08-18 15:20:41 -04:00
Rory Hunter	bd7236cd65	Version bump for 7.9.0 release	2020-08-18 16:07:43 +01:00
Dimitrios Liappis	c870640cbd	[7.x] Introduce 6.8.13 as a version (#61198 ) Introduce version 6.8.13 to branch 7.x	2020-08-18 17:07:16 +03:00
Armin Braun	58d07b2ffc	Remove Unused ByteBufferReference (#61116 ) (#61250 ) We only work with heap byte buffers at this point and those we can and do unwrap the `byte[]` ourselves and use `BytesArray` instead of a needless level of indirection via `ByteBuffer`.	2020-08-18 10:53:40 +02:00
Armin Braun	6ffa7f0737	Fix testConcurrentSnapshotDeleteAndDeleteIndex (#61228 ) (#61249 ) There is a corner case here in which during partial snapshot the index is deleted right between starting the snapshot in the CS and the data node getting to work on it, causing the data node the fail that shard snapshot and making the snapshot `PARTIAL`. Closes #61208	2020-08-18 10:45:30 +02:00
Mark Tozzi	db1df6cc30	[7.x] Remove a bunch of type boilerplate from Aggs (#60852 ) (#61031 )	2020-08-17 12:13:05 -04:00
Nik Everett	1b7bbafd81	Add method to make random DateFormatter pattern (backport of #60613 ) (#61213 ) Adds a method to make a random date `DateFormatter` pattern. We expect this'll be useful for runtime fields to compate their formatting with the standard date field.	2020-08-17 10:57:52 -04:00
Christoph Büscher	6866396e1d	Improve 'ignore_malformed' handling for dates (#60211 ) Currently we occasionally can get ArithmeticException from parsing bad input values on 'date' fields that are passed on even if 'ignore_malformed' is set. This change adds this exception to the ones we already catch for malformed values. Closes #52634	2020-08-17 16:18:08 +02:00
David Turner	b21cb7f466	Reduce allocations when persisting cluster state (#61159 ) Today we allocate a new `byte[]` for each document written to the cluster state. Some of these documents may be quite large. We need a buffer that's at least as large as the largest document, but there's no need to use a fresh buffer for each document. With this commit we re-use the same `byte[]` much more, only allocating it afresh if we need a larger one, and using the buffer needed for one round of persistence as a hint for the size needed for the next one.	2020-08-17 13:45:31 +01:00
David Turner	f3e0c60896	Restrict testing of legacy discovery to tests (#61178 ) The 7.x branch preserves the legacy discovery mechanism from 6.x purely for running internal cluster tests; this mechanism is otherwise completely untested and unsupported. However it is still technically possible to use it outside of the test suite if you dig through the source code to work out what settings need to be set. With this change we make it impossible to use this mechanism in production. Closes #61177	2020-08-17 11:05:27 +01:00
Ryan Ernst	9cb45dafab	Unwrap transport exception when using transport client (#60801 ) The ReloadSecureSettingsIT makes requests to the reload settings apis. In 7.x, the client used from the integ test infrastructure may be a transport client. In that case, the expected exception type, and causes the test to fail (though it will hang indefinitely due to not counting down the latch, see https://github.com/elastic/elasticsearch/pull/60800). This commit adds unwrapping of the remote exception to get the underlying expected exception. closes #51546	2020-08-13 10:24:04 -07:00
Ryan Ernst	c73ab0b16f	Ensure hotthreads do not produce node failures (#61073 ) This commit adds an assertion that no sub-nodes requests within hot threads failed. relates #58842	2020-08-13 10:22:19 -07:00
Armin Braun	3143b5ea47	Stabilize testSnapshotDeleteRelocatingPrimaryIndex (#61088 ) (#61096 ) Use transport blocking to make relocation take forever instead of relying on the relocation to take long enough to clash with the snapshot. Closes #61069	2020-08-13 16:26:56 +02:00
Yannick Welsch	8e775394ac	Fix testNoMasterActionsMetadataWriteMasterBlock (#60605 ) We can't assert on the specific exception, unfortunately.	2020-08-13 10:48:16 +02:00
David Turner	c6276ae177	Fail invalid incremental cluster state writes (#61030 ) It is disastrous if we commit an incremental cluster state update without having written the full state first. We assert that this doesn't happen, but it is hard to fully test the myriad ways that things might fail in a messy production environment. Given the disastrous consequences it is worth erring on the side of caution in this area. This commit fails invalid writes even if assertions are disabled.	2020-08-12 19:46:19 +01:00
Lee Hinman	e3df64a429	[7.x] Add data tiers (hot, warm, cold, frozen) as custom node roles (#60994 ) (#61045 ) This commit adds the `data_hot`, `data_warm`, `data_cold`, and `data_frozen` node roles to the x-pack plugin. These roles are intended to be the base for the formalization of data tiers in Elasticsearch. These roles all act as data nodes (meaning shards can be allocated to them). Nodes with the existing `data` role acts as though they have all of the roles configured (it is a hot, warm, cold, and frozen node). This also includes a custom `AllocationDecider` that allows the user to configure the following settings on a cluster level: - `cluster.routing.allocation.require._tier` - `cluster.routing.allocation.include._tier` - `cluster.routing.allocation.exclude._tier` And in index settings: - `index.routing.allocation.require._tier` - `index.routing.allocation.include._tier` - `index.routing.allocation.exclude._tier` Relates to #60848	2020-08-12 11:06:23 -06:00
Alan Woodward	5b3c10c379	Fix serialization of AllFieldMapper (#61044 ) Converting AllFieldMapper to parametrized form ended up not being run through BWC testing, resulting in an incorrect implementation being committed. This commit fixes the serialization, and adds unit tests as well as unmuting the BWC test that uncovered the bug. Fixes #60986	2020-08-12 17:32:55 +01:00
Yannick Welsch	8c488de576	Gracefully handle null in checkSettingsForTerminalDeprecation Fixes a test failure after backport to 7.x	2020-08-12 18:03:52 +02:00
Yannick Welsch	25404cbe3d	Provide option to allow writes when master is down (#60605 ) Elasticsearch currently blocks writes by default when a master is unavailable. The cluster.no_master_block setting allows a user to change this behavior to also block reads when a master is unavailable. This PR introduces a way to now also still allow writes when a master is offline. Writes will continue to work as long as routing table changes are not needed (as those require the master for consistency), or if dynamic mapping updates are not required (as again, these require the master for consistency). Eventually we should switch the default of cluster.no_master_block to this new mode.	2020-08-12 16:56:45 +02:00
Yannick Welsch	6644f2283d	Do not access snapshot repo on dedicated voting-only master node (#61016 ) Today a snapshot repository verification ensures that all master-eligible and data nodes have write access to the snapshot repository (and can see each other's data) since taking a snapshot requires data nodes and the currently elected master to write to the repository. However, a dedicated voting-only master-eligible node is not a data node and will never be the elected master so we should not require it to have write access to the repository. Closes #59649	2020-08-12 16:56:45 +02:00
Yannick Welsch	af519be9cb	Ensure repo not in use for wildcard repo deletes (#60947 ) Repositories can't be unregistered when they are actively being used for snapshots or restores. Wildcard repository deletes could silently bypass the "repo in use" checks however, which is now fixed.	2020-08-12 16:38:06 +02:00
Dan Hermann	538c93c923	Adding Hit counts and Miss counts for QueryCache exposed through REST api. (#60114 ) (#60993 )	2020-08-12 08:21:09 -05:00
Alan Woodward	c81dc2b8b7	Convert KeywordFieldMapper to parametrized form (#60645 ) This makes KeywordFieldMapper extend ParametrizedFieldMapper, with explicitly defined parameters. In addition, we add a new option to Parameter, restrictedStringParam, which accepts a restricted set of string options.	2020-08-12 11:41:11 +01:00
markharwood	66098e0bf4	Search fix: query_string regex/wildcard searches not working on wildcard fields (#60959 ) (#61010 ) The Query string parser was not delegating the construction of wildcard/regex queries to the underlying field type. The wildcard field has special data structures and queries that operate on them so cannot rely on the basic regex/wildcard queries that were being used for other fields. Closes #60957	2020-08-12 10:44:52 +01:00
Armin Braun	32423a486d	Simplify and Speed up some Compression Usage (#60953 ) (#61008 ) Use thread-local buffers and deflater and inflater instances to speed up compressing and decompressing from in-memory bytes. Not manually invoking `end()` on these should be safe since their off-heap memory will eventually be reclaimed by the finalizer thread which should not be an issue for thread-locals that are not instantiated at a high frequency. This significantly reduces the amount of byte copying and object creation relative to the previous approach which had to create a fresh temporary buffer (that was then resized multiple times during operations), copied bytes out of that buffer to a freshly allocated `byte[]`, used 4k stream buffers needlessly when working with bytes that are already in arrays (`writeTo` handles efficient writing to the compression logic now) etc. Relates #57284 which should be helped by this change to some degree. Also, I expect this change to speed up mapping/template updates a little as those make heavy use of these code paths.	2020-08-12 11:06:23 +02:00
Nik Everett	ce9c5f0e46	Fix diversified sample tests The test assumed that the aggregator only ran once but we turned that off. This turns it back on.	2020-08-11 17:49:43 -04:00
Jay Modi	2fa6448a15	System index reads in separate threadpool (#60927 ) This commit introduces a new thread pool, `system_read`, which is intended for use by system indices for all read operations (get and search). The `system_read` pool is a fixed thread pool with a maximum number of threads equal to lesser of half of the available processors or 5. Given the combination of both get and read operations in this thread pool, the queue size has been set to 2000. The motivation for this change is to allow system read operations to be serviced in spite of the number of user searches. In order to avoid a significant performance hit due to pattern matching on all search requests, a new metadata flag is added to mark indices as system or non-system. Previously created system indices will have flag added to their metadata upon upgrade to a version with this capability. Additionally, this change also introduces a new class, `SystemIndices`, which encapsulates logic around system indices. Currently, the class provides a method to check if an index is a system index and a method to find a matching index descriptor given the name of an index. Relates #50251 Relates #37867 Backport of #57936	2020-08-11 12:16:34 -06:00
Julie Tibshirani	a93be8d577	Handle nested arrays in field retrieval. (#60981 ) We accept _source values with multiple levels of arrays, such as `"field": [[[1, 2]]]`. This PR ensures that field retrieval can handle nested arrays by unwrapping the arrays before parsing the values.	2020-08-11 10:22:16 -07:00
Mark Tozzi	ab8518fb5b	[7.x] Extensibility for Composite Agg #59648 (#60842 )	2020-08-11 09:14:33 -04:00
Alan Woodward	54279212cf	Make MetadataFieldMapper extend ParametrizedFieldMapper (#59847 ) (#60924 ) This commit cuts over all metadata field mappers to parametrized format.	2020-08-11 09:02:28 +01:00
Armin Braun	3e2dfc6eac	Remove GCS Bucket Exists Check (#60899 ) (#60914 ) Same as https://github.com/elastic/elasticsearch/pull/43288 for GCS. We don't need to do the bucket exists check before using the repo, that just needlessly increases the necessary permissions for using the GCS repository.	2020-08-11 09:54:27 +02:00
Julie Tibshirani	d51eae6e9f	Prevent loading 'fields' with stored fields disabled. (#60938 ) Because the 'fields' option loads from _source (which is a stored field), it is not possible to retrieve 'fields' when stored_fields are disabled. This also fixes #60912, where setting stored_fields: _none_ prevented the _ignored fields from being loaded and caused a parsing exception.	2020-08-10 15:40:27 -07:00
Nik Everett	0286d0a769	Move distance_feature query building into MFT (#60614 ) (#60846 ) This moves the `distance_feature` query building out of `DistanceFeatureQueryBuilder` and into subclasses of `MappedFieldType`. Without this we don't have a chance of supporting this for runtime fields. In general I'm not sad to see the `instanceof`s go. Co-authored-by: Elastic Machine <elasticmachine@users.noreply.github.com>	2020-08-10 16:05:17 -04:00
Julie Tibshirani	b216340f50	Make `FetchPhase` logic more readable. (#60779 ) * Factor out FieldsVisitor#postProcess call. * Swap logical order for normal and nested documents. * Extract the method createStoredFieldsVisitor.	2020-08-10 11:04:54 -07:00
Nik Everett	dfd502f9ca	Rework checking if a year is a leap year (#60585 ) (#60790 ) This way is faster, saving about 8% on the microbenchmark that rounds to the nearest month. That is in the hot path for `date_histogram` which is a very popular aggregation so it seems worth it to at least try and speed it up a little. Co-authored-by: Elastic Machine <elasticmachine@users.noreply.github.com>	2020-08-10 12:45:34 -04:00
Jim Ferenczi	f30f1f04e2	Replace AggregatorTestCase#search with AggregatorTestCase#searchAndReduce (#60816 ) This commit removes the ability to test the top level result of an aggregator before it runs the final reduce. All aggregator tests that use AggregatorTestCase#search are rewritten with AggregatorTestCase#searchAndReduce in order to ensure that we test the final output (the one sent to the end user) rather than an intermediary result that could be different. This change also removes spurious commits triggered on top of a random index writer. These commits slow down the tests and are redundant with the commits that the random index writer performs.	2020-08-10 17:23:00 +02:00
David Turner	f44c28b595	Deprecate and ignore join timeout (#60872 ) There is no point in timing out a join attempt any more once a cluster is entirely in 7.x. Timing out and retrying with the same master is pointless, and an in-flight join attempt to one master no longer blocks attempts to join other masters. This commit deprecates this unnecessary setting and removes its effect from the joining process. Relates #60873 which removes this setting in master.	2020-08-10 13:57:41 +01:00
Martijn van Groningen	64bb082f9b	Improve error message for non append-only writes that target data stream (#60874 ) Backport of #60809 to 7.x branch. Closes #60581	2020-08-10 13:18:59 +02:00
Alan Woodward	e8d9185045	Cut over IPFieldMapper to parametrized form (#60602 ) This commit makes IpFieldMapper extend ParametrizedFieldMapper. It also updates the IpFieldMapper docs to add the ignore_malformed parameter, which was not previously documented.	2020-08-10 11:01:10 +01:00
David Turner	1f49e0b9d0	Fix testRerouteOccursOnDiskPassingHighWatermark (#60869 ) Sometimes this test would refresh the disk stats so quickly that it hit the refresh rate limiter even though it was almost completely disabled. This commit allows the rate limiter to be completely disabled. Closes #60587	2020-08-10 09:39:44 +01:00
Ryan Ernst	ddcfbec569	Add assert message for multiple lines in osprobe (#60796 ) Several /proc files are expected to contain a single line. We assert on this in tests, but the contents of the file are lost and the assertion therefore lacks important information to debug why the file appeared to have multiple lines. This commit dumps the contents of the file on assertion failure. relates #59284	2020-08-06 15:53:30 -07:00
Ryan Ernst	fc38af363e	Ensure latch is counted down when assertion trips (#60800 ) The ReloadSecureSettingsIT uses latches to ensure coordination across requests to the underlying in memory cluster. However, in the case of an expected failure, if the assertion fails, the latch will never be counted down, and will cause the test to hang indefinitely. This commit ensures the latch is always counted down with a try/finally. relates #51546	2020-08-06 15:33:46 -07:00
Jim Ferenczi	98119578a1	Disable sort optimization on search collapsing (#60838 ) Collapse search queries that sort by a field can throw an ArrayStoreException due to a bug in the [sort optimization](https://github.com/elastic/elasticsearch/pull/51852) introduced in 7.7.0. Search collapsing were not supposed to be eligible for this sort optimization so this change explicitly filters them from this new feature.	2020-08-06 21:37:12 +02:00
Jim Ferenczi	14980ff97e	Fix AOOBE when setting min_doc_count to 0 in significant_terms (#60823 ) This commit fixes the computation of the subset size on empty buckets (doc count of 0). The aggregator test refactoring in #60683 revealed this bug.	2020-08-06 18:57:09 +02:00
David Turner	721198c29e	Increase logging in testRerouteOccursOnDiskPassingHighWatermark (#60817 ) Relates #60587	2020-08-06 14:08:09 +01:00
Armin Braun	a2c7991e96	Fix CompressibleBytesOutputStreamTests (#60815 ) (#60822 ) Since #60730 the `bytes` field can be `null`. This adds the missing `null` check to the test override. Closes #60814	2020-08-06 15:07:48 +02:00
David Turner	273a6f916d	AwaitsFix for #60814	2020-08-06 12:56:28 +01:00
Tim Brooks	2f76c48ea7	Propagate forceExecution when acquiring permit (#60634 ) Currently the transport replication action does not propagate the force execution parameter when acquiring the indexing permit. The logic to acquire the index permit supports force execution, so this parameter should be propagate. Fixes #60359.	2020-08-05 15:57:40 -06:00
Francisco Fernández Castaño	b4044004aa	Add recovery state tracking for Searchable Snapshots (#60751 ) This pull request adds recovery state tracking for Searchable Snapshots. In order to track recoveries for searchable snapshot backed indices, this pull request adds a new type of RecoveryState. This newRecoveryState instance is able to deal with the small differences that arise during Searchable snapshots recoveries. Those differences can be summarized as follows: - The Directory implementation that's provided by SearchableSnapshots mark the snapshot files as reused during recovery. In order to keep track of the recovery process as the cache is pre-warmed, those files shouldn't be marked as reused. - Once the shard is created, the cache starts its pre-warming phase, meaning that we should keep track of those downloads during that process and tie the recovery to this pre-warming phase. The shard is considered recovered once this pre-warming phase has finished. Backport of #60505	2020-08-05 17:41:49 +02:00
Jake Landis	f3752ba1d5	7.x suport new path for re-index java-api doc (#60319 ) This commit uses the new location for the reindex java-api documentation. Temporary files have been left behind to pacify the docs build. related #60339	2020-08-05 09:05:07 -05:00
Armin Braun	ebfb93ff26	Improve some BytesStreamOutput Usage (#60730 ) (#60736 ) * Stop redundantly creating a `0` length `ByteArray` that is never used * Add efficient way to get a minimal size copy of the bytes in a `BytesStreamOutput` * Avoid multiple redundant `byte[]` copies in search cache key creation	2020-08-05 15:51:06 +02:00
Yannick Welsch	9f6f66f156	Fail searchable snapshot shards on invalid license (#60722 ) Implements license degradation behavior for searchable snapshots. Snapshot-backed shards are failed when the license becomes invalid, and shards won't be reallocated. After valid license is put in place again, shards are allocated again.	2020-08-05 13:14:15 +02:00
Igor Motov	959690a64a	Refactor extendedBounds to use DoubleBounds (#60556 ) (#60681 ) Refactors extendedBounds to use DoubleBounds instead of 2 variables. This is a follow up for #59175	2020-08-04 16:45:47 -04:00
Alan Woodward	b3ae5d26bd	Move mapper validation to the mappers themselves (#60072 ) (#60649 ) Currently, validation of mappers (checking that cross-references are correct, limits on field name lengths and object depths, multiple definitions, etc) is performed by the MapperService. This means that any mapper-specific validation, for example that done on the CompletionFieldMapper, needs to be called specifically from core server code, and so we can't add validation to mappers that live in plugins. This commit reworks the validation framework so that mapper-specific validation is done on the Mapper itself. Mapper gets a new `validate(MappingLookup)` method (already present on `MetadataFieldMapper` and now pulled up to the parent interface), which is called from a new `DocumentMapper.validate()` method. All the validation code currently living on `MapperService` moves either to individual mapper implementations (FieldAliasMapper, CompletionFieldMapper) or into `MappingLookup`, an altered `DocumentFieldMappers` which now knows about object fields and can check for duplicate definitions, or into DocumentMapper which handles soft limit checks.	2020-08-04 14:39:20 +01:00
Armin Braun	212ce22d15	Optimize CS Persistence Stream Use (#60643 ) (#60647 ) In the metadata persistence logic we failed to override the bulk write method on the FilterOutputStream resulting in all the writes to it running byte-by-byte in a loop adding a large number of bounds checks needlessly.	2020-08-04 15:06:57 +02:00
Armin Braun	859ad761bb	Fix Broken Stream Close in writeRawValue (#60625 ) (#60644 ) Small oversight in #56078 that only showed up during backporting where a stream copy was turned from a non-closing to a closing one. Enhanced part of a test in this PR to make it show up in master also even though we practically never use this method with stream targets that actually close.	2020-08-04 13:39:52 +02:00
Armin Braun	7ae9dc2092	Unify Stream Copy Buffer Usage (#56078 ) (#60608 ) We have various ways of copying between two streams and handling thread-local buffers throughout the codebase. This commit unifies a number of them and removes buffer allocations in many spots.	2020-08-04 09:54:52 +02:00
Julie Tibshirani	f99584c6f3	Avoid reloading _source for every inner hit. (#60632 ) Previously if an inner_hits block required _ source, we would reload and parse the root document's source for every hit. This PR adds a shared SourceLookup to the inner hits context that allows inner hits to reuse parsed source if it's already available. This matches our approach for sharing the root document ID. Relates to #32818.	2020-08-03 17:12:27 -07:00
Julie Tibshirani	fc63f8224f	Simplify class hierarchy for ordinals field data. (#60606 ) This PR simplifies the hierarchy for ordinals field data classes: * Remove `AbstractIndexFieldData`, since only `AbstractIndexOrdinalsFieldData` inherits directly from it. * Make `SortedSetOrdinalsIndexFieldData` extend `AbstractIndexOrdinalsFieldData`. This lets us remove some redundant code.	2020-08-03 09:58:29 -07:00
Yannick Welsch	3409e019d2	Ignore shutdown when retrying recoveries (#60586 ) Avoids failures when shutting down a node.	2020-08-03 15:14:38 +02:00
Nik Everett	2cde43b799	Allows nanosecond resolution in search_after (backport of #60328 ) (#60426 ) Allows nanosecond resolution in search_after (#60328) This fixes `search_after` to properly parse string formatted dates that have nanosecond resolution. Closes #52424	2020-08-03 08:17:48 -04:00
David Turner	d2ddf8cd6a	Improve deserialization failure logging (#60577 ) Today when a node fails to properly deserialize a transport message with a parent task we log the following relatively uninformative message: java.lang.IllegalStateException: Message not fully read (response) for requestId [9999], handler [org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler/org.elasticsearch.transport.TransportService$ContextRestoreResponseHandler/org.elasticsearch.transport.TransportService$6@abcdefgh], error [false]; resetting In particular, the wrapping of the listener in the `TransportService` obscures all clues as to the source of the problem, e.g. the action name or the identity of the underlying listener. This commit exposes the inner listener to the logs. Also if the listener is wrapped with `ContextPreservingActionListener` then its identity is similarly hidden. This commit also exposes the wrapped listener in this case. Relates #38939	2020-08-03 11:51:01 +01:00
Armin Braun	3270cb3088	More Efficient Writes for Snapshot Shard Generations (#60458 ) (#60575 ) Same as #59905 but for shard level metadata. Since we wnat to retain the ability to do safe+atomic writes for non-uuid shard generations this PR has to create two separate write paths for both kinds of shard generations.	2020-08-03 11:11:36 +02:00
Armin Braun	204efe9387	Add Repository Setting to Disable Writing index.latest (#60448 ) (#60576 ) Writing the `index.latest` blob is unnecessary unless the contents of the repository are to be used as a URL-repository. Also, in some edge cases, the fact that `index.latest` is the only blob in the repository that regularly gets overwritten was causing compatibility issues with some backing blobstores (Azure no-overwrite policy, Hitachy S3 equivalent). => this commit changes behavior to make snapshots not fail if writing `index.latest` fails and adds a setting to disable writing `index.latest`.	2020-08-03 11:11:24 +02:00
Andrei Dan	ac258f10d6	Data streams: throw ResourceAlreadyExists exception (#60518 ) (#60536 ) For consistency reasons (and reducing the overload of IllegalArgumentException) this changes the exception thrown when trying to create a data stream that already exists. (cherry picked from commit ac2184c4614bba0f3ee377da49aea0daed98bab4) Signed-off-by: Andrei Dan <andrei.dan@elastic.co>	2020-08-01 16:31:09 +01:00
Julie Tibshirani	f1d4fd8c3e	Correct name of IndexFieldData#loadGlobalDirect. (#60492 ) It seems 'localGlobalDirect' was just a typo.	2020-07-31 10:53:21 -07:00
Jim Ferenczi	8db896d290	Fix race condition in SearchPhaseControllerTests#testPartialMergeFailure (#60488 ) This change ensures that we call the listener for partial merge failure before calling the completion listener in order to avoid race condition in tests. Closes #60446	2020-07-31 16:29:20 +02:00
Rene Groeschke	ed4b70190b	Replace immediate task creations by using task avoidance api (#60071 ) (#60504 ) - Replace immediate task creations by using task avoidance api - One step closer to #56610 - Still many tasks are created during configuration phase. Tackled in separate steps	2020-07-31 13:09:04 +02:00
Julie Tibshirani	8ac81a3447	Remove IndexFieldData#clear since it is unused. (#60475 ) This method was never called. It also seemed tricky that calling a method on `IndexFieldData` could clear the contents of a shared cache.	2020-07-30 14:07:55 -07:00
Julie Tibshirani	dfd7f226f0	Clarify SourceLookup sharing across fetch subphases. (#60484 ) The `SourceLookup` class provides access to the _source for a particular document, specified through `SourceLookup#setSegmentAndDocument`. Previously the search context contained a single `SourceLookup` that was shared between different fetch subphases. It was hard to reason about its state: is `SourceLookup` set to the expected document? Is the _source already loaded and available? Instead of using a global source lookup, the fetch hit context now provides access to a lookup that is set to load from the hit document. This refactor closes #31000, since the same `SourceLookup` is no longer shared between the 'fetch _source phase' and script execution.	2020-07-30 13:22:31 -07:00
Dan Hermann	5e5503ac28	Change severity of negative stats messages from WARN to DEBUG (#60375 ) (#60444 )	2020-07-30 06:06:13 -05:00
Armin Braun	3bf4c01d8e	Don't Allocate Redundant Pages in BigArrays (#60201 ) (#60441 ) The oversize algorithm was allocating more pages than necessary to accommodate `minTargetSize`. An example would be that a 16k page size and 15k `minTargetSize` would result in a new size of 32k (2 pages). The difference between the minimum number of necessary pages and the estimated size then keeps growing as sizes increase. I don't think there is much value in preemptively allocating pages by over-sizing aggressively since the behavior of the system is quite different from that of a single array where over-sizing avoids copying once the minimum target size is more than a single page. Relates #60173 which lead me to this when `BytesStreamOutput` would allocate a large number of never used pages during serialization of repository metadata.	2020-07-30 11:09:58 +02:00
Armin Braun	a2c49a4f02	Reduce Heap Use during Shard Snapshot (#60370 ) (#60440 ) Instances of `BlobStoreIndexShardSnapshots` can be of non-trivial size. In case of snapshotting a larger number of shards the previous execution order would lead to memory use proportional to the number of shards for these objects. With this change, the number of these objects on heap is bounded by the size of the snapshot pool (except for in the BwC format path). This PR makes it so that they are written to the repository at the earliest possible point in time so that they can be garbage collected. If shard generations are used, we can safely write these right at the beginning of the shard snapshot. If shard generations are not used we can only write them at the end of the shard snapshot after all other blobs have been written. Closes #60173	2020-07-30 10:45:00 +02:00
Igor Motov	00a1949852	Streamline GeoJSON to map serialization (#60413 ) (#60429 ) Optimizes GeoJSON to map serialization when retrieving spatial data through fields. Closes #60259	2020-07-29 17:56:56 -04:00
Julie Tibshirani	5359417ec3	Minor clean-up around search highlight context. (#60422 ) * Rename SearchContextHighlight -> SearchHighlightContext. * Rename HighlighterContext to FieldHighlightContext. * Make the search highlight context immutable. * Avoid storing SearchHighlightContext on HighlighterContext.	2020-07-29 11:39:17 -07:00
Tim Brooks	85fdf959ad	Add configured indexing memory limit to node stats (#60414 ) This commit adds the configured memory limit to the node stats API.	2020-07-29 12:28:21 -06:00
Nhat Nguyen	9d4a64e749	Allow CCR on nodes with legacy roles only (#60093 ) CCR will stop functioning if the master node is on 7.8, but data nodes are before that version because the master node considers that all data nodes do not have the remote cluster client role. This commit allows CCR work on data nodes with legacy roles only. Relates #54146 Relates #59375	2020-07-29 10:57:31 -04:00
Armin Braun	8429b4ace8	Fix Queued Snapshot Deletes After Finalization Failure (#60285 ) (#60379 ) This fixes the behavior of the snapshot state machine in the following edge case: 1. Snapshot is running 2. Delete/abort for the snapshot is started 3. Snapshot fails to finalize We were not removing the failed snapshot id from the list of snapshots to delete in the delete. This lead to an error in the repository, which throws if we try to delete a non-existing snapshot. This commmit updates the deletions in progress by removing the failed snapshot id. The fact that this could lead to snapshot delete entries without any snapshot ids is not optimized on purpose because it allows for another attempt at writing clean `RepositoryData` and will run basic cleanup on the repository (root level blobs and stale indices) and thus bring the repository back into a clean state after a failed finalization. Closes #60274	2020-07-29 15:54:18 +02:00
Armin Braun	381cec2ba9	Fix ConcurrentSnapshotsIT.testMasterFailOverWithQueuedDeletes (#60307 ) (#60376 ) The test assumed that the master fail-over would always work out as a single step. This is not guaranteed however and we can randomly see master failing over twice, in which case the transport listener will be failed on the node that stops being leader and we have to catch an exception for the deletes as well just like we do for the snapshot. Closes #60262	2020-07-29 15:54:00 +02:00
Armin Braun	0778274b72	Fix IPV6 Scope Id in InetAddressesTests (#60368 ) (#60369 ) Follow up to #60360, turns out at times the name of an interface that isn't loopback is not a valid scope id.	2020-07-29 13:16:12 +02:00
Armin Braun	1f6a3765e4	Fix NPE in SnapshotsInProgress Constructor (#60355 ) Merge oversight between cleanups that removed `null` for `shards` and this corner case spot of no indices in a snapshot. Closes #60330	2020-07-29 10:47:28 +02:00
Armin Braun	4307a45153	Fix IPV6 Scope ID Test (#60360 ) (#60363 ) Use real scope id from first available interface instead of `lo` which might not exist on non-Linux platforms. Closes #60332	2020-07-29 09:55:37 +02:00
Armin Braun	753fd4f6bc	Cleanup and optimize More Serialization Spots (#59959 ) (#60331 ) Same as #59626 for a few more spots.	2020-07-29 07:20:44 +02:00
Zachary Tong	e3d85feecd	Mute testForStringIPv6WithScopeIdInput test Tracking issue: https://github.com/elastic/elasticsearch/issues/60332	2020-07-28 15:05:19 -04:00
Igor Motov	0dd53b76bd	Add aggregation list to node info (#60074 ) (#60256 ) Adds a full list of supported aggregations to the node info API. This list will be used in transform tests and telemetry mapping tests that will be added as follow-up PRs. Fixes #59774	2020-07-28 14:06:12 -04:00
Julie Tibshirani	c7bfb5de41	Add search `fields` parameter to support high-level field retrieval. (#60258 ) This feature adds a new `fields` parameter to the search request, which consults both the document `_source` and the mappings to fetch fields in a consistent way. The PR merges the `field-retrieval` feature branch. Addresses #49028 and #55363.	2020-07-28 10:58:20 -07:00

1 2 3 4 5 ...

5275 Commits