OpenSearch

Commit Graph

Author	SHA1	Message	Date
Tanguy Leroux	55a879ee8d	Align behavior or HDR percentiles iterator with percentile() method (#24206 )	2017-04-20 12:37:33 +02:00
Nik Everett	caf376c8af	Start building analysis-common module (#23614 ) Start moving built in analysis components into the new analysis-common module. The goal of this project is: 1. Remove core's dependency on lucene-analyzers-common.jar which should shrink the dependencies for transport client and high level rest client. 2. Prove that analysis plugins can do all the "built in" things by moving all "built in" behavior to a plugin. 3. Force tests not to depend on any oddball analyzer behavior. If tests need anything more than the standard analyzer they can use the mock analyzer provided by Lucene's test infrastructure.	2017-04-19 18:51:34 -04:00
Jason Tedor	4796557a30	Add primary term to doc write response This commit adds the primary term to the doc write response. Relates #24171	2017-04-19 14:44:22 -04:00
Ryan Ernst	c7e9231a86	Plugins: Remove leniency for missing plugins dir (#24173 ) This leniency was left in after plugin installer refactoring for 2.0 because some tests still relied on it. However, the need for this leniency no longer exists.	2017-04-19 09:09:34 -07:00
Christoph Büscher	a9657a5a09	Add BucketMetricValue interface (#24188 ) Unlike other implementations of InternalNumericMetricsAggregation.SingleValue, the InternalBucketMetricValue aggregation currently doesn't implement a specialized interface that exposes the `keys()` method. This change adds this so that clients can access the keys via the interface.	2017-04-19 16:27:33 +02:00
Jim Ferenczi	f05af0a382	Enable index-time sorting (#24055 ) This change adds an index setting to define how the documents should be sorted inside each Segment. It allows any numeric, date, boolean or keyword field inside a mapping to be used to sort the index on disk. It is not allowed to use a `nested` fields inside an index that defines an index sorting since `nested` fields relies on the original sort of the index. This change does not add early termination capabilities in the search layer. This will be added in a follow up. Relates #6720	2017-04-19 14:36:11 +02:00
Boaz Leskes	8758c541b3	ElectMasterService.hasEnoughMasterNodes should return false if no masters were found This is a regression introduced in #20063	2017-04-19 09:52:06 +02:00
Tanguy Leroux	741c031384	[Test] Add unit tests for InternalHDRPercentilesTests (#24157 ) Related to #22278	2017-04-19 09:37:01 +02:00
Areek Zillur	4f773e2dbb	Replicate write failures (#23314 ) * Replicate write failures Currently, when a primary write operation fails after generating a sequence number, the failure is not communicated to the replicas. Ideally, every operation which generates a sequence number on primary should be recorded in all replicas. In this change, a sequence number is associated with write operation failure. When a failure with an assinged seqence number arrives at a replica, the failure cause and sequence number is recorded in the translog and the sequence number is marked as completed via executing `Engine.noOp` on the replica engine. * use zlong to serialize seq_no * Incorporate feedback * track write failures in translog as a noop in primary * Add tests for replicating write failures. Test that document failure (w/ seq no generated) are recorded as no-op in the translog for primary and replica shards * Update to master * update shouldExecuteOnReplica comment * rename indexshard noop to markSeqNoAsNoOp * remove redundant conditional * Consolidate possible replica action for bulk item request depanding on it's primary execution * remove bulk shard result abstraction * fix failure handling logic for bwc * add more tests * minor fix * cleanup * incorporate feedback * incorporate feedback * add assert to remove handling noop primary response when 5.0 nodes are not supported	2017-04-19 01:23:54 -04:00
Jason Tedor	9e0ebc5965	Rename variable in translog simple commit test This commit renames a variable for clarity in the translog simple commit test.	2017-04-18 23:43:25 -04:00
Jason Tedor	20181dd0ad	Strengthen translog commit with open view test This commit strengthens an assertion in the translog commit with open view test.	2017-04-18 23:41:55 -04:00
Jason Tedor	180d1f2219	Stronger check in translog prepare and commit test This commit strengthens an assertion in the translog prepare commit and commit test.	2017-04-18 23:37:54 -04:00
Jason Tedor	23b224a5a9	Fix translog prepare commit and commit test This test was terribly, horribly, no goodly, and badly broken it's amazing it ever passed so this commit fixes it.	2017-04-18 23:32:47 -04:00
Boaz Leskes	edff30f82a	Engine: store maxUnsafeAutoIdTimestamp in commit (#24149 ) The `maxUnsafeAutoIdTimestamp` timestamp is a safety marker guaranteeing that no retried-indexing operation with a higher auto gen id timestamp was process by the engine. This allows us to safely process documents without checking if they were seen before. Currently this property is maintained in memory and is handed off from the primary to any replica during the recovery process. This commit takes a more natural approach and stores it in the lucene commit, using the same semantics (no retry op with a higher time stamp is part of this commit). This means that the knowledge is transferred during the file copy and also means that we don't need to worry about crazy situations where an original append only request arrives at the engine after a retry was processed and the engine was restarted.	2017-04-18 20:11:32 +02:00
Simon Willnauer	ab9884b2e9	Remove leniency when merging fetched hits in a search response phase (#24158 ) Today when we merge hits we have a hard check to prevent AIOOB exceptions that simply skips an expected search hit. This can only happen if there is a bug in the code which should be turned into a hard exception or an assertion triggered. This change adds an assertion an removes the lenient check for the fetched hits.	2017-04-18 17:19:57 +02:00
Tanguy Leroux	829dd068d6	[Test] Use appropriate DocValueFormats in Aggregations tests (#24155 ) Some aggregations (like Min, Max etc) use a wrong DocValueFormat in tests (like IP or GeoHash). We should not test aggregations that expect a numeric value with a DocValueFormat like IP. Such wrong DocValueFormat can also prevent the aggregation to be rendered as ToXContent, and this will be an issue for the High Level Rest Client tests which expect to be able to parse back aggregations.	2017-04-18 17:03:32 +02:00
Christoph Büscher	8f540346a9	Tests: Fixing typo in class name of InternalGlobalTests Renaming from InternalGlogbalTests -> InternalGlobalTests	2017-04-18 16:27:15 +02:00
Adrien Grand	4632661bc7	Upgrade to a Lucene 7 snapshot (#24089 ) We want to upgrade to Lucene 7 ahead of time in order to be able to check whether it causes any trouble to Elasticsearch before Lucene 7.0 gets released. From a user perspective, the main benefit of this upgrade is the enhanced support for sparse fields, whose resource consumption is now function of the number of docs that have a value rather than the total number of docs in the index. Some notes about the change: - it includes the deprecation of the `disable_coord` parameter of the `bool` and `common_terms` queries: Lucene has removed support for coord factors - it includes the deprecation of the `index.similarity.base` expert setting, since it was only useful to configure coords and query norms, which have both been removed - two tests have been marked with `@AwaitsFix` because of #23966, which we intend to address after the merge	2017-04-18 15:17:21 +02:00
Tanguy Leroux	f217eb8ad8	Merge Percentile class with interface (#24154 ) This commit merges the Percentile interface with the InternalPercentile class, as we don't need to maintain both.	2017-04-18 14:47:18 +02:00
Martijn van Groningen	edada2581e	[TEST] Added unittests for InternalSampler	2017-04-18 14:31:58 +02:00
Yannick Welsch	0b2cb68f6f	[TEST] Randomly add and remove no_master blocks in IndicesClusterStateServiceRandomUpdatesTests Checks that IndicesClusterStateService stays consistent with incoming cluster states that contain no_master blocks (especially discovery.zen.no_master_block=all which disables state persistence). In particular this checks that active shards which have no in-memory data structures on a node are failed.	2017-04-18 14:27:54 +02:00
Martijn van Groningen	ac41fb2c4a	[TEST] Added test for GeoCentroidAggregator and made constructors of GeoCentroidAggregator, GeoCentroidAggregatorFactory and InternalGeoCentroid package protected.	2017-04-18 13:54:31 +02:00
Tanguy Leroux	81dbdb239f	[Test] Add unit tests for InternalTDigestPercentilesTests (#24090 )	2017-04-18 09:48:35 +02:00
Chris Earle	12c8423ec9	Warn on not enough masters during election (#20063 ) This changes the trace level logging to warn, and adds the needed number to the message as well. My fear is that it may get noisy, but this is an issue that you want to be noisy.	2017-04-17 22:18:28 -04:00
Jason Tedor	34eda1a1a8	Do not set path.data in environment if not set When preparing the final settings in the environment, we unconditionally set path.data even if path.data was not explicitly set. This confounds detection for whether or not path.data was explicitly set, and this is trappy. This commit adds logic to only set path.data in the final settings if path.data was explicitly set, and provides a test case that fails without this logic. Relates #24132	2017-04-17 10:43:13 -04:00
Jason Tedor	f7ebe9d18f	Preserve multiple translog generations Today when a flush is performed, the translog is committed and if there are no outstanding views, only the current translog generation is preserved. Yet for the purpose of sequence numbers, we need stronger guarantees than this. This commit migrates the preservation of translog generations to keep the minimum generation that would be needed to recover after the local checkpoint. Relates #24015	2017-04-17 08:51:54 -04:00
Jason Tedor	8033c576b7	Detect remnants of path.data/default.path.data bug In Elasticsearch 5.3.0 a bug was introduced in the merging of default settings when the target setting existed as an array. When this bug concerns path.data and default.path.data, we ended up in a situation where the paths specified in both settings would be used to write index data. Since our packaging sets default.path.data, users that configure multiple data paths via an array and use the packaging are subject to having shards land in paths in default.path.data when that is very likely not what they intended. This commit is an attempt to rectify this situation. If path.data and default.path.data are configured, we check for the presence of indices there. If we find any, we log messages explaining the situation and fail the node. Relates #24099	2017-04-17 07:03:46 -04:00
jaymode	a8be0a5836	Cat APIs should not close the stream obtained from the channel The cat APIs and rest tables would obtain a stream from the RestChannel, which happened to be a ReleasableBytesStreamOutput. These APIs used the stream to write content to, closed the stream, and then tried to send a response. After #23941 was merged, closing the stream meant that the bytes were released for use elsewhere. This caused occasional corruption of the response when the bytes were used prior to the response being sent. This commit changes these two usages to wrap the stream obtained from the channel in a flush on close stream so that the bytes are still reserved until the message is sent.	2017-04-15 14:57:00 -04:00
Jason Tedor	cd8e059885	Do not produce empty IDs in simple versioning test Empty IDs are rejected during indexing, so we should not randomly produce them during tests. This commit modifies the simple versioning tests to no longer produce empty IDs.	2017-04-15 12:15:45 -04:00
Jason Tedor	972bdc09ee	Reject empty IDs When indexing a document via the bulk API where IDs can be explicitly specified, we currently accept an empty ID. This is problematic because such a document can not be obtained via the get API. Instead, we should rejected these requets as accepting them could be a dangerous form of leniency. Additionally, we already have a way of specifying auto-generated IDs and that is to not explicitly specify an ID so we do not need a second way. This commit rejects the individual requests where ID is specified but empty. Relates #24118	2017-04-15 10:36:03 -04:00
Boaz Leskes	ecf81688fb	Use sequence numbers to identify out of order delivery in replicas & recovery (#24060 ) Internal indexing requests in Elasticsearch may be processed out of order and repeatedly. This is important during recovery and due to concurrency in replicating requests between primary and replicas. As such, a replica/recovering shard needs to be able to identify that an incoming request contains information that is old and thus need not be processed. The current logic is based on external version. This is sadly not sufficient. This PR moves the logic to rely on sequences numbers and primary terms which give the semantics we need. Relates to #10708	2017-04-14 21:46:17 +02:00
Jason Tedor	09efdc3151	Improve performance of extracting warning value When building headers for a REST response, we de-duplicate the warning headers based on the actual warning value. The current implementation of this uses a capturing regular expression that is prone to excessive backtracking. In cases a request involves a large number of warnings, this extraction can be a severe performance penalty. An example where this can arise is a bulk indexing request that utilizes a deprecated feature (e.g., using deprecated forms of boolean values). This commit is an attempt to address this performance regression. We already know the format of the warning header, so we do not need to use a regular expression to parse it but rather can parse it by hand to extract the warning value. This gains back the vast majority of the performance lost due to the usage of a deprecated feature. There is still a performance loss due to logging the deprecation message but we do not address that concern in this commit. Relates #24114	2017-04-14 12:18:00 -04:00
Jay Modi	30ab8739a6	Closing a ReleasableBytesStreamOutput closes the underlying BigArray (#23941 ) This commit makes closing a ReleasableBytesStreamOutput release the underlying BigArray so that we can use try-with-resources with these streams and avoid leaking memory by not returning the BigArray. As part of this change, the ReleasableBytesStreamOutput adds protection to only release the BigArray once. In order to make some of the changes cleaner, the ReleasableBytesStream interface has been removed. The BytesStream interface is changed to a abstract class so that we can use it as a useable return type for a new method, Streams#flushOnCloseStream. This new method wraps a given stream and overrides the close method so that the stream is simply flushed and not closed. This behavior is used in the TcpTransport when compression is used with a ReleasableBytesStreamOutput as we need to close the compressed stream to ensure all of the data is written from this stream. Closing the compressed stream will try to close the underlying stream but we only want to flush so that all of the written bytes are available. Additionally, an error message method added in the BytesRestResponse did not use a builder provided by the channel and instead created its own JSON builder. This changes that method to use the channel builder and in turn the bytes stream output that is managed by the channel. Note, this commit differs from `6bfecdf921` in that it updates ReleasableBytesStreamOutput to handle the case of the BigArray decreasing in size, which changes the reference to the BigArray. When the reference is changed, the releasable needs to be updated otherwise there could be a leak of bytes and corruption of data in unrelated streams. This reverts commit `afd45c1432`, which reverted #23572.	2017-04-14 10:50:31 -04:00
Yannick Welsch	e3aa2a89f9	[TEST] Wait in OldIndexBackwardsCompatibilityIT for cluster to be fully initialized There are test failures that suggest that the import of dangling indices is happening too early, before the dangling indices are ready to be consumed. This commit adds an ensureGreen() at the end of cluster initialization to make sure that no cluster state updates are happening while the dangling indices are prepared on-disk.	2017-04-14 11:02:55 +02:00
Ali Beyad	5e54c0261a	[TEST] fixes InternalTopHitsTests test to initialize the SearchHits maxScore to Float.NaN if there is no max score, as that is what Lucene's TopDocs does	2017-04-13 18:27:42 -04:00
Igor Motov	cce321a560	Task Management: Make TaskInfo parsing forwards compatible (#24073 ) TaskInfo is stored as a part of TaskResult and therefore can be read by nodes with an older version. If we add any additional information to TaskInfo (for #23250, for example), nodes with an older version should be able to ignore it, otherwise they will not be able to read TaskResults stored by newer nodes.	2017-04-13 16:16:01 -04:00
Tim Brooks	ffaac5a08a	Simplify BulkProcessor handling and retry logic (#24051 ) This commit collapses the SyncBulkRequestHandler and AsyncBulkRequestHandler into a single BulkRequestHandler. The new handler executes a bulk request and awaits for the completion if the BulkProcessor was configured with a concurrentRequests setting of 0. Otherwise the execution happens asynchronously. As part of this change the Retry class has been refactored. withSyncBackoff and withAsyncBackoff have been replaced with two versions of withBackoff. One method takes a listener that will be called on completion. The other method returns a future that will been complete on request completion.	2017-04-13 14:48:52 -05:00
Jason Tedor	99e0268e0a	Remove support for default settings Today Elasticsearch allows default settings to be used only if the actual setting is not set. These settings are trappy, and the complexity invites bugs. This commit removes support for default settings with the exception of default.path.data, default.path.conf, and default.path.logs which are maintainted to support packaging. A follow-up will remove support for these as well. Relates #24093	2017-04-13 14:25:45 -04:00
Jason Tedor	32b2caad42	Correct handling of default and array settings In Elasticsearch 5.3.0 a bug was introduced in the merging of default settings when the target setting existed as an array. This arose due to the fact that when a target setting is an array, the setting key is broken into key.0, key.1, ..., key.n, one for each element of the array. When settings are replaced by default.key, we are looking for the target key but not the target key.0. This leads to key, and key.0, ..., key.n being present in the constructed settings object. This commit addresses two issues here. The first is that we fix the merging of the keys so that when we try to merge default.key, we also check for the presence of the flattened keys. The second is that when we try to get a setting value as an array from a settings object, we check whether or not the backing map contains the top-level key as well as the flattened keys. This latter check would have caught the first bug. For kicks, we add some tests. Relates #24074	2017-04-13 06:34:58 -04:00
Ryan Ernst	fb3a281755	Build: Switch jna dependency to an elastic version (#24081 ) This new version of jna is rebuilt from the official release of jna, but with native libs linked against older glibc in order to support all platforms elasticsearch supports. closes #23640	2017-04-13 00:17:50 -07:00
Boaz Leskes	215a9b2df9	fix CategoryContextMappingTests compilation bugs	2017-04-13 09:15:10 +02:00
Boaz Leskes	342e745fc7	testConcurrentGetAndSetOnPrimary - fix a race condition between indexing and updating value map Currently the map can be lagging behind what's actually in lucene causes assertions about adding/removing values to fail	2017-04-13 09:03:09 +02:00
Nilabh Sagar	ec421974b9	Allow different data types for category in Context suggester (#23491 ) The "category" in context suggester could be String, Number or Boolean. However with the changes in version 5 this is failing and only accepting String. This will have problem for existing users of Elasticsearch if they choose to migrate to higher version; as their existing Mapping and query will fail as mentioned in a bug #22358 This PR fixes the above mentioned issue and allows user to migrate seamlessly. Closes #22358	2017-04-12 23:43:29 -07:00
Ryan Ernst	c19044ddf6	Restrict build info loading to ES jar, not any jar (#24049 ) This change makes the build info initialization only try to load a jar manifest if it is the elasticsearch jar. Anything else (eg a repackaged ES for use of transport client in an uber jar) will contain "Unknown" for the build info as it does for tests currently. fixes #21955	2017-04-12 23:22:43 -07:00
Jason Tedor	12b46bdbc4	Remove more hidden file leniency from plugins This commit removes one more instance of leniency from the plugin service which skips hidden files in the plugins directory. Relates #23982	2017-04-12 22:23:42 -04:00
Jason Tedor	edd16fa27e	Register error listener in evil logger tests This test needs an error listener registered since we configure logging here.	2017-04-12 21:23:05 -04:00
Jason Tedor	a1c2fe9e3a	Detect using logging before configuration It can easily happen that we touch a logger before logging is configured due to chains of static intializers and other such scenarios. This commit adds detection for this mechanism that will fail startup if we touch a logger before logging is configured. This is a bug that will cause builds to fail. Relates #24076	2017-04-12 21:13:08 -04:00
Nik Everett	31c8903492	Add version constant for 5.5 (#24075 ) This is required in master now that #24071 is in or else we fail during BWC testing because the 5.x branch contains 5.5 but the build thinks it should contain 5.4.	2017-04-12 16:29:47 -04:00
Zachary Tong	1fd50bc54d	Add unit tests for NestedAggregator (#24054 ) Add unit tests for NestedAggregator, change class visibilities Relates to #22278	2017-04-12 15:59:51 -04:00
Nik Everett	e99f90fb46	Add more debugging information to rethrottles I'm still trying to track down failures like: https://elasticsearch-ci.elastic.co/job/elastic+elasticsearch+master+dockeralpine-periodic/1180/console It looks like a task is hanging but I'm not sure why. So this adds more logging for next time.	2017-04-12 08:37:31 -04:00

1 2 3 4 5 ...

7863 Commits