The Shard Had Two Identities: Elasticsearch CVE-2026-103009
The request said logs-public. The shard identifier carried the UUID of restricted-data.
A policy check saw the first name and allowed the operation. The storage path followed the second identity. That small disagreement, not a missing login or an overpowered role, is the center of CVE-2026-103009 and the CWE-639 authorization-bypass class.[6]
This is a reconstructed, authorized-analysis scenario based on Elastic’s advisory and public patch. It is not a claim about a real intrusion. The vulnerability affected Elasticsearch fulfilling clusters using the Remote Cluster Security (RCS) 2.0 model. It was not reachable through the REST API, and clusters that did not expose the remote-cluster transport interface for cross-cluster search were outside the affected configuration.[1]
The interesting lesson is broader than Elasticsearch: when authorization and execution resolve the same object from different attacker-influenced identifiers, both checks can be locally correct while the system is globally wrong.
The boundary was an API key, not a superuser
RCS 2.0 lets a local cluster authenticate to a remote cluster with a cross-cluster API key. The remote administrator defines which indices that key can search or replicate. Local users can be restricted further, but they should never gain more access than the remote API key grants. Cross-cluster traffic uses a dedicated remote-cluster interface, normally TCP 9443, rather than the ordinary REST path.[3]
That architecture creates two authorization layers and a long-lived transport relationship. It also makes the API key a capability: possession is necessary, but its index scope is supposed to be the hard ceiling.
CVE-2026-103009 broke that ceiling for shard-level requests. Elastic describes two independently supplied attributes inside one request. One attribute was used during authorization; another selected the shard actually accessed. An attacker holding a key for one index could make those attributes disagree and cause work against another index.[1]
The preconditions matter:
- the target must be a fulfilling cluster using RCS 2.0 for cross-cluster search;
- the remote-cluster transport interface must be reachable;
- the attacker needs a valid cross-cluster API key with access to at least one index;
- exploitation requires the identifier of a shard belonging to a different index.
That last item is operationally important. The advisory does not describe a universal method for discovering arbitrary shard UUIDs. In a real assessment, I would treat UUID acquisition as a separate hypothesis and look for exposure through logs, diagnostics, snapshots, cluster metadata, monitoring pipelines, or prior access. A low-privilege key alone is not proof of exploitability against a specific hidden index.
Where the identities separated
Elasticsearch represents a shard with more than a display name. A ShardId includes an index name, an index UUID, and a shard number. Names can be reused over time; the UUID binds the name to a specific index instance.
The public patch makes the failed invariant unusually clear. Before the fix, the RBAC path asked whether permissions existed for shardId.getIndexName(). If the name was authorized, the shard passed this part of the decision. The patched code also checks the complete index identity against current cluster metadata. A name paired with another index’s UUID no longer resolves as the authorized object and is denied.[2]
The added regression test constructs a malformed get request whose stated index is idx-a. It then supplies an internal ShardId that keeps an authorized-looking name while carrying another index’s UUID. The test expects an authorization exception. The source comment notes that production code intentionally makes this request difficult to form, so the test writes the transport request manually.[2]

This resembles the identity confusion behind CVE-2026-47849 in Spring Data REST, where the route and mutable body could disagree about the record being changed. It also rhymes with http4s request desynchronization: two components parse one message, then make security decisions about different effective objects. The bug classes differ, but the engineering question is the same: which representation is authoritative at the point of use?
A safe model of the bug
The following lab model contains no Elasticsearch exploit syntax. It isolates the authorization mistake using two fields: an index name and a UUID. I executed it locally during this analysis.
from dataclasses import dataclass
@dataclass(frozen=True)
class ShardId:
name: str
uuid: str
def vulnerable(stated_index, shard, allowed_names):
return stated_index in allowed_names and shard.name in allowed_names
def fixed(stated_index, shard, allowed_names, cluster_metadata):
return (
stated_index in allowed_names
and shard.name in allowed_names
and cluster_metadata.get(shard.name) == shard.uuid
)
metadata = {
"authorized-index": "uuid-A",
"secret-index": "uuid-S",
}
forged = ShardId("authorized-index", "uuid-S")
assert vulnerable("authorized-index", forged, {"authorized-index"}) is True
assert fixed("authorized-index", forged, {"authorized-index"}, metadata) is False
The model returned True for the name-only decision and False after metadata binding. It demonstrates the mechanism, not product exploitability. It deliberately omits transport serialization, request actions, discovery, and data extraction.
A realistic attack chain
Consider a composed scenario in which an analytics cluster is permitted to search only shared-observability-* on a central fulfilling cluster.
First, the operator obtains the analytics cluster’s cross-cluster API key. That could follow compromise of the local cluster, an exposed keystore backup, or excessive administrative access. CVE-2026-103009 does not provide that initial access.
Next, the operator verifies that the remote endpoint speaks RCS 2.0 and that the key can search its intended index. Ordinary success establishes the capability and transport path. The operator then needs a target shard identifier from outside the allowed namespace. A leaked diagnostic bundle or monitoring record could provide it, but that is an environmental dependency, not a property guaranteed by the CVE.
The crafted shard-level request presents the permitted index where RBAC expects a name, while the internal shard identity points at the restricted index. On a vulnerable fulfilling cluster, authorization can be evaluated against the permitted name while execution follows the foreign UUID. The result may expose documents, field mappings, and metadata. Elastic also notes limited modification of retention-lease state, which explains the low integrity impact in the CVSS vector.[1]
This is not a REST-layer BOLA test. Fuzzing _search through port 9200 will not exercise the vulnerable path. An assessment must first prove the RCS 2.0 topology and remote transport exposure. That distinction prevents a common waste pattern: scanning the visible API while the vulnerable parser and authorization code live in a different protocol boundary.
Detection has a before-and-after problem
Elastic reported no vulnerability-specific indicators of compromise.[1] That is not the same as saying the activity is invisible.
After patching, the new RBAC path emits a warning when the request’s stated indices are authorized but the internal shard ID either lacks permission or does not match the authorized index in cluster metadata.[2] Those messages are high-value signals because legitimate clients should not generate contradictory shard identities. Alert on them, retain the complete node context, and correlate the source transport connection with the cross-cluster API key and remote-cluster alias.
Before patching, the dangerous request could pass the flawed check, so defenders should not rely on a denial event that may never have existed. Elastic audit logging can record security events on each node, but it is disabled by default and requires an appropriate subscription. Enable it on every node that handles remote-cluster traffic and centralize the resulting JSON logs.[4] Useful pivots include access_denied, API-key authentication events, the transport action, source address, and bursts of shard-level access that do not match the expected index scope.[5]
Audit data still may not reconstruct the mismatched UUID after the fact. Compensating evidence should include:
- inventory of nodes exposing the remote-cluster interface;
- active and recently invalidated cross-cluster API keys, owners, and index scopes;
- remote-cluster aliases and source networks that used each key;
- access to sensitive indices during the vulnerable window;
- diagnostic exports, snapshots, and observability data that may have disclosed shard UUIDs.
A key authorized to a harmless index is not harmless evidence. The key is the foothold required by this vulnerability. Rotate keys where exposure cannot be excluded, but do not treat rotation as a substitute for upgrading the fulfilling cluster.
Patch the object binding, then reduce the blast radius
Elastic fixed the issue in 8.19.23, 9.4.8, and 9.5.5. Affected ranges were 8.13.0 through 8.19.22, 9.0.0 through 9.4.7, and 9.5.0 through 9.5.4. Elastic Cloud Serverless was not affected.[1]
If an upgrade cannot happen immediately, Elastic’s mitigation is direct: disable remote-cluster connections on the fulfilling cluster when cross-cluster search is not required.[1] Network controls should also restrict TCP 9443 to explicit cluster peers. That does not repair authorization, but it removes opportunistic paths to the transport interface.
Then tighten the capability itself. Give each consuming cluster a distinct API key with narrow index patterns. Avoid sharing one key across environments. The official model already supports further restrictions on local users and, in current documentation, strong identity verification that binds a cross-cluster API key to a certificate identity.[3] Certificate binding does not fix a shard-identity bug, but it makes stolen-key replay from an untrusted cluster harder.
Finally, add invariant tests at the authorization boundary. For every request carrying both a human-readable name and an immutable identifier, mutate them independently. A valid pair should pass. A valid name plus foreign UUID must fail. An unknown name plus valid UUID must fail. A deleted and recreated name with the old UUID must fail. This is the same discipline needed for route/body disagreements, signed-object envelopes, tenant IDs, and confused-deputy workflows such as GitLost’s public-issue to private-repository chain.
Strategic takeaway
The vulnerability is not simply “an API key read too much.” The key’s scope was evaluated against one identity while the storage operation trusted another. The durable fix was to make authorization resolve the complete shard identity against current metadata, then carry that same object into execution.
Security reviews should search for this pattern wherever a request contains aliases, names, IDs, UUIDs, paths, or tenant keys that describe the same resource. Count the representations. Trace which one each layer trusts. If policy and execution can choose independently, the attacker does not need to break either layer. They only need to stand in the gap.
Sources
- Elastic security advisory ESA-2026-199
- Elasticsearch patch: shard-level request consistency
- Elastic documentation: remote clusters with API keys
- Elastic documentation: enable audit logging
- Elastic documentation: audit event types
- MITRE CWE-639
💜 Enjoyed this content? Support the blog with USDT (TRC20):
TX7obcjHQbDUXb4mGqoASEu1QFTKT2CFGG
