Elasticsearch物理层实现及文档索引命令执行流程咨询
Hey folks, let's dive deep into these Elasticsearch technical questions—they're all core concepts that every ES developer should get right. Here's a breakdown for each one:
1. 文档索引命令(PUT /index/type/id {data})的执行流程
Let's walk through this step by step, since understanding the flow helps debug latency or consistency issues:
- Step 1: Request Routing Your client sends the PUT request to any node in the cluster (this becomes the coordinating node). The node hashes the document ID (or uses a custom routing key if you specified one) to calculate which primary shard the document should go to.
- Step 2: Forward to Target Shard The coordinating node forwards the request to the node that holds the target primary shard.
- Step 3: Write to Memory & Translog The target node first validates the request (checks if the index exists, field types match mappings, etc.). Then it writes the document to the in-memory buffer and appends the operation to the translog (a persistent append-only log file on disk).
- Step 4: Return Success As soon as the document is in the memory buffer and translog, the node sends a success response back to the coordinating node, which relays it to your client. Note: At this point, the document isn't yet searchable.
- Step 5: Refresh to Make Searchable By default, every 1 second, the memory buffer is flushed to the filesystem cache, creating a new immutable Lucene segment. This segment is now available for search—this is why ES is near-real-time, not real-time.
- Step 6: Flush to Persist to Disk When the translog reaches a certain size (or every 30 minutes by default), a flush operation triggers:
- Any remaining data in the memory buffer is flushed to the filesystem cache as a new segment
- All segments in the filesystem cache are written to disk
- The translog is cleared, ready to record new operations
2. 文档在物理层面的存储方式
Under the hood, Elasticsearch uses Lucene's storage model, which is all about immutable segments. Here's how documents map to disk:
- Per-Shard Directories Each shard (primary or replica) has its own directory on the node's disk. The path looks something like
data/nodes/0/indices/<index-uuid>/<shard-id>/index/. - Lucene Segment Files Documents are stored in immutable segment files—once a segment is written to disk, it never changes. Each segment is a self-contained inverted index, with files like:
.docs: Stores document IDs and basic metadata.pos: Stores the position of terms within documents (for phrase searches).tim: Stores term frequencies (how often a term appears in a document).cfs: A compound file that bundles smaller segment files to reduce overhead
- Translog Files The translog is a persistent log that records every write operation before it's persisted to disk. If a node crashes, ES uses the translog to replay unflushed operations and recover data.
- Metadata Files Additional files like
.state(stores shard-level state) and.si(stores segment metadata) keep track of the shard's configuration and segment status.
3. index、type、document、shards在硬件层面的含义
Let's map each logical concept to what it looks like on your servers:
- Index Logically, it's a collection of related documents. On hardware, an index corresponds to multiple directories (one per shard) spread across different nodes in the cluster. Each shard's directory holds all the data for that portion of the index.
- Type (Deprecated in ES 7+) Originally a logical grouping within an index, but it never had a separate physical storage layer. Types were just a field (
_type) added to documents, stored alongside other fields in the segment files. - Document The smallest unit of data in ES. On disk, a document is a record within a Lucene segment, stored as binary data with its fields and values. When you update a document, ES doesn't modify the old record—it marks it as deleted and writes a new version in a fresh segment.
- Shards (Primary/Replica) Each shard is an independent Lucene instance. On hardware, a shard is a directory on a node's disk, containing all the segment files, translog, and metadata for that shard. Primary shards are spread across nodes for horizontal scaling, and replica shards are stored on different nodes than their primaries to ensure high availability.
4. Elasticsearch的硬件层具体实现机制
ES is a distributed system that leans heavily on your hardware—here's how it uses each component:
- CPU ES is CPU-intensive for tasks like text analysis (tokenization, stemming), query parsing, and aggregation. Multi-core CPUs shine here because ES can parallelize query processing across shards and handle multiple concurrent requests at once.
- Memory Two critical memory pools:
- JVM Heap Used for storing node state, request processing, aggregation intermediate results, and cached query results. Best practice is to limit this to 50% of your physical memory (max 32GB, since beyond that JVM garbage collection becomes inefficient).
- Filesystem Cache Managed by the OS, this caches frequently accessed segment files. The more memory you leave here, the faster query performance gets—since most queries hit cached segments instead of disk.
- Disk
- SSD vs HDD SSDs are strongly recommended: ES does a lot of random reads (for queries) and sequential writes (for translogs/segments), and SSDs have 10-100x higher IOPS than HDDs, drastically reducing latency.
- RAID Configuration For reliability, use RAID 1 (mirroring) or RAID 5/6 (striping with parity). Avoid RAID 0, as it offers no redundancy.
- Network Cluster nodes communicate constantly (for shard replication, request forwarding, cluster state updates). Low-latency, high-bandwidth networks (like 10Gbps Ethernet) are essential for large clusters to avoid bottlenecks.
- High Availability Primary and replica shards are placed on different nodes (and ideally different racks) to ensure that if one node fails, the replicas take over without data loss.
内容的提问来源于stack exchange,提问作者D vignesh
相关产品推荐
相关产品推荐

