Elasticsearch数据存储文件定位及原始索引数据查看咨询
Hey there! Let's tackle your two Elasticsearch/Lucene questions clearly—they're great ones for understanding how the underlying storage works.
1. Which files store your actual indexed data?
Elasticsearch is built on top of Lucene, so those files are Lucene's segment-specific files. Here's a breakdown of each one you listed:
_0.cfeand_0.cfs: These are Compound File components. Lucene bundles multiple smaller segment files (like inverted indexes, stored document fields, term dictionaries, and more) into a single compound file to cut down on filesystem overhead. Together, these two files hold your actual indexed data—including the content of your documents, the structures that power search, and stored field values._0.si: This is the Segment Info file. It stores metadata about the segment (like document count, segment version, and other administrative details) but not the actual document content.segments_e: This is the segments manifest file. It tracks all active segments in the index, their names, and global index-level metadata—think of it as a table of contents for your index's segments.write.lock: A simple lock file to prevent multiple processes from writing to the index simultaneously. No actual data lives here.
So the short answer: _0.cfe and _0.cfs contain your real indexed data.
2. How to view the actual data if you can't read these files directly?
These are binary files in Lucene's proprietary format, so you can't open them with a regular text editor. Instead, use these reliable methods to access your data:
Use Elasticsearch's REST API
This is the simplest and most recommended approach:
- To search for documents: Run a query like
GET /your_index_name/_search(replaceyour_index_namewith your index's name—from your path, that'satLl0jUNTbuAJKT3OxgpUQ). Addq=*to return all documents, or specify fields to narrow down results. - To fetch a specific document: If you know the document ID, use
GET /your_index_name/_doc/your_doc_id. - To list all documents (for small indexes): Use
GET /your_index_name/_search?size=1000(adjustsizeto match your total document count—note that very largesizevalues aren't ideal for production clusters).
Export data to a readable format
If you want to save your data as human-readable JSON, try these options:
- Use the
_reindexAPI to copy data to a temporary index, then export it with tools likecurl:
You can also use tools likecurl -X POST "localhost:9200/_reindex?pretty" -H 'Content-Type: application/json' -d' { "source": { "index": "your_index_name" }, "dest": { "index": "temp_export_index" } } 'elasticsearch-dumpto directly export the index data to a JSON file on your local machine.
Use Lucene's Luke tool
Luke is a GUI tool designed specifically to inspect Lucene indexes. You'll need to download a version that matches your Elasticsearch's Lucene version (Elasticsearch 6.1.1 uses Lucene 7.1.0, so grab Luke 7.x). Once you open the index folder from your path, you can browse documents, view stored fields, inspect inverted indexes, and more. Just make sure to stop your Elasticsearch cluster before using Luke to avoid lock conflicts.
内容的提问来源于stack exchange,提问作者D vignesh

