如何高效在Elasticsearch存储文本并检索?如何分割特定格式数据按唯一编号检索?
Hey there! Let's tackle your two questions about Elasticsearch storage and retrieval—here's a practical breakdown tailored to your needs.
1. How to Efficiently Store Text in Elasticsearch for Retrieval
To get optimal performance for storing and searching text, follow these key steps:
Preprocess Your Text First
Clean up redundant whitespace, special characters, or irrelevant content before ingestion. For language-specific text (like Chinese), use dedicated analyzers (e.g.,ik_smartfor Chinese) to ensure proper tokenization; for English, the defaultstandardanalyzer works well for most cases.Define a Targeted Mapping
Don’t rely on dynamic mapping for critical fields. Create a custom index mapping where:- Text fields you need to search across use the
texttype with an appropriate analyzer. - Fields used for exact matches (like IDs or categories) use the
keywordtype (this avoids tokenization and enables fast term queries).
Example mapping snippet:
{ "mappings": { "properties": { "content": { "type": "text", "analyzer": "ik_smart" }, "unique_id": { "type": "keyword" } } } }- Text fields you need to search across use the
Use Bulk Ingestion
Avoid indexing documents one by one—Elasticsearch’s_bulkAPI is designed for high-volume ingestion. Group documents into batches (typically 1000-5000 per batch, depending on document size) to minimize network overhead. Example bulk request structure:{"index": {"_index": "your_index", "_id": "1"}} {"unique_id": "SSLEGGU00402-IM", "content": "your text content"} {"index": {"_index": "your_index", "_id": "2"}} {"unique_id": "SSLEGGU00412-IM", "content": "another text content"}Optimize Index Settings
- Temporarily increase
refresh_interval(e.g., to30sor-1during bulk ingestion) to reduce frequent segment refreshes, then reset it to the default1sonce ingestion is done. - Adjust
number_of_shardsbased on your cluster size (1-5 shards per node is a good starting point) and setnumber_of_replicasto at least 1 for high availability.
- Temporarily increase
Optimize Retrieval
Usematchqueries for full-text search across text fields, andterm/termsqueries for exact matches on keyword fields. Combine filters (e.g.,boolquery withfilterclause) to narrow down results without affecting scoring, which boosts performance.
2. Splitting Your Structured File Data & Storing for Unique ID Retrieval
Looking at your sample data, each main record starts with SSLEGGU followed by a unique identifier (like 00402-IM), then includes sub-segments prefixed with numbers (1U, 2U, etc.). Here’s how to process this:
Step 1: Define Data Segmentation Logic
Use regular expressions to split the raw data into individual main records. For example, this regex will capture each SSLEGGU-block until the next SSLEGGU or end of file:
import re raw_data = "your full sample data string here" main_records = re.findall(r'SSLEGGU\w+-\w+.*?(?=SSLEGGU|$)', raw_data, re.DOTALL)
Step 2: Parse Each Record into a Structured Document
For each main record, extract the unique ID, main metadata, and sub-segments. Here’s a Python example to parse and structure the data:
documents = [] for record in main_records: # Split into lines (adjust if your data uses a different separator) segments = record.strip().split() # Find the start of the next main record to split cleanly split_points = [i for i, val in enumerate(segments) if val.startswith('SSLEGGU')][1:] if len(main_records) > 1 else [] # Extract unique ID (first element of the main record) unique_id = segments[0] # Parse main metadata (values right after the unique ID) main_metadata = segments[1:segments.index([s for s in segments if s.startswith('1U')][0])] # Parse sub-segments (group by the 1U, 2U, etc. prefixes) sub_segments = [] current_sub = [] for part in segments: if part.startswith(('1U', '2U', '3U')): if current_sub: sub_segments.append({"sub_id": current_sub[0], "data": current_sub[1:]}) current_sub = [part] else: current_sub.append(part) if current_sub: sub_segments.append({"sub_id": current_sub[0], "data": current_sub[1:]}) # Build the Elasticsearch document documents.append({ "_index": "your_data_index", "_id": unique_id, "_source": { "unique_id": unique_id, "main_metadata": main_metadata, "sub_segments": sub_segments } })
Step 3: Bulk Index the Documents
Use the Elasticsearch Python client to send the bulk request:
from elasticsearch import Elasticsearch from elasticsearch.helpers import bulk # Initialize client (adjust host/port to your cluster) es = Elasticsearch("http://localhost:9200") # Execute bulk ingestion bulk(es, documents)
Step 4: Retrieve by Unique ID
To fetch a record using its unique ID, use a term query on the unique_id keyword field (since we mapped it for exact matches):
GET /your_data_index/_search { "query": { "term": { "unique_id": "SSLEGGU00402-IM" } } }
Alternatively, you can directly get the document by ID for faster access:
GET /your_data_index/_doc/SSLEGGU00402-IM
内容的提问来源于stack exchange,提问作者Shankar

