You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效在Elasticsearch存储文本并检索?如何分割特定格式数据按唯一编号检索?

Answers to Your Elasticsearch Questions

Hey there! Let's tackle your two questions about Elasticsearch storage and retrieval—here's a practical breakdown tailored to your needs.

1. How to Efficiently Store Text in Elasticsearch for Retrieval

To get optimal performance for storing and searching text, follow these key steps:

  • Preprocess Your Text First
    Clean up redundant whitespace, special characters, or irrelevant content before ingestion. For language-specific text (like Chinese), use dedicated analyzers (e.g., ik_smart for Chinese) to ensure proper tokenization; for English, the default standard analyzer works well for most cases.

  • Define a Targeted Mapping
    Don’t rely on dynamic mapping for critical fields. Create a custom index mapping where:

    • Text fields you need to search across use the text type with an appropriate analyzer.
    • Fields used for exact matches (like IDs or categories) use the keyword type (this avoids tokenization and enables fast term queries).
      Example mapping snippet:
    {
      "mappings": {
        "properties": {
          "content": { "type": "text", "analyzer": "ik_smart" },
          "unique_id": { "type": "keyword" }
        }
      }
    }
    
  • Use Bulk Ingestion
    Avoid indexing documents one by one—Elasticsearch’s _bulk API is designed for high-volume ingestion. Group documents into batches (typically 1000-5000 per batch, depending on document size) to minimize network overhead. Example bulk request structure:

    {"index": {"_index": "your_index", "_id": "1"}}
    {"unique_id": "SSLEGGU00402-IM", "content": "your text content"}
    {"index": {"_index": "your_index", "_id": "2"}}
    {"unique_id": "SSLEGGU00412-IM", "content": "another text content"}
    
  • Optimize Index Settings

    • Temporarily increase refresh_interval (e.g., to 30s or -1 during bulk ingestion) to reduce frequent segment refreshes, then reset it to the default 1s once ingestion is done.
    • Adjust number_of_shards based on your cluster size (1-5 shards per node is a good starting point) and set number_of_replicas to at least 1 for high availability.
  • Optimize Retrieval
    Use match queries for full-text search across text fields, and term/terms queries for exact matches on keyword fields. Combine filters (e.g., bool query with filter clause) to narrow down results without affecting scoring, which boosts performance.

2. Splitting Your Structured File Data & Storing for Unique ID Retrieval

Looking at your sample data, each main record starts with SSLEGGU followed by a unique identifier (like 00402-IM), then includes sub-segments prefixed with numbers (1U, 2U, etc.). Here’s how to process this:

Step 1: Define Data Segmentation Logic

Use regular expressions to split the raw data into individual main records. For example, this regex will capture each SSLEGGU-block until the next SSLEGGU or end of file:

import re
raw_data = "your full sample data string here"
main_records = re.findall(r'SSLEGGU\w+-\w+.*?(?=SSLEGGU|$)', raw_data, re.DOTALL)

Step 2: Parse Each Record into a Structured Document

For each main record, extract the unique ID, main metadata, and sub-segments. Here’s a Python example to parse and structure the data:

documents = []
for record in main_records:
    # Split into lines (adjust if your data uses a different separator)
    segments = record.strip().split()
    # Find the start of the next main record to split cleanly
    split_points = [i for i, val in enumerate(segments) if val.startswith('SSLEGGU')][1:] if len(main_records) > 1 else []
    # Extract unique ID (first element of the main record)
    unique_id = segments[0]
    # Parse main metadata (values right after the unique ID)
    main_metadata = segments[1:segments.index([s for s in segments if s.startswith('1U')][0])]
    # Parse sub-segments (group by the 1U, 2U, etc. prefixes)
    sub_segments = []
    current_sub = []
    for part in segments:
        if part.startswith(('1U', '2U', '3U')):
            if current_sub:
                sub_segments.append({"sub_id": current_sub[0], "data": current_sub[1:]})
            current_sub = [part]
        else:
            current_sub.append(part)
    if current_sub:
        sub_segments.append({"sub_id": current_sub[0], "data": current_sub[1:]})
    # Build the Elasticsearch document
    documents.append({
        "_index": "your_data_index",
        "_id": unique_id,
        "_source": {
            "unique_id": unique_id,
            "main_metadata": main_metadata,
            "sub_segments": sub_segments
        }
    })

Step 3: Bulk Index the Documents

Use the Elasticsearch Python client to send the bulk request:

from elasticsearch import Elasticsearch
from elasticsearch.helpers import bulk

# Initialize client (adjust host/port to your cluster)
es = Elasticsearch("http://localhost:9200")
# Execute bulk ingestion
bulk(es, documents)

Step 4: Retrieve by Unique ID

To fetch a record using its unique ID, use a term query on the unique_id keyword field (since we mapped it for exact matches):

GET /your_data_index/_search
{
  "query": {
    "term": {
      "unique_id": "SSLEGGU00402-IM"
    }
  }
}

Alternatively, you can directly get the document by ID for faster access:

GET /your_data_index/_doc/SSLEGGU00402-IM

内容的提问来源于stack exchange,提问作者Shankar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:07:18