You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Rails中通过Scroll从Elasticsearch获取全部记录?

Using Elasticsearch Scroll API to Bypass the 10k Document Limit (Replacing Deprecated Scan)

Hey there! I totally get the confusion since the scan search type got deprecated—let's walk through how to implement the Scroll API correctly in your Ruby code, building off your original attempt.

Step 1: Initialize the Scroll Context

First, instead of using search_type: 'scan', you'll start with a regular search request that includes the scroll parameter (to set how long Elasticsearch keeps the scroll context alive) and your desired size (documents per batch). If you're fetching all documents, use a match_all query:

# Initialize the scroll context
initial_response = client.search(
  index: 'test',
  scroll: '5m',  # Keep the scroll context alive for 5 minutes
  size: 10,      # Number of documents to return per batch
  body: {
    query: {
      match_all: {}  # Fetch all documents in the index
    }
  }
)

# Grab the initial scroll ID
scroll_id = initial_response['_scroll_id']

Step 2: Loop Through Scroll Batches

Next, you'll repeatedly call the scroll API using the latest scroll_id until no more hits are returned. This is similar to your original loop, but we'll handle the scroll ID properly and add cleanup at the end:

begin
  # Process the initial batch of hits
  process_hits(initial_response['hits']['hits'])

  # Continue scrolling until no more hits
  while true
    scroll_response = client.scroll(
      scroll_id: scroll_id,
      scroll: '5m'  # Refresh the scroll context timeout
    )

    hits = scroll_response['hits']['hits']
    break if hits.empty?

    # Process the current batch of hits
    process_hits(hits)

    # Update the scroll ID for the next iteration
    scroll_id = scroll_response['_scroll_id']
  end
ensure
  # Clean up the scroll context to free Elasticsearch resources
  client.clear_scroll(scroll_id: scroll_id)
end

# Helper method to process your hits (customize this as needed)
def process_hits(hits)
  puts hits.map { |r| r['_source']['title'] }
end

Key Notes to Remember

  • Scroll Timeout: The scroll parameter ('5m' here) tells Elasticsearch how long to keep the scroll session active. Set it to a value that gives you enough time to process each batch—don't set it longer than necessary, as it consumes resources.
  • Size Parameter: This controls how many documents are returned per scroll batch, not the total number of documents.
  • Cleanup with clear_scroll: Always call this when you're done (even if there's an error) to avoid leaving orphaned scroll contexts that waste Elasticsearch memory.
  • No More scan: The scan search type was deprecated because the Scroll API alone handles efficient full-index retrieval now—you don't need that extra parameter anymore.

Comparing to Your Original Scan Code

Your original code was close, but the main changes are:

  • Removed search_type: 'scan' (it's no longer needed)
  • Added an explicit match_all query (you can replace this with your own query if needed)
  • Added proper cleanup with clear_scroll
  • Structured the loop with a begin/ensure block to handle errors gracefully

内容的提问来源于stack exchange,提问作者Manoharsinh Rana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 11:37:33