如何在Rails中通过Scroll从Elasticsearch获取全部记录?
Hey there! I totally get the confusion since the scan search type got deprecated—let's walk through how to implement the Scroll API correctly in your Ruby code, building off your original attempt.
Step 1: Initialize the Scroll Context
First, instead of using search_type: 'scan', you'll start with a regular search request that includes the scroll parameter (to set how long Elasticsearch keeps the scroll context alive) and your desired size (documents per batch). If you're fetching all documents, use a match_all query:
# Initialize the scroll context initial_response = client.search( index: 'test', scroll: '5m', # Keep the scroll context alive for 5 minutes size: 10, # Number of documents to return per batch body: { query: { match_all: {} # Fetch all documents in the index } } ) # Grab the initial scroll ID scroll_id = initial_response['_scroll_id']
Step 2: Loop Through Scroll Batches
Next, you'll repeatedly call the scroll API using the latest scroll_id until no more hits are returned. This is similar to your original loop, but we'll handle the scroll ID properly and add cleanup at the end:
begin # Process the initial batch of hits process_hits(initial_response['hits']['hits']) # Continue scrolling until no more hits while true scroll_response = client.scroll( scroll_id: scroll_id, scroll: '5m' # Refresh the scroll context timeout ) hits = scroll_response['hits']['hits'] break if hits.empty? # Process the current batch of hits process_hits(hits) # Update the scroll ID for the next iteration scroll_id = scroll_response['_scroll_id'] end ensure # Clean up the scroll context to free Elasticsearch resources client.clear_scroll(scroll_id: scroll_id) end # Helper method to process your hits (customize this as needed) def process_hits(hits) puts hits.map { |r| r['_source']['title'] } end
Key Notes to Remember
- Scroll Timeout: The
scrollparameter ('5m'here) tells Elasticsearch how long to keep the scroll session active. Set it to a value that gives you enough time to process each batch—don't set it longer than necessary, as it consumes resources. - Size Parameter: This controls how many documents are returned per scroll batch, not the total number of documents.
- Cleanup with
clear_scroll: Always call this when you're done (even if there's an error) to avoid leaving orphaned scroll contexts that waste Elasticsearch memory. - No More
scan: Thescansearch type was deprecated because the Scroll API alone handles efficient full-index retrieval now—you don't need that extra parameter anymore.
Comparing to Your Original Scan Code
Your original code was close, but the main changes are:
- Removed
search_type: 'scan'(it's no longer needed) - Added an explicit
match_allquery (you can replace this with your own query if needed) - Added proper cleanup with
clear_scroll - Structured the loop with a
begin/ensureblock to handle errors gracefully
内容的提问来源于stack exchange,提问作者Manoharsinh Rana

