Windows10下Elasticsearch批量导入报错及导入方式咨询
Hey there, let's start with fixing your current issue first, then walk through several reliable ways to get your Shakespeare dataset into Elasticsearch.
一、报错的根本原因
Looking at your error message:
Warning: Couldn't read data from file "$shakespeare_6.0", this makes an empty Warning: POST. { "error" : { "root_cause" : [ { "type" : "parse_exception", "reason" : "request body is required" } ], "type" : "parse_exception", "reason" : "request body is required" }, "status" : 400 }
The core problem is you're using Linux-style variable syntax ($) in Windows 10. Windows Command Prompt/PowerShell doesn't recognize $shakespeare_6.0 as a file reference—instead, it treats the entire string as a filename, which doesn't exist. That's why curl can't read your data, resulting in an empty POST request that triggers the 400 "request body is required" error.
Quick Fix for Your curl Command
If your data file is directly named shakespeare_6.0, just remove the $ sign. Also, make sure the file is in the same directory where you're running curl, or use the full file path:
# 如果文件在当前目录 curl -H "Content-Type: application/x-ndjson" -XPOST "localhost:9200/shakespeare/doc/_bulk?pretty" --data-binary @shakespeare_6.0 # 如果文件在其他路径,比如下载文件夹 curl -H "Content-Type: application/x-ndjson" -XPOST "localhost:9200/shakespeare/doc/_bulk?pretty" --data-binary @C:\Users\YourUsername\Downloads\shakespeare_6.0
二、向Elasticsearch导入数据的多种方式
Now that we've fixed the immediate issue, here are 4 common methods to import data into ES, each suited for different use cases:
1. curl + NDJSON(适合小批量数据,快速验证)
As you tried originally, curl is great for quick tests. Just ensure your NDJSON format is correct (each index metadata line paired with a data line, no extra spaces) and use the Windows-compatible command above.
2. Kibana Dev Tools(可视化调试,适合中小批量)
If you have Kibana set up, this is the easiest way to debug and import data:
- Open Kibana and navigate to Dev Tools
- Paste your NDJSON data into the editor, prefix it with the bulk API endpoint, and click run:
POST /shakespeare/doc/_bulk {"index":{"_index":"shakespeare","_id":0}} {"type":"act","line_id":1,"play_name":"Henry IV", "speech_number":"","line_number":"","speaker":"","text_entry":"ACT I"} // Add more NDJSON lines here as needed
Kibana will highlight syntax errors and show detailed response messages, making it perfect for troubleshooting.
3. Python Elasticsearch Client(灵活定制,适合中型数据)
For more control over the import process (like filtering data or handling errors), use the official Python client:
First, install the client:
pip install elasticsearch
Then use this script to import your NDJSON file:
import json from elasticsearch import Elasticsearch from elasticsearch.helpers import bulk # Connect to your Elasticsearch instance es = Elasticsearch("http://localhost:9200") # Read and process the NDJSON file actions = [] with open("shakespeare_6.0", "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue # Parse each line as JSON data = json.loads(line) # Handle index metadata lines if "_index" in data.get("index", {}): current_meta = data["index"] else: # Build the bulk action action = { "_index": current_meta["_index"], "_id": current_meta["_id"], "_source": data } actions.append(action) # Bulk insert every 1000 documents to avoid overload if len(actions) == 1000: bulk(es, actions) actions = [] # Insert any remaining documents if actions: bulk(es, actions) print("Dataset imported successfully!")
4. Logstash(适合大规模/生产级数据)
For large datasets (GBs or more), Logstash is the way to go—it supports parallel processing, error handling, and resuming interrupted imports:
- Create a Logstash configuration file (e.g.,
es_shakespeare.conf):
input { file { path => "C:\path\to\shakespeare_6.0" start_position => "beginning" sincedb_path => "NUL" # Disable sincedb on Windows to avoid skipping data } } filter { json { source => "message" } # Extract index metadata from the first line of each pair if [index] { mutate { add_field => { "[@metadata][_index]" => "%{index._index}" } add_field => { "[@metadata][_id]" => "%{index._id}" } remove_field => ["index", "message"] } } else { # Fallback to default index if metadata is missing mutate { add_field => { "[@metadata][_index]" => "shakespeare" } } } } output { elasticsearch { hosts => ["localhost:9200"] index => "%{[@metadata][_index]}" document_id => "%{[@metadata][_id]}" } stdout { codec => rubydebug } # Optional: Print import logs to console }
- Run Logstash with the configuration:
bin\logstash.bat -f es_shakespeare.conf
内容的提问来源于stack exchange,提问作者Umang Bhargava

