Java中不使用Logstash将CSV数据导入Elasticsearch的方法及入门指引
Hey there! Let's break down your questions step by step — first, practical ways to import CSV data into Elasticsearch without relying on Logstash, then a guide to getting started with TransportClient (with an important heads-up: it's deprecated in Elasticsearch 7.0+, but we'll cover both the basics and the modern alternative note).
不用Logstash导入CSV的可行方法
1. 手动使用Elasticsearch _bulk API
This is a lightweight option if you don't want to add extra dependencies:
- First, convert your CSV into the bulk API format: each operation line is followed by the data line. Example:
{"index": {"_index": "your_target_index", "_id": "1"}} {"name": "Alice", "age": 30, "city": "New York"} {"index": {"_index": "your_target_index", "_id": "2"}} {"name": "Bob", "age": 28, "city": "London"} - You can write a simple script (Python/Shell) to parse the CSV and generate this format. For example, in Python, use the
csvmodule to read rows, then construct the bulk lines. - Send the formatted data via a POST request to
http://<es-host>:<port>/_bulk(use tools likecurlor libraries likerequestsin Python). - Pros: No extra tools needed; Cons: Requires handling data conversion and error checking yourself.
2. Use Elasticsearch's official language clients
The official clients (Python, Java, etc.) let you directly read CSV and index data with more flexibility:
Python Client Example
First install the client:
pip install elasticsearch
Then write a script to import:
from elasticsearch import Elasticsearch import csv # Connect to your ES cluster es = Elasticsearch(["http://localhost:9200"]) index_name = "user_data" # Create index (if it doesn't exist) with mappings if not es.indices.exists(index=index_name): es.indices.create( index=index_name, body={ "mappings": { "properties": { "name": {"type": "text"}, "age": {"type": "integer"}, "city": {"type": "keyword"} } } } ) # Read CSV and index rows with open("users.csv", "r") as csv_file: reader = csv.DictReader(csv_file) for row in reader: # Convert data types as needed row["age"] = int(row["age"]) es.index(index=index_name, document=row)
- For large datasets, use the
bulkmethod to batch requests and improve performance. - Pros: Full control over data cleaning and mapping; Cons: Requires basic coding knowledge.
3. Kibana's Visual CSV Importer (if you have Kibana access)
If you prefer a no-code approach:
- Open Kibana → Go to Management → Index Patterns → Click Import data
- Upload your CSV file, follow the wizard to configure field mappings, set the index name, and let Kibana handle the rest.
- Pros: No coding required, visual feedback; Cons: Dependent on Kibana, better for small to medium datasets.
TransportClient入门指南 (Important: Deprecated in 7.0+)
First, a critical note: TransportClient was deprecated in Elasticsearch 7.0 and removed in 8.0. The official recommendation is to use the Rest High Level Client instead. But if you're working with an older ES version (6.x or earlier), here's how to get started:
1. Set up dependencies (Java example)
If using Maven, add these dependencies (match your ES version exactly):
<dependency> <groupId>org.elasticsearch.client</groupId> <artifactId>transport</artifactId> <version>6.8.23</version> </dependency> <dependency> <groupId>org.elasticsearch</groupId> <artifactId>elasticsearch</artifactId> <version>6.8.23</version> </dependency>
2. Initialize the TransportClient
import org.elasticsearch.client.transport.TransportClient; import org.elasticsearch.common.settings.Settings; import org.elasticsearch.common.transport.TransportAddress; import org.elasticsearch.transport.client.PreBuiltTransportClient; import java.net.InetAddress; import java.net.UnknownHostException; public class ESClientSetup { public static void main(String[] args) throws UnknownHostException { // Match your ES cluster's name (from elasticsearch.yml) Settings settings = Settings.builder() .put("cluster.name", "your_cluster_name") .build(); // Connect to the ES node (uses TCP port 9300 by default) TransportClient client = new PreBuiltTransportClient(settings) .addTransportAddress(new TransportAddress(InetAddress.getByName("localhost"), 9300)); // Your operations go here... // Always close the client when done client.close(); } }
3. Import CSV with TransportClient (Java example)
Use a CSV library like OpenCSV to read data, then batch index with BulkRequest:
import com.opencsv.CSVReader; import org.elasticsearch.action.bulk.BulkRequest; import org.elasticsearch.action.bulk.BulkResponse; import org.elasticsearch.action.index.IndexRequest; import java.io.FileReader; import java.util.HashMap; import java.util.Map; // Inside your main method after initializing the client CSVReader reader = new CSVReader(new FileReader("users.csv")); String[] headers = reader.readNext(); // Get CSV headers String[] row; BulkRequest bulkRequest = new BulkRequest(); while ((row = reader.readNext()) != null) { Map<String, Object> document = new HashMap<>(); for (int i = 0; i < headers.length; i++) { // Convert data types (e.g., age to integer) if (headers[i].equals("age")) { document.put(headers[i], Integer.parseInt(row[i])); } else { document.put(headers[i], row[i]); } } // Add index request to bulk batch bulkRequest.add(new IndexRequest("user_data").source(document)); } // Execute the bulk import BulkResponse response = client.bulk(bulkRequest).actionGet(); if (response.hasFailures()) { // Handle errors System.out.println("Bulk import failed: " + response.buildFailureMessage()); }
Key Notes for TransportClient
- Ensure your ES node's
transport.hostsetting allows incoming connections from your client. - Version consistency is critical: The client version must match the ES cluster version exactly to avoid compatibility issues.
- Plan to migrate to the Rest High Level Client if you upgrade to ES 7.x+ — it uses the REST API (port 9200) and is actively maintained.
内容的提问来源于stack exchange,提问作者sai jaswanth kalavendi

