You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中不使用Logstash将CSV数据导入Elasticsearch的方法及入门指引

不用Logstash导入CSV到Elasticsearch的方案 & TransportClient入门指南

Hey there! Let's break down your questions step by step — first, practical ways to import CSV data into Elasticsearch without relying on Logstash, then a guide to getting started with TransportClient (with an important heads-up: it's deprecated in Elasticsearch 7.0+, but we'll cover both the basics and the modern alternative note).


不用Logstash导入CSV的可行方法

1. 手动使用Elasticsearch _bulk API

This is a lightweight option if you don't want to add extra dependencies:

  • First, convert your CSV into the bulk API format: each operation line is followed by the data line. Example:
    {"index": {"_index": "your_target_index", "_id": "1"}}
    {"name": "Alice", "age": 30, "city": "New York"}
    {"index": {"_index": "your_target_index", "_id": "2"}}
    {"name": "Bob", "age": 28, "city": "London"}
    
  • You can write a simple script (Python/Shell) to parse the CSV and generate this format. For example, in Python, use the csv module to read rows, then construct the bulk lines.
  • Send the formatted data via a POST request to http://<es-host>:<port>/_bulk (use tools like curl or libraries like requests in Python).
  • Pros: No extra tools needed; Cons: Requires handling data conversion and error checking yourself.

2. Use Elasticsearch's official language clients

The official clients (Python, Java, etc.) let you directly read CSV and index data with more flexibility:

Python Client Example

First install the client:

pip install elasticsearch

Then write a script to import:

from elasticsearch import Elasticsearch
import csv

# Connect to your ES cluster
es = Elasticsearch(["http://localhost:9200"])
index_name = "user_data"

# Create index (if it doesn't exist) with mappings
if not es.indices.exists(index=index_name):
    es.indices.create(
        index=index_name,
        body={
            "mappings": {
                "properties": {
                    "name": {"type": "text"},
                    "age": {"type": "integer"},
                    "city": {"type": "keyword"}
                }
            }
        }
    )

# Read CSV and index rows
with open("users.csv", "r") as csv_file:
    reader = csv.DictReader(csv_file)
    for row in reader:
        # Convert data types as needed
        row["age"] = int(row["age"])
        es.index(index=index_name, document=row)
  • For large datasets, use the bulk method to batch requests and improve performance.
  • Pros: Full control over data cleaning and mapping; Cons: Requires basic coding knowledge.

3. Kibana's Visual CSV Importer (if you have Kibana access)

If you prefer a no-code approach:

  • Open Kibana → Go to Management → Index Patterns → Click Import data
  • Upload your CSV file, follow the wizard to configure field mappings, set the index name, and let Kibana handle the rest.
  • Pros: No coding required, visual feedback; Cons: Dependent on Kibana, better for small to medium datasets.

TransportClient入门指南 (Important: Deprecated in 7.0+)

First, a critical note: TransportClient was deprecated in Elasticsearch 7.0 and removed in 8.0. The official recommendation is to use the Rest High Level Client instead. But if you're working with an older ES version (6.x or earlier), here's how to get started:

1. Set up dependencies (Java example)

If using Maven, add these dependencies (match your ES version exactly):

<dependency>
    <groupId>org.elasticsearch.client</groupId>
    <artifactId>transport</artifactId>
    <version>6.8.23</version>
</dependency>
<dependency>
    <groupId>org.elasticsearch</groupId>
    <artifactId>elasticsearch</artifactId>
    <version>6.8.23</version>
</dependency>

2. Initialize the TransportClient

import org.elasticsearch.client.transport.TransportClient;
import org.elasticsearch.common.settings.Settings;
import org.elasticsearch.common.transport.TransportAddress;
import org.elasticsearch.transport.client.PreBuiltTransportClient;
import java.net.InetAddress;
import java.net.UnknownHostException;

public class ESClientSetup {
    public static void main(String[] args) throws UnknownHostException {
        // Match your ES cluster's name (from elasticsearch.yml)
        Settings settings = Settings.builder()
                .put("cluster.name", "your_cluster_name")
                .build();

        // Connect to the ES node (uses TCP port 9300 by default)
        TransportClient client = new PreBuiltTransportClient(settings)
                .addTransportAddress(new TransportAddress(InetAddress.getByName("localhost"), 9300));

        // Your operations go here...

        // Always close the client when done
        client.close();
    }
}

3. Import CSV with TransportClient (Java example)

Use a CSV library like OpenCSV to read data, then batch index with BulkRequest:

import com.opencsv.CSVReader;
import org.elasticsearch.action.bulk.BulkRequest;
import org.elasticsearch.action.bulk.BulkResponse;
import org.elasticsearch.action.index.IndexRequest;
import java.io.FileReader;
import java.util.HashMap;
import java.util.Map;

// Inside your main method after initializing the client
CSVReader reader = new CSVReader(new FileReader("users.csv"));
String[] headers = reader.readNext(); // Get CSV headers
String[] row;
BulkRequest bulkRequest = new BulkRequest();

while ((row = reader.readNext()) != null) {
    Map<String, Object> document = new HashMap<>();
    for (int i = 0; i < headers.length; i++) {
        // Convert data types (e.g., age to integer)
        if (headers[i].equals("age")) {
            document.put(headers[i], Integer.parseInt(row[i]));
        } else {
            document.put(headers[i], row[i]);
        }
    }
    // Add index request to bulk batch
    bulkRequest.add(new IndexRequest("user_data").source(document));
}

// Execute the bulk import
BulkResponse response = client.bulk(bulkRequest).actionGet();
if (response.hasFailures()) {
    // Handle errors
    System.out.println("Bulk import failed: " + response.buildFailureMessage());
}

Key Notes for TransportClient

  • Ensure your ES node's transport.host setting allows incoming connections from your client.
  • Version consistency is critical: The client version must match the ES cluster version exactly to avoid compatibility issues.
  • Plan to migrate to the Rest High Level Client if you upgrade to ES 7.x+ — it uses the REST API (port 9200) and is actively maintained.

内容的提问来源于stack exchange,提问作者sai jaswanth kalavendi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:20:50