You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Spring Boot Spring Batch实现大CSV文件的部分数据加载与导出

用Spring Boot + Spring Batch处理大型CSV并导出指定行(第50-100行)

当然可以!Spring Batch本来就是为批量数据处理场景设计的,对付大型CSV完全不在话下,还能精准控制要导出的行范围。下面我一步步给你讲怎么实现:

1. 先准备项目依赖

首先在你的pom.xml(Maven)或者build.gradle里加上Spring Boot和Spring Batch的核心依赖,还有CSV处理相关的包:

Maven示例:

<dependencies>
    <dependency>
        <groupId>org.springframework.boot</groupId>
        <artifactId>spring-boot-starter-batch</artifactId>
    </dependency>
    <dependency>
        <groupId>org.springframework.boot</groupId>
        <artifactId>spring-boot-starter</artifactId>
    </dependency>
    <!-- 用于CSV解析的基础包 -->
    <dependency>
        <groupId>org.springframework.batch</groupId>
        <artifactId>spring-batch-infrastructure</artifactId>
    </dependency>
</dependencies>

2. 定义数据模型

先创建一个实体类来映射CSV的每行数据,假设你的CSV是用户信息,示例如下:

public class User {
    private String id;
    private String name;
    private String email;

    // 生成getter、setter、toString方法,或者用Lombok的@Data注解简化
}

3. 配置Batch Job和核心步骤

核心逻辑就在这里:用FlatFileItemReader控制读取的行范围。要注意两个关键配置:

  • setLinesToSkip:跳过前N行(如果CSV有表头,要先跳过表头,再算目标起始行的偏移量)
  • setMaxItemCount:设置最多读取的行数(目标结束行 - 起始行 + 1)

下面是完整的配置类:

import org.springframework.batch.core.Job;
import org.springframework.batch.core.Step;
import org.springframework.batch.core.configuration.annotation.EnableBatchProcessing;
import org.springframework.batch.core.configuration.annotation.JobBuilderFactory;
import org.springframework.batch.core.configuration.annotation.StepBuilderFactory;
import org.springframework.batch.item.file.FlatFileItemReader;
import org.springframework.batch.item.file.FlatFileItemWriter;
import org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper;
import org.springframework.batch.item.file.mapping.DefaultLineMapper;
import org.springframework.batch.item.file.transform.DelimitedLineAggregator;
import org.springframework.batch.item.file.transform.DelimitedLineTokenizer;
import org.springframework.batch.item.file.transform.BeanWrapperFieldExtractor;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.core.io.FileSystemResource;

@Configuration
@EnableBatchProcessing
public class CsvBatchConfig {

    private final JobBuilderFactory jobBuilderFactory;
    private final StepBuilderFactory stepBuilderFactory;

    // 构造注入工厂类
    public CsvBatchConfig(JobBuilderFactory jobBuilderFactory, StepBuilderFactory stepBuilderFactory) {
        this.jobBuilderFactory = jobBuilderFactory;
        this.stepBuilderFactory = stepBuilderFactory;
    }

    // 读取源CSV的Reader
    @Bean
    public FlatFileItemReader<User> csvItemReader() {
        FlatFileItemReader<User> reader = new FlatFileItemReader<>();
        // 替换成你的大型CSV文件路径
        reader.setResource(new FileSystemResource("src/main/resources/large-data.csv"));
        
        // 先跳过表头(如果你的CSV没有表头,这行删掉)
        int skipHeader = 1;
        reader.setLinesToSkip(skipHeader);
        
        // 重点:跳过前49行(加上表头的1行,刚好从第50行开始读取)
        // 如果没有表头,这里直接设置为49即可
        reader.setLinesToSkip(reader.getLinesToSkip() + 49);
        // 读取51行(第50到100行,共100-50+1=51行)
        reader.setMaxItemCount(51);

        // 配置行映射规则,把CSV行转成User对象
        DefaultLineMapper<User> lineMapper = new DefaultLineMapper<>();
        DelimitedLineTokenizer tokenizer = new DelimitedLineTokenizer();
        // 对应CSV的列名,要和User类的字段一一对应
        tokenizer.setNames("id", "name", "email");
        
        BeanWrapperFieldSetMapper<User> fieldSetMapper = new BeanWrapperFieldSetMapper<>();
        fieldSetMapper.setTargetType(User.class);

        lineMapper.setLineTokenizer(tokenizer);
        lineMapper.setFieldSetMapper(fieldSetMapper);
        reader.setLineMapper(lineMapper);

        return reader;
    }

    // 导出目标CSV的Writer
    @Bean
    public FlatFileItemWriter<User> csvItemWriter() {
        FlatFileItemWriter<User> writer = new FlatFileItemWriter<>();
        // 替换成你要导出的文件路径
        writer.setResource(new FileSystemResource("src/main/resources/exported-data.csv"));
        
        // 配置行转换规则,把User对象转成CSV行
        writer.setLineAggregator(new DelimitedLineAggregator<User>() {{
            setDelimiter(",");
            setFieldExtractor(new BeanWrapperFieldExtractor<User>() {{
                setNames(new String[]{"id", "name", "email"});
            }});
        }});
        
        // 写入前清空目标文件
        writer.setAppendAllowed(false);
        // 写入表头(如果不需要表头,这行删掉)
        writer.setHeaderCallback(writer1 -> writer1.write("id,name,email"));

        return writer;
    }

    // 定义处理步骤
    @Bean
    public Step csvProcessingStep() {
        return stepBuilderFactory.get("csvProcessingStep")
                .<User, User>chunk(10) // 每次批量处理10行,可根据内存情况调整
                .reader(csvItemReader())
                .writer(csvItemWriter())
                .build();
    }

    // 定义整个Job
    @Bean
    public Job csvExportJob() {
        return jobBuilderFactory.get("csvExportJob")
                .start(csvProcessingStep())
                .build();
    }
}

4. 运行Job

你可以在项目启动时自动触发Job,或者用接口触发,这里给个启动时触发的示例:

import org.springframework.batch.core.Job;
import org.springframework.batch.core.JobParameters;
import org.springframework.batch.core.JobParametersBuilder;
import org.springframework.batch.core.launch.JobLauncher;
import org.springframework.boot.CommandLineRunner;
import org.springframework.boot.SpringApplication;
import org.springframework.boot.autoconfigure.SpringBootApplication;
import org.springframework.context.annotation.Bean;

@SpringBootApplication
public class CsvBatchApplication {

    public static void main(String[] args) {
        SpringApplication.run(CsvBatchApplication.class, args);
    }

    @Bean
    public CommandLineRunner runJob(JobLauncher jobLauncher, Job csvExportJob) {
        return args -> {
            // 加时间戳确保每次Job参数唯一,避免重复执行
            JobParameters jobParameters = new JobParametersBuilder()
                    .addLong("time", System.currentTimeMillis())
                    .toJobParameters();
            jobLauncher.run(csvExportJob, jobParameters);
            System.out.println("指定行CSV导出完成!");
        };
    }
}

一些实用注意事项

  • 如果CSV没有表头,记得删掉reader.setLinesToSkip(1),直接设置reader.setLinesToSkip(49)
  • chunk(10)的大小可以根据内存情况调整:值越大处理速度越快,但占用内存也越多
  • Spring Batch是流式处理,不会把整个大型CSV加载到内存,不用担心内存溢出问题
  • 如果需要动态传入起始行和结束行,可以把这两个参数做成JobParameters,在启动Job时传入,然后在Reader里动态设置setLinesToSkip和setMaxItemCount

内容的提问来源于stack exchange,提问作者mimi nette

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:20:28