基于Spring Boot Spring Batch实现大CSV文件的部分数据加载与导出
用Spring Boot + Spring Batch处理大型CSV并导出指定行(第50-100行)
当然可以!Spring Batch本来就是为批量数据处理场景设计的,对付大型CSV完全不在话下,还能精准控制要导出的行范围。下面我一步步给你讲怎么实现:
1. 先准备项目依赖
首先在你的pom.xml(Maven)或者build.gradle里加上Spring Boot和Spring Batch的核心依赖,还有CSV处理相关的包:
Maven示例:
<dependencies> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-batch</artifactId> </dependency> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter</artifactId> </dependency> <!-- 用于CSV解析的基础包 --> <dependency> <groupId>org.springframework.batch</groupId> <artifactId>spring-batch-infrastructure</artifactId> </dependency> </dependencies>
2. 定义数据模型
先创建一个实体类来映射CSV的每行数据,假设你的CSV是用户信息,示例如下:
public class User { private String id; private String name; private String email; // 生成getter、setter、toString方法,或者用Lombok的@Data注解简化 }
3. 配置Batch Job和核心步骤
核心逻辑就在这里:用FlatFileItemReader控制读取的行范围。要注意两个关键配置:
setLinesToSkip:跳过前N行(如果CSV有表头,要先跳过表头,再算目标起始行的偏移量)setMaxItemCount:设置最多读取的行数(目标结束行 - 起始行 + 1)
下面是完整的配置类:
import org.springframework.batch.core.Job; import org.springframework.batch.core.Step; import org.springframework.batch.core.configuration.annotation.EnableBatchProcessing; import org.springframework.batch.core.configuration.annotation.JobBuilderFactory; import org.springframework.batch.core.configuration.annotation.StepBuilderFactory; import org.springframework.batch.item.file.FlatFileItemReader; import org.springframework.batch.item.file.FlatFileItemWriter; import org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper; import org.springframework.batch.item.file.mapping.DefaultLineMapper; import org.springframework.batch.item.file.transform.DelimitedLineAggregator; import org.springframework.batch.item.file.transform.DelimitedLineTokenizer; import org.springframework.batch.item.file.transform.BeanWrapperFieldExtractor; import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.core.io.FileSystemResource; @Configuration @EnableBatchProcessing public class CsvBatchConfig { private final JobBuilderFactory jobBuilderFactory; private final StepBuilderFactory stepBuilderFactory; // 构造注入工厂类 public CsvBatchConfig(JobBuilderFactory jobBuilderFactory, StepBuilderFactory stepBuilderFactory) { this.jobBuilderFactory = jobBuilderFactory; this.stepBuilderFactory = stepBuilderFactory; } // 读取源CSV的Reader @Bean public FlatFileItemReader<User> csvItemReader() { FlatFileItemReader<User> reader = new FlatFileItemReader<>(); // 替换成你的大型CSV文件路径 reader.setResource(new FileSystemResource("src/main/resources/large-data.csv")); // 先跳过表头(如果你的CSV没有表头,这行删掉) int skipHeader = 1; reader.setLinesToSkip(skipHeader); // 重点:跳过前49行(加上表头的1行,刚好从第50行开始读取) // 如果没有表头,这里直接设置为49即可 reader.setLinesToSkip(reader.getLinesToSkip() + 49); // 读取51行(第50到100行,共100-50+1=51行) reader.setMaxItemCount(51); // 配置行映射规则,把CSV行转成User对象 DefaultLineMapper<User> lineMapper = new DefaultLineMapper<>(); DelimitedLineTokenizer tokenizer = new DelimitedLineTokenizer(); // 对应CSV的列名,要和User类的字段一一对应 tokenizer.setNames("id", "name", "email"); BeanWrapperFieldSetMapper<User> fieldSetMapper = new BeanWrapperFieldSetMapper<>(); fieldSetMapper.setTargetType(User.class); lineMapper.setLineTokenizer(tokenizer); lineMapper.setFieldSetMapper(fieldSetMapper); reader.setLineMapper(lineMapper); return reader; } // 导出目标CSV的Writer @Bean public FlatFileItemWriter<User> csvItemWriter() { FlatFileItemWriter<User> writer = new FlatFileItemWriter<>(); // 替换成你要导出的文件路径 writer.setResource(new FileSystemResource("src/main/resources/exported-data.csv")); // 配置行转换规则,把User对象转成CSV行 writer.setLineAggregator(new DelimitedLineAggregator<User>() {{ setDelimiter(","); setFieldExtractor(new BeanWrapperFieldExtractor<User>() {{ setNames(new String[]{"id", "name", "email"}); }}); }}); // 写入前清空目标文件 writer.setAppendAllowed(false); // 写入表头(如果不需要表头,这行删掉) writer.setHeaderCallback(writer1 -> writer1.write("id,name,email")); return writer; } // 定义处理步骤 @Bean public Step csvProcessingStep() { return stepBuilderFactory.get("csvProcessingStep") .<User, User>chunk(10) // 每次批量处理10行,可根据内存情况调整 .reader(csvItemReader()) .writer(csvItemWriter()) .build(); } // 定义整个Job @Bean public Job csvExportJob() { return jobBuilderFactory.get("csvExportJob") .start(csvProcessingStep()) .build(); } }
4. 运行Job
你可以在项目启动时自动触发Job,或者用接口触发,这里给个启动时触发的示例:
import org.springframework.batch.core.Job; import org.springframework.batch.core.JobParameters; import org.springframework.batch.core.JobParametersBuilder; import org.springframework.batch.core.launch.JobLauncher; import org.springframework.boot.CommandLineRunner; import org.springframework.boot.SpringApplication; import org.springframework.boot.autoconfigure.SpringBootApplication; import org.springframework.context.annotation.Bean; @SpringBootApplication public class CsvBatchApplication { public static void main(String[] args) { SpringApplication.run(CsvBatchApplication.class, args); } @Bean public CommandLineRunner runJob(JobLauncher jobLauncher, Job csvExportJob) { return args -> { // 加时间戳确保每次Job参数唯一,避免重复执行 JobParameters jobParameters = new JobParametersBuilder() .addLong("time", System.currentTimeMillis()) .toJobParameters(); jobLauncher.run(csvExportJob, jobParameters); System.out.println("指定行CSV导出完成!"); }; } }
一些实用注意事项
- 如果CSV没有表头,记得删掉
reader.setLinesToSkip(1),直接设置reader.setLinesToSkip(49) chunk(10)的大小可以根据内存情况调整:值越大处理速度越快,但占用内存也越多- Spring Batch是流式处理,不会把整个大型CSV加载到内存,不用担心内存溢出问题
- 如果需要动态传入起始行和结束行,可以把这两个参数做成
JobParameters,在启动Job时传入,然后在Reader里动态设置setLinesToSkip和setMaxItemCount
内容的提问来源于stack exchange,提问作者mimi nette
相关产品推荐
相关产品推荐

