Spring Boot+Spring Batch多源数据写入MongoDB:Reader选型咨询
Hey Jeff, great question! Let's walk through whether AbstractItemStreamItemReader fits your use case, plus how to tackle your Spring Batch job overall.
First, a quick clarification: AbstractItemStreamItemReader is an abstract base class, not a concrete reader you'd use directly. It implements the ItemStream interface, which handles critical batch features like restartability, opening/closing resources at transaction boundaries, and state management.
Good news: The concrete readers you'll need for your scenario are all subclasses of AbstractItemStreamItemReader, so you'll be leveraging its capabilities indirectly (and correctly) without having to implement it yourself.
Let's map your data sources to the right Spring Batch readers:
- Oracle Employee Data: Use
JdbcCursorItemReaderorJdbcPagingItemReader(both subclasses ofAbstractItemStreamItemReader).JdbcCursorItemReaderworks well for small-to-medium datasets, using a JDBC cursor to fetch rows incrementally.JdbcPagingItemReaderis better for large datasets, as it fetches data in pages to avoid memory overload.
- CSV Department Data: Use
FlatFileItemReader(also a subclass ofAbstractItemStreamItemReader). It's purpose-built for delimited files like CSV, supports header skipping, field mapping, and handles resource management out of the box.
The core task here is associating each employee with their corresponding department data from the CSV. Here are two practical approaches:
Approach 1: Load Departments to Memory (Small Dataset)
If your department CSV is small enough to fit in memory, pre-load it into a Map<DeptId, Department> before processing employees. This is simple and efficient:
- Create a
FlatFileItemReaderto read departments. - Load all departments into a HashMap (you can do this in a
@Beanmethod or a pre-processing step). - In an
ItemProcessor, look up the department for each employee using theirdeptIdand embed it into the employee object.
Code Snippets for This Approach
Department Reader
@Bean public FlatFileItemReader<Department> departmentReader() { return new FlatFileItemReaderBuilder<Department>() .name("departmentReader") .resource(new ClassPathResource("departments.csv")) .delimited() .names("deptId", "deptName", "location") .fieldSetMapper(new BeanWrapperFieldSetMapper<Department>() {{ setTargetType(Department.class); }}) .build(); }
Pre-Load Departments to Map
@Bean public Map<String, Department> departmentLookupMap(ItemReader<Department> departmentReader) throws Exception { Map<String, Department> deptMap = new HashMap<>(); Department dept; while ((dept = departmentReader.read()) != null) { deptMap.put(dept.getDeptId(), dept); } return deptMap; }
Employee Reader (Oracle)
@Bean public JdbcCursorItemReader<Employee> employeeReader(DataSource oracleDataSource) { return new JdbcCursorItemReaderBuilder<Employee>() .name("employeeReader") .dataSource(oracleDataSource) .sql("SELECT emp_id, emp_name, dept_id FROM employees") .rowMapper((rs, rowNum) -> { Employee emp = new Employee(); emp.setEmpId(rs.getString("emp_id")); emp.setEmpName(rs.getString("emp_name")); emp.setDeptId(rs.getString("dept_id")); return emp; }) .build(); }
Processor to Embed Department
@Bean public ItemProcessor<Employee, Employee> employeeDepartmentProcessor(Map<String, Department> deptMap) { return employee -> { // Look up department by employee's deptId Department associatedDept = deptMap.get(employee.getDeptId()); if (associatedDept != null) { employee.setDepartment(associatedDept); // Employee class has an embedded Department field } return employee; }; }
MongoDB Writer
@Bean public MongoItemWriter<Employee> mongoEmployeeWriter(MongoTemplate mongoTemplate) { MongoItemWriter<Employee> writer = new MongoItemWriter<>(); writer.setTemplate(mongoTemplate); writer.setCollection("employees"); return writer; }
Approach 2: Persist Departments First (Large Dataset)
If your department CSV is too big for memory:
- Create a separate step to load departments into MongoDB first.
- In the employee processing step, use a
MongoTemplate(or a custom DAO) to look up departments directly from MongoDB as you process each employee. - This avoids loading all departments into memory, but adds a small overhead per employee lookup.
Yes, it's the right foundation for your job! You won't instantiate it directly, but all the readers you'll use inherit from it, giving you the restartability and resource management you need for robust batch processing.
内容的提问来源于stack exchange,提问作者Jeff Cook

