优化重复代码:多标签类别文档匹配代码简化方案问询
Hey, nice work on extracting that repetitive loop into a generic method—you’re already on the right track! Let’s level this up with some more elegant, maintainable optimizations that’ll make your code cleaner and easier to extend later.
1. Encapsulate Tag Categories for Unified Processing
Instead of handling each tag list separately with manual print statements, create a simple structure to pair each tag list with its display name. This lets you loop through all tag categories in one go, eliminating duplicate code for headers and checks.
For Java 16+, use a record (or a simple class if you’re on an older version):
// Define this once outside your processing loop record TagCategory(String displayName, List<String> tags) {} // Initialize all your tag categories in a single list List<TagCategory> tagCategories = List.of( new TagCategory("Keywords", KEYWORDS), new TagCategory("Customers", CUSTOMERS), new TagCategory("System dependencies", SYSTEM_DEPS), new TagCategory("Modules", MODULES), new TagCategory("Drive definitions", DRIVE_DEFS), new TagCategory("Process IDs", PROCESS_IDS) );
Then in your document processing code, replace all those separate loops with:
String documentText = wordExtractor.getText(); // Fetch text once, reuse it! for (TagCategory category : tagCategories) { System.out.printf("\n%s found in the document:\n", category.displayName()); category.tags().stream() .filter(documentText::contains) .forEach(System.out::println); }
2. Preprocess Document Text to Cut Redundant Work
Calling wordExtractor.getText() every time in your loop can be inefficient, especially for large documents. Grab the text once at the start of processing a document and reuse that string everywhere—this reduces unnecessary method calls and improves performance.
3. Optimize Database Queries to Reduce Round-Trips
Right now you’re running separate SQL queries for each tag type. Instead, fetch all tags in one query and group them by their type ID—this cuts down on database calls, which is better for performance and cleaner code.
Replace your multiple queries with this:
// Fetch all tags in one go String sql = "SELECT da_tag_name, da_tag_type_id FROM da_tags"; stmt = conn.prepareStatement(sql); ResultSet resultSet = stmt.executeQuery(); // Group tags by their type ID using a map Map<Integer, List<String>> tagsByTypeId = new HashMap<>(); while (resultSet.next()) { String tagName = resultSet.getString("da_tag_name"); int typeId = resultSet.getInt("da_tag_type_id"); tagsByTypeId.computeIfAbsent(typeId, k -> new ArrayList<>()).add(tagName); } // Populate your lists from the grouped map List<String> CUSTOMERS = tagsByTypeId.getOrDefault(/* insert your customer type ID here */, Collections.emptyList()); List<String> SYSTEM_DEPS = tagsByTypeId.getOrDefault(6, Collections.emptyList()); // Repeat for other tag types using their respective IDs
4. Ditch the Generic Method for Stream Operations
If you prefer a functional style, you can replace your genericForEachLoop method with a concise stream one-liner. This keeps the logic inline without needing a separate helper method, and pairs perfectly with the TagCategory approach above.
Final Optimized Code Snippet
Here’s how your document processing code looks when combining all these improvements:
if(file_name.endsWith(".docx") || file_name.endsWith(".docm")) { System.out.println(file_name); XWPFDocument document = new XWPFDocument(resultSet.getBinaryStream(3)); XWPFWordExtractor wordExtractor = new XWPFWordExtractor(document); String documentText = wordExtractor.getText(); // Process all tag categories in one loop for (TagCategory category : tagCategories) { System.out.printf("\n%s found in the document:\n", category.displayName()); category.tags().stream() .filter(documentText::contains) .forEach(System.out::println); } System.out.println("Log number: " + findLogNumber(wordExtractor)); System.out.println("------------------------------------------"); wordExtractor.close(); }
This approach is:
- Maintainable: Adding a new tag type only requires updating the
TagCategorylist and database grouping—no new loops or print statements needed. - Efficient: Reduces redundant text fetching and database round-trips.
- Clean: Eliminates repetitive boilerplate code that’s prone to typos or missed updates.
内容的提问来源于stack exchange,提问作者dsafas fsafasfsa

