如何利用Apache POI提取Microsoft Word文档中的标题内容?
Alright, since you can already pull full text and individual paragraphs with getText() and getParagraphs(), here's a straightforward way to group content under specific headings like your sample's Title, Background, and Experience.
Step 1: Define How to Identify Headings
First, figure out what makes a paragraph a heading in your documents. From your sample, it looks like headings are short, distinct lines (e.g., Title, Background:, Experience:). You can either:
- Match exact text patterns (check if a paragraph starts with your target headings)
- Use paragraph styles (if your Word docs use built-in heading styles like Heading 1/2—this is more reliable for formatted docs)
Step 2: Traverse Paragraphs & Group Content
Loop through each paragraph, track the current active heading, and append subsequent paragraphs to that heading until you hit the next one. Here's a code example using Apache POI (a go-to library for Word processing in Java):
import org.apache.poi.xwpf.usermodel.XWPFDocument; import org.apache.poi.xwpf.usermodel.XWPFParagraph; import java.io.FileInputStream; import java.util.HashMap; import java.util.List; import java.util.Map; public class HeadingContentExtractor { public static void main(String[] args) throws Exception { // Load your Word document XWPFDocument doc = new XWPFDocument(new FileInputStream("your-document.docx")); List<XWPFParagraph> paragraphs = doc.getParagraphs(); // Map to store heading -> aggregated content Map<String, StringBuilder> headingContentMap = new HashMap<>(); String currentHeading = null; // Define your target headings (tweak this to match your doc's actual headings) String[] targetHeadings = {"Title", "Background:", "Experience:"}; for (XWPFParagraph para : paragraphs) { String paraText = para.getText().trim(); // Check if this paragraph is one of our target headings boolean isHeading = false; for (String heading : targetHeadings) { if (paraText.startsWith(heading)) { currentHeading = heading; headingContentMap.putIfAbsent(currentHeading, new StringBuilder()); isHeading = true; break; } } // If it's not a heading and we have an active heading, add content to it if (!isHeading && currentHeading != null && !paraText.isEmpty()) { headingContentMap.get(currentHeading).append(paraText).append(" "); } } // Output the extracted content per heading for (Map.Entry<String, StringBuilder> entry : headingContentMap.entrySet()) { System.out.println("--- " + entry.getKey() + " ---"); System.out.println(entry.getValue().toString().trim()); } doc.close(); } }
Step 3: Tweak for Your Document's Unique Structure
- If your headings use specific formatting (like bold or built-in Heading styles), replace the text check with a style-based check:
// Check if the paragraph uses a heading style if (para.getStyle() != null && para.getStyle().getName().startsWith("Heading")) { currentHeading = para.getText().trim(); // ... rest of the logic } - If headings have inconsistent formatting (e.g., sometimes with colons, sometimes not), use a more flexible match like
paraText.toLowerCase().contains("background").
This approach will neatly group all content under each of your target headings, making it easy to extract exactly what you need!
内容的提问来源于stack exchange,提问作者dps

