You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Apache POI提取Microsoft Word文档中的标题内容?

Extract Content by Headings from Word Documents

Alright, since you can already pull full text and individual paragraphs with getText() and getParagraphs(), here's a straightforward way to group content under specific headings like your sample's Title, Background, and Experience.

Step 1: Define How to Identify Headings

First, figure out what makes a paragraph a heading in your documents. From your sample, it looks like headings are short, distinct lines (e.g., Title, Background:, Experience:). You can either:

  • Match exact text patterns (check if a paragraph starts with your target headings)
  • Use paragraph styles (if your Word docs use built-in heading styles like Heading 1/2—this is more reliable for formatted docs)

Step 2: Traverse Paragraphs & Group Content

Loop through each paragraph, track the current active heading, and append subsequent paragraphs to that heading until you hit the next one. Here's a code example using Apache POI (a go-to library for Word processing in Java):

import org.apache.poi.xwpf.usermodel.XWPFDocument;
import org.apache.poi.xwpf.usermodel.XWPFParagraph;
import java.io.FileInputStream;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

public class HeadingContentExtractor {
    public static void main(String[] args) throws Exception {
        // Load your Word document
        XWPFDocument doc = new XWPFDocument(new FileInputStream("your-document.docx"));
        List<XWPFParagraph> paragraphs = doc.getParagraphs();
        
        // Map to store heading -> aggregated content
        Map<String, StringBuilder> headingContentMap = new HashMap<>();
        String currentHeading = null;
        
        // Define your target headings (tweak this to match your doc's actual headings)
        String[] targetHeadings = {"Title", "Background:", "Experience:"};
        
        for (XWPFParagraph para : paragraphs) {
            String paraText = para.getText().trim();
            
            // Check if this paragraph is one of our target headings
            boolean isHeading = false;
            for (String heading : targetHeadings) {
                if (paraText.startsWith(heading)) {
                    currentHeading = heading;
                    headingContentMap.putIfAbsent(currentHeading, new StringBuilder());
                    isHeading = true;
                    break;
                }
            }
            
            // If it's not a heading and we have an active heading, add content to it
            if (!isHeading && currentHeading != null && !paraText.isEmpty()) {
                headingContentMap.get(currentHeading).append(paraText).append(" ");
            }
        }
        
        // Output the extracted content per heading
        for (Map.Entry<String, StringBuilder> entry : headingContentMap.entrySet()) {
            System.out.println("--- " + entry.getKey() + " ---");
            System.out.println(entry.getValue().toString().trim());
        }
        
        doc.close();
    }
}

Step 3: Tweak for Your Document's Unique Structure

  • If your headings use specific formatting (like bold or built-in Heading styles), replace the text check with a style-based check:
    // Check if the paragraph uses a heading style
    if (para.getStyle() != null && para.getStyle().getName().startsWith("Heading")) {
        currentHeading = para.getText().trim();
        // ... rest of the logic
    }
    
  • If headings have inconsistent formatting (e.g., sometimes with colons, sometimes not), use a more flexible match like paraText.toLowerCase().contains("background").

This approach will neatly group all content under each of your target headings, making it easy to extract exactly what you need!

内容的提问来源于stack exchange,提问作者dps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:20:49