You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Java/JavaScript中读取Word文件并获取格式与数学表达式信息

读取Word文档中的数学表达式与格式信息问题

我尝试读取一份包含文本格式(粗体、斜体、下划线、上下标)和数学表达式的Word文档,文档内容截图如下:
Word文档内容截图

尝试1:使用XWPFRun读取

我首先尝试用XWPFRun来读取文档内容,代码如下:

public String[] readStringFromFile(String absolutePath) throws Exception {
    FileInputStream fileInputStream = new FileInputStream(absolutePath);
    XWPFDocument document = new XWPFDocument(fileInputStream);
    List<XWPFParagraph> paragraphs = document.getParagraphs();
    List<String> strings = new ArrayList<>();
    
    for (XWPFParagraph paragraph : paragraphs) {
        strings.add(paragraph.getText());
        for (XWPFRun run :
                paragraph.getRuns()) {
            System.out.println("Run: " + run.text());
            System.out.println("Run infos:");
            System.out.println("Bold: " + run.isBold() + " Italic: " + run.isItalic() + " Underlined: " + run.getUnderline());
            System.out.println("Superscript/Subscript: " + run.getVerticalAlignment() + "\n");
        }
    }
    document.close();
    return strings.toArray(new String[0]);
}

输出结果:

Run: This one is a test Docx file.
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: Math: 
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: This is bold
Run infos:
Bold: true Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: This one’s italic
Run infos:
Bold: false Italic: true Underlined: NONE
Superscript/Subscript: baseline

Run: Underlined
Run infos:
Bold: false Italic: false Underlined: SINGLE
Superscript/Subscript: baseline

Run: Bold Italic and underlined
Run infos:
Bold: true Italic: true Underlined: SINGLE
Superscript/Subscript: baseline

Run: Bold and italic
Run infos:
Bold: true Italic: true Underlined: NONE
Superscript/Subscript: baseline

Run: In Same Line: 
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: Bold
Run infos:
Bold: true Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run:  
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: Italic
Run infos:
Bold: false Italic: true Underlined: NONE
Superscript/Subscript: baseline

Run:  
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: Underlined
Run infos:
Bold: false Italic: false Underlined: SINGLE
Superscript/Subscript: baseline

Run:  
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: Bold and Italic
Run infos:
Bold: true Italic: true Underlined: NONE
Superscript/Subscript: baseline

Run: W
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: e have some
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: superscript
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: superscript

Run:  and
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: baseline

Run: subscript
Run infos:
Bold: false Italic: false Underlined: NONE
Superscript/Subscript: subscript

这种方式无法获取任何数学信息。

尝试2:使用Apache Tika读取并生成HTML

之后我尝试用Apache Tika解析文档并生成HTML,代码如下:

private void getHtmlUsingTika(String absolutePath) throws IOException, TikaException, SAXException {
    ContentHandler handler = new ToXMLContentHandler();
    AutoDetectParser parser = new AutoDetectParser();
    Metadata metadata = new Metadata();

    InputStream stream = new FileInputStream(new File(absolutePath));

    parser.parse(stream,handler,metadata);
    System.out.println(handler.toString());
}

输出的HTML内容:

<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<meta name="cp:revision" content="25" />
<meta name="extended-properties:AppVersion" content="15.0000" />
<meta name="meta:paragraph-count" content="1" />
<meta name="meta:word-count" content="33" />
<meta name="extended-properties:Application" content="Microsoft Office Word" />
<meta name="meta:last-author" content="Microsoft account" />
<meta name="extended-properties:Company" content="" />
<meta name="xmpTPg:NPages" content="1" />
<meta name="dcterms:created" content="2022-04-25T09:09:00Z" />
<meta name="meta:line-count" content="1" />
<meta name="dcterms:modified" content="2022-08-28T07:08:00Z" />
<meta name="meta:character-count" content="189" />
<meta name="extended-properties:Template" content="Normal.dotm" />
<meta name="meta:character-count-with-spaces" content="221" />
<meta name="X-TIKA:Parsed-By" content="org.apache.tika.parser.DefaultParser" />
<meta name="X-TIKA:Parsed-By" content="org.apache.tika.parser.microsoft.ooxml.OOXMLParser" />
<meta name="extended-properties:DocSecurityString" content="None" />
<meta name="extended-properties:TotalTime" content="35" />
<meta name="meta:page-count" content="1" />
<meta name="Content-Type" content="application/vnd.openxmlformats-officedocument.wordprocessingml.document" />
<meta name="dc:publisher" content="" />
<title></title>
</head>
<body><p>This one is a test Docx file.</p>
<p>Math: </p>
<p><b>This is bold</b></p>
<p><i>This one’s italic</i></p>
<p><u>Underlined</u></p>
<p><b><i><u>Bold Italic and underlined</u></i></b></p>
<p><b><i>Bold and italic</i></b></p>
<p>In Same Line: <b>Bold</b> <i>Italic</i> <u>Underlined</u> <b><i>Bold and Italic</i></b></p>
<p>We have somesuperscript andsubscript</p>
<p><a name="_GoBack" /></p>
</body></html>

这种方式同样无法获取数学信息,而且连上下标信息也丢失了。

我的问题

我是处理Word文档的新手,已经被这个问题困扰一段时间了。有没有办法在Java中同时获取数学表达式和其他格式信息?另外,使用JavaScript是否也能实现?


内容的提问来源于stack exchange,提问作者Adnan Bin Zahir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 07:27:22