You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Java将Word文档中的项目结构转换为有效文件路径数组?

用Java将Word文档中的项目树形结构转换为文件路径数组

实现思路

  1. 读取Word文档文本:借助Apache POI库提取.docx文档中的纯文本内容,拿到树形结构的原始文本。
  2. 解析树形结构:逐行处理文本,识别每个节点的层级,用栈维护当前路径;遇到文件节点时,拼接完整路径并收集。
  3. 生成结果数组:将收集到的所有文件路径转换为字符串数组输出。

所需依赖

在Maven项目中添加Apache POI相关依赖,用于读取Word文档:

<dependencies>
    <!-- Apache POI 处理 .docx 文件 -->
    <dependency>
        <groupId>org.apache.poi</groupId>
        <artifactId>poi-ooxml</artifactId>
        <version>5.2.5</version>
    </dependency>
</dependencies>

Java代码实现

import org.apache.poi.xwpf.usermodel.XWPFDocument;
import org.apache.poi.xwpf.usermodel.XWPFParagraph;

import java.io.FileInputStream;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class ProjectTreeToPaths {

    // 读取Word文档中的文本内容
    private static String readWordDocument(String filePath) throws IOException {
        StringBuilder content = new StringBuilder();
        try (FileInputStream fis = new FileInputStream(filePath);
             XWPFDocument document = new XWPFDocument(fis)) {

            for (XWPFParagraph paragraph : document.getParagraphs()) {
                content.append(paragraph.getText()).append("\n");
            }
        }
        return content.toString();
    }

    // 解析树形结构文本,生成文件路径数组
    private static String[] parseTreeToPaths(String treeText) {
        List<String> filePaths = new ArrayList<>();
        List<String> pathStack = new ArrayList<>();
        // 匹配树形节点行的正则,提取节点名称
        Pattern pattern = Pattern.compile("^(\\│?\\s*├──\\s*)(.+)$");
        String[] lines = treeText.split("\\n");
        
        if (lines.length == 0) return new String[0];

        // 处理根目录节点
        String rootLine = lines[0].trim();
        if (!rootLine.isEmpty()) {
            pathStack.add(rootLine.replace("/", ""));
        }

        for (int i = 1; i < lines.length; i++) {
            String line = lines[i].trim();
            if (line.isEmpty()) continue;

            Matcher matcher = pattern.matcher(lines[i]);
            if (matcher.find()) {
                String nodeName = matcher.group(2).trim();
                int level = countLevel(lines[i]);

                // 调整路径栈到当前节点层级
                while (pathStack.size() > level) {
                    pathStack.remove(pathStack.size() - 1);
                }

                boolean isDirectory = nodeName.endsWith("/");
                String cleanName = isDirectory ? nodeName.replace("/", "") : nodeName;

                if (isDirectory) {
                    pathStack.add(cleanName);
                } else {
                    // 拼接完整文件路径
                    StringBuilder fullPath = new StringBuilder();
                    for (String pathPart : pathStack) {
                        fullPath.append(pathPart).append("/");
                    }
                    fullPath.append(cleanName);
                    filePaths.add(fullPath.toString());
                }
            }
        }

        return filePaths.toArray(new String[0]);
    }

    // 计算当前节点的层级
    private static int countLevel(String line) {
        // 通过统计行中「│」的数量确定层级,根节点为0层
        int pipeCount = (int) line.chars().filter(c -> c == '│').count();
        return pipeCount + 1;
    }

    public static void main(String[] args) {
        try {
            // 替换为你的Word文档路径
            String wordFilePath = "project-structure.docx";
            String treeText = readWordDocument(wordFilePath);
            String[] filePaths = parseTreeToPaths(treeText);

            // 输出格式化为示例要求的数组形式
            System.out.println("String[] output = [");
            for (int i = 0; i < filePaths.length; i++) {
                System.out.print("    \"" + filePaths[i] + "\"");
                if (i != filePaths.length - 1) {
                    System.out.println(",");
                }
            }
            System.out.println("\n];");
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

代码说明

  • readWordDocument:读取.docx文件的所有段落文本,拼接成树形结构的原始字符串。
  • parseTreeToPaths:用正则匹配每行的节点,通过栈维护当前路径层级;遇到文件节点时,拼接完整路径并加入结果列表。
  • countLevel:统计行中「│」的数量来确定节点层级,以此调整路径栈的长度,保证路径的正确性。
  • 目录与文件区分:通过节点名是否以/结尾判断类型,目录加入路径栈,文件则生成完整路径并收集。

内容的提问来源于stack exchange,提问作者DilliBabu S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 06:45:16