如何用Java将Word文档中的项目结构转换为有效文件路径数组?
用Java将Word文档中的项目树形结构转换为文件路径数组
实现思路
- 读取Word文档文本:借助Apache POI库提取
.docx文档中的纯文本内容,拿到树形结构的原始文本。 - 解析树形结构:逐行处理文本,识别每个节点的层级,用栈维护当前路径;遇到文件节点时,拼接完整路径并收集。
- 生成结果数组:将收集到的所有文件路径转换为字符串数组输出。
所需依赖
在Maven项目中添加Apache POI相关依赖,用于读取Word文档:
<dependencies> <!-- Apache POI 处理 .docx 文件 --> <dependency> <groupId>org.apache.poi</groupId> <artifactId>poi-ooxml</artifactId> <version>5.2.5</version> </dependency> </dependencies>
Java代码实现
import org.apache.poi.xwpf.usermodel.XWPFDocument; import org.apache.poi.xwpf.usermodel.XWPFParagraph; import java.io.FileInputStream; import java.io.IOException; import java.util.ArrayList; import java.util.List; import java.util.regex.Matcher; import java.util.regex.Pattern; public class ProjectTreeToPaths { // 读取Word文档中的文本内容 private static String readWordDocument(String filePath) throws IOException { StringBuilder content = new StringBuilder(); try (FileInputStream fis = new FileInputStream(filePath); XWPFDocument document = new XWPFDocument(fis)) { for (XWPFParagraph paragraph : document.getParagraphs()) { content.append(paragraph.getText()).append("\n"); } } return content.toString(); } // 解析树形结构文本,生成文件路径数组 private static String[] parseTreeToPaths(String treeText) { List<String> filePaths = new ArrayList<>(); List<String> pathStack = new ArrayList<>(); // 匹配树形节点行的正则,提取节点名称 Pattern pattern = Pattern.compile("^(\\│?\\s*├──\\s*)(.+)$"); String[] lines = treeText.split("\\n"); if (lines.length == 0) return new String[0]; // 处理根目录节点 String rootLine = lines[0].trim(); if (!rootLine.isEmpty()) { pathStack.add(rootLine.replace("/", "")); } for (int i = 1; i < lines.length; i++) { String line = lines[i].trim(); if (line.isEmpty()) continue; Matcher matcher = pattern.matcher(lines[i]); if (matcher.find()) { String nodeName = matcher.group(2).trim(); int level = countLevel(lines[i]); // 调整路径栈到当前节点层级 while (pathStack.size() > level) { pathStack.remove(pathStack.size() - 1); } boolean isDirectory = nodeName.endsWith("/"); String cleanName = isDirectory ? nodeName.replace("/", "") : nodeName; if (isDirectory) { pathStack.add(cleanName); } else { // 拼接完整文件路径 StringBuilder fullPath = new StringBuilder(); for (String pathPart : pathStack) { fullPath.append(pathPart).append("/"); } fullPath.append(cleanName); filePaths.add(fullPath.toString()); } } } return filePaths.toArray(new String[0]); } // 计算当前节点的层级 private static int countLevel(String line) { // 通过统计行中「│」的数量确定层级,根节点为0层 int pipeCount = (int) line.chars().filter(c -> c == '│').count(); return pipeCount + 1; } public static void main(String[] args) { try { // 替换为你的Word文档路径 String wordFilePath = "project-structure.docx"; String treeText = readWordDocument(wordFilePath); String[] filePaths = parseTreeToPaths(treeText); // 输出格式化为示例要求的数组形式 System.out.println("String[] output = ["); for (int i = 0; i < filePaths.length; i++) { System.out.print(" \"" + filePaths[i] + "\""); if (i != filePaths.length - 1) { System.out.println(","); } } System.out.println("\n];"); } catch (IOException e) { e.printStackTrace(); } } }
代码说明
readWordDocument:读取.docx文件的所有段落文本,拼接成树形结构的原始字符串。parseTreeToPaths:用正则匹配每行的节点,通过栈维护当前路径层级;遇到文件节点时,拼接完整路径并加入结果列表。countLevel:统计行中「│」的数量来确定节点层级,以此调整路径栈的长度,保证路径的正确性。- 目录与文件区分:通过节点名是否以
/结尾判断类型,目录加入路径栈,文件则生成完整路径并收集。
内容的提问来源于stack exchange,提问作者DilliBabu S
相关产品推荐
相关产品推荐

