求HTML面包屑文件生成目录列表脚本:项目符号文件树转标准目录字符串
HTML项目符号文件树转全路径目录脚本实现
功能说明
可读取包含项目符号格式文件树的HTML文件,自动拼接生成每个文件的完整绝对路径字符串,输出格式符合需求示例。
前置依赖
- 需要安装Python第三方解析库
BeautifulSoup4,安装命令:pip install beautifulsoup4 lxml
实现脚本
from bs4 import BeautifulSoup def generate_full_path(html_path, output_path): # 读取HTML内容 with open(html_path, 'r', encoding='utf-8') as f: html_content = f.read() soup = BeautifulSoup(html_content, 'lxml') result = [] # 递归遍历嵌套ul/li结构拼接路径 def traverse(node, current_path): for li in node.find_all('li', recursive=False): # 提取当前li的直接文本(不含子节点内容) current_name = li.find(string=True, recursive=False).strip() new_path = f"{current_path}\\{current_name}" if current_path else current_name # 存在子ul说明是文件夹,继续遍历;否则为文件加入结果 child_ul = li.find('ul', recursive=False) if child_ul: traverse(child_ul, new_path) else: result.append(new_path) # 从最外层ul开始遍历 root_ul = soup.find('ul') traverse(root_ul, '') # 输出结果文件 with open(output_path, 'w', encoding='utf-8') as f: f.write('\n'.join(result)) # 调用示例 if __name__ == '__main__': generate_full_path('input.html', 'output.txt')
使用说明
- 将待处理的HTML文件命名为
input.html,和脚本放在同一文件夹下 - 运行脚本后,生成的
output.txt文件中每一行对应一条完整文件路径,输出示例如下:
C:\computer list.csv D:\DeptShare\Accounting.DS_Store D:\DeptShare\Accounting\New Folder\8-1-18.xls D:\DeptShare\Accounting\Physical Inventory\Copy of PhysInv Data Entry Dump 2016.xlsx D:\DeptShare\Accounting\Physical Inventory\PhysInv Data Entry Dump 2016.xlsx
内容的提问来源于stack exchange,提问作者mweis
相关产品推荐
相关产品推荐

