You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求HTML面包屑文件生成目录列表脚本:项目符号文件树转标准目录字符串

HTML项目符号文件树转全路径目录脚本实现

功能说明

可读取包含项目符号格式文件树的HTML文件,自动拼接生成每个文件的完整绝对路径字符串,输出格式符合需求示例。

前置依赖

  • 需要安装Python第三方解析库BeautifulSoup4,安装命令:
    pip install beautifulsoup4 lxml

实现脚本

from bs4 import BeautifulSoup

def generate_full_path(html_path, output_path):
    # 读取HTML内容
    with open(html_path, 'r', encoding='utf-8') as f:
        html_content = f.read()
    soup = BeautifulSoup(html_content, 'lxml')
    result = []
    # 递归遍历嵌套ul/li结构拼接路径
    def traverse(node, current_path):
        for li in node.find_all('li', recursive=False):
            # 提取当前li的直接文本(不含子节点内容)
            current_name = li.find(string=True, recursive=False).strip()
            new_path = f"{current_path}\\{current_name}" if current_path else current_name
            # 存在子ul说明是文件夹,继续遍历;否则为文件加入结果
            child_ul = li.find('ul', recursive=False)
            if child_ul:
                traverse(child_ul, new_path)
            else:
                result.append(new_path)
    # 从最外层ul开始遍历
    root_ul = soup.find('ul')
    traverse(root_ul, '')
    # 输出结果文件
    with open(output_path, 'w', encoding='utf-8') as f:
        f.write('\n'.join(result))

# 调用示例
if __name__ == '__main__':
    generate_full_path('input.html', 'output.txt')

使用说明

  • 将待处理的HTML文件命名为input.html,和脚本放在同一文件夹下
  • 运行脚本后,生成的output.txt文件中每一行对应一条完整文件路径,输出示例如下:
C:\computer list.csv
D:\DeptShare\Accounting.DS_Store
D:\DeptShare\Accounting\New Folder\8-1-18.xls
D:\DeptShare\Accounting\Physical Inventory\Copy of PhysInv Data Entry Dump 2016.xlsx
D:\DeptShare\Accounting\Physical Inventory\PhysInv Data Entry Dump 2016.xlsx

内容的提问来源于stack exchange,提问作者mweis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:42:01