You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何精简修改Python脚本实现文件路径唯一输出?

优化HTML链接提取脚本的输出格式

核心思路很简单:用一个变量记录当前已经输出过的文件路径,处理新文件时才输出路径,同一文件下的所有链接直接跟在路径下方,不用重复打印路径。

修改后的脚本

import os
from bs4 import BeautifulSoup

def extract_links(html_path):
    with open(html_path, 'r', encoding='utf-8') as f:
        soup = BeautifulSoup(f, 'html.parser')
    links = []
    for a_tag in soup.find_all('a', href=True):
        href = a_tag['href']
        if href.startswith(('http://', 'https://')):
            links.append(href)
    return links

output_file = 'output.txt'
with open(output_file, 'w', encoding='utf-8') as out_f:
    root_dir = './target_dir'  # 替换成你的目标目录
    current_file = None  # 追踪已输出的文件路径
    for dirpath, _, filenames in os.walk(root_dir):
        for filename in filenames:
            if filename.endswith('.html'):
                file_path = os.path.join(dirpath, filename)
                links = extract_links(file_path)
                if links:  # 跳过无有效链接的文件
                    if file_path != current_file:
                        # 新文件,输出路径
                        out_f.write(f"{file_path}\n")
                        current_file = file_path
                    # 输出所有链接,缩进对齐
                    for link in links:
                        out_f.write(f"  {link}\n")

关键改动点

  • 新增current_file变量,避免重复输出同一文件路径
  • 可选判断if links:,跳过没有有效HTTP/HTTPS链接的文件,减少冗余内容
  • 链接前加两个空格缩进,让层级更清晰(不需要缩进的话直接去掉空格即可)

修改后输出格式示例:

./target_dir/page1.html
https://example.com/link1
https://example.com/link2
./target_dir/sub/page2.html
https://example.com/link3

内容的提问来源于stack exchange,提问作者CluelessDumbo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:42:47