如何精简修改Python脚本实现文件路径唯一输出?
优化HTML链接提取脚本的输出格式
核心思路很简单:用一个变量记录当前已经输出过的文件路径,处理新文件时才输出路径,同一文件下的所有链接直接跟在路径下方,不用重复打印路径。
修改后的脚本
import os from bs4 import BeautifulSoup def extract_links(html_path): with open(html_path, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f, 'html.parser') links = [] for a_tag in soup.find_all('a', href=True): href = a_tag['href'] if href.startswith(('http://', 'https://')): links.append(href) return links output_file = 'output.txt' with open(output_file, 'w', encoding='utf-8') as out_f: root_dir = './target_dir' # 替换成你的目标目录 current_file = None # 追踪已输出的文件路径 for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: if filename.endswith('.html'): file_path = os.path.join(dirpath, filename) links = extract_links(file_path) if links: # 跳过无有效链接的文件 if file_path != current_file: # 新文件,输出路径 out_f.write(f"{file_path}\n") current_file = file_path # 输出所有链接,缩进对齐 for link in links: out_f.write(f" {link}\n")
关键改动点
- 新增
current_file变量,避免重复输出同一文件路径 - 可选判断
if links:,跳过没有有效HTTP/HTTPS链接的文件,减少冗余内容 - 链接前加两个空格缩进,让层级更清晰(不需要缩进的话直接去掉空格即可)
修改后输出格式示例:
./target_dir/page1.html
https://example.com/link1
https://example.com/link2
./target_dir/sub/page2.html
https://example.com/link3
内容的提问来源于stack exchange,提问作者CluelessDumbo
相关产品推荐
相关产品推荐

