Python2.7 BeautifulSoup处理HTML报NameError修复及美化实现问题
问题修复与功能实现方案
1. NameError报错修复
报错原因
- 第一处错误是变量未定义:你在打开文件时传入的
fi没有被赋值,实际需要拼接当前遍历的目录路径和文件名得到完整的文件路径 - 第二处错误是缩进逻辑错误:
content变量仅在匹配到.html文件的分支内定义,而你后续解析HTML的代码缩进层级不对,脱离了文件遍历的内部分支,当目录下没有匹配的.html文件时,content未被定义就会触发报错;即使存在匹配文件,也只会处理遍历到的最后一个文件。
基础修复代码
#!/usr/bin/python from bs4 import BeautifulSoup import os root = '/path/to/directory/test' for dirpath, dirs, files in os.walk(root): for name in files: if name.endswith('.html'): # 拼接完整文件路径 file_path = os.path.join(dirpath, name) # 读取文件内容 with open(file_path, 'r') as f: content = f.read() # 解析HTML放在文件遍历分支内部,每匹配一个文件处理一个 soup = BeautifulSoup(content, 'html.parser') # 删除指定div for div in soup.find_all('div', {'class': 'sidebar'}): div.decompose()
2. HTML原地美化格式化功能实现
直接调用BeautifulSoup内置的prettify()方法即可生成带规范缩进的格式化HTML代码,处理完成后将内容写回原文件即可实现原地修改,Python2.7环境下建议指定编码避免中文乱码,完整代码如下:
#!/usr/bin/python # -*- coding: utf-8 -*- from bs4 import BeautifulSoup import os root = '/path/to/directory/test' for dirpath, dirs, files in os.walk(root): for name in files: if name.endswith('.html'): file_path = os.path.join(dirpath, name) with open(file_path, 'r') as f: content = f.read() soup = BeautifulSoup(content, 'html.parser') # 删除sidebar类的div for div in soup.find_all('div', {'class': 'sidebar'}): div.decompose() # 生成带规范缩进的美化HTML pretty_html = soup.prettify(encoding='utf-8') # 原地写回原文件 with open(file_path, 'w') as f: f.write(pretty_html)
注意事项
操作前建议备份原目录文件,避免误操作导致内容丢失。
内容的提问来源于stack exchange,提问作者Lexx Luxx
相关产品推荐
相关产品推荐

