如何使用BeautifulSoup解析本地HTML文件?
如何用BeautifulSoup加载本地HTML文件
你直接传入文件名字符串给BeautifulSoup是错误的——它会把这个字符串当作HTML内容解析,而不是去读取文件里的实际内容。以下是两种正确的处理方式:
方法一:读取文件内容后传入
先通过open()读取本地HTML文件的内容,再将内容传给BeautifulSoup:
from bs4 import BeautifulSoup # 以只读模式打开文件,指定编码避免乱码 with open('TestOC.html', 'r', encoding='utf-8') as file: html_content = file.read() # 基于文件内容创建BeautifulSoup对象 soup = BeautifulSoup(html_content, 'html.parser')
方法二:直接传入文件对象
BeautifulSoup支持直接接收文件对象,无需提前读取内容:
from bs4 import BeautifulSoup with open('TestOC.html', 'r', encoding='utf-8') as file: soup = BeautifulSoup(file, 'html.parser')
注意事项
- 确保
TestOC.html和你的Python脚本在同一目录下,否则需要传入完整的文件路径(比如Windows下的'C:/Users/xxx/TestOC.html',或Linux/macOS下的'/home/xxx/TestOC.html')。 - 如果文件使用非UTF-8编码(如GBK),请修改
encoding参数为对应编码值。
内容的提问来源于stack exchange,提问作者BatRuby
相关产品推荐
相关产品推荐

