You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup解析本地HTML文件?

如何用BeautifulSoup加载本地HTML文件

你直接传入文件名字符串给BeautifulSoup是错误的——它会把这个字符串当作HTML内容解析,而不是去读取文件里的实际内容。以下是两种正确的处理方式:

方法一:读取文件内容后传入

先通过open()读取本地HTML文件的内容,再将内容传给BeautifulSoup:

from bs4 import BeautifulSoup

# 以只读模式打开文件,指定编码避免乱码
with open('TestOC.html', 'r', encoding='utf-8') as file:
    html_content = file.read()

# 基于文件内容创建BeautifulSoup对象
soup = BeautifulSoup(html_content, 'html.parser')

方法二:直接传入文件对象

BeautifulSoup支持直接接收文件对象,无需提前读取内容:

from bs4 import BeautifulSoup

with open('TestOC.html', 'r', encoding='utf-8') as file:
    soup = BeautifulSoup(file, 'html.parser')

注意事项

  • 确保TestOC.html和你的Python脚本在同一目录下,否则需要传入完整的文件路径(比如Windows下的'C:/Users/xxx/TestOC.html',或Linux/macOS下的'/home/xxx/TestOC.html')。
  • 如果文件使用非UTF-8编码(如GBK),请修改encoding参数为对应编码值。

内容的提问来源于stack exchange,提问作者BatRuby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:25:14