You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML纯文本?解决爬虫结果含标签问题

解决方法

问题出在你直接将BeautifulSoup的标签对象作为字典的键和值,导致输出包含HTML标签。只需提取标签内的纯文本即可,修改代码如下:

from bs4 import BeautifulSoup
import requests

website = 'https://www.klavkarr.com/data-trouble-code-obd2.php?dtc=p0000-p0299#dtc'
result = requests.get(website)
content = result.text

soup = BeautifulSoup(content, 'lxml')

box = soup.find('div', class_='main_article-blog')

table = box.find('table')

# 提取表头文本,strip=True去除多余空格
headers = [header.get_text(strip=True) for header in table.find_all('th')]
# 跳过表头行,提取数据行的文本内容
results = [
    {headers[i]: cell.get_text(strip=True) for i, cell in enumerate(row.find_all('td'))}
    for row in table.find_all('tr')[1:]  # [1:] 跳过第一行表头
]

# 打印结果示例
for item in results[:5]:
    print(item)

关键修改点:

  • 对表头标签header调用get_text(strip=True),提取纯文本并去除首尾空格
  • 对单元格标签cell同样调用get_text(strip=True)获取纯文本内容
  • 使用table.find_all('tr')[1:]跳过第一行表头,避免生成空字典

修改后输出的字典将只包含纯文本内容,不再有HTML标签。

内容的提问来源于stack exchange,提问作者Emiliano Cerasani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 06:33:21