如何用BeautifulSoup提取HTML纯文本?解决爬虫结果含标签问题
解决方法
问题出在你直接将BeautifulSoup的标签对象作为字典的键和值,导致输出包含HTML标签。只需提取标签内的纯文本即可,修改代码如下:
from bs4 import BeautifulSoup import requests website = 'https://www.klavkarr.com/data-trouble-code-obd2.php?dtc=p0000-p0299#dtc' result = requests.get(website) content = result.text soup = BeautifulSoup(content, 'lxml') box = soup.find('div', class_='main_article-blog') table = box.find('table') # 提取表头文本,strip=True去除多余空格 headers = [header.get_text(strip=True) for header in table.find_all('th')] # 跳过表头行,提取数据行的文本内容 results = [ {headers[i]: cell.get_text(strip=True) for i, cell in enumerate(row.find_all('td'))} for row in table.find_all('tr')[1:] # [1:] 跳过第一行表头 ] # 打印结果示例 for item in results[:5]: print(item)
关键修改点:
- 对表头标签
header调用get_text(strip=True),提取纯文本并去除首尾空格 - 对单元格标签
cell同样调用get_text(strip=True)获取纯文本内容 - 使用
table.find_all('tr')[1:]跳过第一行表头,避免生成空字典
修改后输出的字典将只包含纯文本内容,不再有HTML标签。
内容的提问来源于stack exchange,提问作者Emiliano Cerasani
相关产品推荐
相关产品推荐

