You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup网页爬取时如何去除div标签仅保留车型名称?

问题:爬取Jaguar车型页面时输出包含HTML标签,如何提取纯净的车型名称?

你当前代码中直接打印car_model(BeautifulSoup的Tag对象)会输出完整的HTML标签内容,要获取纯净的车型名称,只需提取标签内的文本并清理多余空白字符即可,以下是两种常用解决方案:

方法1:使用.get_text() + .strip()

.get_text()用于提取标签内的所有文本内容,.strip()会移除文本首尾的空格、换行符等空白字符:

修改循环部分代码:

for car_model in soup.find_all(class_='sub-model-with-score-preview__name'):
    # 提取纯文本并清理空白
    clean_model_name = car_model.get_text().strip()
    print(clean_model_name)

方法2:使用.text属性 + .strip()

.text是.get_text()的简写,功能完全一致:

for car_model in soup.find_all(class_='sub-model-with-score-preview__name'):
    clean_model_name = car_model.text.strip()
    print(clean_model_name)

完整可运行代码

注意:你原代码遗漏了BeautifulSoup的导入语句,需补充才能正常运行:

import requests
from bs4 import BeautifulSoup

#Inputs/URLs to scrape:
URL_model = ('https://carbuzz.com/cars/jaguar')
(response := requests.get(URL_model)).raise_for_status()
soup = BeautifulSoup(response.text, 'lxml')

for car_model in soup.find_all(class_='sub-model-with-score-preview__name'):
    clean_model_name = car_model.get_text().strip()
    print(clean_model_name)

执行后会输出纯净的车型名称:

E-Pace
F-Pace SVR
F-Pace
I-Pace
F-Type Coupe

内容的提问来源于stack exchange,提问作者webscrapeartist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 07:09:19