You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup提取网页中的车辆信息?

问题

我用Python做网页爬取,要根据车辆注册号获取车辆信息。目前已成功拿到目标网站的HTML内容,但无法提取品牌、型号、颜色、制造年份、最高时速、变速箱这些信息。当前代码如下:

from bs4 import BeautifulSoup
import requests

url = "https://www.carcheck.co.uk/audi/N18CTN"

r= requests.get(url)

soup = BeautifulSoup(r.text)

print(soup)

目标信息对应的HTML片段:

<td>AUDI</td>
</tr>
<tr>
<th>Model</th>
<td>A3</td>
</tr>
<tr>
<th>Colour</th>
<td>Red</td>
</tr>
<tr>
<th>Year of manufacture</th>
<td>2017</td>
</tr>
<tr>
<th>Top speed</th>
<td>147 mph</td>
</tr>
<tr>
<th>Gearbox</th>
<td>6 speed automatic</td>

解决方案

你可以借助BeautifulSoup的节点查找能力,通过<th>标签的文本定位对应行,再提取<td>中的信息。修改后的代码如下:

from bs4 import BeautifulSoup
import requests

url = "https://www.carcheck.co.uk/audi/N18CTN"

r = requests.get(url)
# 指定解析器,避免默认解析器的警告
soup = BeautifulSoup(r.text, 'html.parser')

# 提取品牌(对应行无<th>标签,直接匹配<td>文本)
brand = soup.find('td', string='AUDI').text.strip()

# 定义要提取的字段与键名的映射
target_fields = {
    'Model': 'model',
    'Colour': 'colour',
    'Year of manufacture': 'manufacture_year',
    'Top speed': 'top_speed',
    'Gearbox': 'gearbox'
}

car_info = {'brand': brand}

# 遍历字段,逐个提取信息
for label, key in target_fields.items():
    th_node = soup.find('th', string=label)
    if th_node:
        # 找到<th>所在的<tr>行,再提取<td>内容
        td_node = th_node.find_parent('tr').find('td')
        car_info[key] = td_node.text.strip()

# 输出结果
for item, value in car_info.items():
    print(f"{item}: {value}")

关键点说明

  • 品牌单独处理:因为它对应的行没有<th>标签,直接通过<td>的文本匹配获取。
  • 其他字段通过<th>的文本精准定位到对应行,再提取该行内的<td>内容,避免无关信息干扰。
  • 指定html.parser解析器,既可以消除BeautifulSoup的警告,也能提升解析的稳定性。

内容的提问来源于stack exchange,提问作者charlie s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 23:05:17