如何用Python+BeautifulSoup提取网页中的车辆信息?
问题
我用Python做网页爬取,要根据车辆注册号获取车辆信息。目前已成功拿到目标网站的HTML内容,但无法提取品牌、型号、颜色、制造年份、最高时速、变速箱这些信息。当前代码如下:
from bs4 import BeautifulSoup import requests url = "https://www.carcheck.co.uk/audi/N18CTN" r= requests.get(url) soup = BeautifulSoup(r.text) print(soup)
目标信息对应的HTML片段:
<td>AUDI</td> </tr> <tr> <th>Model</th> <td>A3</td> </tr> <tr> <th>Colour</th> <td>Red</td> </tr> <tr> <th>Year of manufacture</th> <td>2017</td> </tr> <tr> <th>Top speed</th> <td>147 mph</td> </tr> <tr> <th>Gearbox</th> <td>6 speed automatic</td>
解决方案
你可以借助BeautifulSoup的节点查找能力,通过<th>标签的文本定位对应行,再提取<td>中的信息。修改后的代码如下:
from bs4 import BeautifulSoup import requests url = "https://www.carcheck.co.uk/audi/N18CTN" r = requests.get(url) # 指定解析器,避免默认解析器的警告 soup = BeautifulSoup(r.text, 'html.parser') # 提取品牌(对应行无<th>标签,直接匹配<td>文本) brand = soup.find('td', string='AUDI').text.strip() # 定义要提取的字段与键名的映射 target_fields = { 'Model': 'model', 'Colour': 'colour', 'Year of manufacture': 'manufacture_year', 'Top speed': 'top_speed', 'Gearbox': 'gearbox' } car_info = {'brand': brand} # 遍历字段,逐个提取信息 for label, key in target_fields.items(): th_node = soup.find('th', string=label) if th_node: # 找到<th>所在的<tr>行,再提取<td>内容 td_node = th_node.find_parent('tr').find('td') car_info[key] = td_node.text.strip() # 输出结果 for item, value in car_info.items(): print(f"{item}: {value}")
关键点说明
- 品牌单独处理:因为它对应的行没有
<th>标签,直接通过<td>的文本匹配获取。 - 其他字段通过
<th>的文本精准定位到对应行,再提取该行内的<td>内容,避免无关信息干扰。 - 指定
html.parser解析器,既可以消除BeautifulSoup的警告,也能提升解析的稳定性。
内容的提问来源于stack exchange,提问作者charlie s
相关产品推荐
相关产品推荐

