You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup抓取表格未达预期输出的问题求助

解决Beautiful Soup抓取车辆表格数据缺失问题

问题分析

你的代码存在以下关键错误,导致无法获取车牌号、最后上线时间等预期字段:

  1. 错误复用class选择器:车牌号registration和车辆编号使用了相同的td.number元素,导致输出HTML对象而非车牌号文本。
  2. 未正确提取文本:last_seen变量直接赋值td对象,未调用.text获取内容。
  3. 车型选择错误:model直接取第一个td元素,重复输出车辆编号而非车型信息。
  4. 线路编号与时间的选择逻辑混淆:原代码误将last-seen的内容当作线路编号,导致字段对应错误。

修正后的代码

基于class选择的版本(需匹配实际HTML结构)

from bs4 import BeautifulSoup

# 解析网页HTML
soup = BeautifulSoup(response.text, 'html.parser')

# 定位所有车辆对应的表格
buses = soup.find_all('table', class_='fleet compact')

# 遍历每个车辆表格提取数据
for bus in buses:
    # 提取车辆编号,清理空白字符
    fleet_number = bus.find('td', class_='number').text.strip()
    # 提取车牌号(需根据实际HTML调整class名,示例用'reg')
    registration = bus.find('td', class_='reg').text.strip() if bus.find('td', class_='reg') else 'N/A'
    # 提取线路编号(需根据实际HTML调整class名,示例用'service')
    service_number = bus.find('td', class_='service').text.strip() if bus.find('td', class_='service') else 'N/A'
    # 提取最后上线时间,清理空白字符
    last_seen = bus.find('td', class_='last-seen').text.strip() if bus.find('td', class_='last-seen') else 'N/A'
    # 提取车型(需根据实际HTML调整class名,示例用'model')
    model = bus.find('td', class_='model').text.strip() if bus.find('td', class_='model') else 'N/A'

    # 输出抓取结果
    print(fleet_number)
    print(registration)
    print(service_number)
    print(last_seen)
    print(model)
    print('---')  # 分隔不同车辆信息

基于位置选择的版本(适合class不明确的场景)

如果目标HTML中td元素的顺序固定(车辆编号→车牌号→线路编号→最后上线时间→车型),可以按索引提取:

from bs4 import BeautifulSoup

soup = BeautifulSoup(response.text, 'html.parser')
buses = soup.find_all('table', class_='fleet compact')

for bus in buses:
    tds = bus.find_all('td')
    # 根据td索引提取对应字段,添加容错判断
    fleet_number = tds[0].text.strip() if len(tds) > 0 else 'N/A'
    registration = tds[1].text.strip() if len(tds) > 1 else 'N/A'
    service_number = tds[2].text.strip() if len(tds) > 2 else 'N/A'
    last_seen = tds[3].text.strip() if len(tds) > 3 else 'N/A'
    model = tds[4].text.strip() if len(tds) > 4 else 'N/A'

    print(fleet_number)
    print(registration)
    print(service_number)
    print(last_seen)
    print(model)
    print('---')

关键改动说明

  • 精准定位元素:根据实际HTML结构调整class选择器,或通过td索引匹配字段位置,确保每个变量对应正确的目标数据。
  • 文本提取与清理:所有字段使用.text.strip()获取纯文本并去除多余空白,避免输出HTML对象或无效换行。
  • 容错处理:添加if ... else判断,防止因元素缺失导致代码报错。

内容的提问来源于stack exchange,提问作者Callum Gourlay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 10:10:49