You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何使用BeautifulSoup提取同class名称表格中的文本

问题描述

尝试提取某网站中带有相同class名称的表格内容时,返回结果全为None,无法获取期望的旅行者类型数据(如Couple leisure、Family leisure等)。原代码如下:

import pandas as pd
from bs4 import BeautifulSoup as extractor
import requests
from dateutil import parser
import numpy as np

for page in range(1, 354):
    ba="https://www.airlinequality.com/airline-reviews/british-airways/page/"+str(page)
    Headers={'User-agent': 'Mozilla/5.0'}
    response=requests.get(ba, headers=Headers)
    soup=extractor(response.content, "html.parser")
    for Reviewer in soup.findAll("article", itemprop="review"):
        type_of_traveller=Reviewer.find("tr", class_="review-rating-header type_of_traveller")
解决方案

1. 核心修正代码

import pandas as pd
from bs4 import BeautifulSoup as extractor
import requests
import numpy as np

# 存储结果的列表
traveller_types = []

for page in range(1, 354):
    # 使用f-string简化URL拼接
    ba = f"https://www.airlinequality.com/airline-reviews/british-airways/page/{page}"
    # 完善请求头,提升请求成功率
    headers = {
        'User-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
    }
    response = requests.get(ba, headers=headers)
    
    # 跳过请求失败的页面
    if response.status_code != 200:
        print(f"页面{page}请求失败,状态码:{response.status_code}")
        continue
        
    soup = extractor(response.content, "html.parser")
    # 遍历每一条评论
    for review in soup.find_all("article", itemprop="review"):
        # 用CSS选择器精准定位目标行,兼容多类名顺序变化
        traveller_row = review.select_one("tr.review-rating-header.type_of_traveller")
        if traveller_row:
            # 提取td中的实际文本内容
            traveller_type = traveller_row.find("td", class_="review-value").get_text(strip=True)
            traveller_types.append(traveller_type)
        else:
            # 对缺失数据做标记
            traveller_types.append("N/A")

# 输出结果
print("Type of Traveller")
for t_type in traveller_types:
    print(t_type)

2. 关键修正点

  • 选择器优化:改用select_one("tr.review-rating-header.type_of_traveller"),避免原find方法对多类名顺序的严格依赖,匹配更稳定
  • 内容提取:找到目标tr后,进一步定位td.review-value标签获取实际文本,而非仅获取tr标签本身
  • 请求健壮性:增加状态码检查,跳过请求失败的页面,避免解析无效内容
  • 结果存储:用列表统一收集数据,最后按格式输出

3. 原代码失效原因

  • 原class_参数的字符串写法要求完全匹配class属性的顺序和内容,页面渲染时class顺序可能发生变化,导致匹配失败
  • 未提取td标签中的文本,即使找到tr也无法得到目标内容
  • 缺少请求失败处理,若某页请求失败,解析空内容会生成大量None

内容的提问来源于stack exchange,提问作者Temp4 Watch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 07:52:42