求助:如何使用BeautifulSoup提取同class名称表格中的文本
问题描述
尝试提取某网站中带有相同class名称的表格内容时,返回结果全为None,无法获取期望的旅行者类型数据(如Couple leisure、Family leisure等)。原代码如下:
import pandas as pd from bs4 import BeautifulSoup as extractor import requests from dateutil import parser import numpy as np for page in range(1, 354): ba="https://www.airlinequality.com/airline-reviews/british-airways/page/"+str(page) Headers={'User-agent': 'Mozilla/5.0'} response=requests.get(ba, headers=Headers) soup=extractor(response.content, "html.parser") for Reviewer in soup.findAll("article", itemprop="review"): type_of_traveller=Reviewer.find("tr", class_="review-rating-header type_of_traveller")
解决方案
1. 核心修正代码
import pandas as pd from bs4 import BeautifulSoup as extractor import requests import numpy as np # 存储结果的列表 traveller_types = [] for page in range(1, 354): # 使用f-string简化URL拼接 ba = f"https://www.airlinequality.com/airline-reviews/british-airways/page/{page}" # 完善请求头,提升请求成功率 headers = { 'User-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } response = requests.get(ba, headers=headers) # 跳过请求失败的页面 if response.status_code != 200: print(f"页面{page}请求失败,状态码:{response.status_code}") continue soup = extractor(response.content, "html.parser") # 遍历每一条评论 for review in soup.find_all("article", itemprop="review"): # 用CSS选择器精准定位目标行,兼容多类名顺序变化 traveller_row = review.select_one("tr.review-rating-header.type_of_traveller") if traveller_row: # 提取td中的实际文本内容 traveller_type = traveller_row.find("td", class_="review-value").get_text(strip=True) traveller_types.append(traveller_type) else: # 对缺失数据做标记 traveller_types.append("N/A") # 输出结果 print("Type of Traveller") for t_type in traveller_types: print(t_type)
2. 关键修正点
- 选择器优化:改用
select_one("tr.review-rating-header.type_of_traveller"),避免原find方法对多类名顺序的严格依赖,匹配更稳定 - 内容提取:找到目标
tr后,进一步定位td.review-value标签获取实际文本,而非仅获取tr标签本身 - 请求健壮性:增加状态码检查,跳过请求失败的页面,避免解析无效内容
- 结果存储:用列表统一收集数据,最后按格式输出
3. 原代码失效原因
- 原
class_参数的字符串写法要求完全匹配class属性的顺序和内容,页面渲染时class顺序可能发生变化,导致匹配失败 - 未提取
td标签中的文本,即使找到tr也无法得到目标内容 - 缺少请求失败处理,若某页请求失败,解析空内容会生成大量
None
内容的提问来源于stack exchange,提问作者Temp4 Watch
相关产品推荐
相关产品推荐

