You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取span中的Best Sellers Rank文本?

提取亚马逊页面“Best Sellers Rank”内容的解决方案

问题背景

使用BeautifulSoup爬取亚马逊商品页面时,无法精准提取“Best Sellers Rank”相关内容:

  • 直接获取父div文本会得到包含ASIN、评分、上架日期等的混合内容
  • 使用find_next('span')仅能获取到评分文本,无法定位到目标内容

原爬取代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd

headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:66.0) Gecko/20100101 Firefox/66.0", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}
url = 'https://www.amazon.com/GymCope-Anti-Tear-Cushioning-Non-Slip-Exercise/dp/B0921F1T2P/ref=sr_1_3_sspa?brr=1&pd_rd_r=4b40f0a8-f2d8-44dc-9a98-413c64d3fa34&pd_rd_w=P9ZJI&pd_rd_wg=RS7zW&pf_rd_p=9875e817-188b-48a2-986d-8146749644ac&pf_rd_r=AGWBT5KT04TYKGPZASKA&qid=1642452438&rd=1&rnid=3407731&s=sporting-goods&sr=1-3-spons&psc=1&spLa=ZW5jcnlwdGVkUXVhbGlmaWVyPUExVjhWTk0xQU5WWldPJmVuY3J5cHRlZElkPUEwODE0MzYwMTdMTDZSNDVST08yMiZlbmNyeXB0ZWRBZElkPUEwODQ4MDM0MlE4WEtVUjFKMUdLMiZ3aWRnZXROYW1lPXNwX2F0Zl9icm93c2UmYWN0aW9uPWNsaWNrUmVkaXJlY3QmZG9Ob3RMb2dDbGljaz10cnVl'
response = requests.get(url, headers=headers)
html = response.text
soup = BeautifulSoup(html)
bsr = soup.find("div", class_="a-section table-padding").text

解决方案

以下是三种可行的提取方式,可根据页面结构稳定性选择:

方法1:通过文本匹配定位目标元素

利用“Best Sellers Rank”文本直接定位对应的元素,再提取后续内容:

# 定位父容器
rank_container = soup.find("div", class_="a-section table-padding")
# 找到包含"Best Sellers Rank"的span元素
bsr_label = rank_container.find("span", string=lambda text: text and "Best Sellers Rank" in text)

if bsr_label:
    # 获取完整的排名内容(包含所有排名详情)
    full_rank_content = bsr_label.parent.get_text(strip=True)
    # 拆分出标题和具体排名
    rank_detail = full_rank_content.split("Best Sellers Rank")[-1].split("Date First Available")[0].strip()
    print(f"Best Sellers Rank: {rank_detail}")

方法2:正则表达式提取(适用于结构不稳定场景)

从父容器的混合文本中,用正则匹配目标内容:

import re

# 获取父容器文本
raw_text = soup.find("div", class_="a-section table-padding").text
# 匹配"Best Sellers Rank"到下一个大写开头的字段之间的内容
pattern = re.compile(r'Best Sellers Rank\s*(.*?)\s*(?=[A-Z][a-z]+ First Available|$)')
match_result = pattern.search(raw_text)

if match_result:
    print(f"Best Sellers Rank: {match_result.group(1).strip()}")

方法3:精准CSS选择器定位

利用亚马逊页面的结构特征,直接定位排名所在的行:

# 直接选择包含"Best Sellers Rank"的行元素
bsr_row = soup.select_one('div.a-section.table-padding span:contains("Best Sellers Rank")').parent
# 提取并清理文本内容
cleaned_bsr = bsr_row.get_text(strip=True).replace("Best Sellers Rank", "").split("Date First Available")[0].strip()
print(f"Best Sellers Rank: {cleaned_bsr}")

内容的提问来源于stack exchange,提问作者Bünyamin Erkaya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 11:37:01