You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python eBay乐高已售商品爬虫优化:排除“匹配较少关键词的结果”区域的无关商品列表

Python eBay乐高已售商品爬虫优化:排除“匹配较少关键词的结果”区域的无关商品列表

我完全懂你遇到的麻烦——eBay搜索结果里总会冒出那个“匹配较少关键词的结果”区块,把不相关的商品混进来,直接导致你的平均价格计算完全不准。之前你尝试停止处理这个区域时搞出了#N/A,应该是没找对正确的截断逻辑,咱们来把代码改对。

问题根源

你原来的代码是直接抓取页面上所有带s-item__price的元素,根本没区分“精准匹配的已售商品”和“凑数的低匹配商品”。咱们的核心思路是:先定位到那个无关区块的边界,只处理它之前的有效商品;或者直接锁定精准匹配的商品容器,彻底避开无关内容。

修改后的完整代码

我基于你的原代码做了针对性优化,关键改动都加了注释:

import re
import time
import statistics
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options

# 假设你的CHROME_DRIVER_PATH和chrome_options已经在外部定义好了
def fetch_ebay_average_price(ebay_search_url):
    """ Fetch the average sold price and count of sold listings from eBay sold listings page, excluding low-match results. """
    try:
        service = Service(CHROME_DRIVER_PATH)
        driver = webdriver.Chrome(service=service, options=chrome_options)
        driver.get(ebay_search_url)
        time.sleep(3)  # 可以考虑换成显式等待,比固定sleep更可靠

        soup = BeautifulSoup(driver.page_source, 'html.parser')
        prices = []

        # 第一步:定位到"匹配较少关键词"的标题元素(兼容英文显示)
        low_match_header = soup.find(
            'h3', 
            string=lambda text: text and 'Results Matching Fewer Words' in text
        )

        # 第二步:获取所有有效商品项
        if low_match_header:
            # 如果存在低匹配区域,只处理它之前的所有商品项
            valid_items = low_match_header.find_previous_siblings(class_='s-item')
            # find_previous_siblings是倒序返回,需要反转回正常顺序
            valid_items.reverse()
        else:
            # 没有低匹配区域,取所有商品项
            valid_items = soup.select('div.s-item')

        # 第三步:遍历有效商品,提取价格
        for item in valid_items:
            try:
                # 从当前商品项内找价格元素,避免跨区域误抓
                price_element = item.select_one('span.s-item__price')
                if not price_element:
                    continue
                
                price_text = price_element.text.strip()
                # 处理带"to"的价格区间,只取第一个价格(和原逻辑保持一致)
                price_str = price_text.split()[0].replace(',', '')
                # 移除货币符号,转换为数值
                price_value = float(re.sub(r'[£,$,€]', '', price_str))
                prices.append(price_value)
            except Exception as e:
                print(f"处理单个商品价格出错: {e}")
                continue

        driver.quit()
        if prices:
            average_price = round(statistics.mean(prices), 2)
            sold_count = len(prices)
            return f"£{average_price}", sold_count
        else:
            return "N/A", 0
    except Exception as e:
        print(f"抓取eBay平均价格时出错: {e}")
        # 确保异常场景下浏览器也能正常关闭,避免资源泄漏
        if 'driver' in locals():
            driver.quit()
        return "N/A", 0

关键优化点说明

  • 精准锁定有效商品范围:先检测低匹配区域的标题,存在的话只处理它之前的商品,彻底排除无关内容;同时从单个商品项内提取价格,避免全局抓取的误判。
  • 可靠性提升:新增异常场景下的浏览器关闭逻辑,防止资源泄漏;增加无价格元素的判断,跳过无效商品项。
  • 逻辑兼容:保留了你原有的价格处理逻辑(比如只取价格区间的第一个值),确保计算规则和之前一致。

额外效率建议

把固定的time.sleep(3)换成Selenium的显式等待,能根据页面实际加载情况调整等待时间,更高效也更稳定:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 替换原代码中的time.sleep(3)
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, 's-item'))
)

你可以用目标乐高商品的eBay链接测试,现在应该只会抓取精准匹配的已售商品价格,不会再混入低匹配区域的无关内容了。

备注:内容来源于stack exchange,提问作者click90

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 18:24:36