Hepsiburada商品价格自动排序及信息获取Selenium技术问题
Hepsiburada价格追踪器问题修复方案
以下是你提供的价格追踪器代码:
# libraries import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import matplotlib.pyplot as plt from time import sleep import inspect import os from bs4 import BeautifulSoup import requests # Get the search term and tracking period from the user search_term = input("Please enter the name of the product you want to search: ") months =input("Please enter the number of months you want to track the product: ") # To ensure that the user enters a non-string value while not months.isdigit(): print("Warning: Please enter a valid integer value for the number of months.") months = input("Please enter the number of months you want to track the product: ") months = int(months) # Start the web driver and go to the Hepsiburada homepage chrome_options = webdriver.ChromeOptions() prefs = {"profile.default_content_setting_values.notifications" : 2} chrome_options.add_experimental_option("prefs",prefs) module_path="C:/Users/Desktop/hepsiburada_price_tracker/chromedriver.exe" driver = webdriver.Chrome(executable_path=module_path,options=chrome_options) driver.maximize_window() driver.get("https://www.hepsiburada.com/") # Accept cookies driver.find_element_by_id('onetrust-accept-btn-handler').click() sleep(3) # Enter the search term in the search box and press Enter wait = WebDriverWait(driver, 15) search_box = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, 'theme-IYtZzqYPto8PhOx3ku3c'))) search_box.send_keys(search_term) search_box.send_keys(Keys.RETURN) # Wait for search results and select the first product sleep(3) # Sayfanın yüklenmesi için birkaç saniye bekleyin # Click on the order button driver.find_element_by_class_name('horizontalSortingBar-Ce404X9mUYVCRa5bjV4D').click() sleep(3) # Sort by increasing price driver.find_element_by_class_name('horizontalSortingBar-PkoDOH7UsCwBrQaQx9bn').click() sleep(3) # Get the link, name, and price of the first product in the search results results = driver.find_elements_by_xpath("//h3[@data-test-id='product-card-name']") if not results: print("Sorry, we could not find the product you were looking for.") #driver.quit() else: first_result0 = results[0] first_result= first_result0.text print(first_result) product_link = first_result.find_element(By.XPATH, ".//a[@data-productid]") product_url = product_link.get_attribute("href") product_name = first_result.find_element(By.XPATH, ".//h3").text print("The product selected from the search results is {}: {}".format(product_name, product_url))
问题描述
- 手动操作时按价格升序排序正常,但自动化代码点击排序按钮(排序>升序)后,排序结果错误,是否可修复?
- 无法获取页面商品总数、卖家名称及商品链接,希望在按价格升序排序后批量获取这些信息,并单独提取列表首个商品的信息,如何通过Selenium实现?
解决方案
问题1:排序功能修复
你代码中使用的horizontalSortingBar-Ce404X9mUYVCRa5bjV4D这类带随机字符串的Class Name是动态生成的,页面更新时会变化,导致点击的元素并非预期的排序选项。改用基于文本或data-test-id的稳定定位即可解决:
# 替换原排序相关代码 wait = WebDriverWait(driver, 15) # 点击排序按钮(土耳其语“sırala”对应“排序”) sort_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'sırala')]"))) sort_button.click() # 选择价格升序(土耳其语“En Düşük Fiyat”对应“最低价格”) price_asc_option = wait.until(EC.element_to_be_clickable((By.XPATH, "//div[contains(text(), 'En Düşük Fiyat')]"))) price_asc_option.click() # 等待排序完成(替换固定sleep,用元素状态判断更可靠) wait.until(EC.staleness_of(driver.find_element(By.XPATH, "//div[@data-test-id='product-card']")))
问题2:批量获取商品信息
通过data-test-id定位页面元素,批量提取所需信息:
# 等待商品列表加载完成 wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[@data-test-id='product-card']"))) # 获取商品总数 total_products = wait.until(EC.visibility_of_element_located((By.XPATH, "//span[@data-test-id='total-result']"))).text print(f"商品总数:{total_products}") # 批量提取所有商品信息 product_cards = driver.find_elements(By.XPATH, "//div[@data-test-id='product-card']") product_list = [] for card in product_cards: # 商品名称 name = card.find_element(By.XPATH, ".//h3[@data-test-id='product-card-name']").text # 商品链接 link = card.find_element(By.XPATH, ".//a[@data-test-id='product-card-link']").get_attribute("href") # 卖家名称(部分商品可能无卖家,加异常处理) try: seller = card.find_element(By.XPATH, ".//span[@data-test-id='product-card-seller-name']").text except: seller = "无卖家信息" # 商品价格 price = card.find_element(By.XPATH, ".//div[@data-test-id='current-price']").text product_list.append({ "名称": name, "链接": link, "卖家": seller, "价格": price }) # 打印首个商品信息 if product_list: print("\n首个商品信息:") first_product = product_list[0] for key, value in first_product.items(): print(f"{key}: {value}") # 可选:将数据保存为CSV文件 df = pd.DataFrame(product_list) df.to_csv("hepsiburada_products.csv", index=False, encoding="utf-8-sig")
额外优化建议
- 替换所有固定
sleep()为WebDriverWait显式等待,提升代码稳定性和执行效率。 - 避免使用动态Class Name,优先选择
data-test-id、元素文本或稳定的层级XPath定位。 - 增加异常处理逻辑,应对页面加载超时、元素不存在等场景。
内容的提问来源于stack exchange,提问作者user14178341
相关产品推荐
相关产品推荐

