You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas网页爬取:HSBC产品表格合并及ValueError问题排查

问题解决与实现方案

一、先解决ValueError错误

你遇到的错误是因为pd.read_html()返回的是DataFrame对象的列表,直接把这个列表传给pd.DataFrame()会导致维度不匹配。正确做法是取列表中对应目标表格的DataFrame(页面中只有一个产品表格,所以取第一个元素data1[0]):

def create_dataframe(self):
    # 读取页面所有表格,取第一个目标表格
    data1 = pd.read_html(self.driver.page_source)[0]
    # 直接保存或处理这个DataFrame
    data1.to_csv('Data1.csv', index=False)

二、循环下拉框选项并合并所有DataFrame

以下是完整的实现代码,包含循环遍历客户类型、收集所有产品表格、合并输出的逻辑:

from selenium.webdriver.support.ui import Select
from selenium.common.exceptions import NoSuchElementException
import pandas as pd
import time
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

class HSBCProductScraper:
    def __init__(self, driver):
        self.driver = driver
        self.driver.get("https://intermediaries.hsbc.co.uk/products/product-finder/")
        # 等待页面加载完成
        WebDriverWait(self.driver, 10).until(
            EC.presence_of_element_located((By.ID, 'Availability'))
        )

    def scrape_all_products(self):
        all_product_dfs = []
        select = Select(self.driver.find_element_by_id('Availability'))
        # 获取下拉框所有选项(跳过第一个默认提示选项)
        options = select.options

        for idx in range(1, len(options)):
            customer_type = options[idx].text
            print(f"正在抓取【{customer_type}】的产品数据")
            try:
                # 通过索引选择选项,避免文本匹配的异常
                select.select_by_index(idx)
                # 等待选项加载完成
                time.sleep(2)
                # 点击"查找产品"按钮(用显式等待替代sleep更稳定)
                WebDriverWait(self.driver, 10).until(
                    EC.element_to_be_clickable((By.XPATH, '//button[@type="button" and contains(text(),"Find product")]'))
                ).click()
                # 等待表格加载完成
                WebDriverWait(self.driver, 10).until(
                    EC.presence_of_element_located((By.TAG_NAME, 'table'))
                )
                # 读取当前页面的产品表格
                current_df = pd.read_html(self.driver.page_source)[0]
                # 添加客户类型标记,方便后续区分数据来源
                current_df["客户类型"] = customer_type
                # 将当前表格加入列表
                all_product_dfs.append(current_df)
                # 重新定位下拉框(如果页面刷新,需重新获取元素)
                select = Select(self.driver.find_element_by_id('Availability'))
            except NoSuchElementException as e:
                print(f"处理【{customer_type}】时出错: {str(e)}")
                continue

        # 合并所有DataFrame
        final_df = pd.concat(all_product_dfs, ignore_index=True)
        # 保存到CSV文件(utf-8-sig解决中文乱码)
        final_df.to_csv('HSBC全客户类型产品数据.csv', index=False, encoding='utf-8-sig')
        print(f"数据抓取完成,共{len(final_df)}条记录")
        return final_df

关键说明

  • 循环逻辑:通过遍历下拉框的options索引,避免因文本变更导致的选择失败
  • 数据标记:给每个表格添加客户类型列,便于后续区分不同选项的数据
  • 稳定性优化:用WebDriverWait显式等待替代time.sleep,提升页面元素交互的可靠性
  • 异常处理:捕获单个选项的处理异常,避免流程中断

内容的提问来源于stack exchange,提问作者user19795989

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:15:34