You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium抓取多标签页网站数据并导出CSV的技术求助

问题需求

需要用Python+Selenium抓取AMFI印度官网的基金业绩详情页数据并保存为CSV,要求:

  • 仅处理页面前3个标签页内容
  • 始终选中“ALL”标签
  • 日期设置为当前日期
  • 遍历前3个标签的所有选项组合获取数据

附上的尝试代码未成功,需技术修正。

修正后的代码
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time
import random
from datetime import datetime

def wait_for_element(driver, by, value, timeout=15):
    return WebDriverWait(driver, timeout).until(EC.element_to_be_clickable((by, value)))

def set_current_date(driver):
    # 生成当前日期,匹配页面要求的DD/MM/YYYY格式
    current_date = datetime.now().strftime("%d/%m/%Y")
    date_input = wait_for_element(driver, By.ID, "nav-date")
    date_input.clear()
    date_input.send_keys(current_date)
    time.sleep(random.uniform(0.5, 1))

def scrape_tab_combination(driver, tab_index, end_val, equity_val, cap_val, filename):
    # 切换到目标标签页(仅处理前3个,索引从0开始)
    tabs = wait_for_element(driver, By.CLASS_NAME, "nav-tabs").find_elements(By.TAG_NAME, "li")
    if tab_index >= 3:
        raise ValueError("仅支持前3个标签页")
    wait_for_element(driver, By.LINK_TEXT, tabs[tab_index].text).click()
    time.sleep(random.uniform(1, 2))

    # 强制选中ALL标签
    wait_for_element(driver, By.ID, "all-type").find_element(By.XPATH, ".//option[@value='1']").click()
    time.sleep(random.uniform(0.5, 1))

    # 设置当前日期
    set_current_date(driver)

    # 选择下拉筛选条件
    Select(wait_for_element(driver, By.ID, "end-type")).select_by_value(end_val)
    time.sleep(random.uniform(0.5, 1))
    Select(wait_for_element(driver, By.ID, "equity-type")).select_by_value(equity_val)
    time.sleep(random.uniform(0.5, 1))
    Select(wait_for_element(driver, By.ID, "cap-type")).select_by_value(cap_val)
    time.sleep(random.uniform(0.5, 1))

    # 触发数据加载
    wait_for_element(driver, By.ID, "go-button").click()
    time.sleep(random.uniform(2, 3))

    # 等待表格加载完成并提取数据
    table = wait_for_element(driver, By.ID, "fund-table")
    df = pd.read_html(table.get_attribute('outerHTML'))[0]
    # 新增标识字段,方便后续数据区分
    df["标签页"] = tabs[tab_index].text
    df["产品类型"] = "开放式" if end_val == "1" else "封闭式"
    # 保存CSV(utf-8-sig兼容中文Excel打开)
    df.to_csv(filename, index=False, encoding="utf-8-sig")
    print(f"已保存: {filename}")

# 初始化浏览器
driver = webdriver.Chrome()
driver.maximize_window()
driver.get("https://www.amfiindia.com/research-information/other-data/mf-scheme-performance-details")

# 等待页面核心元素加载
wait_for_element(driver, By.ID, "end-type", timeout=30)
print("页面加载完成")

# 定义需要遍历的维度
target_tabs = [0, 1, 2]  # 前3个标签页索引
end_types = ["1", "2"]  # 开放式/封闭式选项值
equity_types = ["1", "2", "3", "4", "5", "6"]  # 股票类型选项值
cap_types = ["1", "2", "3", "4"]  # 市值类型选项值

# 遍历所有组合
for tab_idx in target_tabs:
    for end_val in end_types:
        for equity_val in equity_types:
            for cap_val in cap_types:
                filename = f"基金数据_标签{tab_idx+1}_{end_val}_{equity_val}_{cap_val}.csv"
                try:
                    scrape_tab_combination(driver, tab_idx, end_val, equity_val, cap_val, filename)
                    time.sleep(random.uniform(2, 4))
                except Exception as e:
                    print(f"组合出错 标签{tab_idx+1}_{end_val}_{equity_val}_{cap_val}: {str(e)}")
                    # 出错后切回初始标签页,避免后续定位异常
                    wait_for_element(driver, By.CLASS_NAME, "nav-tabs").find_elements(By.TAG_NAME, "li")[0].click()
                    time.sleep(2)

driver.quit()
关键修复与优化点
  • 标签页切换逻辑:新增标签页定位与切换代码,明确处理前3个标签页,解决原代码未覆盖标签切换的核心问题
  • 日期自动设置:新增set_current_date函数,自动生成并填入符合格式的当前日期,满足需求
  • 元素等待优化:将原代码的presence_of_element_located改为element_to_be_clickable,确保元素可交互后再操作,避免点击/选择失败
  • ALL标签强制选中:每次切换组合前重新选中ALL标签,确保筛选状态符合要求
  • 异常处理增强:出错后自动切回初始标签页,避免后续遍历因页面状态异常中断
  • 数据可读性优化:在CSV中新增标签页、产品类型等标识字段,方便后续数据区分
  • 反爬友好调整:缩短不必要的长等待,用随机等待降低被网站反爬机制拦截的风险

内容的提问来源于stack exchange,提问作者Starlord22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 08:22:40