You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取Quantalys多页数据并整理为DataFrame

爬取Quantalys多页ETF数据并整理为DataFrame

依赖安装

先安装所需的Python库:

pip install requests beautifulsoup4 pandas

核心实现代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

# 基础请求URL与筛选参数
base_url = "https://www.quantalys.com/Recherche"
params = {
    "Values.lstIdProduits": "1",
    "Values.bExcludeUncommercialized": "true",
    "Values.bETF": "true",
    "Values.nbPerPage": "100",  # 指定每页显示100条数据
    "page": 1  # 初始页码
}

# 模拟浏览器请求头,规避基础反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 存储所有提取到的数据
data_list = []

while True:
    try:
        # 发送GET请求获取页面内容
        response = requests.get(base_url, params=params, headers=headers)
        response.raise_for_status()  # 捕获HTTP请求错误
        soup = BeautifulSoup(response.text, "html.parser")
        
        # 定位表格数据行(跳过表头)
        rows = soup.select("table.table tbody tr")
        if not rows:
            print("已无更多数据,爬取结束")
            break
        
        # 遍历每行提取目标字段
        for row in rows:
            # 提取产品名称与链接
            name_elem = row.select_one("td a")
            product_name = name_elem.get_text(strip=True) if name_elem else None
            product_link = f"https://www.quantalys.com{name_elem['href']}" if name_elem else None
            
            # 提取类别、VL、Dev列数据
            category = row.select_one("td:nth-child(3)").get_text(strip=True) if row.select_one("td:nth-child(3)") else None
            vl_value = row.select_one("td:nth-child(4)").get_text(strip=True) if row.select_one("td:nth-child(4)") else None
            dev_value = row.select_one("td:nth-child(5)").get_text(strip=True) if row.select_one("td:nth-child(5)") else None
            
            # 将单条数据加入列表
            data_list.append({
                "产品名称": product_name,
                "产品链接": product_link,
                "类别": category,
                "VL": vl_value,
                "Dev": dev_value
            })
        
        # 打印爬取进度
        print(f"已完成第{params['page']}页爬取,累计获取{len(data_list)}条数据")
        
        # 页码自增,准备下一页请求
        params["page"] += 1
        
        # 设置请求间隔,避免触发反爬机制
        time.sleep(1)
    
    except Exception as e:
        print(f"爬取第{params['page']}页时出错: {str(e)}")
        break

# 将整理后的数据转换为DataFrame
df = pd.DataFrame(data_list)

# 可选:将数据保存为CSV文件
df.to_csv("quantalys_etf_data.csv", index=False, encoding="utf-8-sig")
print("数据已保存为 quantalys_etf_data.csv")

关键细节说明

  • 每页100条设置:通过Values.nbPerPage=100参数实现,该参数为网站官方提供的分页控制参数,可通过查看页面分页控件的网络请求确认。
  • 翻页逻辑:通过递增page参数实现循环爬取,当请求返回的表格中无数据行时终止循环。
  • 数据定位:使用CSS选择器定位表格元素,若网站结构更新(如类名、列顺序变化),需对应调整选择器索引或类名。
  • 反爬处理:添加浏览器标识头、设置请求延迟,降低被网站封禁IP的风险。

注意事项

  • 若网站更新页面结构,需重新检查并调整BeautifulSoup的选择器规则。
  • 高频次爬取可能导致IP被封禁,建议适当延长请求间隔,或使用代理IP池。
  • 请严格遵守网站的robots.txt协议与使用条款,避免违规爬取。

内容的提问来源于stack exchange,提问作者jacques

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 20:52:45