如何用Python爬取Quantalys多页数据并整理为DataFrame
爬取Quantalys多页ETF数据并整理为DataFrame
依赖安装
先安装所需的Python库:
pip install requests beautifulsoup4 pandas
核心实现代码
import requests from bs4 import BeautifulSoup import pandas as pd import time # 基础请求URL与筛选参数 base_url = "https://www.quantalys.com/Recherche" params = { "Values.lstIdProduits": "1", "Values.bExcludeUncommercialized": "true", "Values.bETF": "true", "Values.nbPerPage": "100", # 指定每页显示100条数据 "page": 1 # 初始页码 } # 模拟浏览器请求头,规避基础反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 存储所有提取到的数据 data_list = [] while True: try: # 发送GET请求获取页面内容 response = requests.get(base_url, params=params, headers=headers) response.raise_for_status() # 捕获HTTP请求错误 soup = BeautifulSoup(response.text, "html.parser") # 定位表格数据行(跳过表头) rows = soup.select("table.table tbody tr") if not rows: print("已无更多数据,爬取结束") break # 遍历每行提取目标字段 for row in rows: # 提取产品名称与链接 name_elem = row.select_one("td a") product_name = name_elem.get_text(strip=True) if name_elem else None product_link = f"https://www.quantalys.com{name_elem['href']}" if name_elem else None # 提取类别、VL、Dev列数据 category = row.select_one("td:nth-child(3)").get_text(strip=True) if row.select_one("td:nth-child(3)") else None vl_value = row.select_one("td:nth-child(4)").get_text(strip=True) if row.select_one("td:nth-child(4)") else None dev_value = row.select_one("td:nth-child(5)").get_text(strip=True) if row.select_one("td:nth-child(5)") else None # 将单条数据加入列表 data_list.append({ "产品名称": product_name, "产品链接": product_link, "类别": category, "VL": vl_value, "Dev": dev_value }) # 打印爬取进度 print(f"已完成第{params['page']}页爬取,累计获取{len(data_list)}条数据") # 页码自增,准备下一页请求 params["page"] += 1 # 设置请求间隔,避免触发反爬机制 time.sleep(1) except Exception as e: print(f"爬取第{params['page']}页时出错: {str(e)}") break # 将整理后的数据转换为DataFrame df = pd.DataFrame(data_list) # 可选:将数据保存为CSV文件 df.to_csv("quantalys_etf_data.csv", index=False, encoding="utf-8-sig") print("数据已保存为 quantalys_etf_data.csv")
关键细节说明
- 每页100条设置:通过
Values.nbPerPage=100参数实现,该参数为网站官方提供的分页控制参数,可通过查看页面分页控件的网络请求确认。 - 翻页逻辑:通过递增
page参数实现循环爬取,当请求返回的表格中无数据行时终止循环。 - 数据定位:使用CSS选择器定位表格元素,若网站结构更新(如类名、列顺序变化),需对应调整选择器索引或类名。
- 反爬处理:添加浏览器标识头、设置请求延迟,降低被网站封禁IP的风险。
注意事项
- 若网站更新页面结构,需重新检查并调整BeautifulSoup的选择器规则。
- 高频次爬取可能导致IP被封禁,建议适当延长请求间隔,或使用代理IP池。
- 请严格遵守网站的
robots.txt协议与使用条款,避免违规爬取。
内容的提问来源于stack exchange,提问作者jacques
相关产品推荐
相关产品推荐

