使用Python+BeautifulSoup爬取基金页面无数据,返回code=0求助
问题解决步骤
1. 核心问题:页面动态加载
目标页面的基金数据是通过JavaScript动态渲染的,requests.get()只能获取初始静态HTML,无法拿到JS加载后的动态内容,所以BeautifulSoup找不到目标元素,自然没有数据输出。
2. 代码修正方案
方案一:用Selenium获取动态页面(能真正拿到数据)
Selenium可以模拟浏览器加载完整页面,适合无编程基础用户操作:
- 先安装依赖:
pip install selenium beautifulsoup4 pandas
- 下载对应浏览器的驱动(比如Chrome的ChromeDriver,需和浏览器版本匹配),放到Python可执行路径下。
- 替换后的完整代码:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import pandas as pd # 配置浏览器无头模式(不弹出窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) url = "https://www.xpi.com.br/investimentos/fundos-de-investimento/lista/#/" driver.get(url) # 等待页面加载完成(可根据网络情况调整等待时间) driver.implicitly_wait(10) # 获取完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 查找所有基金条目 fund_items = soup.find_all('div', class_="xp__card-fund") fund_data = [] for item in fund_items: # 提取基金名称(处理可能找不到元素的情况) title = item.find('div', class_="onclick-fund").get_text(strip=True) if item.find('div', class_="onclick-fund") else "无名称" # 提取基金类型 fund_type = item.find('p', class_="xp__text xp__text--small").get_text(strip=True) if item.find('p', class_="xp__text xp__text--small") else "无类型" fund_data.append({"基金名称": title, "基金类型": fund_type}) # 转为表格并保存为XLS文件 df = pd.DataFrame(fund_data) df.to_excel("xpi_fundos.xlsx", index=False) print("数据已成功保存到xpi_fundos.xlsx") driver.quit()
方案二:原代码的循环错误说明(仅修正语法,仍拿不到数据)
原代码循环里错误使用了lists.find(整个列表对象)而非list.find(当前遍历的单个元素),即使修复这个语法问题,还是拿不到数据,因为页面是动态加载的:
from bs4 import BeautifulSoup import requests url = "https://www.xpi.com.br/investimentos/fundos-de-investimento/lista/#/" page = requests.get(url) soup = BeautifulSoup(page.content,'html.parser') lists = soup.find_all('section', class_="xp-flex-body") for list_item in lists: # 改为用当前遍历的list_item查找子元素 title = list_item.find('div', class_="onclick-fund") fund_type = list_item.find('p', class_="xp__text xp__text--small") info = [title.get_text(strip=True) if title else None, fund_type.get_text(strip=True) if fund_type else None] print(info)
3. 操作提示
运行方案一的代码后,会在当前目录生成xpi_fundos.xlsx文件,直接打开就能看到基金名称和类型数据。
内容的提问来源于stack exchange,提问作者Leonardo Giacomini
相关产品推荐
相关产品推荐

