You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cafef股票数据爬虫返回空DataFrame,请求代码调试与逻辑解析

问题解析与修复方案

问题说明

使用以下代码爬取Cafef网站股票历史交易数据时,返回空DataFrame。核心问题包括:元素定位语法错误、无效操作导致循环提前终止、翻页逻辑依赖脆弱的绝对XPath,以及未导入必要模块。

原代码

from selenium import webdriver
from time import sleep
from selenium.webdriver.common.keys import Keys
import pandas as pd
from selenium.webdriver.support.ui import Select
def crawl(stock):
 date=[]
 price=[]
 volume=[]
 close=[]
 stock_id=[]
 browser = webdriver.Chrome(executable_path="./chromedriver")
 web = browser.get("https://s.cafef.vn/Lich-su-giao-dich-"+stock+"-1.chn")
 sleep(5)
 for count in range (60):
  try:
      date_data=browser.find_elements("Item_DateItem")
      for row in date_data:
        date.append(row.text)
        print(row.text())
      date_data.clear()
      price_data=browser.find_elements_by_class_name("Item_Price1")
      for row in price_data:
        price.append(row.text)
      price_data.clear()
  except:
   break
  try:
    if count == 0:
      next_page = browser.find_element(By.XPATH, "/html/body/form/div[3]/div/div[2]/div[2]/div[1]/div[3]/div/div/div[2]/div[2]/div[2]/div/div/div/div/table/tbody/tr/td[21]/a")
    else:
       try:
          next_page = browser.find_element(By.XPATH, "/html/body/form/div[3]/div/div[2]/div[2]/div[1]/div[3]/div/div/div[2]/div[2]/div[2]/div/div/div/div/table/tbody/tr/td[22]/a")
       except:
          next_page = browser.find_element(By.XPATH, "/html/body/form/div[3]/div/div[2]/div[2]/div[1]/div[3]/div/div/div[2]/div[2]/div[2]/div/div/div/div/table/tbody/tr/td[23]/a")
    next_page.click()
    sleep(5)
  except:
    break
 for i in range (int(len(price)/10)):
  close.append(price[10*i+1].replace(",",""))
  volume.append(price[10*i+2].replace(",",""))
 for i in range (len(date)):
  stock_id.append(stock)
 d = {'Stock': stock_id,'Date': date,'Close': close,'Volume': volume}
 df = pd.DataFrame(data=d)
 df.to_csv(stock+".csv", index=False)
 return df
print(crawl('ABC'))

核心问题分析

  1. 元素定位语法错误

    • browser.find_elements("Item_DateItem") 缺少定位方式,正确写法是browser.find_elements(By.CLASS_NAME, "Item_DateItem")
    • print(row.text()) 错误,text是元素属性而非方法,应改为print(row.text)
    • date_data.clear() 和 price_data.clear() 无效:clear()是输入框的方法,元素列表调用会直接触发异常,导致第一次循环就进入except块终止,数据完全没收集到
    • find_elements_by_class_name 是Selenium旧API,建议使用find_elements(By.CLASS_NAME, ...),且原代码未导入By模块
  2. 翻页逻辑缺陷

    • 绝对XPath(如/html/body/form/.../td[21]/a)极度依赖页面结构,只要页面布局微调就会失效
    • 原逻辑试图通过count判断翻页按钮位置,完全不合理,分页按钮的位置不会随页数变化
  3. 数据匹配风险

    • 依赖len(price)/10推导数据条数,若页面结构变化导致price元素数量不是10的倍数,会出现数据长度不匹配,最终生成空DataFrame

修复后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

def crawl(stock):
    date = []
    close = []
    volume = []
    stock_id = []
    
    # 初始化浏览器(若使用Selenium 4.6+,无需指定executable_path)
    browser = webdriver.Chrome()
    browser.get(f"https://s.cafef.vn/Lich-su-giao-dich-{stock}-1.chn")
    
    for count in range(60):
        try:
            # 等待表格加载完成,避免因页面未加载完全导致元素找不到
            WebDriverWait(browser, 10).until(
                EC.presence_of_element_located((By.CLASS_NAME, "Item_DateItem"))
            )
            
            # 抓取日期数据
            date_data = browser.find_elements(By.CLASS_NAME, "Item_DateItem")
            for row in date_data:
                date_text = row.text.strip()
                if date_text:  # 过滤空文本
                    date.append(date_text)
            
            # 抓取收盘价和成交量数据:直接定位对应列的元素,而非所有Item_Price1
            close_data = browser.find_elements(By.CLASS_NAME, "Item_Price1")
            volume_data = browser.find_elements(By.CLASS_NAME, "Item_Volume")
            
            # 确保日期、收盘价、成交量数量匹配
            for c, v in zip(close_data, volume_data):
                close_text = c.text.strip().replace(",", "")
                volume_text = v.text.strip().replace(",", "")
                if close_text and volume_text:
                    close.append(close_text)
                    volume.append(volume_text)
            
        except Exception as e:
            print(f"第{count+1}页数据抓取失败: {str(e)}")
            break
        
        try:
            # 用相对定位找下一页按钮:分页栏中包含"下一页"文本的链接
            next_page = WebDriverWait(browser, 10).until(
                EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), '下一页')]"))
            )
            next_page.click()
            # 等待页面跳转完成
            WebDriverWait(browser, 10).until(
                EC.staleness_of(date_data[0])  # 等待上一页的日期元素失效
            )
        except Exception as e:
            print(f"翻页失败,已爬取{count+1}页: {str(e)}")
            break
    
    # 生成股票ID列表(确保长度与日期一致)
    stock_id = [stock] * len(date)
    
    # 构建DataFrame并保存
    d = {'Stock': stock_id, 'Date': date, 'Close': close, 'Volume': volume}
    df = pd.DataFrame(data=d)
    df.to_csv(f"{stock}.csv", index=False)
    
    browser.quit()  # 关闭浏览器
    return df

print(crawl('ABC'))

关键修改说明

  1. 导入必要模块:新增By、WebDriverWait和expected_conditions,用于更可靠的元素定位和等待
  2. 修复元素定位:
    • 修正find_elements的语法错误,明确指定定位方式
    • 移除无效的clear()调用,避免循环提前终止
    • 直接定位收盘价(Item_Price1)和成交量(Item_Volume)的元素,无需通过price列表间接提取
  3. 优化翻页逻辑:
    • 用相对XPath定位“下一页”按钮,不再依赖脆弱的绝对路径
    • 使用WebDriverWait等待元素可点击,避免因页面未加载完成导致点击失败
  4. 数据可靠性提升:
    • 过滤空文本,避免无效数据
    • 用zip确保日期、收盘价、成交量一一对应
    • 使用显式等待代替固定sleep,提升爬取效率和稳定性
  5. 资源清理:新增browser.quit()关闭浏览器,避免资源泄漏

内容的提问来源于stack exchange,提问作者Linh Chi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 23:05:02