如何用Python获取股票交易数据?爬取无返回值问题排查
问题排查与修复方案
你的代码无法返回股票数据是以下几个明确错误导致的:
- 核心错误:元素ID拼写错误。CSS选择器中写的目标容器ID是
ctl00_cph1_divSymbols,但BeautifulSoup的find方法里手滑把ID开头的ctl写成了ct1(字母l错写为数字1),完全匹配不到目标DOM节点 - 解析逻辑缺失:就算匹配到了目标div,你也没有继续提取div内部表格的行、单元格数据,仅打印整个div对象无法得到结构化的交易数据
- Selenium驱动写法过时:新版本Selenium已经不支持直接在
webdriver.Chrome()里传executable_path参数,且没有加页面等待逻辑,动态内容未渲染完成就抓取源码会导致拿不到完整内容 - 存在冗余导入:引入的
from xml.dom.minidom import Element全程没有使用,不影响运行但属于无效代码
修复后可运行代码
from selenium import webdriver from bs4 import BeautifulSoup import logging from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') url = "http://eoddata.com/stocklist/NASDAQ/A.htm" # 旧版本Selenium可换回原来的驱动路径写法:webdriver.Chrome(executable_path="C:\Program Files\Chrome\chromedriver") driver = webdriver.Chrome() driver.get(url) # 等待目标表格容器加载完成,最长等待10秒 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, '#ctl00_cph1_divSymbols table')) ) soup = BeautifulSoup(driver.page_source, features="html.parser") # 修正ID拼写错误,定位到目标容器内的表格 table = soup.find('div', {'id': 'ctl00_cph1_divSymbols'}).find('table') stock_data = [] # 跳过表头行,逐行解析交易数据 for row in table.find_all('tr')[1:]: cells = row.find_all('td') if len(cells) < 6: continue stock_info = { "股票代码": cells[0].text.strip(), "股票名称": cells[1].text.strip(), "最高价": cells[2].text.strip(), "最低价": cells[3].text.strip(), "收盘价": cells[4].text.strip(), "成交量": cells[5].text.strip() } stock_data.append(stock_info) logging.info(f"共抓取到{len(stock_data)}条股票交易数据") for item in stock_data[:10]: # 默认打印前10条验证结果,去掉切片即可输出全量数据 logging.info(item) driver.quit()
运行注意事项
- 运行前确保本地Chrome浏览器版本和chromedriver版本完全对应,否则会出现驱动启动失败问题
- 如果访问目标站点网络延迟较高,可以适当调大WebDriverWait的等待时长参数
- 该页面仅展示纳斯达克代码首字母为A的股票列表,需要全量数据可以遍历A-Z的分页路径抓取
内容的提问来源于stack exchange,提问作者Evan Gertis
相关产品推荐
相关产品推荐

