Python循环追加DataFrame爬取股票信息写入Excel报错求助
错误排查及修正方案
需先安装的依赖
如果还没安装对应依赖库,先运行以下命令:pip install selenium pandas beautifulsoup4 xlsxwriter
你代码中的问题汇总
- 路径转义错误:Windows本地路径中的反斜杠
\是转义字符,会导致路径解析失败,需要改为原始字符串(路径前加r)或者用正斜杠/代替 - Chrome初始化参数错误:新版本Selenium要求传入service参数指定chromedriver服务,你已经定义了Service对象但没有使用,同时旧版
headless参数已经失效,需要改为--headless=new - 变量类型错误:你定义的
dn是空列表,列表没有to_excel方法,原本应该是把每只股票的DataFrame合并后再导出 pd.concat用法错误:pd.read_html取[0]之后已经是单个DataFrame,不需要再concat,concat只接受多个DataFrame组成的可迭代对象- ExcelWriter位置错误:你把ExcelWriter放在股票循环内部,每次循环都会新建文件覆盖之前的内容,应该放到循环外部
- 股票代码失效:FB的股票代码已经更改为META,用FB无法爬取到对应页面
修正后的可运行代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import pandas as pd headers= { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:87.0) Gecko/20100101 Firefox/87.0', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1', 'Cache-Control': 'max-age=0' } # 路径前加r转为原始字符串避免转义错误 PATH=r'C:\Program Files (x86)\chromedriver.exe' options = Options() # 使用新版无头模式避免渲染异常 options.add_argument('--headless=new') options.add_argument("--window-size=2550,1440") s = Service(PATH) # 正确传入service参数初始化浏览器 driver = webdriver.Chrome(service=s, options=options) driver.implicitly_wait(10) # 存储所有股票的DataFrame用于后续合并 df_list = [] def accept_cookies(): try: driver.find_element(By.ID, 'accept-choices').click() except: print('无cookie弹窗需要处理') # FB已改名为META,调整股票代码 stocklist=["META","KLIC"] # ExcelWriter放到循环外,避免覆盖之前写入的内容 xlwriter = pd.ExcelWriter(f'financial_statements1.xlsx', engine='xlsxwriter') for x in stocklist: url = f"https://stockanalysis.com/stocks/{x}/financials/" driver.get(url) driver.implicitly_wait(10) accept_cookies() driver.implicitly_wait(10) driver.find_element(By.XPATH, "//span[text()='Quarterly']").click() # 等待表格加载完成 driver.implicitly_wait(10) soup = BeautifulSoup(driver.page_source, 'html.parser') df = pd.read_html(str(soup), attrs={'id': 'financial-table'})[0] # 新增列标记股票代码方便区分 df['股票代码'] = x df_list.append(df) # 每只股票数据单独存入对应名称的sheet df.to_excel(xlwriter, sheet_name=x, index=False) # 所有股票数据合并后存入key sheet all_df = pd.concat(df_list, ignore_index=True) all_df.to_excel(xlwriter, sheet_name='key', index=False) xlwriter.close() driver.quit()
内容的提问来源于stack exchange,提问作者stefan s
相关产品推荐
相关产品推荐

