You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python每日抓取同一股票30天价格并避免列表值被覆盖?

嘿,我来帮你搞定这个问题!你遇到的“新增价格时初始值被替换”的问题,核心原因是每次运行脚本时,price_list和dttm都会被重新初始化为空列表——脚本每次启动都是独立的进程,内存里的列表没法在多次运行间保留数据。另外你的代码还有几个小问题需要调整,比如没精准定位到Catalyst Pharmaceuticals的股票价格,以及文件处理的逻辑不够稳妥。

先说说你的代码里的关键问题

  • 列表无法持久化:每次运行脚本,price_list都是空的,所以append的只有本次抓取的价格,之前的数据根本没地方存。
  • 目标股票定位错误:你当前爬的是Robinhood的“technology集合”页面的第二个a标签,这绝对不是Catalyst Pharmaceuticals的价格,得直接访问它的股票详情页才行。
  • 文件关闭时机不对:在websc函数里直接关闭了CSV文件,要是后续还有写入操作就会报错,最好用with语句自动管理文件。

修正后的完整代码

下面是调整后的代码,解决了上述问题,还增加了错误处理,让脚本更健壮:

import requests
import pytz
from bs4 import BeautifulSoup
import datetime
import csv
import os

# 直接访问Catalyst Pharmaceuticals的股票页面(CPRX是它的股票代码)
TARGET_STOCK_URL = 'https://robinhood.com/stocks/CPRX'

def fetch_current_stock_price():
    """抓取当前目标股票的价格"""
    try:
        response = requests.get(TARGET_STOCK_URL)
        response.raise_for_status()  # 请求失败时抛出异常
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # 定位股票价格元素(Robinhood的页面class可能会变,若失效请检查页面结构)
        price_tag = soup.find('div', {'class': 'QzVHcLdwl2CEuEMpTUFaj'})
        if not price_tag:
            raise ValueError("找不到股票价格元素,请检查页面结构是否更新")
        
        return price_tag.text.strip()
    except Exception as e:
        print(f"抓取价格失败: {str(e)}")
        raise

def save_price_to_csv(price, timestamp):
    """将价格和时间戳追加到CSV,自动处理表头"""
    csv_file_path = 'Catalyst Pharmaceuticals Monthly.csv'
    file_exists = os.path.isfile(csv_file_path)
    
    # 使用with语句自动管理文件,避免手动关闭出错
    with open(csv_file_path, 'a', newline='', encoding='utf-8') as csv_file:
        writer = csv.writer(csv_file)
        # 第一次运行时写入表头
        if not file_exists:
            writer.writerow(['GMT Timestamp', 'Stock Price'])
        # 写入新的价格数据
        writer.writerow([timestamp, price])

def main():
    try:
        # 获取GMT时区的当前时间
        current_gmt_time = datetime.datetime.now(pytz.timezone("GMT")).strftime("%Y-%m-%d %H:%M:%S")
        # 抓取当前价格
        current_price = fetch_current_stock_price()
        print(f"[{current_gmt_time}] 成功抓取价格: {current_price}")
        
        # 保存到CSV
        save_price_to_csv(current_price, current_gmt_time)
        print("数据已成功追加到CSV文件")
        
    except Exception as e:
        print(f"程序出错: {str(e)}")

if __name__ == "__main__":
    main()

实现每日自动运行

要每天自动抓取,有两种靠谱的方式:

  1. 用Python的schedule库:
    先安装库:pip install schedule
    然后修改main部分,添加定时逻辑:
    import schedule
    import time
    
    def daily_scrape_job():
        main()
    
    # 设置每天9点(GMT时区)运行一次
    schedule.every().day.at("09:00").do(daily_scrape_job)
    
    print("脚本已启动,将每天9点自动抓取股票价格...")
    while True:
        schedule.run_pending()
        time.sleep(60)  # 每分钟检查一次任务
    
  2. 用操作系统定时任务(更稳定):
    • Windows:打开「任务计划程序」,创建基本任务,设置每天运行时间,选择执行你的Python脚本。
    • Linux/macOS:编辑crontab(crontab -e),添加一行:0 9 * * * /usr/bin/python3 /path/to/your/script.py(意思是每天GMT时间9点运行脚本)。

如何生成30天的价格列表

因为数据都存在CSV里了,你可以随时读取并筛选最近30天的记录:

import pandas as pd
from datetime import datetime, timedelta
import pytz

csv_path = 'Catalyst Pharmaceuticals Monthly.csv'
# 读取CSV数据
df = pd.read_csv(csv_path)
# 将时间戳列转换为datetime类型
df['GMT Timestamp'] = pd.to_datetime(df['GMT Timestamp'], format="%Y-%m-%d %H:%M:%S")

# 计算30天前的GMT时间
thirty_days_ago = datetime.now(pytz.timezone("GMT")) - timedelta(days=30)
# 筛选最近30天的数据
recent_30_days_data = df[df['GMT Timestamp'] >= thirty_days_ago]

# 转换为你需要的列表格式
price_list = recent_30_days_data['Stock Price'].tolist()
timestamp_list = recent_30_days_data['GMT Timestamp'].dt.strftime("%Y-%m-%d %H:%M:%S").tolist()

print("最近30天的股票价格列表:")
for ts, price in zip(timestamp_list, price_list):
    print(f"{ts}: {price}")

内容的提问来源于stack exchange,提问作者Lucifer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:52:54