Python从Yahoo! Finance下载比特币数据遇日期缺失问题求助
解决Yahoo! Finance比特币数据缺失问题:BeautifulSoup替代方案
我太懂这种糟心的情况了——用fix_yahoo_finance或者pandas_datareader爬Yahoo Finance数据时,时不时就会丢日期,搞得分析数据全是缺口。改用BeautifulSoup直接爬网页确实是个稳当的办法,我给你补全并优化了你的代码,确保能完整拉取比特币(BTC-USD)的历史数据:
import requests import time import pandas as pd from bs4 import BeautifulSoup def time_convert(dt): # 适配Yahoo Finance网页的日期格式,转成时间戳方便后续分析 try: time_struct = time.strptime(dt, '%b %d, %Y') # 网页日期格式为"Oct 05, 2024"这类 return time.mktime(time_struct) except ValueError: return None def get_btc_yahoo_data(): # 构造BTC-USD历史数据页面URL,可调整时间范围参数 url = "https://finance.yahoo.com/quote/BTC-USD/history?p=BTC-USD" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36" } # 发送请求,避免被反爬拦截 response = requests.get(url, headers=headers) if response.status_code != 200: print("请求失败,状态码:", response.status_code) return None soup = BeautifulSoup(response.text, 'html.parser') # 定位历史数据表格 table = soup.find('table', {'data-test': 'historical-prices'}) rows = table.find_all('tr')[1:] # 跳过表头行 data = [] for row in rows: cols = row.find_all('td') if len(cols) < 6: continue # 跳过空行或广告行 date_str = cols[0].text.strip() open_price = cols[1].text.strip().replace(',', '') high_price = cols[2].text.strip().replace(',', '') low_price = cols[3].text.strip().replace(',', '') close_price = cols[4].text.strip().replace(',', '') adj_close = cols[5].text.strip().replace(',', '') volume = cols[6].text.strip().replace(',', '') # 转换数据类型并收集 try: data.append({ 'date': date_str, 'timestamp': time_convert(date_str), 'open': float(open_price), 'high': float(high_price), 'low': float(low_price), 'close': float(close_price), 'adj_close': float(adj_close), 'volume': int(volume) if volume != '-' else 0 }) except ValueError: print(f"跳过无效行:{date_str}") continue # 转成DataFrame并按日期排序 df = pd.DataFrame(data) df = df.sort_values('timestamp').reset_index(drop=True) return df # 调用函数获取并保存数据 btc_data = get_btc_yahoo_data() if btc_data is not None: print(btc_data.head()) btc_data.to_csv('btc_yahoo_data.csv', index=False)
几个关键注意点:
- 请求头(User-Agent):必须添加,否则Yahoo会直接拒绝请求返回403错误
- 日期格式适配:你之前的时间转换格式和网页实际格式不匹配,我调整为适配网页的
%b %d, %Y格式 - 数据清洗:网页表格可能混入广告行,需判断跳过;数值中的逗号要先移除才能转成数字类型
- 分页扩展:如果需要更早的数据,可观察URL中的
offset参数,循环请求多页数据合并
这个方案比API类工具更稳定,只要网页结构不发生大变动,就不会出现数据缺失的问题。
内容的提问来源于stack exchange,提问作者Rookie0007
相关产品推荐
相关产品推荐

