基于BeautifulSoup从无table标签HTML生成CSV及CFTC数据爬取求助
嘿,我来帮你搞定只提取「Random Length Lumber」数据到CSV的问题!下面是一套实用的Python实现方案,从数据提取到自动更新历史CSV都给你梳理清楚:
提取「Random Length Lumber」数据并更新历史CSV的解决方案
1. 先装必备依赖
打开终端运行下面的命令,安装需要的工具库:
pip install requests pandas beautifulsoup4
2. 核心数据提取代码
这段脚本会精准定位目标商品的数据,提取后保存为新CSV,后续你可以基于它扩展自动更新逻辑:
import requests import pandas as pd from bs4 import BeautifulSoup # 目标页面URL url = "https://www.cftc.gov/dea/options/other_lof.htm" # 获取网页内容 response = requests.get(url) response.raise_for_status() # 请求失败时直接抛出错误 soup = BeautifulSoup(response.text, 'html.parser') # 提取页面中的纯文本数据块 data_text = soup.get_text() # 筛选出Random Length Lumber相关的行 target_rows = [] for line in data_text.splitlines(): cleaned_line = line.strip() # 匹配目标商品,用大写避免大小写问题 if "RANDOM LENGTH LUMBER" in cleaned_line.upper(): # 按多空格分割字段(CFTC的数据通常用这种格式) fields = [item.strip() for item in cleaned_line.split() if item.strip()] target_rows.append(fields) # 转为DataFrame并保存 if target_rows: # 替换成CFTC该数据对应的实际表头,比如商品名、交易所、各类持仓数据等 custom_headers = ["商品名称", "交易所", "非商业多头", "非商业空头", "商业多头", "商业空头", "未平仓合约"] new_data_df = pd.DataFrame(target_rows, columns=custom_headers) # 保存新提取的数据 new_data_df.to_csv("random_lumber_new_data.csv", index=False, encoding='utf-8') print("✅ Random Length Lumber数据提取成功,已保存到random_lumber_new_data.csv") else: print("❌ 未找到目标数据,请检查商品名称匹配规则是否正确")
3. 扩展:自动合并更新历史CSV
如果要把新数据和历史CSV合并(避免重复),可以加这段逻辑:
# 读取历史CSV文件 try: historical_df = pd.read_csv("historical_lumber_data.csv") # 按日期字段去重(假设你的数据有日期列,替换成实际唯一标识字段) combined_df = pd.concat([historical_df, new_data_df]).drop_duplicates(subset=["日期"], keep="last") # 保存更新后的历史数据 combined_df.to_csv("historical_lumber_data.csv", index=False, encoding='utf-8') print("✅ 历史CSV已成功更新") except FileNotFoundError: # 如果历史文件不存在,直接把新数据作为初始历史文件 new_data_df.to_csv("historical_lumber_data.csv", index=False, encoding='utf-8') print("ℹ️ 历史文件不存在,已创建初始历史CSV")
4. 实现每周自动运行
- Windows:用「任务计划程序」创建定时任务,设置每周固定时间执行这个Python脚本
- Linux/macOS:用
cron定时任务,比如每周一凌晨3点运行:0 3 * * 1 /usr/bin/python3 /你的脚本路径/extract_lumber_data.py
小提醒:CFTC网站的数据格式偶尔可能变动,记得定期检查提取逻辑是否还能正常工作,比如商品名称的匹配规则、字段分割方式,如果格式变了,及时调整代码里对应的部分就行。
内容的提问来源于stack exchange,提问作者Brian Leonard
相关产品推荐
相关产品推荐

