You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup从无table标签HTML生成CSV及CFTC数据爬取求助

嘿,我来帮你搞定只提取「Random Length Lumber」数据到CSV的问题!下面是一套实用的Python实现方案,从数据提取到自动更新历史CSV都给你梳理清楚:

提取「Random Length Lumber」数据并更新历史CSV的解决方案

1. 先装必备依赖

打开终端运行下面的命令,安装需要的工具库:

pip install requests pandas beautifulsoup4

2. 核心数据提取代码

这段脚本会精准定位目标商品的数据,提取后保存为新CSV,后续你可以基于它扩展自动更新逻辑:

import requests
import pandas as pd
from bs4 import BeautifulSoup

# 目标页面URL
url = "https://www.cftc.gov/dea/options/other_lof.htm"

# 获取网页内容
response = requests.get(url)
response.raise_for_status()  # 请求失败时直接抛出错误
soup = BeautifulSoup(response.text, 'html.parser')

# 提取页面中的纯文本数据块
data_text = soup.get_text()

# 筛选出Random Length Lumber相关的行
target_rows = []
for line in data_text.splitlines():
    cleaned_line = line.strip()
    # 匹配目标商品,用大写避免大小写问题
    if "RANDOM LENGTH LUMBER" in cleaned_line.upper():
        # 按多空格分割字段(CFTC的数据通常用这种格式)
        fields = [item.strip() for item in cleaned_line.split() if item.strip()]
        target_rows.append(fields)

# 转为DataFrame并保存
if target_rows:
    # 替换成CFTC该数据对应的实际表头,比如商品名、交易所、各类持仓数据等
    custom_headers = ["商品名称", "交易所", "非商业多头", "非商业空头", "商业多头", "商业空头", "未平仓合约"]
    new_data_df = pd.DataFrame(target_rows, columns=custom_headers)
    
    # 保存新提取的数据
    new_data_df.to_csv("random_lumber_new_data.csv", index=False, encoding='utf-8')
    print("✅ Random Length Lumber数据提取成功,已保存到random_lumber_new_data.csv")
else:
    print("❌ 未找到目标数据,请检查商品名称匹配规则是否正确")

3. 扩展:自动合并更新历史CSV

如果要把新数据和历史CSV合并(避免重复),可以加这段逻辑:

# 读取历史CSV文件
try:
    historical_df = pd.read_csv("historical_lumber_data.csv")
    # 按日期字段去重(假设你的数据有日期列,替换成实际唯一标识字段)
    combined_df = pd.concat([historical_df, new_data_df]).drop_duplicates(subset=["日期"], keep="last")
    # 保存更新后的历史数据
    combined_df.to_csv("historical_lumber_data.csv", index=False, encoding='utf-8')
    print("✅ 历史CSV已成功更新")
except FileNotFoundError:
    # 如果历史文件不存在,直接把新数据作为初始历史文件
    new_data_df.to_csv("historical_lumber_data.csv", index=False, encoding='utf-8')
    print("ℹ️ 历史文件不存在,已创建初始历史CSV")

4. 实现每周自动运行

  • Windows:用「任务计划程序」创建定时任务,设置每周固定时间执行这个Python脚本
  • Linux/macOS:用cron定时任务,比如每周一凌晨3点运行:
    0 3 * * 1 /usr/bin/python3 /你的脚本路径/extract_lumber_data.py
    

小提醒:CFTC网站的数据格式偶尔可能变动,记得定期检查提取逻辑是否还能正常工作,比如商品名称的匹配规则、字段分割方式,如果格式变了,及时调整代码里对应的部分就行。

内容的提问来源于stack exchange,提问作者Brian Leonard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:50:01