You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Investing.com抓取历史金融报告数据?求示例教程

抓取Investing.com经济日历历史数据指南(以首次申请失业救济金为例)

一、准备依赖工具

先安装Python必备库,打开终端执行:

pip install requests beautifulsoup4 pandas

如果遇到动态加载数据的情况,额外安装selenium:

pip install selenium

二、静态页面数据抓取脚本(适用于直接显示全量历史数据的页面)

直接用requests和BeautifulSoup解析页面,无需模拟浏览器:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 替换成你需要抓取的报告链接
target_url = "https://www.investing.com/economic-calendar/initial-jobless-claims-294"

# 模拟浏览器请求头,避免被反爬拦截
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 获取页面内容
response = requests.get(target_url, headers=request_headers)
response.encoding = "utf-8"

# 解析HTML定位历史数据表格
soup = BeautifulSoup(response.text, "html.parser")
history_table = soup.find("table", class_="history-table")

# 提取表头和数据行
table_headers = [th.text.strip() for th in history_table.find("thead").find_all("th")]
table_rows = []
for row in history_table.find("tbody").find_all("tr"):
    row_content = [td.text.strip() for td in row.find_all("td")]
    table_rows.append(row_content)

# 转换为DataFrame并导出到Excel
data_frame = pd.DataFrame(table_rows, columns=table_headers)
data_frame.to_excel("jobless_claims_history.xlsx", index=False)
print("数据已成功导出至Excel")

三、动态页面数据抓取脚本(适用于需要点击「加载更多」的页面)

如果目标页面需要点击按钮加载更多历史数据,用selenium模拟浏览器操作:

  1. 下载对应浏览器的驱动(比如ChromeDriver,版本要和浏览器匹配),放到Python可执行路径或项目目录下
  2. 执行以下脚本:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time

target_url = "https://www.investing.com/economic-calendar/initial-jobless-claims-294"
driver = webdriver.Chrome()
driver.get(target_url)

# 循环点击「加载更多」直到按钮消失
while True:
    try:
        load_more_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), '加载更多')]"))
        )
        load_more_btn.click()
        time.sleep(2)  # 等待数据加载完成
    except:
        break

# 提取表格数据
history_table = driver.find_element(By.CLASS_NAME, "history-table")
table_headers = [th.text.strip() for th in history_table.find_elements(By.TAG_NAME, "th")]
table_rows = []
# 跳过表头行,遍历数据行
for row in history_table.find_elements(By.TAG_NAME, "tr")[1:]:
    row_content = [td.text.strip() for td in row.find_elements(By.TAG_NAME, "td")]
    table_rows.append(row_content)

# 导出到Excel
data_frame = pd.DataFrame(table_rows, columns=table_headers)
data_frame.to_excel("jobless_claims_full_history.xlsx", index=False)

driver.quit()
print("全量历史数据已导出至Excel")

四、使用说明与注意事项

  • 手动更换数据时,只需修改脚本中的target_url变量为对应报告的链接即可
  • 不要短时间内频繁发送请求,避免被网站封禁IP
  • 若请求被拦截,可更换User-Agent或添加代理IP
  • 导出的Excel文件会保存在脚本执行的目录下,直接打开即可整理数据

内容的提问来源于stack exchange,提问作者Gaming World

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 05:05:25