如何从Investing.com抓取历史金融报告数据?求示例教程
抓取Investing.com经济日历历史数据指南(以首次申请失业救济金为例)
一、准备依赖工具
先安装Python必备库,打开终端执行:
pip install requests beautifulsoup4 pandas
如果遇到动态加载数据的情况,额外安装selenium:
pip install selenium
二、静态页面数据抓取脚本(适用于直接显示全量历史数据的页面)
直接用requests和BeautifulSoup解析页面,无需模拟浏览器:
import requests from bs4 import BeautifulSoup import pandas as pd # 替换成你需要抓取的报告链接 target_url = "https://www.investing.com/economic-calendar/initial-jobless-claims-294" # 模拟浏览器请求头,避免被反爬拦截 request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取页面内容 response = requests.get(target_url, headers=request_headers) response.encoding = "utf-8" # 解析HTML定位历史数据表格 soup = BeautifulSoup(response.text, "html.parser") history_table = soup.find("table", class_="history-table") # 提取表头和数据行 table_headers = [th.text.strip() for th in history_table.find("thead").find_all("th")] table_rows = [] for row in history_table.find("tbody").find_all("tr"): row_content = [td.text.strip() for td in row.find_all("td")] table_rows.append(row_content) # 转换为DataFrame并导出到Excel data_frame = pd.DataFrame(table_rows, columns=table_headers) data_frame.to_excel("jobless_claims_history.xlsx", index=False) print("数据已成功导出至Excel")
三、动态页面数据抓取脚本(适用于需要点击「加载更多」的页面)
如果目标页面需要点击按钮加载更多历史数据,用selenium模拟浏览器操作:
- 下载对应浏览器的驱动(比如ChromeDriver,版本要和浏览器匹配),放到Python可执行路径或项目目录下
- 执行以下脚本:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time target_url = "https://www.investing.com/economic-calendar/initial-jobless-claims-294" driver = webdriver.Chrome() driver.get(target_url) # 循环点击「加载更多」直到按钮消失 while True: try: load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), '加载更多')]")) ) load_more_btn.click() time.sleep(2) # 等待数据加载完成 except: break # 提取表格数据 history_table = driver.find_element(By.CLASS_NAME, "history-table") table_headers = [th.text.strip() for th in history_table.find_elements(By.TAG_NAME, "th")] table_rows = [] # 跳过表头行,遍历数据行 for row in history_table.find_elements(By.TAG_NAME, "tr")[1:]: row_content = [td.text.strip() for td in row.find_elements(By.TAG_NAME, "td")] table_rows.append(row_content) # 导出到Excel data_frame = pd.DataFrame(table_rows, columns=table_headers) data_frame.to_excel("jobless_claims_full_history.xlsx", index=False) driver.quit() print("全量历史数据已导出至Excel")
四、使用说明与注意事项
- 手动更换数据时,只需修改脚本中的
target_url变量为对应报告的链接即可 - 不要短时间内频繁发送请求,避免被网站封禁IP
- 若请求被拦截,可更换
User-Agent或添加代理IP - 导出的Excel文件会保存在脚本执行的目录下,直接打开即可整理数据
内容的提问来源于stack exchange,提问作者Gaming World
相关产品推荐
相关产品推荐

