如何用Python requests模拟点击网页按钮,结合BeautifulSoup提取表格?
解决Investing.com经济日历「Next Week」按钮模拟点击与表格提取问题
问题背景
我想程序化点击http://investing.com/economic-calendar/上的「Next Week」按钮,并用BeautifulSoup提取生成的表格。目前能正常用BeautifulSoup提取表格,但没法模拟点击按钮。
我检查网页后发现目标按钮的类是「newBtn toggleButton LightGray」,想用POST请求模拟点击,但下面的代码没生效,感觉思路接近正确:
import urllib from bs4 import BeautifulSoup import datetime as dt import pandas as pd import requests uri = 'http://investing.com/economic-calendar/' data = {'newBtn toggleButton LightGray' : 'clicked'} page_data = requests.post(uri, data) soup = BeautifulSoup(page_data.content, 'html.parser') table = soup.find('table', {"id": "economicCalendarData"}) tbody = table.find('tbody') rows = tbody.findAll('tr', {"class": "js-event-item"})
从按钮的DOM元素来看,它是一个带有data-range-key="nextWeek"属性的按钮,点击时会触发前端请求加载下周数据。我觉得代码里data = {'newBtn toggleButton LightGray' : 'clicked'}这行有问题,想知道错误原因和正确写法。
错误原因分析
- 你把按钮的class直接作为POST请求的参数键是完全错误的。网页按钮的点击逻辑不是提交自身class作为参数,而是点击时会触发前端发送特定参数的请求,以此告诉服务器要加载的时间范围。
- 另外,该页面加载下周数据的请求是GET请求而非POST,通过URL参数来指定时间范围,和你写的POST请求逻辑完全不符。
正确实现方案
方案1:直接构造时间范围请求(推荐)
观察页面的网络请求可以发现,加载不同时间范围的日历数据时,前端会向特定接口发送带参数的GET请求。比如加载下周数据时,range参数的值为nextWeek。直接构造这个请求就能获取数据,无需模拟点击:
import requests from bs4 import BeautifulSoup import pandas as pd # 数据接口URL api_url = "http://investing.com/economic-calendar/Service/getCalendarFilteredData" # 请求参数:指定时间范围为nextWeek,其他参数可按需调整 request_params = { "timeZone": 8, # 东八区时区,可根据需求修改 "range": "nextWeek", "country": "all", "importance": "3", "category": "all" } # 模拟浏览器请求头,避免被网站拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "X-Requested-With": "XMLHttpRequest" } # 发送请求并解析返回的JSON数据 response = requests.get(api_url, params=request_params, headers=headers) response_data = response.json() # 用BeautifulSoup解析返回的HTML内容 soup = BeautifulSoup(response_data['data'], 'html.parser') table = soup.find('table', {"id": "economicCalendarData"}) # 提取表格数据为DataFrame(可选) calendar_df = pd.read_html(str(table))[0] print(calendar_df)
方案2:用Selenium模拟真实点击(应对复杂反爬)
如果上面的API请求方式失效,可使用Selenium模拟浏览器的真实点击操作,再提取表格数据:
from selenium import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import pandas as pd import time # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get("http://investing.com/economic-calendar/") # 等待页面加载完成,点击「Next Week」按钮 time.sleep(2) next_week_button = driver.find_element(By.CSS_SELECTOR, 'button[data-range-key="nextWeek"]') next_week_button.click() # 等待下周数据加载完成 time.sleep(3) # 解析页面源码提取表格 soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"id": "economicCalendarData"}) calendar_df = pd.read_html(str(table))[0] print(calendar_df) # 关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者Julie Taylor
相关产品推荐
相关产品推荐

