如何用BeautifulSoup爬取美国BLS网站弹窗中的CPI数据表格?
爬取US BLS网站动态加载的CPI区域表格方案
直接用BeautifulSoup单独爬取静态页面源码是拿不到这个表格的——因为点击「Show table」弹出的表格是JavaScript动态渲染的,静态HTML里根本没有这部分内容。不过可以通过以下两种方式解决:
方法1:Selenium模拟交互 + BeautifulSoup解析
用Selenium模拟浏览器操作触发表格加载,再获取渲染后的页面源码,最后用BeautifulSoup提取数据,示例代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 初始化Chrome浏览器(需提前安装对应版本的chromedriver) driver = webdriver.Chrome() driver.get("https://www.bls.gov/charts/consumer-price-index/consumer-price-index-by-region.htm") try: # 等待「Show table」按钮可点击并触发点击 show_table_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Show table']")) ) show_table_btn.click() # 等待表格渲染完成(也可以用等待表格元素出现的方式替代sleep) time.sleep(2) # 获取渲染后的页面源码并解析 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 定位表格并提取数据(根据实际表格class调整选择器) cpi_table = soup.find("table", class_="bls-table") if cpi_table: rows = cpi_table.find_all("tr") for row in rows: cols = row.find_all(["th", "td"]) row_content = [col.get_text(strip=True) for col in cols] print(row_content) finally: driver.quit()
方法2:直接抓取数据API接口(更高效)
打开浏览器开发者工具(F12)切换到「Network」标签,点击「Show table」按钮,观察XHR/fetch请求,能找到加载表格数据的API接口。直接请求这个接口就能拿到结构化JSON数据,无需处理页面渲染:
import requests # 替换为实际抓到的API接口URL api_url = "https://www.bls.gov/charts/consumer-price-index/consumer-price-index-by-region-data.json" response = requests.get(api_url, headers={"User-Agent": "Mozilla/5.0"}) data = response.json() # 解析JSON提取所需数据(结构根据实际返回内容调整) for series in data.get("series", []): print(f"区域:{series.get('region')}") for entry in series.get("values", []): print(f"时间:{entry.get('date')},CPI变动:{entry.get('value')}")
注意事项
- 遵守US BLS网站的使用条款,不要短时间内频繁发送请求,避免IP被封禁。
- 若API接口有验证要求,可复制浏览器请求头中的
User-Agent等参数加入请求。
内容的提问来源于stack exchange,提问作者tacotuesdaydudes
相关产品推荐
相关产品推荐

