爬取USGS地震数据时代码仅执行到print(soup)后停止运行求助
问题原因分析
- 目标页面是Angular单页应用,地震数据通过JavaScript动态加载,
requests.get仅能获取初始静态HTML,其中不存在usgs-event-list这类承载地震数据的标签,导致soup.find_all返回空列表,后续循环无执行。 - 循环内部存在语法错误:遍历
each_eq时,错误使用earthqs.find(整个列表对象)而非each_eq.find(当前遍历的单个元素)。
修复方案
推荐两种实现方式:
方式1:调用USGS官方API(稳定可靠)
USGS提供公开的地震数据API,无需解析动态页面,直接获取结构化数据:
import requests # API端点,可按需调整时间范围、震级等参数 api_url = "https://earthquake.usgs.gov/fdsnws/event/1/query?format=geojson&starttime=2024-01-01&endtime=2024-01-07&minmagnitude=4.5" response = requests.get(api_url) if response.status_code == 200: data = response.json() for eq in data['features']: properties = eq['properties'] magnitude = properties['mag'] location = properties['place'] time = properties['time'] # 时间戳,可自行转换为可读格式 print(f"地震震级: {magnitude}") print(f"地震位置: {location}") print(f"地震时间戳: {time}") print() else: print(f"请求失败,状态码: {response.status_code}")
方式2:用Selenium处理动态渲染页面
若需模拟浏览器加载动态内容,可使用Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup url = "https://earthquake.usgs.gov/earthquakes/map/?extent=-87.0066,-435.23438&extent=86.96966,239.76563&map=false" # 初始化浏览器(需提前安装对应浏览器驱动,如ChromeDriver) driver = webdriver.Chrome() driver.get(url) # 等待动态元素加载完成 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.TAG_NAME, "usgs-event-list"))) # 获取渲染后的页面源码并关闭浏览器 page_content = driver.page_source driver.quit() # 解析页面内容 soup = BeautifulSoup(page_content, "html.parser") earthqs = soup.find_all("usgs-event-list", class_="ng-star-inserted") for each_eq in earthqs: # 需根据页面实际元素结构调整选择器 e_magnitude = each_eq.find("div", class_="magnitude") e_location = each_eq.find("h6", class_="header") e_time = each_eq.find("span", class_="time") if all([e_magnitude, e_location, e_time]): print(f"地震震级: {e_magnitude.text.strip()}") print(f"地震位置: {e_location.text.strip()}") print(f"地震时间: {e_time.text.strip()}") print()
关键提示
- 优先使用官方API,可避免页面结构变更导致的爬虫失效问题,数据获取更稳定。
- 使用Selenium时,需注意页面元素结构可能更新,需根据当前页面实际情况调整CSS选择器。
内容的提问来源于stack exchange,提问作者Lumko Mtengwane
相关产品推荐
相关产品推荐

