You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的BeautifulSoup(BS)从指定CBP网页按年提取数据的实现方法咨询

使用Python的BeautifulSoup(BS)从指定CBP网页按年提取数据的实现方法咨询

嗨,我来帮你梳理下怎么用Python实现这个需求~首先得明确:这个CBP网页的筛选功能是依赖JavaScript交互的,纯BeautifulSoup只能解析静态HTML,搞不定选择年份、点击应用按钮这些动态操作,所以得搭配Selenium来模拟浏览器行为,再用BeautifulSoup来解析提取数据,这样就能完美实现你的需求啦。

我给你整理了一套完整的实现思路和示例代码,你可以参考调整:

第一步:安装依赖包

先把需要的库装上,打开终端运行:

pip install selenium beautifulsoup4 webdriver-manager

webdriver-manager能自动帮你管理Chrome/Edge等浏览器的驱动,不用手动下载配置,省很多事。

第二步:核心代码实现

下面的代码会模拟浏览器操作,循环指定年份、设置每页50条、点击应用,然后提取数据,还处理了分页:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import time
import random

# 初始化Chrome浏览器(无头模式可选,加上就不会弹出浏览器窗口)
options = webdriver.ChromeOptions()
# options.add_argument("--headless=new")  # 可选,启用无头模式
driver = webdriver.Chrome(ChromeDriverManager().install(), options=options)

# 打开目标页面
target_url = "https://www.cbp.gov/newsroom/media-releases/all?field_date_release_value%5Bmin%5D=&field_date_release_value%5Bmax%5D=&field_newsroom_type_target_id_1=54&body_value="
driver.get(target_url)
# 用显式等待替代固定sleep,更稳定
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.ID, "edit-field-date-release-value-min"))
)

# 定义要处理的年份范围,比如2019到2024年
year_range = range(2019, 2025)

for year in year_range:
    print(f"===== 开始提取{year}年的数据 =====")
    
    # 1. 设置年份的起止日期
    min_date_input = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ID, "edit-field-date-release-value-min"))
    )
    max_date_input = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ID, "edit-field-date-release-value-max"))
    )
    
    min_date_input.clear()
    min_date_input.send_keys(f"{year}-01-01")
    max_date_input.clear()
    max_date_input.send_keys(f"{year}-12-31")
    
    # 2. 选择每页显示50条数据
    per_page_select = Select(WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "edit-items-per-page"))
    ))
    # 这里的value值要对应页面里50选项的实际value,你可以F12查看元素确认
    per_page_select.select_by_value("50")
    
    # 3. 点击应用按钮生效筛选
    apply_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ID, "edit-submit-media-releases-all"))
    )
    apply_btn.click()
    # 随机等待2-4秒,避免触发反爬
    time.sleep(random.uniform(2, 4))
    
    # 4. 解析当前页面数据,提取内容
    def extract_news_from_page():
        soup = BeautifulSoup(driver.page_source, "html.parser")
        # 这里的容器class要根据页面实际结构调整,比如新闻条目可能在class为"views-row"的div里
        news_containers = soup.find_all("div", class_="views-row")
        extracted_data = []
        for container in news_containers:
            # 提取标题、发布日期、详情链接等信息,按需调整
            title = container.find("h3", class_="title").get_text(strip=True) if container.find("h3", class_="title") else "无标题"
            pub_date = container.find("span", class_="date-display-single").get_text(strip=True) if container.find("span", class_="date-display-single") else "无日期"
            extracted_data.append({"年份": year, "标题": title, "发布日期": pub_date})
        return extracted_data
    
    current_page_data = extract_news_from_page()
    # 这里可以把数据保存到CSV/JSON,比如用csv模块写入文件
    # 示例:打印提取的内容
    for news in current_page_data:
        print(f" - {news['发布日期']} | {news['标题']}")
    
    # 5. 处理分页:循环点击下一页,直到没有下一页为止
    while True:
        try:
            next_page_btn = WebDriverWait(driver, 5).until(
                EC.element_to_be_clickable((By.LINK_TEXT, "下一页"))
            )
            next_page_btn.click()
            time.sleep(random.uniform(2, 4))
            next_page_data = extract_news_from_page()
            for news in next_page_data:
                print(f" - {news['发布日期']} | {news['标题']}")
        except:
            # 捕获异常说明没有下一页了,跳出循环
            print(f"{year}年数据提取完成")
            break

# 操作完成后关闭浏览器
driver.quit()

关键注意事项

  • 元素定位调整:代码里的ID、class、链接文本(比如"下一页")都是基于当前页面结构假设的,实际使用时一定要用浏览器开发者工具(F12)查看页面的真实元素属性,调整定位方式(比如用XPath替代ID,防止网站更新ID变化)。
  • 反爬规避:不要频繁快速请求,加随机等待时间,必要时可以使用代理IP,避免被网站限制访问。
  • 数据存储:可以把extracted_data里的内容写入CSV文件,比如用csv.DictWriter,这样方便后续分析处理。

备注:内容来源于stack exchange,提问作者emiley mille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 12:59:36