You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python网页爬取下载宾夕法尼亚州选举网站所有数据集

宾夕法尼亚州选举数据集批量下载解决方案

需求说明

  • 目标网站:https://www.electionreturns.pa.gov/ReportCenter/Reports(非美国地区访问会被重定向)
  • 核心任务:遍历页面所有下拉框选项,勾选对应复选框,下载全部数据集
  • 页面元素参考:
    • 下拉框样式:下拉框样式
    • 选中下拉选项后显示的复选框样式:复选框样式

当前遇到的问题

用requests+BeautifulSoup抓取页面时,拿不到下拉框的实际选项,只能得到初始占位内容:

<select class="form-control ng-invalid ng-invalid-required ng-touched bg-border-mandatory" name="ddlElections" ng-change="rpt.electionChagned(rpt.SelectedElection)" ng-model="rpt.SelectedElection" ng-options="value as value.ElectionName for value in rpt.electionList" placeholder="Select Election">
<option value="">Select Election</option>
</select>

当前使用的代码:

html_text=requests.get('https://www.electionreturns.pa.gov/ReportCenter/Reports').content
soup=BeautifulSoup(html_text,'lxml')
options=soup.find_all('option')

问题根源

这个页面用AngularJS做前端渲染,下拉框选项是通过JavaScript动态加载的。requests只能获取页面初始的静态HTML,无法执行JS,自然拿不到动态生成的选项内容。

两种可行解决方案

方案1:直接调用后端API(推荐)

浏览器加载页面时,会通过API请求获取选举列表和对应数据集信息,直接抓这些接口比模拟浏览器更高效:

  1. 打开浏览器F12开发者工具,切换到「Network」标签页,刷新页面
  2. 筛选「XHR/fetch」请求,找到加载选举列表的接口(比如类似/api/Election/GetElections)
  3. 直接请求该接口拿到所有选举数据的JSON,再遍历每个选举ID获取对应数据集的下载链接

示例代码:

import requests

# 替换为你从开发者工具中找到的实际API地址
election_api = "https://www.electionreturns.pa.gov/api/Election/GetElections"
report_api_template = "https://www.electionreturns.pa.gov/api/Report/GetReportsByElectionId?electionId={}"

# 模拟浏览器请求头,避免被拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 获取所有选举列表
election_response = requests.get(election_api, headers=headers)
election_list = election_response.json()

# 遍历每个选举,下载对应数据集
for election in election_list:
    election_id = election["ElectionId"]
    election_name = election["ElectionName"].replace("/", "-")  # 处理文件名特殊字符
    
    # 获取该选举下的所有数据集
    report_response = requests.get(report_api_template.format(election_id), headers=headers)
    report_list = report_response.json()
    
    for report in report_list:
        download_url = report["DownloadUrl"]
        report_name = report["ReportName"].replace("/", "-")
        file_name = f"{election_name}_{report_name}.csv"
        
        # 下载并保存文件
        file_response = requests.get(download_url, headers=headers)
        with open(file_name, "wb") as f:
            f.write(file_response.content)
        print(f"已下载:{file_name}")

方案2:用Selenium模拟浏览器操作

如果找不到API,就用Selenium完全模拟用户操作流程:

from selenium import webdriver
from selenium.webdriver.support.ui import Select
import time

# 初始化浏览器驱动(需提前下载对应浏览器的驱动,比如ChromeDriver)
driver = webdriver.Chrome()
driver.get("https://www.electionreturns.pa.gov/ReportCenter/Reports")
time.sleep(3)  # 等待页面加载完成

# 定位选举下拉框
election_select = Select(driver.find_element("name", "ddlElections"))
# 过滤掉默认的"Select Election"选项
valid_options = [opt for opt in election_select.options if opt.text.strip() != "Select Election"]

for option in valid_options:
    election_name = option.text.strip()
    option.click()
    time.sleep(2)  # 等待复选框加载
    
    # 勾选所有可见的复选框
    checkboxes = driver.find_elements("xpath", "//input[@type='checkbox']")
    for checkbox in checkboxes:
        if not checkbox.is_selected() and checkbox.is_displayed():
            checkbox.click()
    
    # 点击下载按钮(需根据页面实际按钮调整定位方式)
    download_btn = driver.find_element("xpath", "//button[contains(text(), 'Download')]")
    download_btn.click()
    time.sleep(4)  # 等待下载完成

driver.quit()

注意事项

  • 非美国地区访问必须使用代理,否则会被网站重定向
  • 调用API时要保持请求头和浏览器一致,避免被反爬拦截
  • 使用Selenium时,浏览器驱动版本要和本地浏览器版本匹配

内容的提问来源于stack exchange,提问作者user18274957

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 04:05:19