使用Python网页爬取下载宾夕法尼亚州选举网站所有数据集
宾夕法尼亚州选举数据集批量下载解决方案
需求说明
- 目标网站:
https://www.electionreturns.pa.gov/ReportCenter/Reports(非美国地区访问会被重定向) - 核心任务:遍历页面所有下拉框选项,勾选对应复选框,下载全部数据集
- 页面元素参考:
- 下拉框样式:

- 选中下拉选项后显示的复选框样式:

- 下拉框样式:
当前遇到的问题
用requests+BeautifulSoup抓取页面时,拿不到下拉框的实际选项,只能得到初始占位内容:
<select class="form-control ng-invalid ng-invalid-required ng-touched bg-border-mandatory" name="ddlElections" ng-change="rpt.electionChagned(rpt.SelectedElection)" ng-model="rpt.SelectedElection" ng-options="value as value.ElectionName for value in rpt.electionList" placeholder="Select Election"> <option value="">Select Election</option> </select>
当前使用的代码:
html_text=requests.get('https://www.electionreturns.pa.gov/ReportCenter/Reports').content soup=BeautifulSoup(html_text,'lxml') options=soup.find_all('option')
问题根源
这个页面用AngularJS做前端渲染,下拉框选项是通过JavaScript动态加载的。requests只能获取页面初始的静态HTML,无法执行JS,自然拿不到动态生成的选项内容。
两种可行解决方案
方案1:直接调用后端API(推荐)
浏览器加载页面时,会通过API请求获取选举列表和对应数据集信息,直接抓这些接口比模拟浏览器更高效:
- 打开浏览器F12开发者工具,切换到「Network」标签页,刷新页面
- 筛选「XHR/fetch」请求,找到加载选举列表的接口(比如类似
/api/Election/GetElections) - 直接请求该接口拿到所有选举数据的JSON,再遍历每个选举ID获取对应数据集的下载链接
示例代码:
import requests # 替换为你从开发者工具中找到的实际API地址 election_api = "https://www.electionreturns.pa.gov/api/Election/GetElections" report_api_template = "https://www.electionreturns.pa.gov/api/Report/GetReportsByElectionId?electionId={}" # 模拟浏览器请求头,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取所有选举列表 election_response = requests.get(election_api, headers=headers) election_list = election_response.json() # 遍历每个选举,下载对应数据集 for election in election_list: election_id = election["ElectionId"] election_name = election["ElectionName"].replace("/", "-") # 处理文件名特殊字符 # 获取该选举下的所有数据集 report_response = requests.get(report_api_template.format(election_id), headers=headers) report_list = report_response.json() for report in report_list: download_url = report["DownloadUrl"] report_name = report["ReportName"].replace("/", "-") file_name = f"{election_name}_{report_name}.csv" # 下载并保存文件 file_response = requests.get(download_url, headers=headers) with open(file_name, "wb") as f: f.write(file_response.content) print(f"已下载:{file_name}")
方案2:用Selenium模拟浏览器操作
如果找不到API,就用Selenium完全模拟用户操作流程:
from selenium import webdriver from selenium.webdriver.support.ui import Select import time # 初始化浏览器驱动(需提前下载对应浏览器的驱动,比如ChromeDriver) driver = webdriver.Chrome() driver.get("https://www.electionreturns.pa.gov/ReportCenter/Reports") time.sleep(3) # 等待页面加载完成 # 定位选举下拉框 election_select = Select(driver.find_element("name", "ddlElections")) # 过滤掉默认的"Select Election"选项 valid_options = [opt for opt in election_select.options if opt.text.strip() != "Select Election"] for option in valid_options: election_name = option.text.strip() option.click() time.sleep(2) # 等待复选框加载 # 勾选所有可见的复选框 checkboxes = driver.find_elements("xpath", "//input[@type='checkbox']") for checkbox in checkboxes: if not checkbox.is_selected() and checkbox.is_displayed(): checkbox.click() # 点击下载按钮(需根据页面实际按钮调整定位方式) download_btn = driver.find_element("xpath", "//button[contains(text(), 'Download')]") download_btn.click() time.sleep(4) # 等待下载完成 driver.quit()
注意事项
- 非美国地区访问必须使用代理,否则会被网站重定向
- 调用API时要保持请求头和浏览器一致,避免被反爬拦截
- 使用Selenium时,浏览器驱动版本要和本地浏览器版本匹配
内容的提问来源于stack exchange,提问作者user18274957
相关产品推荐
相关产品推荐

