尼泊尔选举网站动态下拉表格爬取失败,求解决方法
如何爬取带动态依赖下拉菜单的选民列表数据?
我尝试爬取尼泊尔选举委员会网站的选民列表表格数据,该网站包含多个依赖前序选择结果的动态下拉菜单。查看网页源码后发现,数据是通过发送HTTP请求渲染的,但我编写的脚本无法正确获取这些请求返回的预期数据。
我的尝试代码如下:
import requests URL = 'https://voterlist.election.gov.np/bbvrs1/index_process_1.php' #API URL payload = 'vdc=5298&ward=1&list_type=reg_centre' #Unique payload fetched from the network request response = requests.post(URL,data=payload,verify=False) #POST request to get the data using URL and Payload information print(response.text)
这段代码未返回预期的表格数据响应,请问这种情况下的最佳解决方法是什么?
问题排查与解决步骤
- 补全请求头信息:服务器可能会校验请求的
User-Agent、Referer、Cookie等头信息,仅传payload容易被拦截。打开浏览器开发者工具的Network面板,复制完整请求头,添加到requests.post的参数中:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Referer': 'https://election.gov.np/np/page/voter-list-db', 'Cookie': '复制浏览器中该域名下的Cookie内容' } response = requests.post(URL, data=payload, headers=headers, verify=False) - 修正Payload格式:手动拼接的字符串payload可能存在编码问题,改用字典格式让requests自动处理编码:
payload = { 'vdc': '5298', 'ward': '1', 'list_type': 'reg_centre' } - 维持会话状态:部分网站需要先访问首页建立会话(获取必要Cookie),再发送请求。使用
requests.Session()保持会话:session = requests.Session() # 先访问目标页面初始化会话 session.get('https://election.gov.np/np/page/voter-list-db', verify=False) # 再发送POST请求 response = session.post(URL, data=payload, headers=headers, verify=False) - 模拟浏览器操作(Selenium):如果上述方法无效,说明网站有JS渲染或反爬机制,用Selenium直接模拟浏览器交互:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import Select import time driver = webdriver.Chrome() driver.get('https://election.gov.np/np/page/voter-list-db') # 模拟选择下拉菜单(需根据页面实际元素ID调整) select_vdc = Select(driver.find_element(By.ID, 'vdc')) select_vdc.select_by_value('5298') time.sleep(1) # 等待依赖选项加载 select_ward = Select(driver.find_element(By.ID, 'ward')) select_ward.select_by_value('1') time.sleep(1) select_list = Select(driver.find_element(By.ID, 'list_type')) select_list.select_by_value('reg_centre') time.sleep(1) # 触发提交(根据页面实际按钮调整) driver.find_element(By.CSS_SELECTOR, 'button[type="submit"]').click() time.sleep(2) # 提取表格数据 table = driver.find_element(By.TAG_NAME, 'table') for row in table.find_elements(By.TAG_NAME, 'tr'): print([col.text.strip() for col in row.find_elements(By.TAG_NAME, 'td')]) driver.quit() - 验证请求有效性:将浏览器Network面板中的请求复制为cURL命令,执行看是否返回数据。如果cURL成功,将其转换为Python requests代码,对比差异修正自己的脚本。
内容的提问来源于stack exchange,提问作者silentcobra
相关产品推荐
相关产品推荐

