使用Selenium抓取NCCPL网页时遭安全拦截,如何绕过?
问题描述
需要从NCCPL的LIPI行业每日数据页面提取数据,使用Selenium编写了抓取代码,但在选择日期环节,即便添加sleep等待,仍被网站安全机制拦截(错误截图:
),代码如下:
from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait as wait from bs4 import BeautifulSoup import pandas as pd import time as t from selenium import webdriver from datetime import datetime from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import Select date = pd.bdate_range('2021-08-05', '2021-08-05') date = date[0].strftime('%d-%m-%y') conv_date = datetime.strptime(date, '%d-%m-%y') month = conv_date.strftime("%b") year = conv_date.strftime("%Y") day_name = conv_date.strftime("%A") day = conv_date.strftime("%d") day = day.lstrip("0") month_full_name = conv_date.strftime("%B") string = 'Select ' + day_name + ', ' + month + ' ' + day + ', ' + year options = Options() driver = webdriver.Chrome(options=options) driver.get('https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily') t.sleep(10) picker = wait(driver, 10).until(EC.presence_of_element_located((By.ID, 'popupDatepicker'))) t.sleep(10) driver.execute_script('arguments[0].scrollIntoView();', picker) picker.click() select_year = Select(wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the year"]')))) select_year.select_by_visible_text(year) select_month = Select( wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the month"]')))) select_month.select_by_visible_text(month_full_name) wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="{}"]'.format(string)))).click() search_button = driver.find_element_by_class_name('search_btn') search_button.click() # addition picker = wait(driver, 10).until(EC.presence_of_element_located((By.ID, 'popupDatepicker1'))) driver.execute_script('arguments[0].scrollIntoView();', picker) picker.click() select_year = Select(wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the year"]')))) select_year.select_by_visible_text(year) select_month = Select( wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the month"]')))) select_month.select_by_visible_text(month_full_name) wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="{}"]'.format(string)))).click() search_button = driver.find_element_by_class_name('search_btn') search_button.click() htmlSource = driver.page_source soup = BeautifulSoup(htmlSource, 'html.parser') data = pd.read_html(htmlSource)
解决方案
针对网站的反爬机制,可通过以下方式优化绕过检测:
伪装浏览器指纹
Selenium默认特征易被识别,需配置Chrome选项隐藏自动化痕迹:options = Options() # 隐藏自动化提示 options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # 设置正常用户代理 options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 禁用自动化控制标识 options.add_argument("--disable-blink-features=AutomationControlled") driver = webdriver.Chrome(options=options) # 清除navigator.webdriver属性 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")模拟真人操作节奏
避免连续快速操作,用随机等待替代固定sleep,增加鼠标交互:from selenium.webdriver.common.action_chains import ActionChains import random # 鼠标移动到元素后再点击 action = ActionChains(driver) action.move_to_element(picker).pause(random.uniform(0.5, 1.5)).click().perform() # 随机等待 t.sleep(random.uniform(2, 4))直接调用数据接口(高效方案)
分析网站网络请求,找到数据接口后用requests直接请求,跳过前端交互:
打开浏览器开发者工具,查看搜索按钮触发的XHR请求,提取URL、参数和请求头,示例:import requests import pandas as pd headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "referer": "https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily" } params = { "date": "05-08-2021", # 其他必要参数从请求中提取 } response = requests.get("目标接口URL", headers=headers, params=params) data = response.json() df = pd.DataFrame(data)使用undetected-chromedriver
该库专门用于绕过反爬检测,替代原生Selenium,无需额外配置指纹:from undetected_chromedriver import Chrome, ChromeOptions options = ChromeOptions() driver = Chrome(options=options) driver.get('https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily') # 后续操作与原代码一致
内容的提问来源于stack exchange,提问作者Lopez
相关产品推荐
相关产品推荐

