You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium抓取NCCPL网页时遭安全拦截,如何绕过?

问题描述

需要从NCCPL的LIPI行业每日数据页面提取数据,使用Selenium编写了抓取代码,但在选择日期环节,即便添加sleep等待,仍被网站安全机制拦截(错误截图:错误截图),代码如下:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait as wait
from bs4 import BeautifulSoup
import pandas as pd
import time as t
from selenium import webdriver
from datetime import datetime
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import Select

date = pd.bdate_range('2021-08-05', '2021-08-05')

date = date[0].strftime('%d-%m-%y')

conv_date = datetime.strptime(date, '%d-%m-%y')

month = conv_date.strftime("%b")
year = conv_date.strftime("%Y")
day_name = conv_date.strftime("%A")
day = conv_date.strftime("%d")
day = day.lstrip("0")
month_full_name = conv_date.strftime("%B")
string = 'Select ' + day_name + ', ' + month + ' ' + day + ', ' + year

options = Options()
driver = webdriver.Chrome(options=options)
driver.get('https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily')
t.sleep(10)
picker = wait(driver, 10).until(EC.presence_of_element_located((By.ID, 'popupDatepicker')))
t.sleep(10)
driver.execute_script('arguments[0].scrollIntoView();', picker)
picker.click()
select_year = Select(wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the year"]'))))
select_year.select_by_visible_text(year)
select_month = Select(
    wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the month"]'))))
select_month.select_by_visible_text(month_full_name)
wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="{}"]'.format(string)))).click()
search_button = driver.find_element_by_class_name('search_btn')
search_button.click()

# addition
picker = wait(driver, 10).until(EC.presence_of_element_located((By.ID, 'popupDatepicker1')))
driver.execute_script('arguments[0].scrollIntoView();', picker)
picker.click()
select_year = Select(wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the year"]'))))
select_year.select_by_visible_text(year)
select_month = Select(
    wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="Change the month"]'))))
select_month.select_by_visible_text(month_full_name)
wait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[title="{}"]'.format(string)))).click()
search_button = driver.find_element_by_class_name('search_btn')
search_button.click()

htmlSource = driver.page_source
soup = BeautifulSoup(htmlSource, 'html.parser')
data = pd.read_html(htmlSource)
解决方案

针对网站的反爬机制,可通过以下方式优化绕过检测:

  • 伪装浏览器指纹
    Selenium默认特征易被识别,需配置Chrome选项隐藏自动化痕迹:

    options = Options()
    # 隐藏自动化提示
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option('useAutomationExtension', False)
    # 设置正常用户代理
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    # 禁用自动化控制标识
    options.add_argument("--disable-blink-features=AutomationControlled")
    driver = webdriver.Chrome(options=options)
    # 清除navigator.webdriver属性
    driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")
    
  • 模拟真人操作节奏
    避免连续快速操作,用随机等待替代固定sleep,增加鼠标交互:

    from selenium.webdriver.common.action_chains import ActionChains
    import random
    
    # 鼠标移动到元素后再点击
    action = ActionChains(driver)
    action.move_to_element(picker).pause(random.uniform(0.5, 1.5)).click().perform()
    
    # 随机等待
    t.sleep(random.uniform(2, 4))
    
  • 直接调用数据接口(高效方案)
    分析网站网络请求,找到数据接口后用requests直接请求,跳过前端交互:
    打开浏览器开发者工具,查看搜索按钮触发的XHR请求,提取URL、参数和请求头,示例:

    import requests
    import pandas as pd
    
    headers = {
        "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "referer": "https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily"
    }
    params = {
        "date": "05-08-2021",
        # 其他必要参数从请求中提取
    }
    response = requests.get("目标接口URL", headers=headers, params=params)
    data = response.json()
    df = pd.DataFrame(data)
    
  • 使用undetected-chromedriver
    该库专门用于绕过反爬检测,替代原生Selenium,无需额外配置指纹:

    from undetected_chromedriver import Chrome, ChromeOptions
    
    options = ChromeOptions()
    driver = Chrome(options=options)
    driver.get('https://www.nccpl.com.pk/en/market-information/fipi-lipi/lipi-sector-wise-daily')
    # 后续操作与原代码一致
    

内容的提问来源于stack exchange,提问作者Lopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 19:54:27