You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫:如何绕过Thefork.co.uk的反爬检测?

解决TheFork网站反爬拦截的实用方案

针对你遇到的TheFork网站反爬拦截问题,除了单纯增加延迟,还有以下更有效的优化方向:

一、替换固定延迟为随机延迟

固定的time.sleep时间过于机械,容易被反爬系统识别。改用随机范围的延迟,模拟人类浏览时的不规则停顿:

import random

# 替换原有的time.sleep(5)、time.sleep(10)等
time.sleep(random.uniform(2, 8))  # 随机等待2-8秒

二、优化undetected_chromedriver启动配置

默认配置可能仍有自动化特征,添加更多模拟真实浏览器的参数:

import undetected_chromedriver as uc

options = uc.ChromeOptions()
# 设置窗口大小,模拟真实浏览器窗口
options.add_argument("--window-size=1920,1080")
# 禁用自动化检测标志
options.add_argument("--disable-blink-features=AutomationControlled")
# 设置真实用户代理(可替换成当前主流浏览器的UA)
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 启用JavaScript(部分网站依赖JS加载内容)
options.add_argument("--enable-javascript")

driver = uc.Chrome(options=options)

三、复用浏览器实例,避免频繁启动

你的代码中get_links和get_data各自创建新的Chrome实例,频繁启动浏览器是典型的爬虫特征。改为复用单个实例:

def get_links(driver):
    # 原get_links逻辑,不再新建driver
    driver.get('https://www.thefork.co.uk/search?coordinates=51.31475930000001%2C-0.5599501')
    # ... 后续逻辑

def get_data(driver):
    # 原get_data逻辑,不再新建driver
    df = pd.read_csv('./Links.csv')
    # ... 后续逻辑

if __name__ == '__main__':
    options = uc.ChromeOptions()
    # 上述配置参数
    driver = uc.Chrome(options=options)
    try:
        # get_links()
        get_data(driver)
    finally:
        driver.quit()  # 最后统一关闭浏览器

四、用WebDriverWait替代固定sleep,精准等待元素加载

固定sleep要么浪费时间,要么因页面加载慢导致元素未出现。改用WebDriverWait等待目标元素加载完成:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 在get_links中替换time.sleep(30)
WebDriverWait(driver, 30).until(
    EC.presence_of_element_located((By.CLASS_NAME, "css-11jucl6"))  # 等待目标链接元素出现
)
content = driver.page_source

五、检测并处理拦截页面

每次页面加载后,检查是否触发了拦截验证(比如你截图中的弹窗),如果出现则暂停等待手动验证:

def check_interception(driver):
    # 根据拦截页面的特征元素调整选择器,比如弹窗的容器ID或类名
    try:
        intercept_popup = driver.find_element(By.CSS_SELECTOR, ".captcha-container")  # 替换为实际拦截元素的选择器
        if intercept_popup.is_displayed():
            print("触发反爬拦截,请手动完成页面验证后按回车继续...")
            input()
    except:
        pass  # 未检测到拦截,继续执行

# 在每次driver.get(url)后调用
driver.get(next_page_full)
check_interception(driver)

六、降低整体请求频率

短时间内大量请求必然触发反爬,可分批次爬取:

  • 每爬取10-20个页面后,暂停1-2分钟
  • 避免连续快速翻页,翻页间隔保持在5-10秒的随机范围

内容的提问来源于stack exchange,提问作者Code Ninja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 10:45:26