You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:解决Instagram自动关注爬虫触发HTTP 429错误的方案

解决Instagram自动关注脚本触发HTTP 429错误的问题

问题描述

我用Scrapy和Selenium编写了Instagram自动关注脚本,通过循环访问目标用户主页执行关注操作。但成功关注100-150个用户后,总会触发HTTP 429错误,提示“Sorry, link maybe broken..”。已尝试添加随机time.sleep(设置不同时长范围)降低请求频率,但问题依旧。

以下是我的代码:

class instaFollowSpider(scrapy.Spider):
    name = 'instaFollowSpider'
    start_urls = ['https://instagram.com']

    def parse(self, response):
        # 登录操作
        chrome_options = Options()
        chrome_options.add_argument("user-data-dir=C:\\Users\\merta\\AppData\\Local\\Google\\Chrome\\User Data\\Default")
        driver = webdriver.Chrome('chromedriver.exe',options=chrome_options)
        driver.get('https://instagram.com')  # 已提前认证登录
        time.sleep(15)

        print("登录成功")
        a=0
        b=0
        baslangic =datetime.now()
        # 关注流程

        for i in followlist:
            try:
                driver.get(f'https://www.instagram.com/{i}/')
                time.sleep(random.randint(15,43))
                if driver.find_element(By.XPATH, "//div[@class='_aacl _aaco _aacw _aad6 _aade']").text == "Takip Et":
                    driver.find_element(By.XPATH, "//div[@class='_aacl _aaco _aacw _aad6 _aade']").click()
                    WebDriverWait(driver, 45).until(EC.visibility_of_element_located((By.XPATH, "//div[@class='_aacl _aaco _aacw _aad6 _aade']")))
                    b += 1
                    a = 0
                    print(i)
                    time.sleep(random.randint(6, 35))


            except NoSuchElementException:
                a += 1
                if a == 3:
                    time.sleep(7500)
                    a=0
                    continue
                else:
                    time.sleep(random.randint(121, 200))
                    continue


            except:
                    driver.get_screenshot_as_file(f"screenshot{i}.png")
                    time.sleep(random.randint(601, 905))
                    continue



        print("总关注人数:", b)
        son =datetime.now()
        print("结束时间:",son," |总运行时长:",son-baslangic)

process = CrawlerProcess()
process.crawl(instaFollowSpider)
process.start()

问题分析

HTTP 429错误本质是Instagram反爬机制触发了限制,除请求频率外,还可能源于以下原因:

  • 行为模式过于机械(固定间隔、无自然操作)
  • 浏览器指纹被识别为自动化工具(如默认ChromeDriver特征)
  • 单账户短时间内关注量超出平台阈值
  • Scrapy与Selenium混用导致请求特征冲突

解决方案

1. 优化行为模拟,贴近真人操作

  • 添加随机浏览动作:进入用户主页后,先模拟滚动页面、随机移动鼠标,避免直接点击关注:
    # 进入主页后随机滚动1-3次
    for _ in range(random.randint(1,3)):
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(random.uniform(2,5))
    
  • 模拟自然点击:用ActionChains模拟鼠标移动到按钮后再点击,避免固定位置直接点击:
    from selenium.webdriver.common.action_chains import ActionChains
    
    follow_btn = driver.find_element(By.XPATH, "//div[text()='Takip Et']")
    ActionChains(driver).move_to_element(follow_btn).pause(random.uniform(0.5,1.5)).click().perform()
    

2. 隐藏浏览器指纹,规避自动化检测

  • 配置ChromeDriver参数:
    chrome_options = Options()
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)
    # 保留原有用户数据目录以维持登录状态
    chrome_options.add_argument("user-data-dir=C:\\Users\\merta\\AppData\\Local\\Google\\Chrome\\User Data\\Default")
    
    driver = webdriver.Chrome('chromedriver.exe', options=chrome_options)
    # 隐藏webdriver标志
    driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")
    
  • 使用高匿代理:固定IP短时间大量请求易被限制,定期更换代理IP可降低风险。

3. 调整关注频率,符合平台限制

  • 控制单时段关注量:Instagram普通账户日关注量通常限制在200-300,1小时内不要超过50个。建议每关注20-30个,主动暂停30-60分钟,而非仅在错误后暂停。
  • 使用更自然的随机间隔:用random.uniform生成浮点型随机时长,避免固定整数间隔的机械感:
    time.sleep(random.uniform(10, 45))  # 替代randint(15,43)
    

4. 修复代码不合理逻辑

  • 移除Scrapy依赖:脚本实际所有操作由Selenium完成,Scrapy的请求队列会与浏览器请求产生特征冲突,建议直接用纯Selenium脚本。
  • 精准捕获异常:避免宽泛的except块,针对具体异常(如TimeoutException、ElementClickInterceptedException)做差异化处理,比如遇到429错误时直接暂停2-4小时。
  • 使用稳定定位方式:Instagram的class名会频繁更新,建议结合文本定位按钮,避免依赖易变的class:
    follow_btn = driver.find_element(By.XPATH, "//div[text()='Takip Et']")
    

5. 账户层面优化

  • 使用老账户:新账户限制更严格,优先使用注册超6个月、有正常互动(点赞、评论、发帖)的老账户。
  • 穿插正常行为:不要仅执行关注操作,穿插点赞、浏览帖子、评论等动作,让账户行为更贴近真人。

内容的提问来源于stack exchange,提问作者freeengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 23:55:28