You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中Instagram机器人状态码异常:返回200却提示429错误

问题分析与解决

你遇到的核心问题是两个请求完全独立,导致状态码显示不一致:

  1. driver.get()是Chrome浏览器发起的请求,带有登录后的Cookie和浏览器标识;而requests.get()是无状态的独立HTTP请求,没有携带任何浏览器的登录信息,属于未授权的匿名请求。这两个请求分属不同会话,状态码自然不同。
  2. 后台提示的429是浏览器会话被Instagram限流(请求频率触发反爬机制),但requests.get()的匿名请求未触发同样的频率限制,或被重定向到登录页面(仍返回200状态码),所以你看到的是200。
  3. 你的异常捕获过于宽泛,except:会捕获所有异常,且如果requests.get()失败,response变量在except块中未定义,会引发额外报错。

修复方案

1. 统一请求会话,复用浏览器Cookie

将浏览器的Cookie传给requests,让它与浏览器保持同一会话:

import requests

# 登录后获取浏览器Cookie
cookies = driver.get_cookies()
session = requests.Session()
for cookie in cookies:
    session.cookies.set(cookie['name'], cookie['value'])

# 用session.get替代requests.get
response = session.get(f'https://www.instagram.com/{i}/')
print(response.status_code)

2. 直接从浏览器捕获真实HTTP状态码

通过Chrome DevTools协议监听网络请求,获取driver.get()的真实状态:

import json
from selenium.webdriver.common.desired_capabilities import DesiredCapabilities

# 初始化浏览器时开启性能日志
caps = DesiredCapabilities.CHROME
caps['goog:loggingPrefs'] = {'performance': 'ALL'}
driver = webdriver.Chrome(executable_path='C:\Program Files\chromedriver.exe', desired_capabilities=caps)

# 在driver.get后解析日志获取状态码
driver.get(f'https://www.instagram.com/{i}/')
logs = driver.get_log('performance')
for log in logs:
    message = json.loads(log['message'])['message']
    if 'Network.responseReceived' in message['method']:
        if message['params']['response']['url'] == f'https://www.instagram.com/{i}/':
            status_code = message['params']['response']['status']
            print(f"真实状态码: {status_code}")
            break

3. 处理限流问题

Instagram反爬机制严格,429是明确的限流警告,可通过以下方式缓解:

  • 替换固定延迟为随机延迟:time.sleep(random.randint(3, 8))
  • 模拟人类行为:偶尔滚动页面、随机调整停留时长
  • 使用代理IP分散请求来源

4. 优化异常捕获

避免宽泛的except:,指定具体异常类型,防止掩盖错误:

from selenium.common.exceptions import TimeoutException, NoSuchElementException

try:
    # 你的业务逻辑代码
except (TimeoutException, NoSuchElementException) as e:
    a +=1
    print(f"错误详情: {e}")
    # 确保response已定义再打印
    if 'response' in locals():
        print(response.headers)
    if a == 3:
        driver.quit()

格式化后的原始代码

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time
from datetime import datetime
import requests

driver = webdriver.Chrome(executable_path='C:\Program Files\chromedriver.exe')
driver.get('https://www.instagram.com/')
time.sleep(5)
baslangic = datetime.now()
yeniliste = []
a = 0
b = 0
for i in followlist:
    try:
        follower = driver.get(f'https://www.instagram.com/{i}/')
        response = requests.get(f'https://www.instagram.com/{i}/')
        print(response.status_code)
        WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.XPATH, '(//div[@class="_aa_c"])')))
        text = driver.find_element(By.XPATH, "//div[@class='_aa_c']").text.upper()
        a = 0
        print(response.headers)
        if ("IEU" in text) or ("İEU" in text) or ("İEÜ" in text) or ("IEÜ" in text) or ("IZMIR UNIVERSITY OF ECONOMICS" in text) or ("İZMIR UNIVERSITY OF ECONOMICS" in text) or ("IUE" in text) or ("İUE" in text) or ("IZMIR EKONOMI ÜNIVERSITESI" in text) or ("İZMIR" in text) or ("IZMIR" in text):
            yeniliste.append(i)
            b += 1

    except:
        a += 1
        print("hata")
        print(response.headers)
        if a == 3:
            driver.quit()

内容的提问来源于stack exchange,提问作者freeengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 01:10:29