Python中Instagram机器人状态码异常:返回200却提示429错误
问题分析与解决
你遇到的核心问题是两个请求完全独立,导致状态码显示不一致:
driver.get()是Chrome浏览器发起的请求,带有登录后的Cookie和浏览器标识;而requests.get()是无状态的独立HTTP请求,没有携带任何浏览器的登录信息,属于未授权的匿名请求。这两个请求分属不同会话,状态码自然不同。- 后台提示的429是浏览器会话被Instagram限流(请求频率触发反爬机制),但
requests.get()的匿名请求未触发同样的频率限制,或被重定向到登录页面(仍返回200状态码),所以你看到的是200。 - 你的异常捕获过于宽泛,
except:会捕获所有异常,且如果requests.get()失败,response变量在except块中未定义,会引发额外报错。
修复方案
1. 统一请求会话,复用浏览器Cookie
将浏览器的Cookie传给requests,让它与浏览器保持同一会话:
import requests # 登录后获取浏览器Cookie cookies = driver.get_cookies() session = requests.Session() for cookie in cookies: session.cookies.set(cookie['name'], cookie['value']) # 用session.get替代requests.get response = session.get(f'https://www.instagram.com/{i}/') print(response.status_code)
2. 直接从浏览器捕获真实HTTP状态码
通过Chrome DevTools协议监听网络请求,获取driver.get()的真实状态:
import json from selenium.webdriver.common.desired_capabilities import DesiredCapabilities # 初始化浏览器时开启性能日志 caps = DesiredCapabilities.CHROME caps['goog:loggingPrefs'] = {'performance': 'ALL'} driver = webdriver.Chrome(executable_path='C:\Program Files\chromedriver.exe', desired_capabilities=caps) # 在driver.get后解析日志获取状态码 driver.get(f'https://www.instagram.com/{i}/') logs = driver.get_log('performance') for log in logs: message = json.loads(log['message'])['message'] if 'Network.responseReceived' in message['method']: if message['params']['response']['url'] == f'https://www.instagram.com/{i}/': status_code = message['params']['response']['status'] print(f"真实状态码: {status_code}") break
3. 处理限流问题
Instagram反爬机制严格,429是明确的限流警告,可通过以下方式缓解:
- 替换固定延迟为随机延迟:
time.sleep(random.randint(3, 8)) - 模拟人类行为:偶尔滚动页面、随机调整停留时长
- 使用代理IP分散请求来源
4. 优化异常捕获
避免宽泛的except:,指定具体异常类型,防止掩盖错误:
from selenium.common.exceptions import TimeoutException, NoSuchElementException try: # 你的业务逻辑代码 except (TimeoutException, NoSuchElementException) as e: a +=1 print(f"错误详情: {e}") # 确保response已定义再打印 if 'response' in locals(): print(response.headers) if a == 3: driver.quit()
格式化后的原始代码
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time from datetime import datetime import requests driver = webdriver.Chrome(executable_path='C:\Program Files\chromedriver.exe') driver.get('https://www.instagram.com/') time.sleep(5) baslangic = datetime.now() yeniliste = [] a = 0 b = 0 for i in followlist: try: follower = driver.get(f'https://www.instagram.com/{i}/') response = requests.get(f'https://www.instagram.com/{i}/') print(response.status_code) WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.XPATH, '(//div[@class="_aa_c"])'))) text = driver.find_element(By.XPATH, "//div[@class='_aa_c']").text.upper() a = 0 print(response.headers) if ("IEU" in text) or ("İEU" in text) or ("İEÜ" in text) or ("IEÜ" in text) or ("IZMIR UNIVERSITY OF ECONOMICS" in text) or ("İZMIR UNIVERSITY OF ECONOMICS" in text) or ("IUE" in text) or ("İUE" in text) or ("IZMIR EKONOMI ÜNIVERSITESI" in text) or ("İZMIR" in text) or ("IZMIR" in text): yeniliste.append(i) b += 1 except: a += 1 print("hata") print(response.headers) if a == 3: driver.quit()
内容的提问来源于stack exchange,提问作者freeengineer
相关产品推荐
相关产品推荐

