如何用Scrapy爬取Instagram查询?及绕过403/301响应爬取数据求助
Hey there, let's tackle your Instagram scraping questions one by one—Instagram's anti-bot measures are pretty strict, but with the right tweaks, you can get past those issues.
Here's a practical step-by-step approach to build a working Scrapy spider for Instagram:
Handle User Authentication First
Instagram locks down almost all non-public content behind a login wall, so simulating a real user login is non-negotiable. In Scrapy, send a POST request to their AJAX login endpoint with your credentials (pro tip: use environment variables instead of hardcoding your username/password for security). Scrapy automatically persists session cookies, so once you're authenticated, all follow-up requests will carry valid credentials.
Example login code snippet:import os import time import scrapy class InstagramSpider(scrapy.Spider): name = "instagram_scraper" def start_requests(self): yield scrapy.Request( url='https://www.instagram.com/accounts/login/ajax/', method='POST', formdata={ 'username': os.getenv('IG_USERNAME'), 'enc_password': f'#PWD_INSTAGRAM_BROWSER:0:{int(time.time())}:{os.getenv("IG_PASSWORD")}' }, headers={ 'X-Requested-With': 'XMLHttpRequest', 'Referer': 'https://www.instagram.com/accounts/login/' }, callback=self.after_login ) def after_login(self, response): data = response.json() if data.get('authenticated'): # Proceed to scrape target content yield scrapy.Request(url='https://www.instagram.com/target_username/', callback=self.parse_user_page) else: self.logger.error("Login failed—check credentials or encryption format")Note: The
enc_passwordfollows Instagram's browser-side encryption format. You can grab this by inspecting the login request in your browser's dev tools, or use helper libraries (but understanding the logic yourself avoids dependency headaches).Locate the Correct GraphQL Query Endpoint
Thegraphql/queryendpoint you mentioned is valid, but double-check these key parameters:query_id: A fixed value Instagram uses to identify the query type (it changes occasionally, so grab the latest from your browser's dev tools)id: The target user's numeric ID (not their username—fetch this by scraping the user's profile page or using a small GraphQL query for username-to-ID conversion)first: The number of posts to return per page
Build the Scraping Logic
After logging in, construct your GraphQL request with the correct parameters, parse the JSON response to extract your target plain text content, and handle pagination using theend_cursorfrom the response to fetch subsequent pages.Add Basic Anti-Scraping Guards
Start with a reasonableDOWNLOAD_DELAY(2-5 seconds) to avoid triggering rate limits, and use a rotating User-Agent to mimic real browser traffic.
These errors usually come from Instagram detecting bot behavior or missing critical request metadata. Here's how to fix them:
Ensure Valid Session Authentication
A 301 redirect often means your session is invalid, forcing you back to the login page. Verify your login succeeds by checking forauthenticated: truein the JSON response. If login fails, double-check your password encryption format—Instagram sometimes updates this, or you might be hitting a captcha (you'll need to handle captchas manually or use a third-party service for automation).Mirror Browser Request Headers Exactly
Instagram checks multiple header fields to confirm the request comes from a real browser. Copy these headers directly from your browser's dev tools when inspecting a valid GraphQL request:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'X-IG-App-ID': '936619743392459', 'X-Requested-With': 'XMLHttpRequest', 'Referer': 'https://www.instagram.com/target_username/' }Missing or incorrect headers like
X-IG-App-IDare a common cause of 403 errors.Update the
query_id
Instagram periodically rotatesquery_idvalues. If yourquery_idis outdated, you'll get a 403. Refresh it by inspecting a valid GraphQL request in your browser's Network tab, then replace the old value in your code.Use Proxy IPs
If your IP gets blocked by Instagram, you'll see 403 responses. Set up a proxy pool and configure Scrapy to use rotating proxies viaDOWNLOADER_MIDDLEWARESto avoid IP bans.Throttle Your Requests
Even with a valid session, too many requests in a short time will trigger rate limits. Enable Scrapy's auto-throttle withAUTOTHROTTLE_ENABLED = Trueinsettings.py, or set a higherDOWNLOAD_DELAY(3-5 seconds) to mimic human browsing speed.
内容的提问来源于stack exchange,提问作者Ron

