使用Scrapy或Selenium无法向网站发送Cookie的技术求助
Hey there! Let's break down how to fix your cookie issue and successfully scrape that Sanal Market page—since you're new to web scraping and Scrapy, I'll keep this clear and actionable.
Common Reasons Cookie Injection Fails
First, let's cover why your current approach might not be working:
- Cookies are often time-sensitive (session IDs/csrf tokens expire quickly)
- You might be missing critical cookie fields (like
sessionidorcsrfmiddlewaretoken) - The site checks for matching
User-Agentheaders between your cookie source and scraping tool - Domain mismatches (e.g., using
www.sanalmarket.com.trinstead of.sanalmarket.com.trfor cookies)
Solution 1: Scrapy (Cookie Injection + Fallback to Simulated Login)
Step 1: Grab Valid Cookies Manually
- Log into Sanal Market in your browser, select your city/region, and navigate to the fruit page
- Open DevTools (F12) → Go to Application tab → Cookies → Copy all cookie key-value pairs (pay special attention to
sessionidandcsrfmiddlewaretoken)
Step 2: Inject Cookies in Scrapy
Use the start_requests method to pass cookies and matching headers:
import scrapy class SanalMarketSpider(scrapy.Spider): name = 'sanalmarket_fruits' allowed_domains = ['sanalmarket.com.tr'] target_url = 'https://www.sanalmarket.com.tr/kweb/sclist/30011-tum-meyveler' def start_requests(self): # Paste your copied cookies here cookies = { 'sessionid': 'your_valid_session_id', 'csrfmiddlewaretoken': 'your_csrf_token', 'city_selection': 'your_saved_city_code', # Add all other cookies from your browser } # Match the User-Agent from your logged-in browser headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } yield scrapy.Request( url=self.target_url, cookies=cookies, headers=headers, callback=self.parse_products ) def parse_products(self, response): # Replace with actual selectors for product name/price for product in response.css('div.product-item'): yield { 'name': product.css('h3.product-name::text').get().strip(), 'price': product.css('span.product-price::text').get().strip(), }
Fallback: Simulate Login in Scrapy
If cookie injection still fails, mimic the actual login flow with FormRequest:
- Find the login POST endpoint (check DevTools → Network tab when logging in)
- Extract the login form fields (including
csrfmiddlewaretokenfrom the login page) - Submit the form, and Scrapy will automatically persist cookies for subsequent requests
Solution 2: Selenium (More Reliable for Dynamic Sites)
Selenium handles dynamic content and cookie persistence more seamlessly. Here's how to fix your cookie issue:
Step 1: Correct Cookie Injection
Always visit the site first before adding cookies to avoid domain mismatches:
from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() # First visit the site to set the correct domain context driver.get('https://www.sanalmarket.com.tr') # Add cookies (copy these from your logged-in browser) cookies = [ {'name': 'sessionid', 'value': 'your_session_id', 'domain': '.sanalmarket.com.tr'}, {'name': 'csrfmiddlewaretoken', 'value': 'your_csrf_token', 'domain': '.sanalmarket.com.tr'}, {'name': 'selected_district', 'value': 'your_district_code', 'domain': '.sanalmarket.com.tr'}, ] for cookie in cookies: driver.add_cookie(cookie) # Refresh or navigate directly to the fruit page driver.get('https://www.sanalmarket.com.tr/kweb/sclist/30011-tum-meyveler') time.sleep(3) # Wait for dynamic content to load # Scrape products for product in driver.find_elements(By.CSS_SELECTOR, 'div.product-item'): name = product.find_element(By.CSS_SELECTOR, 'h3.product-name').text.strip() price = product.find_element(By.CSS_SELECTOR, 'span.product-price').text.strip() print(f"Product: {name} | Price: {price}") driver.quit()
Fallback: Simulate Full User Flow
If cookie injection doesn't work, just automate the login and city selection directly:
- Use
driver.find_elementto locate username/password fields - Submit the login form
- Click through the city/region selection prompts
- The browser will automatically save all necessary cookies for you
Extra Tips for Avoiding Anti-Scrape Blocks
- Use a realistic
User-Agentthat matches your browser - Add random delays between requests (don't hammer the site)
- For Scrapy, consider using
scrapy-playwrightto handle JavaScript-heavy pages - Never share your personal account credentials in public code
内容的提问来源于stack exchange,提问作者Muharrem Akkaya

