You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy或Selenium无法向网站发送Cookie的技术求助

Hey there! Let's break down how to fix your cookie issue and successfully scrape that Sanal Market page—since you're new to web scraping and Scrapy, I'll keep this clear and actionable.

First, let's cover why your current approach might not be working:

  • Cookies are often time-sensitive (session IDs/csrf tokens expire quickly)
  • You might be missing critical cookie fields (like sessionid or csrfmiddlewaretoken)
  • The site checks for matching User-Agent headers between your cookie source and scraping tool
  • Domain mismatches (e.g., using www.sanalmarket.com.tr instead of .sanalmarket.com.tr for cookies)

Step 1: Grab Valid Cookies Manually

  1. Log into Sanal Market in your browser, select your city/region, and navigate to the fruit page
  2. Open DevTools (F12) → Go to Application tab → Cookies → Copy all cookie key-value pairs (pay special attention to sessionid and csrfmiddlewaretoken)

Step 2: Inject Cookies in Scrapy

Use the start_requests method to pass cookies and matching headers:

import scrapy

class SanalMarketSpider(scrapy.Spider):
    name = 'sanalmarket_fruits'
    allowed_domains = ['sanalmarket.com.tr']
    target_url = 'https://www.sanalmarket.com.tr/kweb/sclist/30011-tum-meyveler'

    def start_requests(self):
        # Paste your copied cookies here
        cookies = {
            'sessionid': 'your_valid_session_id',
            'csrfmiddlewaretoken': 'your_csrf_token',
            'city_selection': 'your_saved_city_code',
            # Add all other cookies from your browser
        }

        # Match the User-Agent from your logged-in browser
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        }

        yield scrapy.Request(
            url=self.target_url,
            cookies=cookies,
            headers=headers,
            callback=self.parse_products
        )

    def parse_products(self, response):
        # Replace with actual selectors for product name/price
        for product in response.css('div.product-item'):
            yield {
                'name': product.css('h3.product-name::text').get().strip(),
                'price': product.css('span.product-price::text').get().strip(),
            }

Fallback: Simulate Login in Scrapy

If cookie injection still fails, mimic the actual login flow with FormRequest:

  1. Find the login POST endpoint (check DevTools → Network tab when logging in)
  2. Extract the login form fields (including csrfmiddlewaretoken from the login page)
  3. Submit the form, and Scrapy will automatically persist cookies for subsequent requests

Solution 2: Selenium (More Reliable for Dynamic Sites)

Selenium handles dynamic content and cookie persistence more seamlessly. Here's how to fix your cookie issue:

Step 1: Correct Cookie Injection

Always visit the site first before adding cookies to avoid domain mismatches:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

driver = webdriver.Chrome()
# First visit the site to set the correct domain context
driver.get('https://www.sanalmarket.com.tr')

# Add cookies (copy these from your logged-in browser)
cookies = [
    {'name': 'sessionid', 'value': 'your_session_id', 'domain': '.sanalmarket.com.tr'},
    {'name': 'csrfmiddlewaretoken', 'value': 'your_csrf_token', 'domain': '.sanalmarket.com.tr'},
    {'name': 'selected_district', 'value': 'your_district_code', 'domain': '.sanalmarket.com.tr'},
]

for cookie in cookies:
    driver.add_cookie(cookie)

# Refresh or navigate directly to the fruit page
driver.get('https://www.sanalmarket.com.tr/kweb/sclist/30011-tum-meyveler')
time.sleep(3)  # Wait for dynamic content to load

# Scrape products
for product in driver.find_elements(By.CSS_SELECTOR, 'div.product-item'):
    name = product.find_element(By.CSS_SELECTOR, 'h3.product-name').text.strip()
    price = product.find_element(By.CSS_SELECTOR, 'span.product-price').text.strip()
    print(f"Product: {name} | Price: {price}")

driver.quit()

Fallback: Simulate Full User Flow

If cookie injection doesn't work, just automate the login and city selection directly:

  • Use driver.find_element to locate username/password fields
  • Submit the login form
  • Click through the city/region selection prompts
  • The browser will automatically save all necessary cookies for you

Extra Tips for Avoiding Anti-Scrape Blocks

  • Use a realistic User-Agent that matches your browser
  • Add random delays between requests (don't hammer the site)
  • For Scrapy, consider using scrapy-playwright to handle JavaScript-heavy pages
  • Never share your personal account credentials in public code

内容的提问来源于stack exchange,提问作者Muharrem Akkaya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:31:21