You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Scrapy从搜索引擎抓取数据?以1mg.com为例求助

Hey there! I’ve worked with Scrapy on similar e-commerce scraping tasks, so let’s break down how you can pull data from 1mg.com step by step—super doable once you get the hang of it:

Step 1: Spin Up Your Scrapy Project

First, create a new Scrapy project if you haven’t already. Open your terminal and run:

scrapy startproject one_mg_scraper
cd one_mg_scraper

This sets up the basic project structure you’ll need.

Step 2: Create a Custom Spider

Next, generate a spider specifically for 1mg.com:

scrapy genspider one_mg www.1mg.com

This will create a file one_mg.py inside the spiders folder of your project—this is where you’ll write your scraping logic.

Step 3: Write the Scraping Logic

Open the one_mg.py file and replace the default code with something like this (we’ll use a diabetes category page as an example, but you can adjust the URL to your target page):

import scrapy

class OneMgSpider(scrapy.Spider):
    name = "one_mg"
    allowed_domains = ["1mg.com"]
    # Replace this URL with the specific page you want to scrape
    start_urls = ["https://www.1mg.com/categories/diabetes-100"]

    def parse(self, response):
        # Target each product card on the page
        product_cards = response.css('div.style__product-box___liepi')
        
        for card in product_cards:
            # Extract key data points using CSS selectors
            yield {
                'product_name': card.css('div.style__pro-title___2QwJy::text').get().strip(),
                'price': card.css('div.style__price-tag___cOxY8::text').get(),
                'rating': card.css('div.style__rating___1jYO_::text').get(),
                'product_link': response.urljoin(card.css('a::attr(href)').get())
            }
        
        # Handle pagination to scrape multiple pages
        next_page = response.css('a.style__next-arrow___2wMSq::attr(href)').get()
        if next_page:
            # Follow the next page link and run parse again
            yield response.follow(next_page, callback=self.parse)

Note: CSS selectors might change if 1mg updates their site design. If your scraper stops working, use your browser’s dev tools (F12) to re-inspect the page and update the selectors.

Step 4: Beat Basic Anti-Scraping Measures

1mg has basic anti-scraping checks, so tweak your settings.py file to avoid getting blocked:

  • Add a realistic User-Agent to mimic a real browser:
    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    
  • Add a delay between requests to avoid overwhelming the server:
    DOWNLOAD_DELAY = 2  # 2 seconds between requests
    
  • Check the site’s robots.txt to make sure you’re allowed to scrape the pages you’re targeting. By default, Scrapy respects robots.txt, but if you need to bypass it (only if allowed by the site’s terms), set ROBOTSTXT_OBEY = False.
Step 5: Run the Spider & Save Your Data

Finally, run the spider and save the output to a file (JSON, CSV, etc.):

# Save to JSON
scrapy crawl one_mg -o 1mg_products.json

# Or save to CSV
scrapy crawl one_mg -o 1mg_products.csv

You’ll find the output file in your project folder with all the scraped data!

Quick Heads-Up
  • Always follow the site’s Terms of Service—don’t scrape aggressively or use the data for anything that violates their policies.
  • If you run into dynamically loaded content (data that only appears when scrolling or clicking), you’ll need tools like Scrapy Splash or Playwright to render JavaScript.

内容的提问来源于stack exchange,提问作者Naveenkumar Jain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:39:07