如何通过Scrapy从搜索引擎抓取数据?以1mg.com为例求助
Hey there! I’ve worked with Scrapy on similar e-commerce scraping tasks, so let’s break down how you can pull data from 1mg.com step by step—super doable once you get the hang of it:
First, create a new Scrapy project if you haven’t already. Open your terminal and run:
scrapy startproject one_mg_scraper cd one_mg_scraper
This sets up the basic project structure you’ll need.
Next, generate a spider specifically for 1mg.com:
scrapy genspider one_mg www.1mg.com
This will create a file one_mg.py inside the spiders folder of your project—this is where you’ll write your scraping logic.
Open the one_mg.py file and replace the default code with something like this (we’ll use a diabetes category page as an example, but you can adjust the URL to your target page):
import scrapy class OneMgSpider(scrapy.Spider): name = "one_mg" allowed_domains = ["1mg.com"] # Replace this URL with the specific page you want to scrape start_urls = ["https://www.1mg.com/categories/diabetes-100"] def parse(self, response): # Target each product card on the page product_cards = response.css('div.style__product-box___liepi') for card in product_cards: # Extract key data points using CSS selectors yield { 'product_name': card.css('div.style__pro-title___2QwJy::text').get().strip(), 'price': card.css('div.style__price-tag___cOxY8::text').get(), 'rating': card.css('div.style__rating___1jYO_::text').get(), 'product_link': response.urljoin(card.css('a::attr(href)').get()) } # Handle pagination to scrape multiple pages next_page = response.css('a.style__next-arrow___2wMSq::attr(href)').get() if next_page: # Follow the next page link and run parse again yield response.follow(next_page, callback=self.parse)
Note: CSS selectors might change if 1mg updates their site design. If your scraper stops working, use your browser’s dev tools (F12) to re-inspect the page and update the selectors.
1mg has basic anti-scraping checks, so tweak your settings.py file to avoid getting blocked:
- Add a realistic User-Agent to mimic a real browser:
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' - Add a delay between requests to avoid overwhelming the server:
DOWNLOAD_DELAY = 2 # 2 seconds between requests - Check the site’s
robots.txtto make sure you’re allowed to scrape the pages you’re targeting. By default, Scrapy respectsrobots.txt, but if you need to bypass it (only if allowed by the site’s terms), setROBOTSTXT_OBEY = False.
Finally, run the spider and save the output to a file (JSON, CSV, etc.):
# Save to JSON scrapy crawl one_mg -o 1mg_products.json # Or save to CSV scrapy crawl one_mg -o 1mg_products.csv
You’ll find the output file in your project folder with all the scraped data!
- Always follow the site’s Terms of Service—don’t scrape aggressively or use the data for anything that violates their policies.
- If you run into dynamically loaded content (data that only appears when scrolling or clicking), you’ll need tools like Scrapy Splash or Playwright to render JavaScript.
内容的提问来源于stack exchange,提问作者Naveenkumar Jain

