You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python抓取pricerunner.dk产品数据并导出至Excel的技术求助

Hey Rick, let's tackle your scraping problem step by step. I've looked at your code and the target website, so here's a breakdown of fixes and solutions to get you the product URL, name, and EAN code you need, plus exporting everything to Excel.

1. Why Your Current Code Isn't Working

First, the CSS classes you're using (like mIkxpLfxgo css-183umi2) are dynamically generated by the site's frontend framework (likely React), which means they change regularly. That's why your soup.find calls are failing—those classes don't exist anymore. Also, you're not handling pagination, so you'd only get the first page of products, and you haven't added logic to fetch EAN codes (which aren't present on category pages).

2. Better Approach: Use the Site's Internal API

Pricerunner loads products via XHR requests to an internal API, which is way more reliable than scraping raw HTML. If you open your browser's DevTools > Network tab and refresh the category page, you'll see a request to a URL like https://www.pricerunner.dk/_next/data/<hash>/cl/1424/OEl-Spiritus.json. This JSON response contains all the product data for the page (names, URLs) without messy HTML parsing.

3. How to Get the EAN Code

EAN codes aren't in the category API response, so you'll need to scrape each product's detail page. On product pages, EAN is usually stored in:

  • A hidden <script type="application/ld+json"> tag (structured JSON-LD data)
  • The product specifications section

We can parse either of these with BeautifulSoup.

4. Full Working Code

Here's a revised version of your code that fixes all these issues, handles pagination, fetches EAN codes, and exports everything to Excel:

import openpyxl
import requests
from bs4 import BeautifulSoup
import random
import time
import json
from urllib.parse import urlparse

# Rotating User-Agent list to avoid detection
USER_AGENTS = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36",
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:126.0) Gecko/20100101 Firefox/126.0",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 14_5) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.5 Safari/605.1.15",
    "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
]

def get_random_headers():
    return {"User-Agent": random.choice(USER_AGENTS)}

def fetch_product_ean(product_url):
    """Extract EAN code from a product's detail page"""
    try:
        headers = get_random_headers()
        response = requests.get(product_url, headers=headers)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # First try: Get EAN from JSON-LD structured data (most reliable)
        json_ld_script = soup.find('script', type='application/ld+json')
        if json_ld_script:
            product_data = json.loads(json_ld_script.string)
            if 'gtin13' in product_data:
                return product_data['gtin13']
        
        # Fallback: Check product specifications section
        spec_items = soup.find_all('div', class_='product-spec-item')
        for item in spec_items:
            label = item.find('div', class_='product-spec-label')
            value = item.find('div', class_='product-spec-value')
            if label and 'EAN' in label.text.strip():
                return value.text.strip()
        
        return "N/A"
    except Exception as e:
        print(f"Failed to fetch EAN for {product_url}: {str(e)}")
        return "N/A"

def scrape_category(category_url, sheet):
    """Scrape all products from a category, including pagination"""
    page = 1
    while True:
        try:
            # Get initial page to extract Next.js build ID (for API URL)
            headers = get_random_headers()
            initial_response = requests.get(category_url, headers=headers)
            initial_response.raise_for_status()
            soup = BeautifulSoup(initial_response.text, 'html.parser')
            
            # Extract Next.js build ID to construct the API URL
            next_data_script = soup.find('script', id='__NEXT_DATA__')
            if not next_data_script:
                print("Could not find site's data script—stopping scraping for this category")
                break
            
            next_data = json.loads(next_data_script.string)
            build_id = next_data['buildId']
            api_url = f"https://www.pricerunner.dk/_next/data/{build_id}{urlparse(category_url).path}.json?page={page}"
            
            # Fetch product data from API
            api_response = requests.get(api_url, headers=headers)
            api_response.raise_for_status()
            product_data = api_response.json()
            
            products = product_data['pageProps']['products']
            if not products:
                print(f"No more products on page {page}—moving to next category")
                break
            
            print(f"Scraping page {page} of {category_url}...")
            for product in products:
                product_name = product.get('name', 'N/A')
                product_full_url = f"https://www.pricerunner.dk{product.get('url', '')}"
                ean = fetch_product_ean(product_full_url)
                
                # Append to Excel sheet
                sheet.append([product_name, product_full_url, ean])
                print(f"Added: {product_name} | EAN: {ean}")
            
            # Add delay to avoid being blocked
            time.sleep(random.uniform(1.5, 3))
            page += 1
            
        except Exception as e:
            print(f"Error scraping page {page}: {str(e)}")
            time.sleep(random.uniform(2, 4))
            continue

# Initialize Excel workbook
wb = openpyxl.Workbook()

# Scrape Beer & Spirits category
sheet_spirits = wb.active
sheet_spirits.title = "Beer & Spirits"
sheet_spirits.append(["Product Name", "Product URL", "EAN"])
scrape_category("https://www.pricerunner.dk/cl/1424/OEl-Spiritus", sheet_spirits)

# Scrape Wine category
sheet_wine = wb.create_sheet(title="Wine")
sheet_wine.append(["Product Name", "Product URL", "EAN"])
scrape_category("https://www.pricerunner.dk/cl/465/Vin", sheet_wine)

# Save final Excel file
wb.save("Pricerunner_Product_Data.xlsx")
print("Scraping complete! Data saved to Pricerunner_Product_Data.xlsx")

5. Key Improvements & Notes

  • Dynamic API Detection: Extracts the site's Next.js build ID automatically, so you don't have to hardcode changing API URLs.
  • EAN Extraction: Checks two reliable sources for EAN codes, with a fallback to "N/A" if none is found.
  • Pagination Handling: Automatically iterates through all pages until no more products are available.
  • Anti-Blocking: Uses rotating User-Agents and random delays to avoid being flagged as a bot.
  • Error Resilience: Try/except blocks handle network errors gracefully, so a single failure doesn't stop the entire scrape.

Important Reminders

  • Check the site's robots.txt (https://www.pricerunner.dk/robots.txt) to ensure scraping is allowed for your use case.
  • If you get blocked, increase the delay between requests or add more User-Agents to the list.
  • The site's API structure may change over time—you'll need to adjust the JSON parsing logic if the response format shifts.

内容的提问来源于stack exchange,提问作者Rick Prins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 21:42:35