使用Python抓取pricerunner.dk产品数据并导出至Excel的技术求助
Hey Rick, let's tackle your scraping problem step by step. I've looked at your code and the target website, so here's a breakdown of fixes and solutions to get you the product URL, name, and EAN code you need, plus exporting everything to Excel.
1. Why Your Current Code Isn't Working
First, the CSS classes you're using (like mIkxpLfxgo css-183umi2) are dynamically generated by the site's frontend framework (likely React), which means they change regularly. That's why your soup.find calls are failing—those classes don't exist anymore. Also, you're not handling pagination, so you'd only get the first page of products, and you haven't added logic to fetch EAN codes (which aren't present on category pages).
2. Better Approach: Use the Site's Internal API
Pricerunner loads products via XHR requests to an internal API, which is way more reliable than scraping raw HTML. If you open your browser's DevTools > Network tab and refresh the category page, you'll see a request to a URL like https://www.pricerunner.dk/_next/data/<hash>/cl/1424/OEl-Spiritus.json. This JSON response contains all the product data for the page (names, URLs) without messy HTML parsing.
3. How to Get the EAN Code
EAN codes aren't in the category API response, so you'll need to scrape each product's detail page. On product pages, EAN is usually stored in:
- A hidden
<script type="application/ld+json">tag (structured JSON-LD data) - The product specifications section
We can parse either of these with BeautifulSoup.
4. Full Working Code
Here's a revised version of your code that fixes all these issues, handles pagination, fetches EAN codes, and exports everything to Excel:
import openpyxl import requests from bs4 import BeautifulSoup import random import time import json from urllib.parse import urlparse # Rotating User-Agent list to avoid detection USER_AGENTS = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36", "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:126.0) Gecko/20100101 Firefox/126.0", "Mozilla/5.0 (Macintosh; Intel Mac OS X 14_5) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.5 Safari/605.1.15", "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" ] def get_random_headers(): return {"User-Agent": random.choice(USER_AGENTS)} def fetch_product_ean(product_url): """Extract EAN code from a product's detail page""" try: headers = get_random_headers() response = requests.get(product_url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # First try: Get EAN from JSON-LD structured data (most reliable) json_ld_script = soup.find('script', type='application/ld+json') if json_ld_script: product_data = json.loads(json_ld_script.string) if 'gtin13' in product_data: return product_data['gtin13'] # Fallback: Check product specifications section spec_items = soup.find_all('div', class_='product-spec-item') for item in spec_items: label = item.find('div', class_='product-spec-label') value = item.find('div', class_='product-spec-value') if label and 'EAN' in label.text.strip(): return value.text.strip() return "N/A" except Exception as e: print(f"Failed to fetch EAN for {product_url}: {str(e)}") return "N/A" def scrape_category(category_url, sheet): """Scrape all products from a category, including pagination""" page = 1 while True: try: # Get initial page to extract Next.js build ID (for API URL) headers = get_random_headers() initial_response = requests.get(category_url, headers=headers) initial_response.raise_for_status() soup = BeautifulSoup(initial_response.text, 'html.parser') # Extract Next.js build ID to construct the API URL next_data_script = soup.find('script', id='__NEXT_DATA__') if not next_data_script: print("Could not find site's data script—stopping scraping for this category") break next_data = json.loads(next_data_script.string) build_id = next_data['buildId'] api_url = f"https://www.pricerunner.dk/_next/data/{build_id}{urlparse(category_url).path}.json?page={page}" # Fetch product data from API api_response = requests.get(api_url, headers=headers) api_response.raise_for_status() product_data = api_response.json() products = product_data['pageProps']['products'] if not products: print(f"No more products on page {page}—moving to next category") break print(f"Scraping page {page} of {category_url}...") for product in products: product_name = product.get('name', 'N/A') product_full_url = f"https://www.pricerunner.dk{product.get('url', '')}" ean = fetch_product_ean(product_full_url) # Append to Excel sheet sheet.append([product_name, product_full_url, ean]) print(f"Added: {product_name} | EAN: {ean}") # Add delay to avoid being blocked time.sleep(random.uniform(1.5, 3)) page += 1 except Exception as e: print(f"Error scraping page {page}: {str(e)}") time.sleep(random.uniform(2, 4)) continue # Initialize Excel workbook wb = openpyxl.Workbook() # Scrape Beer & Spirits category sheet_spirits = wb.active sheet_spirits.title = "Beer & Spirits" sheet_spirits.append(["Product Name", "Product URL", "EAN"]) scrape_category("https://www.pricerunner.dk/cl/1424/OEl-Spiritus", sheet_spirits) # Scrape Wine category sheet_wine = wb.create_sheet(title="Wine") sheet_wine.append(["Product Name", "Product URL", "EAN"]) scrape_category("https://www.pricerunner.dk/cl/465/Vin", sheet_wine) # Save final Excel file wb.save("Pricerunner_Product_Data.xlsx") print("Scraping complete! Data saved to Pricerunner_Product_Data.xlsx")
5. Key Improvements & Notes
- Dynamic API Detection: Extracts the site's Next.js build ID automatically, so you don't have to hardcode changing API URLs.
- EAN Extraction: Checks two reliable sources for EAN codes, with a fallback to "N/A" if none is found.
- Pagination Handling: Automatically iterates through all pages until no more products are available.
- Anti-Blocking: Uses rotating User-Agents and random delays to avoid being flagged as a bot.
- Error Resilience: Try/except blocks handle network errors gracefully, so a single failure doesn't stop the entire scrape.
Important Reminders
- Check the site's
robots.txt(https://www.pricerunner.dk/robots.txt) to ensure scraping is allowed for your use case. - If you get blocked, increase the delay between requests or add more User-Agents to the list.
- The site's API structure may change over time—you'll need to adjust the JSON parsing logic if the response format shifts.
内容的提问来源于stack exchange,提问作者Rick Prins

