You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python requests爬取超市营养数据?解决页面HTML不一致问题

Fixing Mismatched HTML When Scraping Mysupermarket with Requests

Hey there, I’ve dealt with this exact issue dozens of times—let’s break down why it’s happening and how to fix it:

The core problem is that the page uses JavaScript to dynamically load content (like those nutrition facts) after the initial raw HTML is sent to your browser. The requests library only fetches that initial HTML, not the fully rendered version you see in Chrome DevTools. Browsers automatically execute the page’s JS to populate dynamic data like nutrition info, but requests doesn’t do that.

Let’s walk through two solutions, starting with the simplest:

1. First, Try Adding Proper Request Headers

Many sites block or serve stripped-down content to non-browser clients. Spoofing a real browser’s request headers can often trick the server into sending the same content you see when you visit the page manually.

Here’s how to adjust your code:

import requests
from bs4 import BeautifulSoup

url = "http://www.mysupermarket.co.uk/tesco-price-comparison/Fruit/Tesco_Gala_Apple_Approx_160g.html"

# Mimic a real Chrome browser's request headers
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
    "Accept-Encoding": "gzip, deflate, br"
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# Try targeting the nutrition section (adjust the selector to match what you see in DevTools)
nutrition_section = soup.find("div", class_="nutrition")
if nutrition_section:
    print(nutrition_section.get_text(strip=True, separator="\n"))
else:
    print("Still no luck? Move on to the next solution below.")

2. Use Browser Automation for Fully Rendered HTML

If headers don’t work (which is common for sites with heavy dynamic content), you’ll need a tool that simulates a real browser. Selenium or Playwright are perfect for this—they launch a browser, execute all the page’s JavaScript, and let you grab the fully loaded page source.

Let’s go with Selenium (it’s widely used and straightforward to set up):

First, install the required packages and browser driver:

pip install selenium
# Download ChromeDriver matching your Chrome version, then add it to your system PATH

Then write the automation code:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "http://www.mysupermarket.co.uk/tesco-price-comparison/Fruit/Tesco_Gala_Apple_Approx_160g.html"

# Launch a Chrome browser instance
driver = webdriver.Chrome()
driver.get(url)

try:
    # Wait up to 10 seconds for the nutrition section to load (adjust the selector to match your target element)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "nutrition"))
    )
    
    # Grab the fully rendered page source
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    
    # Extract and print the nutrition data
    nutrition_data = soup.find("div", class_="nutrition").get_text(strip=True, separator="\n")
    print(nutrition_data)
finally:
    # Always close the browser to avoid hanging processes
    driver.quit()

Quick Pro Tips:

  • Use Chrome DevTools to copy the exact selector for the nutrition element (right-click the element > Copy > Copy Selector) to ensure your code targets the right part of the page.
  • If you hit anti-bot blocks, try adding small delays or using undetected-chromedriver (a modified version of ChromeDriver that bypasses common anti-scraping tools).

内容的提问来源于stack exchange,提问作者user9409939

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:29:25