You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Requests无法获取完整HTML,BeautifulSoup无法定位Woolworths商品元素的爬取方案咨询

Troubleshooting Woolworths Product Scraping with BeautifulSoup

Hey Jacob, I’ve run into this exact issue with dynamic e-commerce sites like Woolworths before—let’s walk through why your current setup isn’t working and what you can do instead, without leaning on Selenium if possible.

Why Your Current Code Fails

Woolworths loads most of its product content dynamically via JavaScript after the initial page loads. When you use requests.get(), you’re only fetching the static "shell" of the HTML page—none of the product tiles or pricing elements exist in that initial response. Tools like cloudscraper help bypass anti-bot measures like Cloudflare, but they don’t execute JavaScript, so you still end up with an incomplete HTML structure. Switching parsers (lxml vs html.parser) won’t fix this because the elements simply aren’t present in the raw response.

Solutions to Get the Data You Need

1. Find and Call the Backend API (Best Option)

Nearly all dynamic e-commerce sites pull product data from a REST API instead of embedding it directly in the HTML. This is the most reliable and efficient way to get the data:

  • Open your browser’s DevTools (F12) and go to the Network tab.
  • Refresh the Woolworths search page, then filter requests by XHR/Fetch.
  • Look for requests with URLs containing keywords like products or search—you’ll likely find a JSON response that includes all product details (names, prices, weights, etc.).
  • Instead of scraping HTML, send a GET request directly to this API endpoint. You can copy the headers (including any auth tokens or cookies) from the DevTools request to mimic a browser request.

Example snippet for API scraping (adjust the endpoint to match what you find):

import requests

api_url = "YOUR_FOUND_API_ENDPOINT_HERE"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "application/json",
    # Add any other headers copied from DevTools here
}

response = requests.get(api_url, headers=headers)
product_data = response.json()

# Extract the fields you need
for item in product_data["products"]:
    print(item["name"], item["price"], item["pricePerUnit"])

2. Use a JS-Rendering Library with BeautifulSoup

If you can’t find the API or need to stick with HTML parsing, use a library that can execute JavaScript to render the full page before parsing. requests-html is a great lightweight alternative to Selenium:

  • Install it first: pip install requests-html
  • Then modify your code to render the page:
from requests_html import HTMLSession
from bs4 import BeautifulSoup

session = HTMLSession()
item_url = "https://www.woolworths.com.au/shop/search/products?searchTerm=uncle%20tobys%20oats%20500g&sortBy=TraderRelevance"

# Fetch and render the page (executes JS to load dynamic content)
response = session.get(item_url)
response.html.render()  # This waits for JS to finish loading elements

# Now parse the fully rendered HTML with BeautifulSoup
soup = BeautifulSoup(response.html.html, 'lxml')

product = soup.find_all('a', class_='shelfProductTile-descriptionLink')
price_per_weight = soup.find_all('div', class_='shelfProductTile-cupPrice ng-star-inserted')

print(product)
print(price_per_weight)

3. Scrapy with JS-Rendering (For Large-Scale Scraping)

If you’re planning to scrape large volumes of data, Scrapy is a robust choice. To handle dynamic content, you can integrate scrapy-playwright (a lightweight headless browser tool) to render pages before parsing. This avoids Selenium’s overhead while still handling JavaScript-driven content.

Final Notes

  • Avoid Selenium if possible: It’s slower and more likely to trigger anti-bot detection. APIs or JS-rendering libraries like requests-html are better for most use cases.
  • Check Woolworths’ Terms of Service: Ensure your scraping activity complies with their rules to avoid getting blocked.

内容的提问来源于stack exchange,提问作者Jacob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 23:42:37