使用Python从网站获取链接(入门):BeautifulSoup4无法抓取产品网格链接问题排查
Hey there! Let's break down why your productlist is coming up empty and get your web scraping working properly.
最可能的原因:动态内容加载
Modern websites often use JavaScript to load content after the initial page HTML is downloaded. When you use requests.get(), you're only fetching the raw, unrendered HTML sent by the server—and the product grid you're targeting is probably loaded later via AJAX calls. That's why your soup can't find the products product-grid div.
解决方案1:用Selenium模拟浏览器渲染
Selenium acts like a real browser, so it waits for JavaScript to load all content before you scrape. Here's how to adjust your code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd baseurl = "https://www.xl-byg.dk/" url = 'https://www.xl-byg.dk/produkter/trae-byggematerialer/gulve/laminatgulve' # Initialize Chrome browser (make sure you have chromedriver matching your Chrome version) driver = webdriver.Chrome() driver.get(url) # Wait up to 10 seconds for the product grid to load try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".products.product-grid")) ) finally: # Grab the fully rendered page source soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() productlist = soup.find_all("div", class_="products product-grid") print(f"Found {len(productlist)} product grid blocks") productlinks = [] for product in productlist: # Find all product links in the grid links = product.find_all("a", class_="reset-style") for link in links: href = link.get('href') if href: productlinks.append(baseurl + href) print(f"Scraped {len(productlinks)} product links total")
解决方案2:直接调用网站的API(更高效)
Instead of scraping HTML, you can often find the API that the website uses to load product data:
- Open your browser's DevTools (F12) and go to the Network tab
- Refresh the page and look for XHR/fetch requests (filter by "XHR")
- Look for requests that return product data (usually JSON format)
- You can then use
requests.get()to call this API directly and parse the JSON, which is faster and more reliable than HTML scraping.
快速调试技巧
- First, print
soup.prettify()ork.textto check if theproducts product-griddiv exists in the raw HTML. If it's not there, dynamic loading is definitely the issue. - Double-check your selector: using
class_="products product-grid"is equivalent to your original{"class": "..."}syntax, so that's not the problem here—but it's good to verify the class names haven't changed on the website.
内容的提问来源于stack exchange,提问作者awi1100

