You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup解析的HTML Soup中无表格数据标签?

Troubleshooting Missing Table Data in Your Web Scraper

Hey there! I totally get the frustration when you think you've got your soup ready but the data you need is nowhere to be found. Let's break down why this is happening and how to fix it.

Why Your Current Code Isn't Working

The Grainger website loads its product table dynamically using JavaScript. When you use requests.get(), you're only fetching the initial static HTML that gets sent to the browser—this doesn't include the product rows, because those are loaded later via API calls once the page renders in a browser. That's why your soup doesn't have any <tr> elements with the class search-table-view__web-parent-table-row.

Solution 1: Use Selenium to Render JavaScript

Selenium simulates a real browser, so it waits for the JavaScript to load all the content before you grab the HTML. Here's how to adjust your code:

First, install Selenium and make sure you have the ChromeDriver (or your browser's equivalent driver) installed:

pip install selenium

Then update your scraper code:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

main_url = 'https://www.grainger.com/category/pipe-hose-tube-fittings/hose-products/hose-fittings-couplings/cam-groove-fittings-gaskets/metal-cam-groove-fittings/stainless-steel-cam-groove-fittings'

# Set up Chrome options (headless mode runs without a visible window)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.97 Safari/537.36")

# Initialize the browser driver
driver = webdriver.Chrome(options=chrome_options)
driver.get(main_url)

# Wait for the page to fully load (implicit wait for 10 seconds)
driver.implicitly_wait(10)

# Get the fully rendered HTML
page_source = driver.page_source
soup = BeautifulSoup(page_source, "lxml")

# Now you should be able to find the table rows
rows = soup.find_all("tr", {"class":"search-table-view__web-parent-table-row"})

# Verify the results
print(f"Found {len(rows)} product rows!")

# Don't forget to close the browser when done
driver.quit()

Solution 2: Directly Call the Website's API (More Efficient)

Instead of simulating a browser, you can find the API endpoint that Grainger uses to load product data. Here's how to do that:

  • Open your browser's DevTools (F12) and go to the Network tab.
  • Refresh the page and look for XHR/Fetch requests (filter by "XHR" to narrow it down).
  • Look for requests that return JSON data with product details—you'll likely see a URL like /api/search/products or similar.
  • Copy the request headers and parameters, then use requests to call this API directly. This is faster than using Selenium because you're skipping the browser rendering step.

Important Notes

  • Always check Grainger's robots.txt and terms of service before scraping—make sure you're allowed to access their data programmatically.
  • Add delays between requests (using time.sleep()) to avoid getting your IP blocked for too many rapid requests.

内容的提问来源于stack exchange,提问作者bknoll16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 13:32:55