为何BeautifulSoup解析的HTML Soup中无表格数据标签?
Hey there! I totally get the frustration when you think you've got your soup ready but the data you need is nowhere to be found. Let's break down why this is happening and how to fix it.
Why Your Current Code Isn't Working
The Grainger website loads its product table dynamically using JavaScript. When you use requests.get(), you're only fetching the initial static HTML that gets sent to the browser—this doesn't include the product rows, because those are loaded later via API calls once the page renders in a browser. That's why your soup doesn't have any <tr> elements with the class search-table-view__web-parent-table-row.
Solution 1: Use Selenium to Render JavaScript
Selenium simulates a real browser, so it waits for the JavaScript to load all the content before you grab the HTML. Here's how to adjust your code:
First, install Selenium and make sure you have the ChromeDriver (or your browser's equivalent driver) installed:
pip install selenium
Then update your scraper code:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup main_url = 'https://www.grainger.com/category/pipe-hose-tube-fittings/hose-products/hose-fittings-couplings/cam-groove-fittings-gaskets/metal-cam-groove-fittings/stainless-steel-cam-groove-fittings' # Set up Chrome options (headless mode runs without a visible window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.97 Safari/537.36") # Initialize the browser driver driver = webdriver.Chrome(options=chrome_options) driver.get(main_url) # Wait for the page to fully load (implicit wait for 10 seconds) driver.implicitly_wait(10) # Get the fully rendered HTML page_source = driver.page_source soup = BeautifulSoup(page_source, "lxml") # Now you should be able to find the table rows rows = soup.find_all("tr", {"class":"search-table-view__web-parent-table-row"}) # Verify the results print(f"Found {len(rows)} product rows!") # Don't forget to close the browser when done driver.quit()
Solution 2: Directly Call the Website's API (More Efficient)
Instead of simulating a browser, you can find the API endpoint that Grainger uses to load product data. Here's how to do that:
- Open your browser's DevTools (F12) and go to the Network tab.
- Refresh the page and look for XHR/Fetch requests (filter by "XHR" to narrow it down).
- Look for requests that return JSON data with product details—you'll likely see a URL like
/api/search/productsor similar. - Copy the request headers and parameters, then use
requeststo call this API directly. This is faster than using Selenium because you're skipping the browser rendering step.
Important Notes
- Always check Grainger's
robots.txtand terms of service before scraping—make sure you're allowed to access their data programmatically. - Add delays between requests (using
time.sleep()) to avoid getting your IP blocked for too many rapid requests.
内容的提问来源于stack exchange,提问作者bknoll16

