如何用Python 3.6抓取JQuery渲染内容?SAP站点爬取遇空列表求助
The issue here is that the SAP Partner Finder page loads its search results dynamically using JavaScript. When you use requests.get(), you’re only fetching the initial static HTML of the page—those search result headers you’re targeting aren’t present in that initial response. They get added to the page’s DOM after the browser executes the site’s JavaScript, which requests and BeautifulSoup can’t handle on their own.
BeautifulSoup only parses the raw HTML received from the request, so it can’t access content that loads client-side after the initial page load.
Solution: Use a Headless Browser to Render JavaScript
To scrape dynamically loaded content, you need a tool that mimics a real browser and executes JavaScript. Selenium is a popular, straightforward choice for this. Here’s how to adjust your code:
Step 1: Install Required Tools
First, install Selenium and set up a browser driver (we’ll use Chrome for this example):
pip install selenium
Newer versions of Selenium can automatically manage the ChromeDriver for you, but if you run into issues, download the driver matching your Chrome version and ensure it’s in your system path.
Step 2: Updated Scraping Code
This code uses Selenium to load the page, wait for the target elements to appear, then extract the content:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize Chrome browser (you can use Firefox or Edge too) driver = webdriver.Chrome() try: # Navigate to the partner finder page driver.get('https://partneredge.sap.com/content/partnerfinder/search.html#/') # Wait up to 10 seconds for the search result headers to load wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'header.search-result__head'))) # Get the fully rendered page source (including JS-loaded content) page_source = driver.page_source page_soup = BeautifulSoup(page_source, "lxml") # Find and print the target elements trazenje = page_soup.find_all('header', class_='search-result__head') for result in trazenje: print(result.get_text(strip=True)) finally: # Ensure the browser closes even if an error occurs driver.quit()
Key Tips:
- The
WebDriverWaitis critical—it prevents your code from trying to parse the page before the dynamic content loads. Adjust the timeout (10 seconds here) if the page takes longer to load. - Using
class_in BeautifulSoup is more explicit than combining the tag and class in a single string selector. - Always wrap the driver in a
try/finallyblock to avoid leaving browser windows open accidentally.
Alternative: Playwright (Modern Alternative)
If you prefer a more streamlined tool, Playwright is a great choice that handles headless browsing and dynamic content with minimal setup:
from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup with sync_playwright() as p: browser = p.chromium.launch() page = browser.new_page() page.goto('https://partneredge.sap.com/content/partnerfinder/search.html#/') # Wait for the target elements to exist page.wait_for_selector('header.search-result__head') # Extract the fully rendered page content page_source = page.content() page_soup = BeautifulSoup(page_source, "lxml") # Process and print results trazenje = page_soup.find_all('header', class_='search-result__head') for result in trazenje: print(result.get_text(strip=True)) browser.close()
To use Playwright, install it first:
pip install playwright playwright install
内容的提问来源于stack exchange,提问作者mr. M

