You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3.6抓取JQuery渲染内容?SAP站点爬取遇空列表求助

Why Your Scraping Code Returns Empty Results

The issue here is that the SAP Partner Finder page loads its search results dynamically using JavaScript. When you use requests.get(), you’re only fetching the initial static HTML of the page—those search result headers you’re targeting aren’t present in that initial response. They get added to the page’s DOM after the browser executes the site’s JavaScript, which requests and BeautifulSoup can’t handle on their own.

BeautifulSoup only parses the raw HTML received from the request, so it can’t access content that loads client-side after the initial page load.

Solution: Use a Headless Browser to Render JavaScript

To scrape dynamically loaded content, you need a tool that mimics a real browser and executes JavaScript. Selenium is a popular, straightforward choice for this. Here’s how to adjust your code:

Step 1: Install Required Tools

First, install Selenium and set up a browser driver (we’ll use Chrome for this example):

pip install selenium

Newer versions of Selenium can automatically manage the ChromeDriver for you, but if you run into issues, download the driver matching your Chrome version and ensure it’s in your system path.

Step 2: Updated Scraping Code

This code uses Selenium to load the page, wait for the target elements to appear, then extract the content:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize Chrome browser (you can use Firefox or Edge too)
driver = webdriver.Chrome()

try:
    # Navigate to the partner finder page
    driver.get('https://partneredge.sap.com/content/partnerfinder/search.html#/')
    
    # Wait up to 10 seconds for the search result headers to load
    wait = WebDriverWait(driver, 10)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'header.search-result__head')))
    
    # Get the fully rendered page source (including JS-loaded content)
    page_source = driver.page_source
    page_soup = BeautifulSoup(page_source, "lxml")
    
    # Find and print the target elements
    trazenje = page_soup.find_all('header', class_='search-result__head')
    for result in trazenje:
        print(result.get_text(strip=True))
        
finally:
    # Ensure the browser closes even if an error occurs
    driver.quit()

Key Tips:

  • The WebDriverWait is critical—it prevents your code from trying to parse the page before the dynamic content loads. Adjust the timeout (10 seconds here) if the page takes longer to load.
  • Using class_ in BeautifulSoup is more explicit than combining the tag and class in a single string selector.
  • Always wrap the driver in a try/finally block to avoid leaving browser windows open accidentally.

Alternative: Playwright (Modern Alternative)

If you prefer a more streamlined tool, Playwright is a great choice that handles headless browsing and dynamic content with minimal setup:

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto('https://partneredge.sap.com/content/partnerfinder/search.html#/')
    # Wait for the target elements to exist
    page.wait_for_selector('header.search-result__head')
    
    # Extract the fully rendered page content
    page_source = page.content()
    page_soup = BeautifulSoup(page_source, "lxml")
    
    # Process and print results
    trazenje = page_soup.find_all('header', class_='search-result__head')
    for result in trazenje:
        print(result.get_text(strip=True))
        
    browser.close()

To use Playwright, install it first:

pip install playwright
playwright install

内容的提问来源于stack exchange,提问作者mr. M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:18:40