使用BeautifulSoup提取kununu大众页面全部Pro信息的求助
Great question! The issue you’re hitting is super common with modern websites—Kununu loads comments dynamically instead of serving everything in the initial HTML. Your current requests + BeautifulSoup setup only grabs the static initial page (which only includes the top 10 reviews), so it can’t access the hidden comments that load when you scroll or click "Mehr anzeigen" (Load More).
Here’s how to fix this using Selenium, which simulates a real browser to load all content:
Step 1: Set Up Dependencies
First, install Selenium and grab the browser driver matching your browser version (e.g., ChromeDriver for Google Chrome):
pip install selenium
Step 2: Full Code to Scrape All "Pro" Comments
This code will launch a browser, load every comment by clicking the "Load More" button repeatedly, then parse the complete page:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time url = 'https://www.kununu.com/de/volkswagen/kommentare' # Initialize Chrome driver (adjust the path to your driver if needed) driver = webdriver.Chrome() driver.get(url) # Wait for the initial comment section to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "col-xs-12.col-lg-12")) ) # Keep clicking "Mehr anzeigen" until no more comments are left to load while True: try: # Wait for the load more button to be clickable, then click it load_more_btn = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Mehr anzeigen')]")) ) load_more_btn.click() time.sleep(2) # Give time for new comments to load except: # No more load button means all comments are loaded break # Grab the full page source after all content is loaded page_source = driver.page_source driver.quit() # Parse the complete page with BeautifulSoup soup = BeautifulSoup(page_source, 'html.parser') # Extract and print all "Pro" sections for h2 in soup.find_all('h2', class_='h3', text=['Pro']): pro_content = h2.find_next_sibling('p').get_text(strip=True) print(pro_content)
Quick Tips:
- Headless Mode: If you don’t want a visible browser window, add these options when initializing the driver:
options = webdriver.ChromeOptions() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options) - Wait Times: Adjust the
time.sleep(2)value if your internet is slower—this gives the page time to load new comments after clicking the button.
内容的提问来源于stack exchange,提问作者codecodecode

