使用BeautifulSoup开发德国黄页爬虫:如何处理无URL变化的Load More按钮?
Hey there! I totally get where you're coming from—dynamic content that loads via button clicks without changing the URL is a common roadblock when you're new to scraping. Since requests only fetches the initial static HTML, we need to use Selenium to mimic real browser behavior and trigger those "Load More" clicks. Let's walk through how to adapt your code to do this.
First, Let's Get Set Up
First, you'll need to install Selenium and a browser driver. For Chrome:
- Install Selenium via pip:
pip install selenium - Download the ChromeDriver that matches your Chrome browser version, and make sure it's in your system PATH (or specify its full path in the code later).
Modified Code with Selenium Integration
Here's how to update your script to handle the "Load More" button, with a customizable number of clicks:
import pandas as pd from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import ElementClickInterceptedException, NoSuchElementException main_list = [] # Customize how many times you want to click "Load More" LOAD_MORE_CLICKS = 5 # Change this number to get more results def extract_with_selenium(url): # Optional: Add a custom user-agent to avoid being blocked from selenium.webdriver.chrome.options import Options options = Options() options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # Initialize Chrome browser (swap with webdriver.Firefox() if using Firefox) driver = webdriver.Chrome(options=options) driver.get(url) driver.implicitly_wait(10) # Wait up to 10 seconds for elements to load initially for _ in range(LOAD_MORE_CLICKS): try: # Wait until the "Load More" button is clickable load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Load More')]")) ) # Scroll to the button to ensure it's visible (avoids click errors) driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", load_more_btn) load_more_btn.click() # Wait for new content to load (adjust time if your page is slower) driver.implicitly_wait(5) except (ElementClickInterceptedException, NoSuchElementException): # If the button is gone or unclickable, stop the loop early print("No more 'Load More' button available. Stopping clicks.") break # Grab the full page source after all clicks are done soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # Close the browser to free up resources return soup.find_all('article', class_='mod mod-Treffer') def transform(articles): for item in articles: name = item.find('h2', {'data-wipe-name': 'Titel'}).text.strip() tel = item.find('p', class_='mod-AdresseKompakt__phoneNumber').text.strip() try: website = item.find('a', class_='contains-icon-homepage gc-btn gc-btn--s')['href'] except: website = '' business = { 'Name': name, 'Website': website, 'Telefonnummer': tel, } main_list.append(business) return def load(): df = pd.DataFrame(main_list) df.to_csv('stb_fürth.csv', index=False) # Replace with your actual German Yellow Pages URL articles = extract_with_selenium('YOUR_GERMAN_YELLOW_PAGES_URL_HERE') transform(articles) load() print('CSV Datei erstellt')
Key Changes Explained
- Selenium Browser Setup: We launch a real Chrome browser (with a custom user-agent to avoid blocking) that can execute JavaScript and interact with dynamic elements.
- Custom Click Control: The
LOAD_MORE_CLICKSvariable lets you set exactly how many times you want to trigger the button—tweak this number to get the amount of results you need. - Smart Waiting: We use
WebDriverWaitto ensure the button is fully loaded and clickable before trying to interact with it, which prevents random errors. - Error Handling: If the button disappears (no more results to load), the loop stops gracefully instead of crashing your script.
- Cleanup: We close the browser once we're done scraping to avoid leaving unnecessary processes running.
Quick Tips
- Double-check that your ChromeDriver version matches your installed Chrome browser—this is a super common source of bugs.
- Adjust the
implicitly_waittimes if your page loads slower (increase the number) or faster (decrease it) than expected. - If the "Load More" button has a different text (like in German), update the XPath to match—e.g.,
//button[contains(text(), 'Mehr laden')].
That should do it! This script will now click the "Load More" button as many times as you specify, scrape all loaded results, and save everything to your CSV file.
内容的提问来源于stack exchange,提问作者newprogrammingdude

