使用Selenium与Beautiful Soup爬取论坛评论数量失败求助
Solution to LIHKG Comment Number Scraping Issue
Common Causes of the Problem
- Dynamic Content Loading: The forum loads comments asynchronously, so a 2-second sleep might not be enough for all elements to render.
- Incorrect Element Selection: The parent div selector (
_36ZEkSvpdj_igmog0nluzh) might target non-comment elements, leading toNonewhen searching for the span inside them. - Lazy Loading: Comments below the viewport may not load until you scroll down the page.
Fixes with Code Examples
1. Use WebDriverWait Instead of Fixed Sleep
Replace static sleep(2) with explicit waits to ensure elements are fully loaded before parsing:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup URL = "https://lihkg.com/thread/3717611/page/1" num_list = [] driver = webdriver.Chrome() # Or your preferred browser driver driver.get(URL) # Wait up to 10 seconds for comment number spans to appear wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0"))) # Parse the fully loaded page source html = BeautifulSoup(driver.page_source, 'html.parser') span_list = html.find_all("span", {"class": "_3SqN3KZ8m8vCsD9FNcxcki _208tAU6LsyjP5LKTdcPXD0"}) # Extract and clean comment numbers for span in span_list: num = span.text.strip().lstrip('#') num_list.append(num) print(num_list) driver.quit()
2. Handle Lazy Loading by Scrolling
For threads with comments loaded on scroll, add code to scroll to the bottom and wait for new content:
# After initial page load, scroll to bottom to trigger lazy loading driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new comments to load wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0"))) # Repeat scroll for very long threads (uncomment if needed) # for _ in range(3): # driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0")))
3. Use Text Pattern Matching (Robust to Class Changes)
If span class names are dynamic, target spans starting with # instead:
# Replace the span_list line with this span_list = html.find_all("span", text=lambda text: text and text.startswith('#')) for span in span_list: num = span.text.strip().lstrip('#') num_list.append(num)
Key Notes
- Explicit waits (
WebDriverWait) are far more reliable than fixed sleeps for dynamic content. - Directly targeting comment number spans avoids irrelevant parent divs causing
Nonevalues. - For long threads, implement scrolling to ensure all comments are loaded before parsing.
内容的提问来源于stack exchange,提问作者pak hei wong
相关产品推荐
相关产品推荐

