You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium与Beautiful Soup爬取论坛评论数量失败求助

Solution to LIHKG Comment Number Scraping Issue

Common Causes of the Problem

  • Dynamic Content Loading: The forum loads comments asynchronously, so a 2-second sleep might not be enough for all elements to render.
  • Incorrect Element Selection: The parent div selector (_36ZEkSvpdj_igmog0nluzh) might target non-comment elements, leading to None when searching for the span inside them.
  • Lazy Loading: Comments below the viewport may not load until you scroll down the page.

Fixes with Code Examples

1. Use WebDriverWait Instead of Fixed Sleep

Replace static sleep(2) with explicit waits to ensure elements are fully loaded before parsing:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

URL = "https://lihkg.com/thread/3717611/page/1"
num_list = []

driver = webdriver.Chrome()  # Or your preferred browser driver
driver.get(URL)

# Wait up to 10 seconds for comment number spans to appear
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0")))

# Parse the fully loaded page source
html = BeautifulSoup(driver.page_source, 'html.parser')
span_list = html.find_all("span", {"class": "_3SqN3KZ8m8vCsD9FNcxcki _208tAU6LsyjP5LKTdcPXD0"})

# Extract and clean comment numbers
for span in span_list:
    num = span.text.strip().lstrip('#')
    num_list.append(num)

print(num_list)
driver.quit()

2. Handle Lazy Loading by Scrolling

For threads with comments loaded on scroll, add code to scroll to the bottom and wait for new content:

# After initial page load, scroll to bottom to trigger lazy loading
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
# Wait for new comments to load
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0")))

# Repeat scroll for very long threads (uncomment if needed)
# for _ in range(3):
#     driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
#     wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span._3SqN3KZ8m8vCsD9FNcxcki._208tAU6LsyjP5LKTdcPXD0")))

3. Use Text Pattern Matching (Robust to Class Changes)

If span class names are dynamic, target spans starting with # instead:

# Replace the span_list line with this
span_list = html.find_all("span", text=lambda text: text and text.startswith('#'))

for span in span_list:
    num = span.text.strip().lstrip('#')
    num_list.append(num)

Key Notes

  • Explicit waits (WebDriverWait) are far more reliable than fixed sleeps for dynamic content.
  • Directly targeting comment number spans avoids irrelevant parent divs causing None values.
  • For long threads, implement scrolling to ensure all comments are loaded before parsing.

内容的提问来源于stack exchange,提问作者pak hei wong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 13:02:35