You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium WebDriver为何无法定位全部指定class的HTML元素?

Why Selenium Only Finds 2 cmt__title Elements Instead of 12?

I’ve run into this exact kind of discrepancy before—here’s what’s going on and how to fix it:

1. Dynamic Content is the #1 Culprit

When you view the page in your regular browser, you’re probably scrolling down, which triggers the site to load more comments (and their cmt__title elements). Your current Selenium code grabs elements immediately after the page loads, so it only catches the first 2 that load right away, not the 10 that load dynamically later.

2. Headless Mode Quirks

Headless Chrome has a tiny default window size, and some sites serve mobile-optimized content (with fewer initial elements) when they detect a small viewport. Plus, older headless modes can be flagged by sites, leading to simplified content being served.

Let’s Fix This

Option 1: Wait for All Elements to Load

Use explicit waits to tell Selenium to hold off until all 12 elements are present. This is way more reliable than grabbing elements immediately:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--ignore-certificate-errors')
options.add_argument('--incognito')
options.add_argument('--headless=new')  # Newer headless mode matches regular Chrome better
options.add_argument('--window-size=1920,1080')  # Force desktop viewport

driver = webdriver.Chrome("/home/matan/chrome-driver/chromedriver", options=options)
driver.get("https://www.themarker.com/law/1.7254050")

# Wait up to 10 seconds for at least 12 elements to show up
wait = WebDriverWait(driver, 10)
titles = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'cmt__title')))

print(f"Found {len(titles)} elements!")
for title in titles:
    print(title.text)

driver.quit()

Option 2: Simulate Scrolling to Load More

If comments load as you scroll, mimic that behavior in Selenium:

# After loading the page
# Scroll to the bottom a few times to trigger dynamic content loads
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    driver.implicitly_wait(2)  # Give the site time to load new content

# Now fetch all titles
titles = driver.find_elements(By.CLASS_NAME, 'cmt__title')
print(f"Found {len(titles)} elements")

Option 3: Make Headless Chrome Less Detectable

Some sites block or alter content for headless browsers. Add these arguments to make your headless instance look like a regular desktop browser:

options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
options.add_argument('--disable-blink-features=AutomationControlled')

Pro Tips

  • Explicit waits are always better than implicit waits for dynamic content—they target specific elements instead of waiting arbitrarily.
  • Check your browser’s DevTools Network tab: if comments load via API calls, you could skip scraping the page entirely and call those APIs directly. It’s faster and more stable.
  • Use the newer --headless=new flag (available in Chrome 112+)—it behaves almost identically to regular Chrome, avoiding many headless-related issues.

内容的提问来源于stack exchange,提问作者CrazySynthax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:51:00