You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium采集网页链接遇加载更多按钮失效问题

解决Selenium无法触发「加载更多新闻」按钮的问题

原脚本存在的核心问题

  • 页面加载新内容后,之前获取的load_more_button会变成陈旧元素,无法再次点击
  • 仅判断按钮是否显示,未等待按钮处于可点击状态,可能按钮已显示但还未绑定点击事件
  • 固定sleep(10)不够灵活,网络慢时可能加载未完成,网络快时浪费时间
  • 循环逻辑中,若新加载的链接未及时渲染,会误判为没有新链接而提前跳出循环

修改后的脚本及说明

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import StaleElementReferenceException, ElementClickInterceptedException, TimeoutException
import json

options = webdriver.ChromeOptions()
# 可按需添加配置,比如无头模式
# options.add_argument('--headless=new')
# options.add_argument('--disable-gpu')

driver = webdriver.Chrome(options=options)

url = "..."
base_url = "..."

driver.get(url)
outlinks = []
wait = WebDriverWait(driver, 90)

while True:
    # 获取当前页面所有目标链接
    current_links = driver.find_elements(By.CSS_SELECTOR, 'a.text-truncate')
    current_link_count = len(current_links)
    
    # 滚动至页面底部,确保加载更多按钮进入视野
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    
    try:
        # 等待加载更多按钮可点击,每次循环重新获取最新元素
        load_more_button = wait.until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.listing-load-more-btn[title="Load More News"]'))
        )
        # 用JS点击规避元素遮挡问题
        driver.execute_script("arguments[0].click();", load_more_button)
        
        # 等待新内容加载完成:直到链接数量增加,最多等待30秒
        wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, 'a.text-truncate')) > current_link_count)
        
    except (StaleElementReferenceException, ElementClickInterceptedException):
        # 元素陈旧或被遮挡,重新尝试
        continue
    except TimeoutException:
        # 超时说明无更多内容,退出循环
        break

# 收集最终所有有效链接
final_links = driver.find_elements(By.CSS_SELECTOR, 'a.text-truncate')
for link in final_links:
    href = link.get_attribute('href')
    if href:  # 过滤空链接
        # 处理相对/绝对链接
        full_url = base_url + href if not href.startswith('http') else href
        en_url = full_url.replace("ar-ae", "en")
        outlinks.append(en_url)

# 保存结果到JSON文件
with open('outlinks.json', 'w', encoding='utf-8') as f:
    json.dump(outlinks, f, ensure_ascii=False, indent=2)

print(f"共采集到 {len(outlinks)} 条链接")
driver.quit()

关键优化点

  • 每次循环重新获取加载更多按钮,彻底避免陈旧元素异常
  • 使用element_to_be_clickable确保按钮真正具备交互能力
  • 用JS点击方式解决元素被遮挡无法点击的常见问题
  • 以链接数量变化作为加载完成的判断依据,替代固定sleep,稳定性更强
  • 增加异常捕获,处理元素陈旧、点击拦截等场景
  • 增加空链接过滤,避免无效数据

内容的提问来源于stack exchange,提问作者kutsalzamazingo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 10:20:47