You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取华尔街日报世界新闻失败,请求技术协助

解决WSJ世界板块爬取失效的问题

核心问题分析

  • 硬编码XPATH易失效:你使用的XPATH依赖页面固定ID和层级结构,WSJ页面可能动态更新,导致选择器无法定位元素。
  • 未处理隐私弹窗:WSJ加载后会弹出隐私政策同意框,未关闭会干扰元素定位。
  • 低效的滚动加载:固定滚动到指定像素的方式不稳定,无法确保内容完全加载。
  • 无有效错误排查:宽泛的异常捕获直接跳过错误,无法定位具体失败原因。
  • 缺少浏览器伪装:Selenium默认的浏览器标识可能被WSJ反爬机制识别。

修正后的代码

from selenium import webdriver
from selenium.webdriver.edge.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.edge.options import Options
import os

# 设置浏览器选项,模拟真实用户
edge_options = Options()
edge_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36 Edg/118.0.2088.69")
edge_options.add_argument("--disable-blink-features=AutomationControlled")  # 隐藏自动化标识

# 设置驱动路径
desktop_path = os.path.join(os.environ['USERPROFILE'], 'Desktop')
e_driver_path = os.path.join(desktop_path, "msedgedriver.exe")

def wsj_world():
    website_link = 'https://www.wsj.com/news/world'
    
    # 初始化浏览器
    s = Service(e_driver_path)
    driver = webdriver.Edge(service=s, options=edge_options)
    driver.maximize_window()
    driver.get(website_link)
    
    try:
        # 处理隐私政策弹窗(等待并点击同意)
        accept_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Accept All")]'))
        )
        accept_btn.click()
    except:
        # 若没有弹窗则继续
        pass
    
    # 滚动加载更多内容(滚动到底部,重复几次确保加载)
    for _ in range(3):
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'article[data-testid="article"]'))
        )
    
    # 使用通用选择器定位新闻元素,替代硬编码XPATH
    news_elements = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'article[data-testid="article"] a[data-testid="link"]'))
    )
    
    nt = []
    nu = []
    for elem in news_elements:
        # 获取标题文本
        title = elem.find_element(By.CSS_SELECTOR, 'span[data-testid="headline-text"]').text
        # 获取链接
        url = elem.get_attribute('href')
        if title and url:
            nt.append(title)
            nu.append(url)
    
    # 打印结果
    for title, url in zip(nt, nu):
        print(f"{title}\n{url}\n---")
    
    driver.quit()
    return nt, nu

wsj_world()

关键改进说明

  • 浏览器伪装:添加User-Agent并隐藏自动化标识,避免被反爬机制识别。
  • 处理隐私弹窗:显式等待并点击同意按钮,确保页面正常交互。
  • 通用选择器:使用data-testid属性定位元素,这类属性比固定ID更稳定,不易随页面更新失效。
  • 高效滚动加载:滚动到页面底部并等待新闻元素加载,确保内容完全加载。
  • 优化错误处理:仅针对弹窗做异常捕获,保留其他错误的排查可能性。

内容的提问来源于stack exchange,提问作者turtle rough

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 04:10:42