You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium抓取半岛电视台文章遇阻:DevTools可定位XPath无法获内容

解决Al Jazeera新闻段落抓取的Selenium XPath问题

以下是针对你遇到的DevTools中XPath有效但Selenium抓取失败问题的具体解决方案:

核心原因及对应解决办法

1. 内容动态加载延迟

Selenium执行定位操作时,页面可能尚未完成渲染,导致元素还未加载。必须使用显式等待确保元素可见后再抓取:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 等待目标段落可见,超时时间设为15秒
wait = WebDriverWait(driver, 15)
first_paragraph = wait.until(EC.visibility_of_element_located((By.XPATH, '//main[@id="main-content-area"]//div[@class="wysiwyg wysiwyg--all-content css-1ck9wyi"]/p[1]')))
print(first_paragraph.text)

2. 避免依赖不稳定的索引定位

你原XPath中的div[2]属于索引定位,页面结构可能因加载状态或布局变化失效。改用更稳定的class属性定位文章内容容器:

  • 替换原XPath为://div[@class="wysiwyg wysiwyg--all-content css-1ck9wyi"]//p
  • 这个容器是Al Jazeera文章内容的固定父节点,直接定位其下的所有<p>标签即可覆盖所有段落。

3. 模拟真实浏览器环境

Al Jazeera可能存在反爬机制,需设置正确的User-Agent模拟真实访问:

from selenium import webdriver

options = webdriver.ChromeOptions()
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
driver = webdriver.Chrome(options=options)

4. 检查iframe嵌套

若目标元素被嵌套在iframe中,需先切换到iframe再定位:

# 按iframe的id或name切换,若没有则用索引
driver.switch_to.frame("iframe_id")
# 执行定位操作
# 操作完成后切回主文档
driver.switch_to.default_content()

完整抓取示例代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化浏览器
options = webdriver.ChromeOptions()
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
driver = webdriver.Chrome(options=options)

# 访问目标新闻页
url = "https://www.aljazeera.com/economy/2023/2/6/who-is-gautam-adani-and-why-is-he-controversial"
driver.get(url)

try:
    # 等待文章内容容器加载完成
    wait = WebDriverWait(driver, 15)
    content_container = wait.until(EC.presence_of_element_located((By.XPATH, '//div[@class="wysiwyg wysiwyg--all-content css-1ck9wyi"]')))
    
    # 抓取所有段落文本
    paragraphs = content_container.find_elements(By.TAG_NAME, 'p')
    for idx, p in enumerate(paragraphs, 1):
        text = p.text.strip()
        if text:
            print(f"段落{idx}: {text}")
finally:
    driver.quit()

内容的提问来源于stack exchange,提问作者turtle rough

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:20:33