You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium解析网页XML并提取所有链接?

使用Selenium解析网页XML并提取标签链接

实现思路

浏览器会自动把XML文档渲染成可操作的DOM节点,直接用Selenium的元素定位方法抓取<loc>标签即可,无需额外解析XML字符串。

Python代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化Chrome驱动(需确保chromedriver已配置到环境变量或指定路径)
driver = webdriver.Chrome()
target_url = "https://www.thetutorsdirectory.com/usa/sitemap/sitemap_l1.xml"

try:
    # 加载目标XML页面
    driver.get(target_url)
    
    # 等待页面加载完成,确保<loc>标签可被定位
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "loc"))
    )
    
    # 获取所有<loc>标签元素
    loc_tags = driver.find_elements(By.TAG_NAME, "loc")
    
    # 提取并输出每个标签内的链接
    for tag in loc_tags:
        print(tag.text)
        
finally:
    # 关闭浏览器
    driver.quit()

常见问题排查

  • 加载超时:检查网络状态,或者把WebDriverWait的超时时间调整到15秒以上。
  • 定位不到元素:打开浏览器开发者工具(F12)确认页面DOM里确实存在<loc>标签,排除页面加载异常。
  • 浏览器兼容:如果用Firefox等其他浏览器,替换对应的驱动(比如Firefox用geckodriver)。

内容的提问来源于stack exchange,提问作者Andrej Stomnaroski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:05:17