You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium与BeautifulSoup解析XML文件失败,求解决方法

问题

尝试使用Selenium结合BeautifulSoup解析Propublica非营利组织页面的XML文件,但无法提取<PhoneNum>元素,输出为None。

代码示例

import time
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

options = Options()
# options.add_argument('--headless=new')  
options.add_argument("start-maximized")
options.add_argument('--log-level=3')  
options.add_experimental_option("prefs", {"profile.default_content_setting_values.notifications": 1})    
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('excludeSwitches', ['enable-logging'])
options.add_experimental_option('useAutomationExtension', False)
options.add_argument('--disable-blink-features=AutomationControlled') 
srv=Service()
driver = webdriver.Chrome (service=srv, options=options)    
waitWD = WebDriverWait (driver, 10)  

wLink = "https://projects.propublica.org/nonprofits/organizations/830370609"
driver.get(wLink) 
driver.execute_script("arguments[0].click();", waitWD.until(EC.element_to_be_clickable((By.XPATH, '(//a[text()="XML"])[1]'))))  
driver.switch_to.window(driver.window_handles[1])    
time.sleep(3) 
print(driver.current_url)
soup = BeautifulSoup (driver.page_source, 'lxml')   
worker = soup.find("PhoneNum")
print(worker)

错误输出

(selenium) C:\DEV\Fiverr2025\TRY\austibn>python test.py
https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?response-content-disposition=inline&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA266MJEJYTM5WAG5Y%2F20250423%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20250423T152903Z&X-Amz-Expires=1800&X-Amz-SignedHeaders=host&X-Amz-Signature=9743a63b41a906fac65c397a2bba7208938ca5b865f1e5a33c4f711769c815a4
None

解决方法

核心原因

原代码失败有两个关键问题:

  1. 使用**HTML解析器(lxml)**处理XML,导致XML结构被错误解析;
  2. 目标XML包含命名空间(xmlns="http://www.irs.gov/efile"),直接通过标签名"PhoneNum"无法匹配到实际元素。

方案1:改进Selenium+BeautifulSoup代码

修改解析器为XML专用的lxml-xml,并处理命名空间:

# 替换原代码中BeautifulSoup初始化和查找部分
soup = BeautifulSoup(driver.page_source, 'lxml-xml')  # 用XML解析器
# 目标XML的命名空间为http://www.irs.gov/efile,需带命名空间查找
phone_num = soup.find("{http://www.irs.gov/efile}PhoneNum")
print(phone_num.text if phone_num else "未找到PhoneNum元素")

另外,建议用WebDriverWait替代time.sleep(3),等待页面加载完成:

# 替换time.sleep(3)
waitWD.until(lambda d: '<?xml' in d.page_source)  # 等待XML内容加载

方案2:直接请求XML(更高效,无需Selenium)

既然已经能获取到XML的直接URL,跳过Selenium直接用requests请求,节省资源:

import requests
from bs4 import BeautifulSoup

xml_url = "https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?response-content-disposition=inline&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA266MJEJYTM5WAG5Y%2F20250423%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20250423T152903Z&X-Amz-Expires=1800&X-Amz-SignedHeaders=host&X-Amz-Signature=9743a63b41a906fac65c397a2bba7208938ca5b865f1e5a33c4f711769c815a4"

response = requests.get(xml_url)
soup = BeautifulSoup(response.content, 'lxml-xml')
phone_num = soup.find("{http://www.irs.gov/efile}PhoneNum")
print(phone_num.text if phone_num else "未找到PhoneNum元素")

如果需要批量处理,也可以从原页面提取XML链接,不用点击:

import requests
from bs4 import BeautifulSoup

base_url = "https://projects.propublica.org/nonprofits/organizations/830370609"
response = requests.get(base_url)
soup = BeautifulSoup(response.text, 'lxml')
# 提取第一个XML链接
xml_link = soup.find('a', text='XML')['href']
# 请求XML并解析
xml_response = requests.get(xml_link)
xml_soup = BeautifulSoup(xml_response.content, 'lxml-xml')
phone_num = xml_soup.find("{http://www.irs.gov/efile}PhoneNum")
print(phone_num.text if phone_num else "未找到PhoneNum元素")

方案3:用原生XML解析库(更专业)

对于XML解析,Python原生的xml.etree.ElementTree更适合:

import requests
import xml.etree.ElementTree as ET

xml_url = "https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?..."
response = requests.get(xml_url)
root = ET.fromstring(response.content)
# 注册命名空间前缀
ns = {'irs': 'http://www.irs.gov/efile'}
# 用XPath查找元素
phone_num = root.find('.//irs:PhoneNum', ns)
print(phone_num.text if phone_num else "未找到PhoneNum元素")

内容的提问来源于stack exchange,提问作者Rapid1898

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 07:42:32