使用Selenium与BeautifulSoup解析XML文件失败,求解决方法
问题
尝试使用Selenium结合BeautifulSoup解析Propublica非营利组织页面的XML文件,但无法提取<PhoneNum>元素,输出为None。
代码示例
import time from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By options = Options() # options.add_argument('--headless=new') options.add_argument("start-maximized") options.add_argument('--log-level=3') options.add_experimental_option("prefs", {"profile.default_content_setting_values.notifications": 1}) options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('excludeSwitches', ['enable-logging']) options.add_experimental_option('useAutomationExtension', False) options.add_argument('--disable-blink-features=AutomationControlled') srv=Service() driver = webdriver.Chrome (service=srv, options=options) waitWD = WebDriverWait (driver, 10) wLink = "https://projects.propublica.org/nonprofits/organizations/830370609" driver.get(wLink) driver.execute_script("arguments[0].click();", waitWD.until(EC.element_to_be_clickable((By.XPATH, '(//a[text()="XML"])[1]')))) driver.switch_to.window(driver.window_handles[1]) time.sleep(3) print(driver.current_url) soup = BeautifulSoup (driver.page_source, 'lxml') worker = soup.find("PhoneNum") print(worker)
错误输出
(selenium) C:\DEV\Fiverr2025\TRY\austibn>python test.py https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?response-content-disposition=inline&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA266MJEJYTM5WAG5Y%2F20250423%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20250423T152903Z&X-Amz-Expires=1800&X-Amz-SignedHeaders=host&X-Amz-Signature=9743a63b41a906fac65c397a2bba7208938ca5b865f1e5a33c4f711769c815a4 None
解决方法
核心原因
原代码失败有两个关键问题:
- 使用**HTML解析器(
lxml)**处理XML,导致XML结构被错误解析; - 目标XML包含命名空间(
xmlns="http://www.irs.gov/efile"),直接通过标签名"PhoneNum"无法匹配到实际元素。
方案1:改进Selenium+BeautifulSoup代码
修改解析器为XML专用的lxml-xml,并处理命名空间:
# 替换原代码中BeautifulSoup初始化和查找部分 soup = BeautifulSoup(driver.page_source, 'lxml-xml') # 用XML解析器 # 目标XML的命名空间为http://www.irs.gov/efile,需带命名空间查找 phone_num = soup.find("{http://www.irs.gov/efile}PhoneNum") print(phone_num.text if phone_num else "未找到PhoneNum元素")
另外,建议用WebDriverWait替代time.sleep(3),等待页面加载完成:
# 替换time.sleep(3) waitWD.until(lambda d: '<?xml' in d.page_source) # 等待XML内容加载
方案2:直接请求XML(更高效,无需Selenium)
既然已经能获取到XML的直接URL,跳过Selenium直接用requests请求,节省资源:
import requests from bs4 import BeautifulSoup xml_url = "https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?response-content-disposition=inline&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA266MJEJYTM5WAG5Y%2F20250423%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20250423T152903Z&X-Amz-Expires=1800&X-Amz-SignedHeaders=host&X-Amz-Signature=9743a63b41a906fac65c397a2bba7208938ca5b865f1e5a33c4f711769c815a4" response = requests.get(xml_url) soup = BeautifulSoup(response.content, 'lxml-xml') phone_num = soup.find("{http://www.irs.gov/efile}PhoneNum") print(phone_num.text if phone_num else "未找到PhoneNum元素")
如果需要批量处理,也可以从原页面提取XML链接,不用点击:
import requests from bs4 import BeautifulSoup base_url = "https://projects.propublica.org/nonprofits/organizations/830370609" response = requests.get(base_url) soup = BeautifulSoup(response.text, 'lxml') # 提取第一个XML链接 xml_link = soup.find('a', text='XML')['href'] # 请求XML并解析 xml_response = requests.get(xml_link) xml_soup = BeautifulSoup(xml_response.content, 'lxml-xml') phone_num = xml_soup.find("{http://www.irs.gov/efile}PhoneNum") print(phone_num.text if phone_num else "未找到PhoneNum元素")
方案3:用原生XML解析库(更专业)
对于XML解析,Python原生的xml.etree.ElementTree更适合:
import requests import xml.etree.ElementTree as ET xml_url = "https://pp-990-xml.s3.us-east-1.amazonaws.com/202403189349311780_public.xml?..." response = requests.get(xml_url) root = ET.fromstring(response.content) # 注册命名空间前缀 ns = {'irs': 'http://www.irs.gov/efile'} # 用XPath查找元素 phone_num = root.find('.//irs:PhoneNum', ns) print(phone_num.text if phone_num else "未找到PhoneNum元素")
内容的提问来源于stack exchange,提问作者Rapid1898
相关产品推荐
相关产品推荐

