Python3解析XML遇xml.parsers.expat.ExpatError错误的解决方法
解决方法
目标网站启用了反爬机制,会先返回带有JS加密逻辑的HTML页面,要求浏览器执行JS生成__test Cookie后,才能获取真实的XML内容。你需要模拟这个JS加密过程,生成合法的Cookie再发起请求。
步骤说明
- 提取加密参数:从返回的HTML中解析出AES解密所需的三个十六进制参数(
a、b、c) - 模拟AES解密:用Python实现JS中
slowAES.decrypt的逻辑,生成__testCookie的值 - 带Cookie请求:携带生成的Cookie再次请求目标URL(需带上
?i=1参数) - 解析XML内容:对第二次请求返回的真实XML进行解析
具体代码实现
首先安装依赖库:
pip install pycryptodome requests
替换你的代码为以下内容:
from xml.dom import minidom import requests import re from Crypto.Cipher import AES from Crypto.Util.Padding import unpad selectedserverurl = 'http://fairbird.liveblog365.com/TSpanel/TSipanel.xml' def hex_to_bytes(hex_str): # 实现JS中的toNumbers逻辑,转换为字节 return bytes.fromhex(hex_str) def decrypt_test_cookie(a_hex, b_hex, c_hex): # 转换参数为字节 key = hex_to_bytes(a_hex) iv = hex_to_bytes(b_hex) ciphertext = hex_to_bytes(c_hex) # AES-CBC解密,对应JS中的slowAES.decrypt(c, 2, a, b),mode=2即CBC cipher = AES.new(key, AES.MODE_CBC, iv) plaintext = unpad(cipher.decrypt(ciphertext), AES.block_size) # 转换为十六进制字符串,对应JS中的toHex return plaintext.hex() def downloadxmlpage(): # 第一次请求,获取带JS的HTML session = requests.Session() response = session.get(selectedserverurl) html_content = response.text # 用正则提取a、b、c的十六进制值 pattern = r'var a=toNumbers\("([0-9a-f]+)"\),b=toNumbers\("([0-9a-f]+)"\),c=toNumbers\("([0-9a-f]+)"\);' match = re.search(pattern, html_content) if not match: print("无法提取加密参数") return a_hex, b_hex, c_hex = match.groups() # 生成__test cookie的值 test_cookie_value = decrypt_test_cookie(a_hex, b_hex, c_hex) # 设置cookie session.cookies.set('__test', test_cookie_value, domain='fairbird.liveblog365.com', path='/', expires=1735689355) # 第二次请求,带上?i=1参数 target_url = f"{selectedserverurl}?i=1" xml_response = session.get(target_url) # 解析XML if xml_response.headers.get('Content-Type', '').startswith('application/xml'): gotPageLoad(xml_response.content) else: print("仍未获取到XML内容,可能反爬机制已更新") def gotPageLoad(data = None): if data != None: xmlparse = minidom.parseString(data) for plugins in xmlparse.getElementsByTagName('plugins'): item = plugins.getAttribute('cont') if 'TSpanel' in item: for plugin in plugins.getElementsByTagName('plugin'): tsitem = plugin.getAttribute('name') print("tsitem:", tsitem) downloadxmlpage()
关键点说明
- 使用
requests.Session保持会话,自动处理Cookie - 用
pycryptodome库模拟JS中的AES-CBC解密逻辑,注意填充方式为PKCS7(JS的slowAES默认使用该填充) - 严格按照HTML中的JS逻辑生成Cookie,包括过期时间(对应JS中的
expires=Thu, 31-Dec-37 23:55:55 GMT,转换为时间戳是1735689355)
内容的提问来源于stack exchange,提问作者user18418127
相关产品推荐
相关产品推荐

