为何urllib.parse无法解析某URL的查询参数?
问题分析与解决
问题原因
两个URL的参数位置不同,导致解析结果差异:
- 正常运行的URL参数在
?之后,属于标准查询参数(query),parse_qs正是用来解析这部分内容的。 - 无法解析的URL参数在
#之后,这部分属于URL的锚点片段(fragment),urlparse会将其存入parsed_url.fragment属性,而非parsed_url.query,因此直接解析query会得到空字典。
解决方案
将解析对象从parsed_url.query替换为parsed_url.fragment即可:
from urllib.parse import urlparse from urllib.parse import parse_qs url = 'https://newassets.hcaptcha.com/captcha/v1/000919d/static/hcaptcha.html#frame=checkbox&id=0d4abkdnbvpa&host=2captcha.com&sentry=true&reportapi=https%3A%2F%2Faccounts.hcaptcha.com&recaptchacompat=off&custom=false&hl=ru&tplinks=on&sitekey=41b778e7-8f20-45cc-a804-1f1ebb45c579&theme=light&origin=https%3A%2F%2F2captcha.com' parsed_url = urlparse(url) # 解析锚点片段而非查询参数 captured_value = parse_qs(parsed_url.fragment) print(captured_value)
输出结果
{'frame': ['checkbox'], 'id': ['0d4abkdnbvpa'], 'host': ['2captcha.com'], 'sentry': ['true'], 'reportapi': ['https://accounts.hcaptcha.com'], 'recaptchacompat': ['off'], 'custom': ['false'], 'hl': ['ru'], 'tplinks': ['on'], 'sitekey': ['41b778e7-8f20-45cc-a804-1f1ebb45c579'], 'theme': ['light'], 'origin': ['https://2captcha.com']}
补充说明
urlparse会将URL拆分为scheme、netloc、path、params、query、fragment几个部分:query对应?到#之间的内容(若存在#)fragment对应#之后的所有内容
parse_qs仅负责解析key1=value1&key2=value2格式的键值对字符串,无论该字符串原本是query还是fragment,只要格式匹配就能正常解析。
内容的提问来源于stack exchange,提问作者Eugene
相关产品推荐
相关产品推荐

