Python中如何保留URL中的'+'并将其编码为%2B?
URL中保留
+并转成%2B的修复方案 问题描述
我有一个解码后的URL输入,其中包含未正确转义的+符号,例如http://example.com/path/tofile?query=param+with+questionmark,需要将这些+转换为%2B。我编写了以下安全编码URL的代码:
# Safely encode the path encoded_path = quote(parsed.path) # Safely encode query parameters query_params = parse_qsl(parsed.query, keep_blank_values=True) encoded_query = urlencode(query_params, quote_via=quote, safe='')
但发现parse_qsl会把+转换成空格,导致最终编码后的查询参数中出现%20而非预期的%2B,该如何修复?
可复现代码如下:
import urllib from urllib.parse import urlparse, urlunparse, quote, urlencode, parse_qsl, quote_plus def safe_presigned_url(raw_url): parsed = urlparse(raw_url) # Safely encode the path encoded_path = quote(parsed.path) # Safely encode query parameters query_params = parse_qsl(parsed.query, keep_blank_values=True) encoded_query = urlencode(query_params, quote_via=quote, safe='') # Rebuild and return the full URL return urlunparse(( parsed.scheme, parsed.netloc, encoded_path, parsed.params, encoded_query, parsed.fragment )) url = 'http://example.com/path/tofile?query=param+with+questionmark' #decoded url print(f'BEFORE {url}') # BEFORE http://example.com/path/tofile?query=param+with+questionmark url = safe_presigned_url(url) print(f'AFTER {url}') # AFTER http://example.com/path/tofile?query=param%20with%20questionmark # EXPECTED http://example.com/path/tofile?query=param%2Bwith%2Bquestionmark
修复方案
只需修改parse_qsl的调用参数,添加plus=False即可阻止它将+解析为空格,再配合原有的urlencode设置就能得到预期结果。
修改后的完整代码:
import urllib from urllib.parse import urlparse, urlunparse, quote, urlencode, parse_qsl, quote_plus def safe_presigned_url(raw_url): parsed = urlparse(raw_url) # Safely encode the path encoded_path = quote(parsed.path) # 保留原始+符号,不解析为空格 query_params = parse_qsl(parsed.query, keep_blank_values=True, plus=False) encoded_query = urlencode(query_params, quote_via=quote, safe='') # Rebuild and return the full URL return urlunparse(( parsed.scheme, parsed.netloc, encoded_path, parsed.params, encoded_query, parsed.fragment )) url = 'http://example.com/path/tofile?query=param+with+questionmark' #decoded url print(f'BEFORE {url}') url = safe_presigned_url(url) print(f'AFTER {url}') # 输出:AFTER http://example.com/path/tofile?query=param%2Bwith%2Bquestionmark
原理说明
parse_qsl默认参数plus=True,会遵循URL查询字符串的编码规范,将+解析为空格(URL查询中+是空格的替代编码)。但在该场景中,+是原始内容的一部分,设置plus=False可让parse_qsl直接保留原始的+字符。urlencode使用quote_via=quote时,会将所有非安全字符(包括+)编码为对应的百分号形式,因此保留下来的+会被转换成%2B,符合预期。
内容的提问来源于stack exchange,提问作者Prithvi Raj
相关产品推荐
相关产品推荐

