Python urllib urlparse解析差异疑问及RFC 3986统一解析方法求助
Is Your Understanding Correct?
Yes, your observation and reasoning are totally on point! Here’s the breakdown of why this happens in Python’s urllib.parse.urlparse:
- For unknown/custom schemes (like
some://), Python defaults to strict RFC 3986 parsing. RFC 3986 treats semicolons (;) as valid path characters, so they stay in the.pathattribute—hence your result of'/path;'. - For well-known schemes (like
http://,https://,ftp://), Python uses legacy parsing rules aligned with RFC 2396. In RFC 2396, semicolons acted as delimiters for path parameters, sourlparsestrips them from the.pathfield (moving any trailing parameters to.params; in your example, since there’s nothing after the semicolon,.pathbecomes'/path').
You can confirm this by checking the .params attribute for the HTTP case:
from urllib.parse import urlparse result = urlparse('http://foo.bar/path;') print(result.path) # Output: '/path' print(result.params) # Output: '' (empty string, since no parameters follow the semicolon)
How to Parse Any String Strictly According to RFC 3986
If you want consistent RFC 3986 parsing no matter the scheme, here are two reliable options:
Option 1: Use the rfc3986 Third-Party Library
This library is built specifically to strictly implement RFC 3986, so it handles all schemes identically. First install it:
pip install rfc3986
Then use it like this:
from rfc3986 import urlparse # Parse HTTP URL with RFC 3986 rules result = urlparse('http://foo.bar/path;') print(result.path) # Output: '/path;' # Parse custom scheme URL (consistent behavior) result = urlparse('some://foo.bar/path;') print(result.path) # Output: '/path;'
Option 2: Force Standard Library urlparse to Use RFC 3986 Rules
If you prefer sticking to Python’s standard library, you can trick urlparse into treating all schemes as "unknown" (which triggers RFC 3986 parsing) by temporarily swapping the scheme, parsing, then restoring it. Here’s a helper function to do this:
from urllib.parse import urlparse def rfc3986_urlparse(url): parsed = urlparse(url) # List of well-known schemes that use legacy parsing legacy_schemes = ('http', 'https', 'ftp', 'ftps', 'file', 'mailto') if parsed.scheme in legacy_schemes: # Replace scheme with a dummy one to trigger RFC 3986 parsing dummy_url = url.replace(f"{parsed.scheme}://", "rfc3986://", 1) dummy_parsed = urlparse(dummy_url) # Restore original scheme and return the adjusted result return dummy_parsed._replace(scheme=parsed.scheme) # For unknown schemes, return the original parsed result return parsed # Test with HTTP URL result = rfc3986_urlparse('http://foo.bar/path;') print(result.path) # Output: '/path;' # Test with custom scheme result = rfc3986_urlparse('some://foo.bar/path;') print(result.path) # Output: '/path;'
Note: This relies on knowing which schemes Python treats as legacy. The list above covers most common ones, but you can extend it if working with other niche schemes.
内容的提问来源于stack exchange,提问作者Ildar Gafurov

