如何用Python构建符合指定规则的自定义Canonical URL?
构建my.com网站的Canonical URL解决方案
需求说明
- 主域名
my.com始终包含www子域名,子域名(如sub.my.com)保持原样 - 强制使用HTTPS协议
- 移除默认的80、443端口
- 移除路径末尾的斜杠
转换示例
- http://my.com => https://www.my.com
- http://my.com/ => https://www.my.com
- https://my.com:80/ => https://www.my.com
- https://sub.my.com/ => https://sub.my.com
- https://sub.my.com?term=t1 => https://sub.my.com?term=t1
原代码问题分析
你写的代码存在两个关键问题:
- 修改
parsed对象属性时误用了比较运算符==,应该用赋值运算符= - 未利用
urllib.parse提供的urlunparse方法来规范拼接URL组件
修正后的完整代码
from urllib.parse import urlparse, urlunparse def build_canonical_url(request): # 获取请求的完整绝对URL absolute_url = request.build_absolute_uri() parsed = urlparse(absolute_url) # 1. 强制切换为HTTPS协议 scheme = 'https' # 2. 为主域名添加www子域名,子域名不做修改 hostname = parsed.hostname if hostname == 'my.com': hostname = 'www.my.com' # 3. 移除默认端口,非默认端口保留 port = None if parsed.port in (80, 443) else parsed.port # 4. 移除路径末尾的斜杠,空路径保持为空 path = parsed.path.rstrip('/') if parsed.path else '' # 拼接网络位置(主机名+端口) netloc = hostname if port is not None: netloc = f"{hostname}:{port}" # 用urlunparse重新组合所有URL组件 canonical_url = urlunparse(( scheme, netloc, path, parsed.params, parsed.query, parsed.fragment )) return canonical_url
关键代码说明
- 用
urlunparse替代手动拼接,它能自动处理URL各组件的格式规范,避免手动拼接出现的格式错误 - 精确匹配主域名
my.com添加www,避免误修改子域名 - 用
rstrip('/')安全移除路径末尾斜杠,空路径不会被处理成多余字符 - 仅保留非默认端口,80和443端口直接省略
内容的提问来源于stack exchange,提问作者Hooman Bahreini
相关产品推荐
相关产品推荐

