如何在Python中拼接任意URL片段?urljoin路径丢失问题解析
解决Python中urljoin拼接URL丢失路径的问题
你遇到的问题是urljoin的规则导致的:当第二个参数(要拼接的片段)以斜杠/开头时,它会替换base URL的最后一级路径,而不是追加。你的base URL是https://dummy.restapiexample.com/api/v1,因为它没有以斜杠结尾,api/v1会被当作最后一个路径段(类似“文件名”的角色),所以被/employees替换掉了。
下面是几种正确拼接任意URL片段的方法:
方法1:给base URL添加结尾斜杠
这是最符合urljoin设计逻辑的方式,给base加上结尾斜杠后,api/v1会被识别为一个目录路径,此时以斜杠开头的片段就会追加到这个目录下:
from urllib.parse import urljoin base = "https://dummy.restapiexample.com/api/v1/" tail = "/employees" print(urljoin(base, tail)) # 输出结果: 'https://dummy.restapiexample.com/api/v1/employees'
方法2:手动处理路径后拼接
如果不想修改base URL,或者需要处理各种格式的片段(比如有的带斜杠有的不带),可以拆分base的结构,手动拼接路径:
from urllib.parse import urlparse, urlunparse base = "https://dummy.restapiexample.com/api/v1" tail = "/employees" # 解析base URL的各个部分 parsed_base = urlparse(base) # 处理路径,避免重复或缺失斜杠 new_path = f"{parsed_base.path.rstrip('/')}/{tail.lstrip('/')}" # 重新组装URL new_url = urlunparse(parsed_base._replace(path=new_path)) print(new_url) # 输出结果: 'https://dummy.restapiexample.com/api/v1/employees'
方法3:简单字符串拼接(需处理斜杠)
如果场景简单,直接用字符串拼接,同时处理首尾的斜杠,避免出现//或者缺失的情况:
base = "https://dummy.restapiexample.com/api/v1" tail = "/employees" new_url = f"{base.rstrip('/')}/{tail.lstrip('/')}" print(new_url) # 输出结果: 'https://dummy.restapiexample.com/api/v1/employees'
内容的提问来源于stack exchange,提问作者Dims
相关产品推荐
相关产品推荐

