Python如何将类URL字符串转换为符合规范的标准URL格式
Python实现字符串转标准URL格式
核心处理逻辑
按照需求拆分三个核心处理环节:
- 非法URL识别:通过正则校验输入是否符合合法域名结构,要求域名必须包含多级后缀(即必须包含
.分隔符),如果带端口需要校验端口在0-65535的合法范围内,不符合规则直接返回invalid url - 协议自动补全:端口为443时自动匹配
https协议,无端口、端口为80或其他自定义端口时自动匹配http协议 - 默认端口省略:当显式声明的端口为当前协议的默认端口(http对应80、https对应443)时,拼接结果时省略端口部分
实现代码
import re def standardize_url(input_str: str) -> str: # 预处理:去除首尾空白字符 content = input_str.strip() # 匹配域名+可选端口的正则规则 url_pattern = r'^([a-zA-Z0-9-]+\.)+[a-zA-Z]{2,}(:(\d+))?$' match_res = re.match(url_pattern, content) # 格式不匹配直接返回非法 if not match_res: return "invalid url" full_match = match_res.group() # 拆分域名和端口 if ':' in full_match: domain, port_str = full_match.split(':', 1) port = int(port_str) # 校验端口合法范围 if not (0 <= port <= 65535): return "invalid url" # 按端口匹配协议、处理默认端口省略 if port == 443: return f"https://{domain}" elif port == 80: return f"http://{domain}" else: return f"http://{domain}:{port}" else: # 无显式端口默认走http return f"http://{full_match}"
效果验证
对应需求给出的示例,运行结果完全符合预期:
print(standardize_url("example.com")) # 输出 http://example.com print(standardize_url("example.com:80")) # 输出 http://example.com print(standardize_url("example.com:443")) # 输出 https://example.com print(standardize_url("example.com:8000")) # 输出 http://example.com:8000 print(standardize_url("examplecom")) # 输出 invalid url
如果需要支持带路径、查询参数的URL场景,可以在现有正则基础上扩展对应匹配规则即可。
内容的提问来源于stack exchange,提问作者Huong
相关产品推荐
相关产品推荐

