You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何将类URL字符串转换为符合规范的标准URL格式

Python实现字符串转标准URL格式

核心处理逻辑

按照需求拆分三个核心处理环节:

  • 非法URL识别:通过正则校验输入是否符合合法域名结构,要求域名必须包含多级后缀(即必须包含.分隔符),如果带端口需要校验端口在0-65535的合法范围内,不符合规则直接返回invalid url
  • 协议自动补全:端口为443时自动匹配https协议,无端口、端口为80或其他自定义端口时自动匹配http协议
  • 默认端口省略:当显式声明的端口为当前协议的默认端口(http对应80、https对应443)时,拼接结果时省略端口部分

实现代码

import re

def standardize_url(input_str: str) -> str:
    # 预处理:去除首尾空白字符
    content = input_str.strip()
    # 匹配域名+可选端口的正则规则
    url_pattern = r'^([a-zA-Z0-9-]+\.)+[a-zA-Z]{2,}(:(\d+))?$'
    match_res = re.match(url_pattern, content)
    
    # 格式不匹配直接返回非法
    if not match_res:
        return "invalid url"
    
    full_match = match_res.group()
    # 拆分域名和端口
    if ':' in full_match:
        domain, port_str = full_match.split(':', 1)
        port = int(port_str)
        # 校验端口合法范围
        if not (0 <= port <= 65535):
            return "invalid url"
        # 按端口匹配协议、处理默认端口省略
        if port == 443:
            return f"https://{domain}"
        elif port == 80:
            return f"http://{domain}"
        else:
            return f"http://{domain}:{port}"
    else:
        # 无显式端口默认走http
        return f"http://{full_match}"

效果验证

对应需求给出的示例,运行结果完全符合预期:

print(standardize_url("example.com"))      # 输出 http://example.com
print(standardize_url("example.com:80"))   # 输出 http://example.com
print(standardize_url("example.com:443"))  # 输出 https://example.com
print(standardize_url("example.com:8000")) # 输出 http://example.com:8000
print(standardize_url("examplecom"))       # 输出 invalid url

如果需要支持带路径、查询参数的URL场景,可以在现有正则基础上扩展对应匹配规则即可。

内容的提问来源于stack exchange,提问作者Huong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 14:39:53