You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式解析HTTP URL及长格式字符串处理技术问询

嘿,针对你提出的两个技术需求,我整理了实用的落地方案,直接上干货:

需求1:使用正则表达式解析HTTP URL

HTTP URL的标准结构一般是 scheme://netloc/path?query#fragment,咱们可以用一个覆盖大部分常见场景的正则来提取各个核心部分。

推荐正则表达式

^(https?)://([^:/\s]+)(?::(\d+))?(/[^\s?#]*)?(?:\?([^\s#]*))?(?:#([^\s]*))?$

每个分组对应的含义:

  • 第1组:https? - 协议(仅匹配http或https)
  • 第2组:[^:/\s]+ - 域名或IP地址
  • 第3组:\d+ - 端口号(可选,未指定时为空)
  • 第4组:/[^\s?#]* - 请求路径(可选,默认根路径/)
  • 第5组:[^\s#]* - 查询参数(可选)
  • 第6组:[^\s]* - 锚点片段(可选)

代码示例(Python)

import re

# 预编译正则,提升匹配效率
url_regex = re.compile(r'^(https?)://([^:/\s]+)(?::(\d+))?(/[^\s?#]*)?(?:\?([^\s#]*))?(?:#([^\s]*))?$')
test_url = "https://example.com:8080/api/user?id=123&name=foo#profile"

match_result = url_regex.match(test_url)
if match_result:
    scheme = match_result.group(1)
    netloc = match_result.group(2)
    port = match_result.group(3) or ("443" if scheme == "https" else "80")
    path = match_result.group(4) or "/"
    query_params = match_result.group(5) or ""
    fragment = match_result.group(6) or ""
    
    print(f"协议: {scheme}")
    print(f"域名: {netloc}")
    print(f"端口: {port}")
    print(f"路径: {path}")
    print(f"查询参数: {query_params}")
    print(f"锚点: {fragment}")
需求2:处理格式异常的JSON长字符串

先拆解你给出的示例字符串,里面存在3处明显的格式错误:

  1. "tu":"http"://bus.mapit.me/iot/pipe/ - URL内部错误拆分了引号,应该合并为"tu":"http://bus.mapit.me/iot/pipe/"
  2. "uu":"http"://bus.mapit.me/firmware/ - 同上,需修正为"uu":"http://bus.mapit.me/firmware/"
  3. "sw":mapit2_v245 - 字符串值未添加双引号,符合JSON规范的写法是"sw":"mapit2_v245"

处理方案:先批量修复格式,再解析JSON

我们可以用正则表达式快速修复这些共性错误,再用标准JSON库完成解析:

代码示例(Python)

import json
import re

# 原始格式异常的字符串
raw_json_str = '{ "method": "POST", "url": "/iot/pipe/", "query": {}, "body": { "d": {"l": 1523737659, "n": "861359030665564", "b": 100, "v": "02.45", "t": 3, "dev": {"vr":7, "ae":1, "at":5, "ad":2, "as":4, "al":60, "tp":60, "tr":3, "tu":"http"://bus.mapit.me/iot/pipe/, "gt":50, "gm":120, "gh":400, "gs":3, "gr":2, "gg":1, "ua":0, "uu":"http"://bus.mapit.me/firmware/, "le":0, "lt":0, "sw":mapit2_v245, "sp":240, "rt":0, "sa":1}}}, "headers": { "host": "nod...'

# 修复错误1和2:替换拆分的URL引号
fixed_str = re.sub(r'"http"://', r'"http://', raw_json_str)
# 修复错误3:给sw的字符串值添加双引号(假设值仅包含字母、数字和下划线)
fixed_str = re.sub(r'"sw":(\w+_\w+)', r'"sw":"\1"', fixed_str)

# 注意:示例字符串末尾的headers部分是截断的,实际场景需要补全闭合符号(这里模拟补全)
fixed_str += '}}}'

try:
    parsed_data = json.loads(fixed_str)
    print("JSON解析成功!")
    # 示例:提取设备核心信息
    print(f"设备IMEI: {parsed_data['body']['d']['n']}")
    print(f"设备固件版本: {parsed_data['body']['d']['v']}")
    print(f"设备上报的服务地址: {parsed_data['body']['d']['dev']['tu']}")
except json.JSONDecodeError as e:
    print(f"解析失败,错误详情: {e}")

额外容错方案

如果你的原始数据经常出现这类非标准格式问题,可以使用容错性更强的第三方库(比如demjson),它能自动兼容部分不严格的JSON写法:

import demjson

# 无需手动修复部分小错误,直接解析
parsed_data = demjson.decode(raw_json_str)

不过还是建议优先修复原始数据的格式问题,标准JSON解析的稳定性和性能会更优。


内容的提问来源于stack exchange,提问作者Flame_Phoenix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:12:45