Java:如何按冒号正确分割HTTP请求头的名称与值对
HTTP头字段拆分解决方案
你之前使用str.Split(":", 2)失效的原因是,HTTP/2伪头的开头冒号属于头名称的一部分,该方法会把开头的冒号识别为分割点,导致头名称拆分错误。
核心实现逻辑
HTTP头名称本身遵循RFC规范,除了HTTP/2伪头会以单个冒号开头外,头名称的其他位置不允许出现冒号,因此可以按以下规则匹配分隔用的冒号:
- 若行首字符不是冒号,第一个出现的冒号即为分隔符
- 若行首字符是冒号,第二个出现的冒号即为分隔符
匹配到分隔符后,前半部分作为头名,后半部分去除开头空白后作为头值即可。
代码示例
遍历实现(Python为例)
def split_header(line: str) -> tuple[str, str]: colon_count = 0 split_idx = -1 for idx, char in enumerate(line): if char == ":": colon_count += 1 # 普通头取第一个冒号,伪头取第二个冒号作为分割点 if (line[0] != ":" and colon_count == 1) or (line[0] == ":" and colon_count == 2): split_idx = idx break if split_idx == -1: # 非法格式行可根据业务需求调整处理逻辑 return (line.strip(), "") header_name = line[:split_idx].strip() header_value = line[split_idx+1:].strip() return (header_name, header_value) # 测试 test_lines = [ "accept-encoding: gzip, deflate, br", ":authority: stackoverflow.com", "something: some:value" ] for line in test_lines: name, value = split_header(line) print(f"{name} => {value}")
运行输出完全符合预期:
accept-encoding => gzip, deflate, br :authority => stackoverflow.com something => some:value
正则实现(通用性更强)
可以用正则表达式一次性匹配头名和头值,匹配规则为^(:?[^:]+):\s*(.*),第一分组为头名,第二分组为头值,Python示例如下:
import re header_pattern = re.compile(r'^(:?[^:]+):\s*(.*)') def split_header_by_re(line: str) -> tuple[str, str]: match_res = header_pattern.match(line) if not match_res: return (line.strip(), "") return (match_res.group(1), match_res.group(2))
内容的提问来源于stack exchange,提问作者HelloWorld
相关产品推荐
相关产品推荐

