如何使用C#解析批量HTTP请求文本并转换为对象及属性?
解析HTTP Multipart Batch 请求体
这是标准的HTTP Multipart Batch 请求,核心是用指定的boundary(这里是batch_f1d4b121-35c3-40a0-bbbd-cb1a9f1a5f13)分隔多个独立的HTTP请求。下面给两种可行的解析方案:
方案一:正则表达式解析(解决你的正则问题)
先提取boundary,再用它拆分每个请求块,最后解析每个块里的HTTP内容。以Python为例:
import re # 你的原始请求体文本 batch_content = """--batch_f1d4b121-35c3-40a0-bbbd-cb1a9f1a5f13 Content-Type: application/http; msgtype=request POST /api/values HTTP/1.1 Host: localhost:5102 Content-Type: application/json; charset=utf-8 {"value": "Hello World"} --batch_f1d4b121-35c3-40a0-bbbd-cb1a9f1a5f13 Content-Type: application/http; msgtype=request PUT /api/values/5 HTTP/1.1 Host: localhost:5102 Content-Type: application/json; charset=utf-8 {"value": "Hello World"} --batch_f1d4b121-35c3-40a0-bbbd-cb1a9f1a5f13 Content-Type: application/http; msgtype=request DELETE /api/values/5 HTTP/1.1 Host: localhost:5102 --batch_f1d4b121-35c3-40a0-bbbd-cb1a9f1a5f13--""" # 第一步:提取boundary boundary_match = re.match(r'--(batch_.+)\n', batch_content) if boundary_match: boundary = boundary_match.group(1) else: raise ValueError("无法找到batch boundary") # 第二步:拆分每个请求块 # 正则匹配:跳过boundary,捕获到下一个boundary或结束符之间的内容 pattern = re.compile(rf'--{re.escape(boundary)}\n(.*?)(?=\n--{re.escape(boundary)}|\n--{re.escape(boundary)}--)', re.DOTALL) request_blocks = pattern.findall(batch_content) # 第三步:解析每个请求块中的HTTP内容 parsed_requests = [] for block in request_blocks: # 拆分块的头部和HTTP请求内容(跳过Content-Type行,然后分割HTTP头和体) _, http_part = block.split('\n\n', 1) # 拆分HTTP请求行、头、体 request_line, rest = http_part.split('\n', 1) method, path, version = request_line.split() headers_part, *body_parts = rest.split('\n\n', 1) headers = {} for line in headers_part.split('\n'): if line.strip(): key, value = line.split(': ', 1) headers[key] = value body = body_parts[0].strip() if body_parts else '' parsed_requests.append({ 'method': method, 'path': path, 'version': version, 'headers': headers, 'body': body }) # 输出解析结果 for idx, req in enumerate(parsed_requests): print(f"=== 请求 {idx+1} ===") print(f"方法: {req['method']}") print(f"路径: {req['path']}") print(f"头部: {req['headers']}") print(f"请求体: {req['body']}\n")
这个正则用了re.DOTALL让.匹配换行,同时用正向预查确保只捕获到下一个boundary前的内容,避免匹配错误。
方案二:用专门的Multipart解析库(更可靠)
正则容易在复杂场景(比如请求体里包含boundary字符串、换行不一致)出错,推荐用成熟的库处理:
Python 示例(用requests-toolbelt)
from requests_toolbelt.multipart import decoder # 构造multipart内容的headers(需要指定Content-Type) content_type = f'multipart/mixed; boundary={boundary}' # 解析multipart内容 multipart_data = decoder.MultipartDecoder(batch_content.encode('utf-8'), content_type) parsed_requests = [] for part in multipart_data.parts: # 读取part的内容,解析HTTP请求 http_content = part.content.decode('utf-8') request_line, rest = http_content.split('\n', 1) method, path, version = request_line.split() headers_part, *body_parts = rest.split('\n\n', 1) headers = {} for line in headers_part.split('\n'): if line.strip(): key, value = line.split(': ', 1) headers[key] = value body = body_parts[0].strip() if body_parts else '' parsed_requests.append({ 'method': method, 'path': path, 'version': version, 'headers': headers, 'body': body }) # 输出结果 for idx, req in enumerate(parsed_requests): print(f"=== 请求 {idx+1} ===") print(f"方法: {req['method']}") print(f"路径: {req['path']}") print(f"头部: {req['headers']}") print(f"请求体: {req['body']}\n")
这种方法能处理各种边界情况,比如不同的换行符、请求体里的特殊字符,比正则更稳定。
内容的提问来源于stack exchange,提问作者Mohamed Magdy
相关产品推荐
相关产品推荐

