You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析含单双引号及特殊字符文本的分割处理问题咨询

问题说明

需要解析同时包含双引号(")、单引号(')及特殊字符的文本,拆分规则如下:

  • 预处理阶段移除所有换行符,将制表符替换为空格
  • 按空格为分隔符拆分字段,但双引号包裹的内容作为整体不拆分
  • 拆分后需要标记原本被双引号包裹的字段,不能丢失引号标识(shlex.split()默认会剥离包裹的双引号,无法区分普通字段和引号包裹字段,不符合要求)
  • 兼容引号内容与非引号内容出现在同一行的场景,且引号内部的单引号、空格不影响拆分逻辑

示例输入文本:

"""
Brian: "I am not the messiah" Arthur:\n\t "I say you are Lord and I should know I've followed a few"
"""

期望拆分结果:

['Brian:', '"I am not the messiah"', 'Arthur:', '"I say you are Lord and I should know I\'ve followed a few"']
实现方案

直接用状态机实现轻量解析,不需要依赖第三方库,逻辑完全可控,不会出现shlex自动转义、剥离字符的问题。

版本1:保留双引号作为包裹标记

直接在拆分结果中保留包裹字段的双引号,和给出的预期输出完全一致:

def split_quoted_text(raw_line: str):
    # 预处理:移除末尾换行,制表符替换为空格
    content = raw_line.removesuffix("\n").replace("\t", " ").strip()
    segments = []
    buffer = []
    inside_quotes = False
    for c in content:
        if c == '"':
            buffer.append(c)
            inside_quotes = not inside_quotes
            continue
        if c == " " and not inside_quotes:
            if buffer:
                segments.append("".join(buffer))
                buffer.clear()
            continue
        buffer.append(c)
    # 追加最后一个字段
    if buffer:
        segments.append("".join(buffer))
    return segments

测试效果:

test_text = 'Brian: "I am not the messiah" Arthur:  "I say you are Lord and I should know I\'ve followed a few"'
print(split_quoted_text(test_text))
# 输出:['Brian:', '"I am not the messiah"', 'Arthur:', '"I say you are Lord and I should know I\'ve followed a few"']

版本2:独立标记引号包裹状态

如果不需要保留双引号字符,只需要单独标记字段是否为引号包裹,可以用以下版本,返回结构为(字段内容, 是否为引号包裹)的元组列表,方便后续流程直接判断:

def split_quoted_text_with_flag(raw_line: str):
    content = raw_line.removesuffix("\n").replace("\t", " ").strip()
    segments = []
    buffer = []
    inside_quotes = False
    is_quoted = False
    for c in content:
        if c == '"':
            if not inside_quotes:
                is_quoted = True
            inside_quotes = not inside_quotes
            continue
        if c == " " and not inside_quotes:
            if buffer:
                segments.append(("".join(buffer), is_quoted))
                buffer.clear()
                is_quoted = False
            continue
        buffer.append(c)
    if buffer:
        segments.append(("".join(buffer), is_quoted))
    return segments

该版本测试输出:

[
    ('Brian:', False),
    ('I am not the messiah', True),
    ('Arthur:', False),
    ("I say you are Lord and I should know I've followed a few", True)
]

从文件逐行读取时,直接对每一行调用上述函数即可,自动兼容同一行内混合普通字段和引号字段的场景,也不会被内容里的单引号、特殊字符干扰。

内容的提问来源于stack exchange,提问作者Pioneer_11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 23:12:19