Python解析含单双引号及特殊字符文本的分割处理问题咨询
问题说明
需要解析同时包含双引号(")、单引号(')及特殊字符的文本,拆分规则如下:
- 预处理阶段移除所有换行符,将制表符替换为空格
- 按空格为分隔符拆分字段,但双引号包裹的内容作为整体不拆分
- 拆分后需要标记原本被双引号包裹的字段,不能丢失引号标识(
shlex.split()默认会剥离包裹的双引号,无法区分普通字段和引号包裹字段,不符合要求) - 兼容引号内容与非引号内容出现在同一行的场景,且引号内部的单引号、空格不影响拆分逻辑
示例输入文本:
""" Brian: "I am not the messiah" Arthur:\n\t "I say you are Lord and I should know I've followed a few" """
期望拆分结果:
['Brian:', '"I am not the messiah"', 'Arthur:', '"I say you are Lord and I should know I\'ve followed a few"']
实现方案
直接用状态机实现轻量解析,不需要依赖第三方库,逻辑完全可控,不会出现shlex自动转义、剥离字符的问题。
版本1:保留双引号作为包裹标记
直接在拆分结果中保留包裹字段的双引号,和给出的预期输出完全一致:
def split_quoted_text(raw_line: str): # 预处理:移除末尾换行,制表符替换为空格 content = raw_line.removesuffix("\n").replace("\t", " ").strip() segments = [] buffer = [] inside_quotes = False for c in content: if c == '"': buffer.append(c) inside_quotes = not inside_quotes continue if c == " " and not inside_quotes: if buffer: segments.append("".join(buffer)) buffer.clear() continue buffer.append(c) # 追加最后一个字段 if buffer: segments.append("".join(buffer)) return segments
测试效果:
test_text = 'Brian: "I am not the messiah" Arthur: "I say you are Lord and I should know I\'ve followed a few"' print(split_quoted_text(test_text)) # 输出:['Brian:', '"I am not the messiah"', 'Arthur:', '"I say you are Lord and I should know I\'ve followed a few"']
版本2:独立标记引号包裹状态
如果不需要保留双引号字符,只需要单独标记字段是否为引号包裹,可以用以下版本,返回结构为(字段内容, 是否为引号包裹)的元组列表,方便后续流程直接判断:
def split_quoted_text_with_flag(raw_line: str): content = raw_line.removesuffix("\n").replace("\t", " ").strip() segments = [] buffer = [] inside_quotes = False is_quoted = False for c in content: if c == '"': if not inside_quotes: is_quoted = True inside_quotes = not inside_quotes continue if c == " " and not inside_quotes: if buffer: segments.append(("".join(buffer), is_quoted)) buffer.clear() is_quoted = False continue buffer.append(c) if buffer: segments.append(("".join(buffer), is_quoted)) return segments
该版本测试输出:
[ ('Brian:', False), ('I am not the messiah', True), ('Arthur:', False), ("I say you are Lord and I should know I've followed a few", True) ]
从文件逐行读取时,直接对每一行调用上述函数即可,自动兼容同一行内混合普通字段和引号字段的场景,也不会被内容里的单引号、特殊字符干扰。
内容的提问来源于stack exchange,提问作者Pioneer_11
相关产品推荐
相关产品推荐

