PyParsing中runTests处理硬编码与文件字符串的解析差异
PyParsing解析行为不一致问题排查
问题描述
使用PyParsing的runTests方法时,发现硬编码字符串与从文件读取的字符串解析结果存在差异。运行环境为Python 3.9.6(Mac系统)。
测试代码
import pyparsing as pp import sys from pathlib import Path # 基于多行块捕获示例构建 pp.enable_diag(pp.Diagnostics.enable_debug_on_named_expressions) EOL = pp.LineEnd() EmptyLine = pp.Suppress(pp.LineStart() + EOL) englishLine_StopSeparator = pp.LineStart() + "OM" englishLines = pp.Optional(EOL) + pp.Group( pp.OneOrMore(pp.SkipTo(pp.LineEnd()) + EOL, stopOn=englishLine_StopSeparator) ) titleLine = pp.Combine( pp.Word(pp.nums) + (pp.SkipTo(EOL)) + pp.Suppress(EOL) ).setResultsName("titleLine_Section*") invocationLine_StopSeparator = pp.LineStart() + titleLine invocationLines = pp.Optional(EOL) + pp.Group( pp.OneOrMore(pp.SkipTo(pp.LineEnd()) + EOL, stopOn=invocationLine_StopSeparator) ) Separator1 = pp.Keyword("*************") prasna_Separator = pp.Keyword("============").suppress() Separators = Separator1 | prasna_Separator prasna_StopSeparator = pp.LineStart() + prasna_Separator englishPreface = pp.Combine(pp.Keyword("Notes") + englishLines).setResultsName( "englishPreface_Section*" ) invocation = pp.Combine(pp.Keyword("OM") + invocationLines).setResultsName( "invocation_Section*" ) sectiontitleLine = pp.Combine( pp.Word(pp.nums) + pp.Literal(".") + pp.Word(pp.nums) ).setResultsName("sectiontitleLine_Section*") prasnaLines = pp.Optional(EOL) + pp.Combine( pp.OneOrMore(pp.SkipTo(pp.LineEnd()) + EOL, stopOn=prasna_StopSeparator) ).setResultsName("prasna_Section*") prasna = pp.Optional(prasna_Separator + pp.SkipTo(EOL)).suppress() + prasnaLines prasna.setName("prasna") parser = englishPreface + invocation + titleLine + pp.OneOrMore(prasna) if args_count := len(sys.argv) < 2: print("Provide a file name ") raise SystemExit(2) file_name = sys.argv[1] text = Path(file_name).read_text(encoding="utf-8") text="""\ Notes for the users of this document EnglishPrefaceNotes EnglishPrefaceNotes OM invocationFirstLine invocationFirstLine 3 Title Title Title Title 3.1 SectionTitle SectionTitle AnnexureText AnnexureText ======================== 3.2 SectionTitle SectionTitle SectionTitle PrasnaEndLine2_2 ======================== """ #text = Path(file_name).read_text(encoding="utf-8") texts=[text] parser.runTests(texts)
硬编码字符串解析结果
Match prasna at loc 147(6,1) 3.1 SectionTitle SectionTitle ^ Matched prasna -> ['3.1 SectionTitle SectionTitle\nAnnexureText AnnexureText \n'] Match prasna at loc 204(8,1) ======================== ^ Matched prasna -> ['\n', '3.2 SectionTitle SectionTitle SectionTitle\nPrasnaEndLine2_2\n'] Match prasna at loc 290(12,1) ======================== ^ Matched prasna -> ['\n', ''] Match prasna at loc 316(13,2) ^ Match prasna failed, ParseException raised: , found end of text (at char 316), (line:13, col:2) Notes for the users of this document EnglishPrefaceNotes EnglishPrefaceNotes OM invocationFirstLine invocationFirstLine 3 Title Title Title Title 3.1 SectionTitle SectionTitle AnnexureText AnnexureText ======================== 3.2 SectionTitle SectionTitle SectionTitle PrasnaEndLine2_2 ======================== ['Notes for the users of this document\n\nEnglishPrefaceNotes EnglishPrefaceNotes\n', 'OM invocationFirstLine invocationFirstLine\n', '3 Title Title Title Title', '3.1 SectionTitle SectionTitle\nAnnexureText AnnexureText \n', '\n', '3.2 SectionTitle SectionTitle SectionTitle\nPrasnaEndLine2_2\n', '\n', ''] - englishPreface_Section: ['Notes for the users of this document\n\nEnglishPrefaceNotes EnglishPrefaceNotes\n'] - invocation_Section: ['OM invocationFirstLine invocationFirstLine\n'] - prasna_Section: ['3.1 SectionTitle SectionTitle\nAnnexureText AnnexureText \n', '3.2 SectionTitle SectionTitle SectionTitle\nPrasnaEndLine2_2\n', ''] - titleLine_Section: ['3 Title Title Title Title']
文件读取字符串解析结果
Match prasna at loc 147(6,1) 3.1 SectionTitle SectionTitle ^ Matched prasna -> ['3.1 SectionTitle SectionTitle\nAnnexureText AnnexureText \n'] Match prasna at loc 204(8,1) ======================== ^ Matched prasna -> ['\n', '3.2 SectionTitle SectionTitle SectionTitle\nPrasnaEndLine2_2\n'] Match prasna at loc 290(12,1) ======================== ^ Match prasna failed, ParseException raised: , found end of text (at char 315), (line:12, col:26) Notes for the users of this document EnglishPrefaceNotes EnglishPrefaceNotes OM invocationFirstLine invocationFirstLine 3 Title Title Title Title 3.1 SectionTitle SectionTitle AnnexureText AnnexureText ======================== 3.2 SectionTitle SectionTitle SectionTitle PrasnaEndLine2_2 ======================== ======================== ^ ParseException: Expected end of text, found '=' (at char 290), (line:12, col:1) FAIL: Expected end of text, found '=' (at char 290), (line:12, col:1)
原因分析与解决
核心问题
- 换行符格式差异:Mac系统原生换行符为
\r(CR),而硬编码字符串使用\n(LF)。PyParsing的LineEnd()默认匹配平台原生换行符,导致两种场景下换行符匹配逻辑不一致。 - 文件内容不一致:从解析结果可见,文件末尾多了一行
========================分隔符,而硬编码字符串中没有这一行,直接触发了解析失败。
解决方案
- 统一换行符处理:读取文件后将所有换行符转换为
\n,消除平台差异:
text = Path(file_name).read_text(encoding="utf-8").replace('\r', '\n')
- 显式指定换行符:将
LineEnd()替换为显式匹配\n,避免依赖平台默认行为:
EOL = pp.Literal('\n')
- 对齐文件内容:确保文件内容与硬编码字符串完全一致,删除末尾多余的分隔符行。
内容的提问来源于stack exchange,提问作者Harihara Vinayakaram
相关产品推荐
相关产品推荐

