You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pyparsing技术问题:ParseResults无法获取predictors与covariates全部值

问题原因与解决方案

这个问题我之前也碰到过,核心原因是pyparsing处理重复结构中命名匹配的默认行为:当你在OneOrMore()这类重复解析器里直接给元素命名(比如var('covariates')),pyparsing每次匹配到一个var时,都会用新匹配的值覆盖之前存储的covariates条目,而不是把它追加到列表里。

虽然parseString返回的结果字典看起来包含完整的列表,但那是pyparsing内部保留的所有匹配记录,默认的属性访问(results.covariates)或字典索引(results['covariates'])只会返回最后一次匹配的值。

解决方法

要让pyparsing把所有匹配到的元素收集成列表,有两种常用方式:

方式1:使用setResultsName并指定listAllMatches=True

显式告诉pyparsing要收集所有匹配项:

from pyparsing import Word, alphanums, OneOrMore, Optional, Suppress

var = Word(alphanums)
# 给重复匹配的元素设置命名,并指定收集所有匹配
predictor_part = var.setResultsName('predictors', listAllMatches=True) + Optional(Suppress('+'))
covariate_part = var.setResultsName('covariates', listAllMatches=True) + Optional(Suppress('+'))

reg = OneOrMore(predictor_part) + '~' + OneOrMore(covariate_part)
string = 'y1 ~ f1 + f2 + f3'
results = reg.parseString(string)

print(results.predictors)  # 输出: ['y1']
print(results.covariates)  # 输出: ['f1', 'f2', 'f3']

方式2:使用*前缀简化命名(pyparsing 3.0+支持)

在命名前加*,等价于listAllMatches=True,写法更简洁:

from pyparsing import Word, alphanums, OneOrMore, Optional, Suppress

var = Word(alphanums)
# 用*前缀表示收集所有匹配项
reg = OneOrMore(var('*predictors') + Optional(Suppress('+'))) + '~' + OneOrMore(var('*covariates') + Optional(Suppress('+')))

string = 'y1 ~ f1 + f2 + f3'
results = reg.parseString(string)

print(results.predictors)  # 输出: ['y1']
print(results.covariates)  # 输出: ['f1', 'f2', 'f3']

额外说明

如果你的表达式里有多个预测变量(比如y1 + y2 ~ f1 + f2),原来的写法同样会只返回最后一个预测变量y2,用上面的方法也能解决这个问题,确保收集到完整的列表。

内容的提问来源于stack exchange,提问作者Ankur Ankan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:21:21