You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中‘first argument must be string or compiled pattern’错误排查求助

问题分析:正则匹配抛出"first argument must be string or compiled pattern"异常

问题重现

运行以下代码时触发指定异常:

df1 = pd.read_csv('syslog.csv')
with open('data.json',encoding="utf8") as data_file:
  file_contents = data_file.read()
data = json.loads(file_contents)
df2 = pd.json_normalize(data['results'])

count = []

for index, row in df1.iterrows():
    try:
        df3 = df2["sample_syslog_message"].str.match(row['Regex'])
        if df3.any() == True:
            count.append(index)
    except Exception as e:
        print(e)
        print(index)
        break

触发异常的正则表达式为:%PLATFORM-PWRMON-1-IDPROM_INVALID\s*:.*,且该正则已通过多个在线平台验证格式正确。

失败原因

这个异常的核心原因是传入str.match()的row['Regex']并非字符串或编译后的正则对象,常见两种触发场景:

  • 场景1:CSV中对应正则的单元格为空:pd.read_csv()会将空单元格解析为NaN(浮点类型),而非字符串。当循环到该行时,把NaN传给str.match()就会触发类型错误。
  • 场景2:正则表达式的类型被意外转换:比如CSV读取时,某些正则被误解析为数值类型,导致不符合str.match()的参数要求。

注:若CSV中存储正则时反斜杠未正确转义(比如直接写\s而非\\s),会导致读取后的正则丢失转义字符,但这种情况通常触发正则语法错误,而非当前的类型异常,优先级低于类型问题。

快速验证方法

在循环中加入类型检查,确认出错行的row['Regex']类型:

for index, row in df1.iterrows():
    try:
        print(f"Index {index}, Regex type: {type(row['Regex'])}")
        df3 = df2["sample_syslog_message"].str.match(row['Regex'])
        if df3.any() == True:
            count.append(index)
    except Exception as e:
        print(e)
        print(index)
        break

内容的提问来源于stack exchange,提问作者Amankumar Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 06:28:10