You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本报错TypeError:首个参数需为字符串或编译模式的解决求助

修复按Excel唯一值拆分表格的TypeError问题

我用网上找到的Python脚本,根据Excel表格中的唯一值将其拆分为多个独立表格,代码如下:

from pathlib import Path
import pandas as pd  # pip install pandas
import os

# Define & create output directory
output_dir = Path(__file__).parent / "output"
output_dir.mkdir(parents=True, exist_ok=True)


# Define Excel file path
excel_file = Path(__file__).parent / "partnercontacts2.xlsx"

df = pd.read_excel(excel_file)
column_name = "Company"
df[column_name] = df[column_name].str.strip().str.title()
unique_values = df[column_name].unique()

for unique_value in unique_values:
    df_output = df[df[column_name].str.fullmatch(unique_value).fillna(False)]
    output_path = output_dir / f"{unique_value}.xlsx"
    df_output.to_excel(output_path, sheet_name=unique_value[:31], index=False)

运行几千次拆分操作后,出现如下错误:

TypeError: first argument must be string or compiled pattern


问题原因

这个错误是因为unique_values中存在非字符串类型的值(比如NaN、数字或其他非str类型),而pandas的str.fullmatch方法要求传入的第一个参数必须是字符串或编译后的正则表达式,非字符串类型传入后就会触发这个报错。

另外,这里用str.fullmatch属于过度设计——你只是需要精确匹配列值,完全不需要正则匹配功能。

修复方案

方案1:替换为直接相等判断(推荐)

把str.fullmatch替换成普通的相等判断,既解决类型问题,又提升运行效率:

修改后的核心代码:

df = pd.read_excel(excel_file)
column_name = "Company"
# 先处理空值,转成空字符串避免后续问题
df[column_name] = df[column_name].fillna("").str.strip().str.title()
unique_values = df[column_name].unique()

for unique_value in unique_values:
    # 直接用相等判断,无需正则
    df_output = df[df[column_name] == unique_value]
    # 处理空字符串的文件名,避免生成无效文件名
    output_path = output_dir / f"{unique_value}.xlsx" if unique_value else output_dir / "Empty_Company.xlsx"
    # 处理空字符串的工作表名
    sheet_name = unique_value[:31] if unique_value else "Empty"
    df_output.to_excel(output_path, sheet_name=sheet_name, index=False)

方案2:确保unique_value为字符串类型

如果坚持要用str.fullmatch,可以先把所有unique_value转成字符串:

修改后的循环部分:

for unique_value in unique_values:
    # 确保传入的匹配值是字符串类型
    match_str = str(unique_value) if not isinstance(unique_value, str) else unique_value
    # 使用na=False直接忽略空值,无需额外fillna处理
    df_output = df[df[column_name].str.fullmatch(match_str, na=False)]
    output_path = output_dir / f"{match_str}.xlsx"
    df_output.to_excel(output_path, sheet_name=match_str[:31], index=False)

内容的提问来源于stack exchange,提问作者Mike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:12:48