You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas zip()提取数据时遇ValueError解包值不足问题排查

问题:从Pandas DataFrame提取信息时出现ValueError:预期3个值但只得到2个

我用Python和Pandas清洗CSV数据,要从DataFrame的'Notes'列提取社会保障号码(SSN)、出生日期(DOB)、亲属关系(Relationship)等结构化信息,但一直碰到以下错误:

PS C:\Users\hokop\Documents\GitHub\Tina-Agency-of-Texas-Data> python test2.py
Traceback (most recent call last):
  File "C:\Users\hokop\Documents\GitHub\Tina-Agency-of-Texas-Data\test2.py", line 80, in <module>
    df['SSN'],df['DOB'],df['Relationship'] = zip(*df['Notes'].apply(extract_info))
    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ValueError: not enough values to unpack (expected 3, got 2)

我原本以为extract_info函数总会返回SSN、DOB、Relationship三个值,测试时也能看到这三个变量,但错误提示偶尔只返回两个值。以下是简化后的代码:

import re
import pandas as pd

# Sample input data
df = pd.read_csv('contacts.csv')

# Define regex patterns for DOB and SSN
dob_pattern = r'\b(?:DOB:|DOB;|DOB: |DOB;)\s*:? ?([0-9]{2}/[0-9]{2}/[0-9]{4})\b'
ssn_pattern = r'\b(?:SS|SS |SS#|SS:|SS: |SS;|SS; |SS# |SS#:|SS#: )\s*:? ?([0-9]{3}-[0-9]{2}-[0-9]{4}|[0-9]{9})\b'
name_pattern3 = r'(?P<first>[A-Za-z]+)(?:\s+(?P<middle>[A-Za-z]+))?\s+(?P<last>[A-Za-z]+)'
name_pattern2 = r'(?P<first>[A-Za-z\'-]+)\s+(?P<last>[A-Za-z\'-]+)'

# Define a list of relationship keywords
relationship_keywords = [
    "father",
    "mother",
    "brother",
    "sister",
    "friend",
    "spouse",
    "partner",
    "child",
    "aunt",
    "uncle",
    "cousin"
]

# Compile a regex pattern for the relationships
relationship_pattern = r'\b(?:' + '|'.join(relationship_keywords) + r')\b'

# Function to extract structured information
def extract_info(entry):
    if not isinstance(entry, str):  # Check if the entry is a string
        return '',''  # Return empty values for non-strings

    
    # Initialize variables
    name = ""
    dob = ""
    ssn = ""
    relationship = "asd"
    
    # Split entry into lines
    lines = entry.splitlines()
    for line in lines:
        line = line.strip()
        
        # if re.match(relationship_pattern, line): 
        #     relationship = re.search(relationship_pattern, line).group(1)
            
        #     if re.match(name_pattern3, line): 
        #         name = re.search(name_pattern3, line).group(1)
        #     if re.match(name_pattern2, line):
        #         name = re.search(name_pattern2, line).group(1)
        # elif not relationship:
        #     relationship = 'asd'
        if re.match(name_pattern3, line): 
            
            name = re.search(name_pattern3, line).group(1)
        elif re.match(name_pattern2, line):
            name = re.search(name_pattern2, line).group(1)
        elif re.match(ssn_pattern, line):
            # Extract SSN
            ssn = re.search(ssn_pattern, line).group(1)
        elif re.match(dob_pattern, line):
            # Extract DOB
            dob = re.search(dob_pattern, line).group(1)
        else:
            # Assume the remaining line is the name
            if line.strip() != '':
                name = line
            else:
                name = ''
    relationship = "asd"


    return ssn, dob, relationship
# Process each entry and create a list of dictionaries

df['SSN'],df['DOB'],df['Relationship'] = zip(*df['Notes'].apply(extract_info))

# Convert structured data to a DataFrame for better visualization
df.to_csv('ssn.csv', index=False)

# Display the DataFrame
print(df)

问题原因

核心问题出在extract_info函数的开头判断逻辑:当entry不是字符串类型(比如CSV里的空值NaN、None,或者数字类型)时,函数返回的是('', '')——只有两个值,但后续代码zip(*...)期望每个函数调用都返回3个值,所以只要有一条非字符串的Notes条目,就会触发这个错误。你之前测试可能没覆盖到非字符串的情况,所以误以为函数始终返回3个值。

修复与调试建议

  • 修复返回值一致性:把非字符串情况的返回语句改成返回三个空值,确保无论输入是什么,函数都返回长度为3的元组:
    if not isinstance(entry, str):  # Check if the entry is a string
        return '', '', ''  # 返回三个空值,保持元组长度统一
    
  • 定位非字符串条目:可以先排查Notes列里哪些行不是字符串,方便确认问题来源:
    # 筛选出Notes列中非字符串的行
    non_string_entries = df[~df['Notes'].apply(lambda x: isinstance(x, str))]
    print("非字符串的Notes条目:")
    print(non_string_entries)
    
  • 优化正则匹配效率:已经用re.match判断过匹配结果,无需再重复调用re.search,直接使用匹配对象即可:
    match = re.match(name_pattern3, line)
    if match: 
        name = match.group(1)
    
  • 完善亲属关系提取:目前relationship固定返回"asd",后续可以恢复注释掉的关系匹配逻辑,调试正则确保能正确识别亲属关系关键词。

内容的提问来源于stack exchange,提问作者Hoko L

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 21:57:34