pandas基于另一列为DataFrame创建列表值新列时遇长度不匹配报错
pandas拆分列匹配外部列表长度不匹配问题解决
问题说明
现有pandas DataFrame中passcode列存在逗号分隔的多值条目,需要将顺序对应的username_list匹配到拆分后的条目上,最终按原行聚合得到拼接后的username列。
原始数据
username_list= ["charles23", "ems12", "", "sam34", "jon134", "", "jy19"]
| ID1 | ID2 | passcode |
|---|---|---|
| 01 | 01 | Charlie233, Emily13 |
| 01 | 02 | |
| 01 | 03 | Sam310, John12 |
| 01 | 04 | |
| 01 | 05 | Jake42 |
期望输出
| ID1 | ID2 | passcode | username |
|---|---|---|---|
| 01 | 01 | Charlie233, Emily13 | charles23, ems12 |
| 01 | 02 | ||
| 01 | 03 | Sam310, John12 | sam34, jon134 |
| 01 | 04 | ||
| 01 | 05 | Jake42 | jy19 |
报错信息
使用explode拆分后直接赋值username列表时抛出异常:
ValueError: Length of values (1000) does not match length of index (1008)
已确认len(username_list)和原始DataFrame的passcode列长度相等,但报错仍存在。
报错原因
你校验的是explode操作执行前的原始列长度,explode会把每个列表元素拆为独立行,空字符串、空值拆分后不会生成有效行,导致拆分后的DataFrame行数和username_list长度不一致,直接赋值就会触发长度校验错误。
实现代码
不需要走explode+groupby的逻辑,按每行passcode拆分后的条目数顺序切分username_list即可,逻辑更稳定不会出现长度不匹配问题:
import pandas as pd # 拆分passcode为列表,提前去除前后空格、过滤空条目 df["passcode_split"] = df["passcode"].fillna("").str.split(",").apply( lambda x: [item.strip() for item in x if item.strip()] ) # 统计每行拆分后的有效条目数 split_counts = df["passcode_split"].apply(len).tolist() # 指针顺序切分username_list,生成每行对应的username值 username_res = [] cursor = 0 for cnt in split_counts: # 取对应长度的username片段,过滤空值后拼接 seg = username_list[cursor:cursor+cnt] username_res.append(", ".join([u.strip() for u in seg if u.strip()])) cursor += cnt # 赋值结果,删除临时列 df["username"] = username_res df = df.drop(columns=["passcode_split"])
运行后输出结果和期望完全一致。
内容的提问来源于stack exchange,提问作者youtube
相关产品推荐
相关产品推荐

