You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用apply为pandas DataFrame添加新列并保持索引对齐

问题解决方法

问题根源

你之前的写法将未匹配到结果的行直接丢弃,flattenedWeights、flattenedCounts仅保留了匹配成功的内容,长度和原DataFrame的8000行不一致,赋值时pandas会自动在末尾补空值,导致有效内容全部集中在列顶部,和原索引无法对齐。

正确实现方案

推荐直接使用pandas自带的字符串匹配方法,自动对齐原DataFrame索引,无需手动处理列表,代码更简洁效率也更高:

import pandas as pd
import re

# 1. 匹配重量列:支持ml/g/gm单位,自动忽略大小写和数字与单位之间的空格
weight_pattern = r'(\d+\s*(?:ml|g|gm))'
df['flattenedWeights'] = df['item_name'].str.extract(weight_pattern, flags=re.IGNORECASE)

# 2. 匹配包装规格列:合并两种匹配规则,支持x/pk/pack/packs、pack of 等格式
count_pattern = r'(\d+\s*(?:x|pk|pack|packs)|packs? of \d+)'
df['flattenedCounts'] = df['item_name'].str.extract(count_pattern, flags=re.IGNORECASE)

# 若某行可能匹配到多个结果,需要保留所有结果时,改用findall+join的写法:
# df['flattenedWeights'] = df['item_name'].str.findall(weight_pattern, flags=re.IGNORECASE).str.join(',')

# 导出结果
df.to_csv('newColumns.csv', index=False)

之前写法报错原因

  • 自定义fab函数没有接收行参数也没有返回值,apply调用时每一行都会重复执行全量匹配,既拿不到对应行的结果,效率也极低
  • 直接把列表flattenedWeights传入apply方法,pandas无法将列表识别为可执行函数,会直接报错

内容的提问来源于stack exchange,提问作者Osiris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 02:27:02