You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame列名拼接至对应行值且空值跳过的实现方法

Pandas DataFrame列名拼接至对应行取值前的实现方法
  • 首先筛选出需要拼接处理的列,排除不需要改动的列(本例中为fruit_tag)
  • 对每个待处理列的取值做判断:仅当取值非空(排除空值、空字符串、全空格内容)时,将列名+: 拼接在取值前方,否则保留空值

小数据量易读实现(循环写法)

import pandas as pd

# 构造示例数据
data = {'fruit_tag': {0: 'apple', 1: 'apple', 2: 'banana', 3: 'apple', 4: 'watermelon'}, 'location': {0: 'Hong Kong', 1: 'Tokyo', 2: '', 3: '', 4: ''}, 'rating': {0: 'bad', 1: 'good', 2: 'good', 3: 'bad', 4: 'good'}, 'measure_score': {0: 0.9529434442520142, 1: 0.952498733997345, 2: 0.9080725312232971, 3: 0.8847543001174927, 4: 0.8679852485656738}}
dat = pd.DataFrame.from_dict(data)

# 配置不需要处理的列
exclude_cols = ['fruit_tag']
process_cols = [col for col in dat.columns if col not in exclude_cols]

# 循环处理每一列
for col in process_cols:
    dat[col] = dat[col].apply(
        lambda x: f"{col}: {x}" if pd.notna(x) and str(x).strip() != '' else ''
    )

# 输出结果
print(dat)

大数据量高性能实现(向量化写法)

万行以上数据推荐使用该写法,性能远高于逐行apply:

import pandas as pd

# 构造示例数据
data = {'fruit_tag': {0: 'apple', 1: 'apple', 2: 'banana', 3: 'apple', 4: 'watermelon'}, 'location': {0: 'Hong Kong', 1: 'Tokyo', 2: '', 3: '', 4: ''}, 'rating': {0: 'bad', 1: 'good', 2: 'good', 3: 'bad', 4: 'good'}, 'measure_score': {0: 0.9529434442520142, 1: 0.952498733997345, 2: 0.9080725312232971, 3: 0.8847543001174927, 4: 0.8679852485656738}}
dat = pd.DataFrame.from_dict(data)

exclude_cols = ['fruit_tag']
process_cols = [col for col in dat.columns if col not in exclude_cols]

# 向量化批量处理
dat[process_cols] = dat[process_cols].apply(
    lambda s: f"{s.name}: " + s.astype(str)
).where(dat[process_cols].astype(str).str.strip() != '', '')

如果需要保留原始数据不被修改,只要在处理前先执行dat_processed = dat.copy(),后续对dat_processed做处理即可。

内容的提问来源于stack exchange,提问作者codedancer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 04:15:03