You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用原DataFrame的值迭代替换所有单元格?

解决Pandas中逐行逐列调用自定义函数替换单元格的问题

需求说明

遍历DataFrame的每一行,将每行的Date(数字字符串类型)与title1、title2、title3列的短语传入do_calculations(date, phrase)函数,用函数返回值替换原短语所在的单元格,完成所有行的处理。

原代码问题分析

你之前的代码存在以下问题:

  • 仅处理了Title1列(注意大小写,示例DataFrame中列名为小写title1),未循环覆盖title2、title3列
  • 索引逻辑混乱,row.index[1]、row.index[2]这种写法依赖列的顺序,极易出错且无法适配多列处理
  • iterrows结合at的手动循环效率较低,数据量大时性能劣势明显

高效解决方案(优先选择)

方案1:批量列处理(性能最优)

利用Pandas的apply结合广播特性,对每个title列单独处理,代码简洁且效率更高:

import pandas as pd
import random

# 生成示例DataFrame(已修正Date为数字字符串格式)
date_rng = pd.date_range(start='1/1/2023', end='1/10/2023', freq='D')
phrases = ['Hello world', 'Python is awesome', None, 'Data science is fun', 'I love coding', 'Pandas is powerful', 'Pineapples', 'Pizza', 'Krusty', 'krab', 'Is the pizza']
df = pd.DataFrame(columns=['Date', 'title1', 'title2', 'title3'])

for date in date_rng:
    # 将Date转为数字字符串,符合需求
    row = [date.strftime('%Y-%m-%d')]
    row.extend(random.sample(phrases, 3))
    df = df.append(pd.Series(row, index=df.columns), ignore_index=True)

# 自定义计算函数(示例逻辑,替换为你的业务代码)
def do_calculations(date_str, phrase):
    if phrase is None:
        return None
    return f"[{date_str}] {phrase.upper()}"

# 遍历所有title列,批量处理
title_cols = ['title1', 'title2', 'title3']
for col in title_cols:
    df[col] = df.apply(lambda x: do_calculations(x['Date'], x[col]), axis=1)

# 查看结果
print(df)

方案2:向量化改造(极致效率)

如果do_calculations可以改造成向量化函数,性能会进一步提升(适合大数据量场景):

import numpy as np

# 向量化版本的计算函数
def do_calculations_vec(date_arr, phrase_arr):
    # 过滤空值
    mask = phrase_arr.notna()
    # 对非空值应用逻辑,空值保留None
    result = np.where(mask, "[" + date_arr[mask] + "] " + phrase_arr[mask].str.upper(), None)
    return result

# 批量处理列
for col in title_cols:
    df[col] = do_calculations_vec(df['Date'], df[col])

修正后的循环方案(小数据量适用)

如果必须使用逐行循环,修正索引和循环逻辑即可:

for i, row in df.iterrows():
    current_date = row['Date']
    # 遍历所有title列
    for col in title_cols:
        phrase = row[col]
        if phrase is not None:
            df.at[i, col] = do_calculations(current_date, phrase)

总结

优先选择批量列处理或向量化改造的方案,数据量越大性能优势越明显;原代码的核心问题是未覆盖所有目标列,且索引依赖列顺序导致逻辑错误。

内容的提问来源于stack exchange,提问作者Jay Jung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:15:19