You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas/Vaex复制day_of_year=140行替换148行?

用Pandas和Vaex实现替换DataFrame行的方案

没问题,我来给你详细拆解在Pandas和Vaex里怎么实现这个需求——把day_of_year=148的1000条记录,全部替换成从day_of_year=140的记录中复制的内容:

一、Pandas实现方案

Pandas处理中小规模数据很灵活,这里分两种常见场景实现:

场景1:删除原148行,添加1000条140行(保持总行数不变)

import pandas as pd

# 1. 筛选出day_of_year=140的行,用copy()避免视图修改问题
df_140 = df[df['day_of_year'] == 140].copy()

# 2. 生成刚好1000条140的行(通用写法:即使140的行不足1000,也能重复凑够数量)
needed_count = len(df[df['day_of_year'] == 148])  # 这里就是1000
repeat_times = (needed_count // len(df_140)) + 1
repeated_140 = pd.concat([df_140]*repeat_times, ignore_index=True).head(needed_count)

# 3. 删除原148的行,再添加新的140行
df = df[df['day_of_year'] != 148].reset_index(drop=True)
df = pd.concat([df, repeated_140], ignore_index=True)

场景2:直接替换原148行的内容(保持行索引位置不变)

如果需要保留原DataFrame的行结构,只是把148行的内容换成140的,可以这样:

import pandas as pd

# 1. 筛选140的行和148行的索引
df_140 = df[df['day_of_year'] == 140].copy()
idx_148 = df[df['day_of_year'] == 148].index
needed_count = len(idx_148)

# 2. 生成足够的140行数据
repeat_times = (needed_count // len(df_140)) + 1
repeated_140 = pd.concat([df_140]*repeat_times, ignore_index=True).head(needed_count)

# 3. 直接替换对应索引的行内容
df.loc[idx_148] = repeated_140.values

二、Vaex实现方案

Vaex适合处理大规模数据集(比如你的20万条140行),它的懒加载特性不会占用过多内存,实现步骤如下:

import vaex

# 1. 筛选出day_of_year=140的行
df_140 = df[df.day_of_year == 140]

# 2. 生成1000条140的行(Vaex的repeat方法高效且懒执行)
needed_count = len(df[df.day_of_year == 148])
repeat_times = (needed_count // len(df_140)) + 1
repeated_140 = df_140.repeat(repeat_times).head(needed_count)

# 3. 过滤掉原148的行,再合并新的140行
df_filtered = df[df.day_of_year != 148]
df_final = df_filtered.concat(repeated_140)

补充说明

因为你的day_of_year=140有20万条,远多于需要的1000条,所以上面代码里的repeat_times实际会是1,直接取前1000条140的行就够了,写通用写法是为了兼容140行不足的情况。

内容的提问来源于stack exchange,提问作者Valdemar sousa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 18:17:32