如何用Pandas/Vaex复制day_of_year=140行替换148行?
用Pandas和Vaex实现替换DataFrame行的方案
没问题,我来给你详细拆解在Pandas和Vaex里怎么实现这个需求——把day_of_year=148的1000条记录,全部替换成从day_of_year=140的记录中复制的内容:
一、Pandas实现方案
Pandas处理中小规模数据很灵活,这里分两种常见场景实现:
场景1:删除原148行,添加1000条140行(保持总行数不变)
import pandas as pd # 1. 筛选出day_of_year=140的行,用copy()避免视图修改问题 df_140 = df[df['day_of_year'] == 140].copy() # 2. 生成刚好1000条140的行(通用写法:即使140的行不足1000,也能重复凑够数量) needed_count = len(df[df['day_of_year'] == 148]) # 这里就是1000 repeat_times = (needed_count // len(df_140)) + 1 repeated_140 = pd.concat([df_140]*repeat_times, ignore_index=True).head(needed_count) # 3. 删除原148的行,再添加新的140行 df = df[df['day_of_year'] != 148].reset_index(drop=True) df = pd.concat([df, repeated_140], ignore_index=True)
场景2:直接替换原148行的内容(保持行索引位置不变)
如果需要保留原DataFrame的行结构,只是把148行的内容换成140的,可以这样:
import pandas as pd # 1. 筛选140的行和148行的索引 df_140 = df[df['day_of_year'] == 140].copy() idx_148 = df[df['day_of_year'] == 148].index needed_count = len(idx_148) # 2. 生成足够的140行数据 repeat_times = (needed_count // len(df_140)) + 1 repeated_140 = pd.concat([df_140]*repeat_times, ignore_index=True).head(needed_count) # 3. 直接替换对应索引的行内容 df.loc[idx_148] = repeated_140.values
二、Vaex实现方案
Vaex适合处理大规模数据集(比如你的20万条140行),它的懒加载特性不会占用过多内存,实现步骤如下:
import vaex # 1. 筛选出day_of_year=140的行 df_140 = df[df.day_of_year == 140] # 2. 生成1000条140的行(Vaex的repeat方法高效且懒执行) needed_count = len(df[df.day_of_year == 148]) repeat_times = (needed_count // len(df_140)) + 1 repeated_140 = df_140.repeat(repeat_times).head(needed_count) # 3. 过滤掉原148的行,再合并新的140行 df_filtered = df[df.day_of_year != 148] df_final = df_filtered.concat(repeated_140)
补充说明
因为你的day_of_year=140有20万条,远多于需要的1000条,所以上面代码里的repeat_times实际会是1,直接取前1000条140的行就够了,写通用写法是为了兼容140行不足的情况。
内容的提问来源于stack exchange,提问作者Valdemar sousa
相关产品推荐
相关产品推荐

