如何在Pandas DataFrame中标记指定行并添加新列
Pandas按行范围新增标记列解决方案
方法一:直接赋值(最直观)
先初始化新列为默认值3,再按索引范围修改对应值:
import pandas as pd # 构造示例DataFrame(可替换为你的数据) df = pd.DataFrame({ 'no.': range(1, 10), 'col1': ['abc', 'bcd', 'cde', 'def', 'efg', 'fgh', 'ghi', 'hij', 'ijk'], 'col2': [123, 234, 345, 456, 567, 678, 789, 890, 901] }) # 新增列并设置默认值 df['new_col'] = 3 # 前4行(索引0-3,对应数据里的第1-4行)设为1 df.loc[:3, 'new_col'] = 1 # 第5-7行(索引4-6,对应数据里的第5-7行)设为2 df.loc[4:6, 'new_col'] = 2
方法二:用numpy.select(多条件场景更灵活)
通过定义条件列表和对应值列表,批量赋值:
import pandas as pd import numpy as np df = pd.DataFrame({ 'no.': range(1, 10), 'col1': ['abc', 'bcd', 'cde', 'def', 'efg', 'fgh', 'ghi', 'hij', 'ijk'], 'col2': [123, 234, 345, 456, 567, 678, 789, 890, 901] }) # 定义条件和对应值 conditions = [ df.index < 4, # 前4行 df.index.between(4,6) # 第5-7行 ] values = [1, 2] # 生成新列,不满足条件的默认设为3 df['new_col'] = np.select(conditions, values, default=3)
方法三:用pd.cut(区间划分场景适用)
把索引按区间分段,直接映射到对应标记:
import pandas as pd import numpy as np df = pd.DataFrame({ 'no.': range(1, 10), 'col1': ['abc', 'bcd', 'cde', 'def', 'efg', 'fgh', 'ghi', 'hij', 'ijk'], 'col2': [123, 234, 345, 456, 567, 678, 789, 890, 901] }) # 按索引区间划分,bins参数对应区间边界,labels对应标记值 df['new_col'] = pd.cut( df.index, bins=[-1, 4, 7, np.inf], # 区间:[-1,4), [4,7), [7,无穷) labels=[1, 2, 3], include_lowest=True # 包含左边界 ).astype(int) # 转为整数类型
执行任意一种方法后,都会得到符合需求的结果:前4行标记1,第5-7行标记2,剩余行标记3。
内容的提问来源于stack exchange,提问作者user18334254
相关产品推荐
相关产品推荐

