如何在Python中统计行内指定列的"y"并新增计数列
实现Python中统计指定列内"y"的出现次数并新增列存储结果
方法一:使用Pandas处理结构化表格数据(推荐)
适用于CSV、Excel等结构化数据,代码简洁高效:
import pandas as pd # 读取数据,替换为你的文件路径(支持csv/xlsx等格式) df = pd.read_csv("your_input_data.csv") # 核心操作:统计每行第3-7列(位置索引从0开始,对应索引2到6)中"y"的总次数 # 若列中存在非字符串类型数据,需先转为字符串:row.astype(str).str.count('y') df['y_count'] = df.iloc[:, 2:7].apply(lambda row: row.str.count('y').sum(), axis=1) # 保存结果到新文件,不保留索引列 df.to_csv("output_with_y_count.csv", index=False)
关键代码说明
iloc[:, 2:7]:选中所有行,以及第3到第7列(位置索引左闭右开,包含索引2到6)apply(..., axis=1):对每一行执行统计操作row.str.count('y').sum():对该行每个目标列的元素统计"y"的数量,再求和得到该行总次数df['y_count']:新增一列存储统计结果,列名可自行修改
方法二:纯Python处理文本格式数据
若不想依赖第三方库,可直接处理文本文件(以逗号分隔的CSV为例):
# 读取原始数据文件 with open("your_input_data.txt", "r", encoding="utf-8") as infile: lines = infile.readlines() processed_lines = [] for line in lines: # 分割行数据,按实际分隔符修改(比如制表符用'\t') columns = line.strip().split(',') # 截取第3到第7列(索引2到6) target_columns = columns[2:7] # 统计"y"的出现次数(忽略前后空格) y_total = sum(1 for col in target_columns if col.strip() == "y") # 将统计结果追加为新列 columns.append(str(y_total)) # 重新拼接为行字符串 processed_lines.append(','.join(columns) + '\n') # 写入处理后的结果文件 with open("output_with_y_count.txt", "w", encoding="utf-8") as outfile: outfile.writelines(processed_lines)
内容的提问来源于stack exchange,提问作者Virendra Patel
相关产品推荐
相关产品推荐

