如何基于Source列提取CSV中6个唯一源的最新行?
提取CSV中每个Source对应的最新行
用Python pandas实现的步骤:
- 先导入工具并读取CSV文件:
import pandas as pd # 替换成你的CSV文件路径 df = pd.read_csv('your_data.csv')
- 处理时间列(用时间判断“最新”的前提):
假设你的数据里有记录时间的列(比如叫record_time),先把它转成时间格式,避免排序逻辑出错:
df['record_time'] = pd.to_datetime(df['record_time'])
- 分组提取每个Source的最新行:
先按时间从新到旧排序,再按Source分组取每组第一行,结果直接存在latest_rows变量里,刚好对应6个唯一Source的最新数据:
# 按时间倒序排序后,分组取每个Source的第一条(即最新行) latest_rows = df.sort_values('record_time', ascending=False).groupby('Source').first().reset_index()
- 高亮指定行:
如果要标记某一行(比如Source为WebAPI的行),可以这样操作:
# 遍历打印时高亮目标行 for _, row in latest_rows.iterrows(): if row['Source'] == 'WebAPI': print(f"★ 高亮行:{row.to_dict()}") else: print(row.to_dict())
要是在Jupyter环境里想给表格加背景色高亮:
def highlight_target(row): return ['background-color: #ffffcc' if row['Source'] == 'WebAPI' else '' for _ in row] # 生成带高亮样式的表格 styled_df = latest_rows.style.apply(highlight_target, axis=1) display(styled_df)
内容的提问来源于stack exchange,提问作者iceman
相关产品推荐
相关产品推荐

