如何批量截取Pandas DataFrame字符串并保留两位小数?
Pandas批量处理字符串数值:截取并四舍五入
问题描述
我有一个名为country的Pandas DataFrame,原始数据如下:
Booking date Country1 Country2 Country3 Country 4 2023-07-08T00:00:00.000 NaN NaN 129.6119.7449.3519.7439.4819 13.018.614 2023-07-89T00:00:00.000 NaN NaN 19.7439.4849.3516.09.8739.4834.4819.7419 67.518.616.557.629
希望处理后得到如下格式:
Booking date Country1 Country2 Country3 Country4 2023-07-08T00:00:00.000 NaN NaN 129.61 13.02 2023-07-89T00:00:00.000 NaN NaN 19.74 67.52
需求:将DataFrame中每个字符串截取到第一个小数点后三位,再四舍五入保留两位小数。之前尝试过country['Country1'].str[:5]实现单列操作,但无法对整个DataFrame批量处理,请问该如何实现?
解决方案
要实现批量处理,不用对每一列单独编写代码,我们可以结合Pandas的批量操作和正则表达式来完成:
1. 统一列名(可选但推荐)
原始数据中Country 4列名包含空格,先统一列名格式避免后续操作出错:
country.columns = country.columns.str.replace(' ', '')
2. 批量处理所有目标列
这里提供两种实用方法:
方法一:自定义函数 + applymap
通过自定义处理函数,用applymap批量应用到所有目标列:
import pandas as pd def clean_numeric_string(val): # 保留NaN值 if pd.isna(val): return val # 提取第一个"整数部分.三位小数"的片段 extracted_segment = pd.Series(val).str.extract(r'^(\d+\.\d{3})')[0].iloc[0] if extracted_segment: # 转换为浮点数后四舍五入保留两位小数 return round(float(extracted_segment), 2) return val # 选择除Bookingdate外的所有列进行处理 target_cols = country.columns.drop('Bookingdate') country[target_cols] = country[target_cols].applymap(clean_numeric_string)
方法二:正则提取 + 批量数值转换(更简洁)
直接遍历目标列,用正则提取片段后转换为数值并四舍五入:
# 遍历所有需要处理的列 for col in country.columns.drop('Bookingdate'): # 提取第一个有效数值片段,转浮点数后四舍五入 country[col] = country[col].str.extract(r'^(\d+\.\d{3})')[0].astype(float).round(2)
两种方法都能满足需求:保留原始NaN值,将每个字符串截取到第一个小数点后三位,最终四舍五入保留两位小数。
内容的提问来源于stack exchange,提问作者Victor Roos
相关产品推荐
相关产品推荐

