You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量截取Pandas DataFrame字符串并保留两位小数?

Pandas批量处理字符串数值:截取并四舍五入

问题描述

我有一个名为country的Pandas DataFrame,原始数据如下:

Booking date            Country1 Country2 Country3                                 Country 4
2023-07-08T00:00:00.000 NaN      NaN      129.6119.7449.3519.7439.4819             13.018.614
2023-07-89T00:00:00.000 NaN      NaN      19.7439.4849.3516.09.8739.4834.4819.7419 67.518.616.557.629

希望处理后得到如下格式:

Booking date            Country1 Country2 Country3 Country4
2023-07-08T00:00:00.000 NaN      NaN      129.61   13.02
2023-07-89T00:00:00.000 NaN      NaN      19.74    67.52

需求:将DataFrame中每个字符串截取到第一个小数点后三位,再四舍五入保留两位小数。之前尝试过country['Country1'].str[:5]实现单列操作,但无法对整个DataFrame批量处理,请问该如何实现?

解决方案

要实现批量处理,不用对每一列单独编写代码,我们可以结合Pandas的批量操作和正则表达式来完成:

1. 统一列名(可选但推荐)

原始数据中Country 4列名包含空格,先统一列名格式避免后续操作出错:

country.columns = country.columns.str.replace(' ', '')

2. 批量处理所有目标列

这里提供两种实用方法:

方法一:自定义函数 + applymap

通过自定义处理函数,用applymap批量应用到所有目标列:

import pandas as pd

def clean_numeric_string(val):
    # 保留NaN值
    if pd.isna(val):
        return val
    # 提取第一个"整数部分.三位小数"的片段
    extracted_segment = pd.Series(val).str.extract(r'^(\d+\.\d{3})')[0].iloc[0]
    if extracted_segment:
        # 转换为浮点数后四舍五入保留两位小数
        return round(float(extracted_segment), 2)
    return val

# 选择除Bookingdate外的所有列进行处理
target_cols = country.columns.drop('Bookingdate')
country[target_cols] = country[target_cols].applymap(clean_numeric_string)

方法二:正则提取 + 批量数值转换(更简洁)

直接遍历目标列,用正则提取片段后转换为数值并四舍五入:

# 遍历所有需要处理的列
for col in country.columns.drop('Bookingdate'):
    # 提取第一个有效数值片段,转浮点数后四舍五入
    country[col] = country[col].str.extract(r'^(\d+\.\d{3})')[0].astype(float).round(2)

两种方法都能满足需求:保留原始NaN值,将每个字符串截取到第一个小数点后三位,最终四舍五入保留两位小数。


内容的提问来源于stack exchange,提问作者Victor Roos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 12:25:36