如何在含字符串列的Pandas DataFrame中将数字转为两位浮点数?
搞定Pandas含字符串列的浮点数格式化问题
嗨,我来帮你解决这个格式化难题~你说的没错,applymap()没法直接全表用,因为有Names这个字符串列,但我们可以针对性地用它,同时我也会给你更高效的实现方案,优化你现有的代码。
一、用applymap()的可行方案
既然applymap()是作用于每个元素,那我们可以先筛选出DataFrame里的数值列,只对这些列应用格式化逻辑:
# 先筛选出所有数值类型的列(int/float都包含) numeric_cols = df.select_dtypes(include=['int64', 'float64']).columns # 对数值列统一格式化两位小数 df[numeric_cols] = df[numeric_cols].applymap(lambda x: "{:.2f}".format(x))
⚠️ 注意:如果你已经给金额列加了$符号(比如你代码里把Subtotal列转为带$的字符串),这些列会变成object类型,就没法用数值格式化了。所以正确顺序是先格式化数值,再添加$符号。
二、更优的实现方式(推荐)
你现有代码里用循环concat构建DataFrame,效率偏低,格式化步骤也比较零散。这里给你重构后的代码,逻辑更清晰,运行也更高效:
import pandas as pd # 1. 用列表收集用户输入,比循环concat高效多了 people_count = int(input('How many people ordered? ')) order_data = [] cider_unit_price = 5.50 juice_unit_price = 4.50 for i in range(people_count): name = input(f"Enter the name of Person #{i+1}: ") cider_qty = int(input(f"How many orders of cider did {name} have? ")) juice_qty = int(input(f"How many orders of juice did {name} have? ")) # 计算各项金额,直接保留两位小数 cider_sub = round(cider_qty * cider_unit_price, 2) juice_sub = round(juice_qty * juice_unit_price, 2) total = round(cider_sub + juice_sub, 2) order_data.append([name, cider_qty, juice_qty, cider_sub, juice_sub, total]) # 2. 一次性创建完整的DataFrame df = pd.DataFrame( order_data, columns=["Names", "Cider", "Juice", "Subtotal(Cider)", "Subtotal(Juice)", "Total"] ) # 3. 添加Total和Average行,单独构建更直观 total_row = pd.DataFrame({ "Names": "Total", "Cider": df["Cider"].sum(), "Juice": df["Juice"].sum(), "Subtotal(Cider)": df["Subtotal(Cider)"].sum(), "Subtotal(Juice)": df["Subtotal(Juice)"].sum(), "Total": df["Total"].sum() }, index=[people_count]) average_row = pd.DataFrame({ "Names": "Average", "Cider": df["Cider"].mean(), "Juice": df["Juice"].mean(), "Subtotal(Cider)": df["Subtotal(Cider)"].mean(), "Subtotal(Juice)": df["Subtotal(Juice)"].mean(), "Total": df["Total"].mean() }, index=[people_count+1]) df = pd.concat([df, total_row, average_row]) # 4. 统一格式化:先处理金额列(加$+两位小数),再处理数量列 for col in ["Subtotal(Cider)", "Subtotal(Juice)", "Total"]: df[col] = df[col].apply(lambda x: f"$ {x:.2f}") df[["Cider", "Juice"]] = df[["Cider", "Juice"]].applymap(lambda x: f"{x:.2f}") # 5. 设置索引 df.set_index("Names", inplace=True) print(df)
优化点说明:
- 用列表收集数据再一次性创建DataFrame,避免循环
concat的性能损耗 - 统一处理数值格式化,逻辑更集中,不容易出错
- 用
f-string简化字符串拼接和格式化,代码更简洁 - 单独构建Total和Average行,可读性更强
三、可选方案:用Styler做展示层格式化
如果只是需要把DataFrame格式化后展示(比如在Jupyter Notebook里),可以用Pandas的Styler功能,不会修改原数据的类型,只做展示层的美化:
# 定义展示用的格式化规则 styled_df = df.style.format({ "Subtotal(Cider)": lambda x: f"$ {x:.2f}", "Subtotal(Juice)": lambda x: f"$ {x:.2f}", "Total": lambda x: f"$ {x:.2f}", "Cider": "{:.2f}", "Juice": "{:.2f}" }) display(styled_df) # 在Notebook中显示格式化后的表格
这个方案适合只需要美观展示,不需要修改原数据的场景。
内容的提问来源于stack exchange,提问作者Meruemu
相关产品推荐
相关产品推荐

