如何在Pandas列中实现句号后新句子首字母大写?
Pandas DataFrame文本句子首字母批量大写方案
不用手动遍历行,直接用Pandas的向量化字符串操作配合正则表达式就能搞定,高效且简洁。
实现代码
假设你的DataFrame是df,目标列为Description:
import pandas as pd # 一步到位处理:先转全小写,再用正则替换实现句子首字母大写 df['Description'] = df['Description'].str.lower().str.replace( r'(^|[.!?]\s)([a-z])', lambda match: match.group(1) + match.group(2).upper(), regex=True )
逻辑解释
str.lower():先把整段文本转为小写,解决原文本中存在的无意义大写(比如示例里的VERY会变成very)。- 正则匹配规则:
(^|[.!?]\s):匹配两种句子起始位置——文本开头,或者句号/感叹号/问号加空格之后的位置,这部分会原封不动保留。([a-z]):匹配起始位置后的第一个小写字母,作为要转换的目标。
- lambda替换函数:把匹配到的第一个小写字母转为大写,和前面保留的部分拼接,实现每个句子首字母大写的效果。
测试验证
输入示例文本:
"This product has several different features. it is also VERY cost effective. it is one of my favorite products."
处理后输出:
"This product has several different features. It is also very cost effective. It is one of my favorite products."
完全符合需求,而且全程是Pandas的向量化操作,不需要手动循环行,性能更优。
内容的提问来源于stack exchange,提问作者periclesrocha
相关产品推荐
相关产品推荐

