Python中如何从DataFrame整数列提取前4位数字及月份?求助解决提取结果错误问题
问题分析与解决方案
你遇到的问题核心是错误地将整个Series对象转为字符串,而不是对Series中的每个元素单独转换。str(dvd['CalendarYearMonth'])会把整个Series变成类似"0 202108\n1 202109\n..."的字符串,再切片自然得不到每个元素的年份和月份。
下面提供两种可靠的解决方法:
方法1:字符串切片法(直观易懂)
先把整数列转为字符串类型的Series,再用Pandas的str访问器对每个元素做切片:
# 提取年份(前4位) dvd['yy'] = dvd['CalendarYearMonth'].astype(str).str[:4] # 提取月份(后2位) dvd['mon'] = dvd['CalendarYearMonth'].astype(str).str[-2:]
如果需要月份保持两位数字格式(比如"08"而不是"8"),这个方法天生就能满足,因为原始整数转字符串后是6位,切片后直接是两位。
方法2:数学运算法(效率更高)
因为你的CalendarYearMonth是整数格式(比如202108 = 2021*100 + 8),可以用整除和取模运算直接提取,不需要转字符串,大数据量下性能更好:
# 提取年份:整除100,得到前4位的整数 dvd['yy'] = dvd['CalendarYearMonth'] // 100 # 提取月份:取模100,得到后2位的整数 dvd['mon'] = dvd['CalendarYearMonth'] % 100
如果需要月份是两位字符串格式,可以再补零:
dvd['mon'] = (dvd['CalendarYearMonth'] % 100).astype(str).str.zfill(2)
验证结果
用上述两种方法处理你的示例数据后,会得到符合预期的输出:
| CalendarYearMonth | CountryCode | Dividends | yy | mon |
|---|---|---|---|---|
| 202108 | CN | 196.0 | 2021 | 08 |
| 202109 | CN | 380.0 | 2021 | 09 |
| 202108 | IN | NaN | 2021 | 08 |
| 202109 | IN | 115.0 | 2021 | 09 |
内容的提问来源于stack exchange,提问作者Ashish
相关产品推荐
相关产品推荐

