如何将DataFrame中5或6位字符串列拆分为3-2或3-3的两列
Pandas按字符串长度拆分列为两列(适配5/6位)
方法1:直接字符串切片(高效简洁)
核心逻辑:不管原字符串是5位还是6位,前3位统一拆分到第一列,从第4位开始的剩余部分拆分到第二列,刚好匹配你的需求:
- 5位字符串:
str[:3]取前3位,str[3:]取后2位 - 6位字符串:
str[:3]取前3位,str[3:]取后3位
代码示例:
import pandas as pd # 示例DataFrame df = pd.DataFrame({'code': ['12345', '456789', '78901', '123456']}) # 先将列转为字符串类型(若原列是数字类型必须做这一步) df['code'] = df['code'].astype(str) # 拆分生成新列 df['col1'] = df['code'].str[:3] df['col2'] = df['code'].str[3:]
执行后结果:
code col1 col2 0 12345 123 45 1 456789 456 789 2 78901 789 01 3 123456 123 456
方法2:正则表达式提取(更灵活)
如果后续需要适配更多长度规则,用正则提取更方便,匹配开头3位数字+结尾2-3位数字的模式:
import pandas as pd df = pd.DataFrame({'code': ['12345', '456789', '78901', '123456']}) # 直接拆分并生成两列 df[['col1', 'col2']] = df['code'].astype(str).str.extract(r'^(\d{3})(\d{2,3})$')
注意事项
- 必须确保原列是字符串类型,否则无法使用
.str相关方法;如果原列是数值类型,一定要先通过astype(str)转换。 - 如果数据中存在非5/6位的字符串,需要先过滤(比如
df = df[df['code'].str.len().isin([5,6])]),避免拆分出空值或不符合预期的结果。
内容的提问来源于stack exchange,提问作者datagolfer
相关产品推荐
相关产品推荐

