如何拆分DataFrame列中带枚举编号的文本并生成新行?
拆分DataFrame中的枚举文本到多行
要实现你要的效果,不能用普通字符串拆分,得用正则表达式匹配枚举编号的分隔符,具体操作如下:
核心拆分规则
在str.split()中填写正则表达式:r'\d+[.)] ',这个规则的含义是:
\d+:匹配1个或多个数字[.)]:匹配右括号)或者点号.- 末尾的空格:匹配编号后的空格分隔符
完整代码示例
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ 'col': [ '1) wake up 2) brush your teeth 3) go to school', '1. save 2. download', 'this is a great day' ] }) # 拆分并展开为多行 df = df['col'].str.split(r'\d+[.)] ', expand=True).stack().reset_index(drop=True).to_frame('col') # 清理空内容和多余空格 df = df[df['col'].str.strip() != ''].reset_index(drop=True) df['col'] = df['col'].str.strip() print(df)
运行后就能得到你期望的结果:
col 0 wake up 1 brush your teeth 2 go to school 3 save 4 download 5 this is a great day
内容的提问来源于stack exchange,提问作者gh1222
相关产品推荐
相关产品推荐

