You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按数字块拆分DataFrame列并转换为多行?

拆分DataFrame列并展开多行的实现方案

我们可以用pandas结合正则表达式来实现你的需求,步骤清晰易懂,适合新手操作:

步骤1:导入必要的库

import pandas as pd
import re

步骤2:构造示例DataFrame

先把你提供的测试数据转换成DataFrame,方便后续操作:

data = {
    "Column A": ["Cell 1", "Cell 3"],
    "Column B": ["1234 abcd 667 randomthings", "4455 abcd abc 847 other randomthings 1 endings"]
}
df = pd.DataFrame(data)

步骤3:定义拆分函数

用正则表达式提取所有以数字开头的片段,直到下一个数字出现或字符串结束:

def split_col_b(text):
    # 正则匹配:数字开头,后续内容直到下一个空格+数字或结尾
    parts = re.findall(r'\d+.*?(?=\s+\d+|$)', text)
    # 去除每个片段前后的多余空格
    return [part.strip() for part in parts]

步骤4:应用拆分并展开多行

对Column B应用拆分函数,再用explode把列表拆成多行:

# 生成拆分后的列表列
df['Column B'] = df['Column B'].apply(split_col_b)
# 展开列表为多行
result_df = df.explode('Column B').reset_index(drop=True)

查看结果

运行后result_df就是你想要的格式:

print(result_df)

输出结果:

Column A               Column B
0   Cell 1            1234 abcd
1   Cell 1        667 randomthings
2   Cell 3       4455 abcd abc
3   Cell 3  847 other randomthings
4   Cell 3              1 endings

关键说明

  • 正则表达式r'\d+.*?(?=\s+\d+|$)':\d+匹配连续数字,.*?非贪婪匹配后续内容,(?=\s+\d+|$)是正向预查,确保在遇到空格+数字或字符串结尾时停止匹配,精准拆分每个数字开头的块。
  • explode方法是pandas专门用来把列表类型的列展开成多行的工具,会自动保留对应行的Column A值。

内容的提问来源于stack exchange,提问作者Samiiir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 20:22:14