You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将长字符串按随机长度均分填充至空行?

需求实现:Pandas长字符串拆分填充空值行

原始DataFrame

string  some_col
0  But were so TESSA tell me a little bit more t ...        10
1                                                           15
2                                                           14
3  Some other text xxxxxxxxxx                               20

需求说明

将string列中的长字符串拆分为随机长度的片段,均等填充至后续的空值行中,处理后效果如下:

处理后效果示例

string  some_col
0   But were so TESSA tell me        10
1        little bit more t a        15
2              you pretty upset        14

可复现代码

import pandas as pd
data = [['But were so TESSA tell me a  you pretty upset.', 10], ['', 15], ['', 14]]
df = pd.DataFrame(data, columns=['string', 'some_col']) 
print(df)

实现步骤及代码

核心步骤

  • 标记分组:将每个非空字符串行与其后续的空值行划分为一个处理组
  • 拆分字符串:把长字符串按随机长度拆分为对应行数的片段(保证每个片段非空)
  • 填充片段:将拆分后的片段替换组内的空值行
  • 整理结果:合并处理后的组并清理临时列

完整实现代码

import pandas as pd
import random

# 初始化测试数据
data = [['But were so TESSA tell me a you pretty upset.', 10], ['', 15], ['', 14], ['Some other text xxxxxxxxxx', 20]]
df = pd.DataFrame(data, columns=['string', 'some_col']) 

# 1. 生成分组标识:将非空行和后续空行归为一组
df['group_id'] = df['string'].ne('').cumsum()
# 过滤掉无后续空行的组(比如示例中的第4行)
target_groups = df.groupby('group_id').filter(lambda x: len(x) > 1)

# 2. 逐个处理分组
processed_groups = []
for _, group in target_groups.groupby('group_id'):
    source_str = group['string'].iloc[0]
    required_parts = len(group)
    
    # 将字符串拆分为单词列表,避免拆分到单词中间
    word_list = source_str.split()
    # 随机生成分割点,确保每个片段至少有一个单词
    split_positions = sorted(random.sample(range(1, len(word_list)), required_parts - 1))
    
    # 分割单词列表并拼接为字符串片段
    segments = []
    start_idx = 0
    for pos in split_positions:
        segments.append(' '.join(word_list[start_idx:pos]))
        start_idx = pos
    segments.append(' '.join(word_list[start_idx:]))
    
    # 将片段赋值给组内的string列
    group['string'] = segments
    processed_groups.append(group)

# 3. 合并结果并清理临时列
final_df = pd.concat(processed_groups).drop('group_id', axis=1)
print(final_df)

关键知识点说明

  • df['string'].ne('').cumsum():通过判断字符串是否非空生成累加分组ID,实现非空行与后续空行的分组绑定
  • random.sample(range(1, len(word_list)), required_parts - 1):从单词索引中随机选取分割点,保证每个拆分片段至少包含一个完整单词
  • groupby+filter:快速筛选出需要处理的(非空行+后续空行)组合

内容的提问来源于stack exchange,提问作者Bhargav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 20:25:16