如何用Pandas将长字符串按随机长度均分填充至空行?
需求实现:Pandas长字符串拆分填充空值行
原始DataFrame
string some_col 0 But were so TESSA tell me a little bit more t ... 10 1 15 2 14 3 Some other text xxxxxxxxxx 20
需求说明
将string列中的长字符串拆分为随机长度的片段,均等填充至后续的空值行中,处理后效果如下:
处理后效果示例
string some_col 0 But were so TESSA tell me 10 1 little bit more t a 15 2 you pretty upset 14
可复现代码
import pandas as pd data = [['But were so TESSA tell me a you pretty upset.', 10], ['', 15], ['', 14]] df = pd.DataFrame(data, columns=['string', 'some_col']) print(df)
实现步骤及代码
核心步骤
- 标记分组:将每个非空字符串行与其后续的空值行划分为一个处理组
- 拆分字符串:把长字符串按随机长度拆分为对应行数的片段(保证每个片段非空)
- 填充片段:将拆分后的片段替换组内的空值行
- 整理结果:合并处理后的组并清理临时列
完整实现代码
import pandas as pd import random # 初始化测试数据 data = [['But were so TESSA tell me a you pretty upset.', 10], ['', 15], ['', 14], ['Some other text xxxxxxxxxx', 20]] df = pd.DataFrame(data, columns=['string', 'some_col']) # 1. 生成分组标识:将非空行和后续空行归为一组 df['group_id'] = df['string'].ne('').cumsum() # 过滤掉无后续空行的组(比如示例中的第4行) target_groups = df.groupby('group_id').filter(lambda x: len(x) > 1) # 2. 逐个处理分组 processed_groups = [] for _, group in target_groups.groupby('group_id'): source_str = group['string'].iloc[0] required_parts = len(group) # 将字符串拆分为单词列表,避免拆分到单词中间 word_list = source_str.split() # 随机生成分割点,确保每个片段至少有一个单词 split_positions = sorted(random.sample(range(1, len(word_list)), required_parts - 1)) # 分割单词列表并拼接为字符串片段 segments = [] start_idx = 0 for pos in split_positions: segments.append(' '.join(word_list[start_idx:pos])) start_idx = pos segments.append(' '.join(word_list[start_idx:])) # 将片段赋值给组内的string列 group['string'] = segments processed_groups.append(group) # 3. 合并结果并清理临时列 final_df = pd.concat(processed_groups).drop('group_id', axis=1) print(final_df)
关键知识点说明
df['string'].ne('').cumsum():通过判断字符串是否非空生成累加分组ID,实现非空行与后续空行的分组绑定random.sample(range(1, len(word_list)), required_parts - 1):从单词索引中随机选取分割点,保证每个拆分片段至少包含一个完整单词groupby+filter:快速筛选出需要处理的(非空行+后续空行)组合
内容的提问来源于stack exchange,提问作者Bhargav
相关产品推荐
相关产品推荐

