如何在DataFrame中按指定长度拆分长字符串并生成新行(保留完整单词)
解决方案
下面分两种阈值场景给出实现方法,均保证在完整单词处拆分,不会截断单词:
一、按单词数(9个单词)拆分
先写一个拆分函数,判断字符串的单词数量,超过9个则在第9个单词后拆分为两段;不足阈值则保持原字符串:
import pandas as pd def split_by_word_count(s, threshold=9): words = s.strip("'").split() # 去除示例中的单引号并拆分单词 if len(words) <= threshold: return [f"'{' '.join(words)}'"] # 拆分为前9个单词和剩余部分 part1 = f"'{' '.join(words[:threshold])}'" part2 = f"'{' '.join(words[threshold:])}'" return [part1, part2] # 构建示例数据 df = pd.DataFrame({ 'var1': ["'this is a long string that should break here but it keeps going'", "'this string is a fine length'"], 'var2': [1, 2], 'var3': ['a', 'b'] }) # 执行拆分并展开结果 result = df.assign(var1=df['var1'].apply(split_by_word_count)).explode('var1').reset_index(drop=True) print(result)
运行后输出与示例一致:
var1 var2 var3 0 'this is a long string that should break' 1 a 1 'here but it keeps going' 1 a 2 'this string is a fine length' 2 b
二、按字符数(80字符)拆分
编写函数找到80字符范围内最后一个空格的位置进行拆分,避免截断单词:
def split_by_char_length(s, threshold=80): s_clean = s.strip("'") if len(s_clean) <= threshold: return [f"'{s_clean}'"] # 定位阈值内最后一个空格的位置 split_pos = s_clean.rfind(' ', 0, threshold) # 极端情况:无空格则直接按阈值拆分(可按需调整逻辑) if split_pos == -1: split_pos = threshold part1 = f"'{s_clean[:split_pos]}'" part2 = f"'{s_clean[split_pos+1:]}'" return [part1, part2] # 应用函数处理 result_char = df.assign(var1=df['var1'].apply(split_by_char_length)).explode('var1').reset_index(drop=True) print(result_char)
对你现有代码的修改思路
你之前的代码是按特定字符(逗号)拆分,核心逻辑是str.split()+explode(),只需把拆分逻辑替换成自定义函数即可,比如:
result = df.set_index(['var2', 'var3']) .assign(var1=lambda x: x['var1'].apply(split_by_word_count)) .explode('var1') .reset_index()
这样就能复用你原有的索引操作逻辑,同时实现按单词/字符拆分的需求。
内容的提问来源于stack exchange,提问作者joules
相关产品推荐
相关产品推荐

