Python实现CSV按组将多行描述合并为单个字符串
嘿,作为Python新手碰到这种CSV格式的“碎行”问题太正常啦!我帮你搞定这个把多行description合并到单个单元格的需求~
问题分析
你的CSV里,同一记录的description被拆分成了多行,而且还有空行间隔——比如bundy这条记录的description其实是good dog smells kind of weird needs to be washed,我们要把这些分散的行合并成一个完整的字段,同时忽略空行,让每条记录只占一行。
解决方案代码
下面是直接能用的Python代码,我会逐段解释逻辑:
import csv # 替换成你的输入输出文件路径 input_path = "your_raw_data.csv" output_path = "cleaned_data.csv" current_entry = None cleaned_data = [] with open(input_path, "r", newline="", encoding="utf-8") as infile: reader = csv.reader(infile) # 先读取表头,加入清理后的数据 header = next(reader) cleaned_data.append(header) for row in reader: # 跳过完全空的行(所有字段都无内容) if not any(field.strip() for field in row): continue name, date, desc = row # 判断当前行是不是新记录的开头:name或date有内容 if name.strip() or date.strip(): # 如果之前有正在构建的记录,先把它整理好加入结果 if current_entry is not None: # 把多行description拼接成一个字符串(这里用空格分隔,可换成'\n'保留换行) merged_desc = " ".join(current_entry[2]) cleaned_data.append([current_entry[0], current_entry[1], merged_desc]) # 初始化新的记录,用列表存description的多行内容 current_entry = [name.strip(), date.strip(), [desc.strip()]] else: # 这行是当前记录的description续行,追加到列表里 if current_entry is not None and desc.strip(): current_entry[2].append(desc.strip()) # 别忘了处理最后一条记录(循环结束后可能还有未保存的) if current_entry is not None: merged_desc = " ".join(current_entry[2]) cleaned_data.append([current_entry[0], current_entry[1], merged_desc]) # 把清理后的数据写入新CSV with open(output_path, "w", newline="", encoding="utf-8") as outfile: writer = csv.writer(outfile) writer.writerows(cleaned_data)
关键细节说明
- 跳过空行:用
if not any(field.strip() for field in row)判断行是否完全为空,直接跳过,避免干扰。 - 记录拼接逻辑:用
current_entry变量跟踪正在构建的记录,把分散的description先存在列表里,最后统一拼接。 - 自定义分隔符:如果想保留description里的换行(而不是空格),把代码里的
" ".join(...)改成"\n".join(...)就行。 - 边界处理:循环结束后要单独处理最后一条记录,不然会漏掉它。
运行这段代码后,你的输出CSV就会变成这样:
name,date,description bundy,12-12-2017,good dog smells kind of weird needs to be washed
内容的提问来源于stack exchange,提问作者Dylan Smith
相关产品推荐
相关产品推荐

