You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理RTF文件:去除每行末尾字符导出CSV为空的解决方法

问题分析

你的代码输出为空主要是这几个原因:

  1. RTF不是纯文本格式:直接用open()读取会拿到大量RTF格式标记(比如{\rtf1\ansi...),而非你需要的有效字符串,导致后续逻辑无法匹配内容。
  2. 逻辑与需求不符:你需要去除每行最后一个字符,但代码仅在最后一个字符是.时才处理,若行尾不是.就直接原封不动写入,和需求矛盾。
  3. 行尾换行符干扰:读取的每行末尾默认带换行符\n,此时line[-1]取到的是换行符而非目标字符,导致判断条件永远不成立。
  4. 冗余的文件关闭操作:with语句会自动管理文件的打开和关闭,手动调用close()属于多余操作,甚至可能引发异常。
解决方案

根据文件实际类型,分两种场景处理:

场景1:文件是误标为.rtf的纯文本文件

如果文件本质是纯文本(只是后缀改成了rtf),用以下修正后的代码直接处理:

input_file = "usernames.rtf"
output_file = "outputfile.csv"

with open(input_file, "r", encoding="utf-8") as input_f, open(output_file, "w", encoding="utf-8") as output_f:
    for line in input_f:
        stripped_line = line.strip()  # 去除换行符和首尾空白
        if stripped_line:  # 跳过空行
            # 删除最后一个字符,处理单行只有一个字符的情况
            modified_line = stripped_line[:-1] if len(stripped_line) > 1 else ""
            output_f.write(f"{modified_line}\n")

print("处理完成,查看输出文件:", output_file)

场景2:真正的RTF格式文件

如果是标准RTF文件,需要用专门工具提取纯文本,这里使用striprtf库(先执行pip install striprtf安装):

from striprtf.striprtf import rtf_to_text

input_file = "usernames.rtf"
output_file = "outputfile.csv"

# 读取RTF并转换为纯文本
with open(input_file, "r", encoding="utf-8") as f:
    rtf_content = f.read()
plain_text = rtf_to_text(rtf_content)

# 处理每行并写入CSV
with open(output_file, "w", encoding="utf-8") as output_f:
    for line in plain_text.splitlines():
        stripped_line = line.strip()
        if stripped_line:
            modified_line = stripped_line[:-1] if len(stripped_line) > 1 else ""
            output_f.write(f"{modified_line}\n")

print("处理完成,查看输出文件:", output_file)
关键说明
  • 先执行strip()处理,避免换行符、空格等干扰最后一个字符的判断。
  • 加入空行和单行字符的判断,防止出现索引越界错误。
  • 指定utf-8编码,避免中文或特殊字符乱码。

内容的提问来源于stack exchange,提问作者David Guri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 13:33:26