You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法将FTP制表符分隔文本文件保存为UTF-8编码的技术求助

老旧嵌入式FTP服务器文件下载与编码处理问题

背景信息

  • 无法修改FTP设置,服务器属于嵌入式设备,极为老旧,使用WS_FTP 4.8(距今约20年),不支持PASV、TLS等现代功能;执行set PASV命令返回502,确认不支持被动模式
  • FTP登录提示:

Connection established, waiting for welcome message...
Status: Insecure server, it does not support FTP over TLS.
Status: Server does not support non-ASCII characters.

尝试的解决方法及报错

方法1:二进制模式下载+文本模式写入

代码:

with open(local_temp_file, 'wb', encoding='UTF-8', errors='replace') as local_file:
    conn.retrbinary('RETR ' + filename_convention
                    + yesterday + '.txt', local_file.write)

FTP日志:

*resp* '200 Type set to I.'
*resp* '200 PORT command successful.'
*cmd* 'RETR Data Log Trend_Ops_Data_Log_230804.txt'
*resp* '150 Opening BINARY mode data connection for Data Log Trend_Ops_Data_Log_230804.txt.'

报错:

{'TypeError'}
Traceback (most recent call last):
  File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 123, in get_log_data
    conn.retrbinary('RETR ' + filename_convention
  File "D:\Anaconda\Lib\ftplib.py", line 441, in retrbinary
    callback(data)
TypeError: write() argument must be str, not bytes

提示:write()参数应为str而非bytes

方法2:ASCII模式下载+文本模式写入

代码:

with open(local_temp_file, 'w', encoding='UTF-8', errors='replace') as local_file:
    conn.retrlines('RETR ' + filename_convention
                   + yesterday + '.txt', local_file.write)

FTP日志:

*resp* '200 Type set to A.'
*resp* '200 PORT command successful.'
*cmd* 'RETR Data Log Trend_Ops_Data_Log_230804.txt'
*resp* '150 Opening ASCII mode data connection for Data Log Trend_Ops_Data_Log_230804.txt.'

报错:

{'UnicodeDecodeError'}
Traceback (most recent call last):
  File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 123, in get_log_data
    conn.retrlines('RETR ' + filename_convention
  File "D:\Anaconda\Lib\ftplib.py", line 465, in retrlines
    line = fp.readline(self.maxline + 1)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "<frozen codecs>", line 322, in decode
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte

提示:UTF-8编码无法解码位置0的0xff字节

方法3:尝试编码检测

编码检测函数:

def detect_encoding(file):
    detector = chardet.universaldetector.UniversalDetector()
    with open(file, "rb") as f:
        for line in f:
            detector.feed(line)
            if detector.done:
                break
        detector.close()
    return detector.result

错误尝试1:

f = open(local_temp_file, 'wb')
conn.retrbinary('RETR ' + filename_convention
                + yesterday + '.txt', f.write)
f.close()
f.encode('utf-8')
print(detect_encoding(f))

报错:

{'AttributeError'}
Traceback (most recent call last):
  File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 139, in get_log_data
    f.encode('utf-8')
    ^^^^^^^^
AttributeError: '_io.BufferedWriter' object has no attribute 'encode'

错误尝试2:

f = open(local_temp_file, 'wb')
conn.retrbinary('RETR ' + filename_convention
                + yesterday + '.txt', f.write)
f.close()
print(detect_encoding(f))

报错:

{'TypeError'}
Traceback (most recent call last):
  File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 140, in get_log_data
    print(detect_encoding(f))
          ^^^^^^^^^^^^^^^^^^
  File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 69, in detect_encoding
    with open(file, "rb") as f:
         ^^^^^^^^^^^^^^^^
TypeError: expected str, bytes or os.PathLike object, not BufferedWriter

实际问题影响

  • 制表符分隔文件上传至S3后,Glue爬虫无法正确识别列,仅显示单列(文件开头有菱形??符号,已设置对应分类器)
  • 需要为文件每行追加站点名称和时间戳,但当前无法对文件进行编辑操作

内容的提问来源于stack exchange,提问作者Shenanigator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 05:08:10