无法将FTP制表符分隔文本文件保存为UTF-8编码的技术求助
老旧嵌入式FTP服务器文件下载与编码处理问题
背景信息
- 无法修改FTP设置,服务器属于嵌入式设备,极为老旧,使用WS_FTP 4.8(距今约20年),不支持PASV、TLS等现代功能;执行
set PASV命令返回502,确认不支持被动模式 - FTP登录提示:
Connection established, waiting for welcome message...
Status: Insecure server, it does not support FTP over TLS.
Status: Server does not support non-ASCII characters.
尝试的解决方法及报错
方法1:二进制模式下载+文本模式写入
代码:
with open(local_temp_file, 'wb', encoding='UTF-8', errors='replace') as local_file: conn.retrbinary('RETR ' + filename_convention + yesterday + '.txt', local_file.write)
FTP日志:
*resp* '200 Type set to I.' *resp* '200 PORT command successful.' *cmd* 'RETR Data Log Trend_Ops_Data_Log_230804.txt' *resp* '150 Opening BINARY mode data connection for Data Log Trend_Ops_Data_Log_230804.txt.'
报错:
{'TypeError'} Traceback (most recent call last): File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 123, in get_log_data conn.retrbinary('RETR ' + filename_convention File "D:\Anaconda\Lib\ftplib.py", line 441, in retrbinary callback(data) TypeError: write() argument must be str, not bytes
提示:write()参数应为str而非bytes
方法2:ASCII模式下载+文本模式写入
代码:
with open(local_temp_file, 'w', encoding='UTF-8', errors='replace') as local_file: conn.retrlines('RETR ' + filename_convention + yesterday + '.txt', local_file.write)
FTP日志:
*resp* '200 Type set to A.' *resp* '200 PORT command successful.' *cmd* 'RETR Data Log Trend_Ops_Data_Log_230804.txt' *resp* '150 Opening ASCII mode data connection for Data Log Trend_Ops_Data_Log_230804.txt.'
报错:
{'UnicodeDecodeError'} Traceback (most recent call last): File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 123, in get_log_data conn.retrlines('RETR ' + filename_convention File "D:\Anaconda\Lib\ftplib.py", line 465, in retrlines line = fp.readline(self.maxline + 1) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "<frozen codecs>", line 322, in decode UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
提示:UTF-8编码无法解码位置0的0xff字节
方法3:尝试编码检测
编码检测函数:
def detect_encoding(file): detector = chardet.universaldetector.UniversalDetector() with open(file, "rb") as f: for line in f: detector.feed(line) if detector.done: break detector.close() return detector.result
错误尝试1:
f = open(local_temp_file, 'wb') conn.retrbinary('RETR ' + filename_convention + yesterday + '.txt', f.write) f.close() f.encode('utf-8') print(detect_encoding(f))
报错:
{'AttributeError'} Traceback (most recent call last): File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 139, in get_log_data f.encode('utf-8') ^^^^^^^^ AttributeError: '_io.BufferedWriter' object has no attribute 'encode'
错误尝试2:
f = open(local_temp_file, 'wb') conn.retrbinary('RETR ' + filename_convention + yesterday + '.txt', f.write) f.close() print(detect_encoding(f))
报错:
{'TypeError'} Traceback (most recent call last): File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 140, in get_log_data print(detect_encoding(f)) ^^^^^^^^^^^^^^^^^^ File "c:\users\justin\onedrive\documents\epic_cleantec_work\ftp log retriever\batch_data_get.py", line 69, in detect_encoding with open(file, "rb") as f: ^^^^^^^^^^^^^^^^ TypeError: expected str, bytes or os.PathLike object, not BufferedWriter
实际问题影响
- 制表符分隔文件上传至S3后,Glue爬虫无法正确识别列,仅显示单列(文件开头有菱形??符号,已设置对应分类器)
- 需要为文件每行追加站点名称和时间戳,但当前无法对文件进行编辑操作
内容的提问来源于stack exchange,提问作者Shenanigator
相关产品推荐
相关产品推荐

