Python脚本转换WXR为TXT时触发TypeError:write()需str而非bytes
解决WXR转TXT脚本的字节写入错误及非Post类型文章问题
让我们一步步解决你遇到的问题:
首先修复TypeError错误
你遇到的TypeError: write() argument must be str, not bytes是因为:
- 用
'w'模式打开文件时,文件操作对象只接受字符串类型,但你调用了content.encode('utf8')把字符串转成了字节(bytes),导致类型不匹配。
解决这个问题有两种更合理的方式:
方式1(推荐):直接写入字符串,指定文件编码
打开文件时明确指定encoding='utf-8',直接写入content字符串即可,不需要手动encode:
with open(os.path.join(output_dir, post_filename), 'w', encoding='utf-8') as post_file: post_file.write(content if content else "")
方式2:用二进制模式写入字节
如果一定要用encode,就把打开模式改成'wb'(二进制写入):
with open(os.path.join(output_dir, post_filename), 'wb') as post_file: post_file.write(content.encode('utf8') if content else b"")
修复非Post类型文章的运行问题
你的脚本目前还有几个影响非Post类型的bug:
- 重复计数错误:在else分支里错误地再次执行了
post_counter += 1,导致计数混乱 - 未处理空内容/标题:如果某些非Post类型的文章没有content或title,会抛出AttributeError
- 文件名生成逻辑不够健壮:title可能为None,需要处理这种情况
完整修改后的代码
#!/usr/bin/env python """This script converts WXR file to a number of plain text files. WXR stands for "WordPress eXtended RSS", which basically is just a regular XML file. This script extracts entries from the WXR file into plain text files. Output format: article name prefixed by date for posts, article name for pages. Usage: wxr2txt.py filename [-o output_dir]""" import os import re import sys from xml.etree import ElementTree NAMESPACES = { 'content': 'http://purl.org/rss/1.0/modules/content/', 'wp': 'http://wordpress.org/export/1.2/', } USAGE_STRING = "Usage: wxr2txt.py filename [-o output_dir]" def main(argv): filename, output_dir = _parse_and_validate_output(argv) try: data = ElementTree.parse(filename).getroot() except ElementTree.ParseError: _error("Invalid input file format. Can not parse the input.") page_counter, post_counter = 0, 0 for post in data.find('channel').findall('item'): post_type = post.find('wp:post_type', namespaces=NAMESPACES).text content = post.find('content:encoded', namespaces=NAMESPACES).text date = post.find('wp:post_date', namespaces=NAMESPACES).text title = post.find('title').text # 处理空日期、标题的边界情况 date_str = date.split(' ')[0].replace('-', '') if date else "no_date" # 处理标题为空的情况,生成默认文件名 if title: title_clean = re.sub(r'[_]+', '_', re.sub(r'[^a-z0-9+]', '_', title.lower())) else: title_clean = f"untitled_{post_type}_{post_counter + page_counter + 1}" if post_type == 'post': post_filename = f"{date_str}_{title_clean}.txt" post_counter += 1 else: post_filename = f"{title_clean}.txt" page_counter += 1 # 写入文件:用utf-8编码,处理空content的情况 with open(os.path.join(output_dir, post_filename), 'w', encoding='utf-8') as post_file: post_file.write(content if content else "") print("Saved {} posts and {} pages in directory '{}'.".format( post_counter, page_counter, output_dir)) def _parse_and_validate_output(argv): if len(argv) not in (2, 4): _error("Wrong number of arguments.") filename = argv[1] if not os.path.isfile(filename): _error("Input file does not exist (or not enough permissions).") output_dir = argv[3] if len(argv) == 4 and argv[2] == '-o' else os.getcwd() if not os.path.isdir(output_dir): _error("Output directory does not exist (or not enough permissions).") return filename, output_dir def _error(text): print(text) print(USAGE_STRING) sys.exit(1) if __name__ == "__main__": main(sys.argv)
关键修改点说明:
- 修复了写入时的类型错误,用
'w'模式+指定encoding='utf-8'直接写入字符串 - 移除了else分支里错误的
post_counter += 1,确保计数准确 - 增加了对空日期、空标题、空内容的边界处理,避免脚本崩溃
- 优化了标题为空时的默认文件名生成逻辑,避免重名
内容的提问来源于stack exchange,提问作者Jordi
相关产品推荐
相关产品推荐

