You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本转换WXR为TXT时触发TypeError:write()需str而非bytes

解决WXR转TXT脚本的字节写入错误及非Post类型文章问题

让我们一步步解决你遇到的问题:

首先修复TypeError错误

你遇到的TypeError: write() argument must be str, not bytes是因为:

  • 用'w'模式打开文件时,文件操作对象只接受字符串类型,但你调用了content.encode('utf8')把字符串转成了字节(bytes),导致类型不匹配。

解决这个问题有两种更合理的方式:

方式1(推荐):直接写入字符串,指定文件编码

打开文件时明确指定encoding='utf-8',直接写入content字符串即可,不需要手动encode:

with open(os.path.join(output_dir, post_filename), 'w', encoding='utf-8') as post_file:
    post_file.write(content if content else "")

方式2:用二进制模式写入字节

如果一定要用encode,就把打开模式改成'wb'(二进制写入):

with open(os.path.join(output_dir, post_filename), 'wb') as post_file:
    post_file.write(content.encode('utf8') if content else b"")

修复非Post类型文章的运行问题

你的脚本目前还有几个影响非Post类型的bug:

  1. 重复计数错误:在else分支里错误地再次执行了post_counter += 1,导致计数混乱
  2. 未处理空内容/标题:如果某些非Post类型的文章没有content或title,会抛出AttributeError
  3. 文件名生成逻辑不够健壮:title可能为None,需要处理这种情况

完整修改后的代码

#!/usr/bin/env python
"""This script converts WXR file to a number of plain text files. WXR stands for "WordPress eXtended RSS", which basically is just a regular XML file. This script extracts entries from the WXR file into plain text files. Output format: article name prefixed by date for posts, article name for pages. Usage: wxr2txt.py filename [-o output_dir]"""
import os
import re
import sys
from xml.etree import ElementTree

NAMESPACES = {
    'content': 'http://purl.org/rss/1.0/modules/content/',
    'wp': 'http://wordpress.org/export/1.2/',
}

USAGE_STRING = "Usage: wxr2txt.py filename [-o output_dir]"

def main(argv):
    filename, output_dir = _parse_and_validate_output(argv)
    try:
        data = ElementTree.parse(filename).getroot()
    except ElementTree.ParseError:
        _error("Invalid input file format. Can not parse the input.")
    page_counter, post_counter = 0, 0
    for post in data.find('channel').findall('item'):
        post_type = post.find('wp:post_type', namespaces=NAMESPACES).text
        content = post.find('content:encoded', namespaces=NAMESPACES).text
        date = post.find('wp:post_date', namespaces=NAMESPACES).text
        title = post.find('title').text

        # 处理空日期、标题的边界情况
        date_str = date.split(' ')[0].replace('-', '') if date else "no_date"
        # 处理标题为空的情况,生成默认文件名
        if title:
            title_clean = re.sub(r'[_]+', '_', re.sub(r'[^a-z0-9+]', '_', title.lower()))
        else:
            title_clean = f"untitled_{post_type}_{post_counter + page_counter + 1}"

        if post_type == 'post':
            post_filename = f"{date_str}_{title_clean}.txt"
            post_counter += 1
        else:
            post_filename = f"{title_clean}.txt"
            page_counter += 1

        # 写入文件:用utf-8编码,处理空content的情况
        with open(os.path.join(output_dir, post_filename), 'w', encoding='utf-8') as post_file:
            post_file.write(content if content else "")

    print("Saved {} posts and {} pages in directory '{}'.".format(
        post_counter, page_counter, output_dir))

def _parse_and_validate_output(argv):
    if len(argv) not in (2, 4):
        _error("Wrong number of arguments.")
    filename = argv[1]
    if not os.path.isfile(filename):
        _error("Input file does not exist (or not enough permissions).")
    output_dir = argv[3] if len(argv) == 4 and argv[2] == '-o' else os.getcwd()
    if not os.path.isdir(output_dir):
        _error("Output directory does not exist (or not enough permissions).")
    return filename, output_dir

def _error(text):
    print(text)
    print(USAGE_STRING)
    sys.exit(1)

if __name__ == "__main__":
    main(sys.argv)

关键修改点说明:

  • 修复了写入时的类型错误,用'w'模式+指定encoding='utf-8'直接写入字符串
  • 移除了else分支里错误的post_counter += 1,确保计数准确
  • 增加了对空日期、空标题、空内容的边界处理,避免脚本崩溃
  • 优化了标题为空时的默认文件名生成逻辑,避免重名

内容的提问来源于stack exchange,提问作者Jordi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:47:28