You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何将分词得到的tokens列表写入文本文件?

How to Write Full Token Lists to a Text File (Instead of Individual Items) in Python

Problem Description

I’ve written this Python code to fetch text from a URL and tokenize it using CrazyTokenizer:

from urllib import request
from redditscore.tokenizer import CrazyTokenizer
tokenizer = CrazyTokenizer()
url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt"
for line in request.urlopen(url):
    tokens = tokenizer.tokenize(line.decode('utf-8'))
    #print(tokens)
    with open('your_file.txt', 'a') as f:
        print(tokens)
        for item in tokens:
            f.write("%s\n" % item)

Right now, each token is written to a new line in the output file. But I want to save the complete list object (e.g., ['\ufeffsave', 'bbc', 'world', 'service', 'from', 'savage', 'cuts']) to the file, with one list per line instead. How can I adjust my code to achieve this?


Solution

The simplest way to write the full list to your file is to convert the tokens list directly to its string representation, then write that string plus a newline. Here are a few straightforward approaches:

1. Use str() for a human-readable list string

Replace the loop that writes individual tokens with a single write call that converts the list to a string:

from urllib import request
from redditscore.tokenizer import CrazyTokenizer
tokenizer = CrazyTokenizer()
url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt"
for line in request.urlopen(url):
    tokens = tokenizer.tokenize(line.decode('utf-8'))
    with open('your_file.txt', 'a') as f:
        print(tokens)
        # Convert list to string and write it
        f.write(f"{tokens}\n")

This will write each list exactly as it appears when you run print(tokens)—complete with square brackets, quotes, and commas.

2. Use repr() for a precise, reproducible representation

If you need an exact string version of the list (useful for debugging or later parsing back into a Python list), use repr() instead:

f.write(f"{repr(tokens)}\n")

repr() ensures special characters (like the BOM \ufeff in your first sample list) are preserved as escape sequences, matching exactly what you see in the console.

3. Optimize file handling (optional but recommended)

Opening the file inside the loop can be inefficient for large datasets. Instead, open the file once before processing lines:

from urllib import request
from redditscore.tokenizer import CrazyTokenizer
tokenizer = CrazyTokenizer()
url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt"

# Open file once outside the loop
with open('your_file.txt', 'w') as f:
    for line in request.urlopen(url):
        tokens = tokenizer.tokenize(line.decode('utf-8'))
        print(tokens)
        f.write(f"{tokens}\n")

Use 'w' mode to overwrite the file each run, or keep 'a' if you want to append to an existing file.


内容的提问来源于stack exchange,提问作者john

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 14:09:05