Python:如何将分词得到的tokens列表写入文本文件?
Problem Description
I’ve written this Python code to fetch text from a URL and tokenize it using CrazyTokenizer:
from urllib import request from redditscore.tokenizer import CrazyTokenizer tokenizer = CrazyTokenizer() url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt" for line in request.urlopen(url): tokens = tokenizer.tokenize(line.decode('utf-8')) #print(tokens) with open('your_file.txt', 'a') as f: print(tokens) for item in tokens: f.write("%s\n" % item)
Right now, each token is written to a new line in the output file. But I want to save the complete list object (e.g., ['\ufeffsave', 'bbc', 'world', 'service', 'from', 'savage', 'cuts']) to the file, with one list per line instead. How can I adjust my code to achieve this?
Solution
The simplest way to write the full list to your file is to convert the tokens list directly to its string representation, then write that string plus a newline. Here are a few straightforward approaches:
1. Use str() for a human-readable list string
Replace the loop that writes individual tokens with a single write call that converts the list to a string:
from urllib import request from redditscore.tokenizer import CrazyTokenizer tokenizer = CrazyTokenizer() url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt" for line in request.urlopen(url): tokens = tokenizer.tokenize(line.decode('utf-8')) with open('your_file.txt', 'a') as f: print(tokens) # Convert list to string and write it f.write(f"{tokens}\n")
This will write each list exactly as it appears when you run print(tokens)—complete with square brackets, quotes, and commas.
2. Use repr() for a precise, reproducible representation
If you need an exact string version of the list (useful for debugging or later parsing back into a Python list), use repr() instead:
f.write(f"{repr(tokens)}\n")
repr() ensures special characters (like the BOM \ufeff in your first sample list) are preserved as escape sequences, matching exactly what you see in the console.
3. Optimize file handling (optional but recommended)
Opening the file inside the loop can be inefficient for large datasets. Instead, open the file once before processing lines:
from urllib import request from redditscore.tokenizer import CrazyTokenizer tokenizer = CrazyTokenizer() url = "http://www.site.uottawa.ca/~diana/csi5386/A1_2020/microblog2011.txt" # Open file once outside the loop with open('your_file.txt', 'w') as f: for line in request.urlopen(url): tokens = tokenizer.tokenize(line.decode('utf-8')) print(tokens) f.write(f"{tokens}\n")
Use 'w' mode to overwrite the file each run, or keep 'a' if you want to append to an existing file.
内容的提问来源于stack exchange,提问作者john

