Python读取文本为字典替换缩写报错:replace()参数2需为字符串而非列表
Hey there! Let's break down why you're hitting that replace() argument 2 must be str, not list error and fix it up.
What's Causing the Error?
Take a look at how you're building your dictionary:
d = { row[0]:row[1:] for row in rows }
When you split each line with split('\t'), row[1:] returns a list—even if there's only one full word after the tab, it becomes something like ['awesome']. When you try to use this list as the replacement value in word.replace(), Python throws an error because replace() expects a string, not a list.
How to Fix It
Store Strings in the Dictionary: Instead of
row[1:], userow[1].strip()to grab the full word as a string. Thestrip()removes any extra whitespace or newlines from the end of the line.- This assumes your
acronym.txtfollows the formatacronym\tfull_word(one tab-separated pair per line). If some lines have multiple tabs, you could use'\t'.join(row[1:]).strip()to join all parts after the first tab into a single string.
- This assumes your
Simplify the Replacement Logic: Your nested loops can be streamlined. For exact word matches (the most common use case), check if the word exists directly in the dictionary instead of iterating through all keys for every word.
Corrected Code
Here's the fixed version of your function:
from nltk.tokenize import word_tokenize # Don't forget this import! def acronym(text): # Load and format the acronym dictionary correctly with open('acronym.txt') as fin: # Strip whitespace from each line first, then split rows = (line.strip().split('\t') for line in fin) # Filter out empty lines and ensure each row has at least two elements d = {row[0].strip(): row[1].strip() for row in rows if len(row) >= 2} words = word_tokenize(text) replaced_words = [] for word in words: # Replace if the word is in our acronym dict; keep it as-is if not replaced_words.append(d.get(word, word)) # Join the words back into a single string return ' '.join(replaced_words)
Bonus: Handling Partial Matches
If you need to replace acronyms that are part of longer words (e.g., replacing "gr8" in "gr8day" with "great"), you can adjust the loop like this—just be aware it might lead to unintended replacements:
for word in words: modified_word = word for acronym_key, full_word in d.items(): if acronym_key in modified_word: modified_word = modified_word.replace(acronym_key, full_word) replaced_words.append(modified_word)
Quick Checks
- Make sure
acronym.txtis properly formatted (e.g.,awsm\tawesomeon one line,gr8\tgreaton another). - If you haven't already, install NLTK and download the tokenizer data with
nltk.download('punkt').
内容的提问来源于stack exchange,提问作者Asmaa M. Elmohamady

