如何在处理TXT文件与字典时去除字符串末尾的'\n'?
问题描述
目前代码里除了列表打印功能外,其余逻辑运行都正常。我试过用rstrip()、splitline()、replace()这些方法来处理文本,但可能用法不对。刚接触字典的使用,希望能针对这个打印问题得到具体指导。
购物清单TXT内容
Shopping List ITEM 1000: Frozen Burgers ITEM 1001: Bread ITEM 1002: Frozen Burgers ITEM 1003: Mayonnaise ITEM 1004: Bread ITEM 1005: Frozen Burgers ITEM 1006: Mustard ITEM 1007: Tomato ITEM 1008: Tomato
现有代码
from collections import Counter def get_items(): print("""Items that repeat twice or more in the lists: Item Name Repeat Times ------------------------------""") item_count = Counter() with open('Shopping_List.txt') as f: for line in f.readlines()[1:]: item_count.update([line[11:]]) print('\r',*item_count.most_common(3),sep='\n') def get_dictionaries(textfile): items = {} numbers = {} FIRST_NUM = 1000 with open(textfile) as f: for lines in f.readlines()[1:]: numbers[FIRST_NUM]=lines.lstrip()[11:] items[lines.lstrip()[11:]]=items.setdefault(lines.lstrip()[11:],0)+1 FIRST_NUM += 1 return items,numbers get_items()
解决方案
你的问题主要出在文本处理时未清除换行符和打印格式未对齐表头这两点上,以下是针对性修复:
1. 修复统计准确性
读取每行时,line[11:]会保留行尾的换行符\n,导致Counter把带换行和不带换行的同名称物品当成不同条目;同时固定索引11提取物品名不够健壮(如果ITEM编号位数变化就会出错),建议用字符串分割处理。
2. 修复打印格式
item_count.most_common()返回元组列表,直接打印会破坏预设的表头对齐格式,需要手动格式化输出,同时添加判断只显示重复次数≥2的项。
修正后的完整代码:
from collections import Counter def get_items(): print("""Items that repeat twice or more in the lists: Item Name Repeat Times ------------------------------""") item_count = Counter() with open('Shopping_List.txt') as f: # 跳过标题行,遍历所有物品行 for line in f.readlines()[1:]: # 按": "分割提取物品名,清除首尾空白(含换行) item = line.split(': ', 1)[1].strip() item_count.update([item]) # 按统计次数排序,格式化输出符合条件的项 for item, count in item_count.most_common(): if count >= 2: # 左对齐物品名(占16位),右对齐次数(占13位),匹配表头格式 print(f"{item:<16} {count:>13}") def get_dictionaries(textfile): items = {} numbers = {} FIRST_NUM = 1000 with open(textfile) as f: for line in f.readlines()[1:]: item = line.split(': ', 1)[1].strip() numbers[FIRST_NUM] = item # 用get方法简化字典计数逻辑,更直观 items[item] = items.get(item, 0) + 1 FIRST_NUM += 1 return items, numbers get_items()
额外优化
- 统一文本处理逻辑,确保
get_dictionaries里的物品名也清除了换行符,避免字典存储脏数据 - 用
items.get(item, 0) + 1替代items.setdefault(item,0)+1,计数逻辑更清晰
内容的提问来源于stack exchange,提问作者Strider
相关产品推荐
相关产品推荐

