如何在Python中为字典添加嵌套键以按政党分类统计推特数据
实现按政党分类的嵌套字典统计
问题说明
需要编写代码统计推特用户的发推数量,最终返回的字典需包含Democrats和Republicans两个顶层键,每个键对应的值是该政党用户的发推数字典。现有代码已能正确统计用户发推数,但未构建出目标嵌套结构。
原代码
def getPartyUserTwitterUsage(tweetFile): import csv myFile = open(tweetFile,"r") # opening file in read csvReader = csv.reader(myFile,delimiter=",") # splitting for ',' next(csvReader) # skipping header tweetList = {} repTweet = 0 demoTweet = 0 for row in csvReader: # iterating through file if (row[1] == 'R'): if (row[0] not in tweetList): tweetList[row[0]] = 1 else: tweetList[row[0]] += 1 if (row[1] == 'D'): if (row[0] not in tweetList): tweetList[row[0]] = 1 else: tweetList[row[0]] += 1 return tweetList
当前输出
{'ChrisMurphyCT': 1000, 'SenBlumenthal': 1000, 'SenatorCarper': 1000, 'ChrisCoons': 1000, 'brianschatz': 1000, 'maziehirono': 1000, 'SenatorDurbin': 1000, 'SenatorHarkin': 1000, 'lisamurkowski': 1000, 'JeffFlake': 1000, 'marcorubio': 1000, 'MikeCrapo': 958, 'SenatorRisch': 1000, 'ChuckGrassley': 1000, 'SenPatRoberts': 1000, 'JerryMoran': 1000}
期望输出
{'Republicans': {'lisamurkowski': 1000, 'JeffFlake': 1000, 'marcorubio': 1000, 'MikeCrapo': 958, 'SenatorRisch': 1000, 'ChuckGrassley': 1000, 'SenPatRoberts': 1000, 'JerryMoran': 1000}, 'Democrats': {'ChrisMurphyCT': 1000, 'SenBlumenthal': 1000, 'SenatorCarper': 1000, 'ChrisCoons': 1000, 'brianschatz': 1000, 'maziehirono': 1000, 'SenatorDurbin': 1000, 'SenatorHarkin': 1000}}
修改后的代码
def getPartyUserTwitterUsage(tweetFile): import csv # 初始化嵌套字典,顶层键为目标政党 party_tweets = { 'Republicans': {}, 'Democrats': {} } # 使用with语句自动管理文件资源,避免手动关闭 with open(tweetFile, "r") as myFile: csvReader = csv.reader(myFile, delimiter=",") next(csvReader) # 跳过表头 for row in csvReader: username = row[0] party_code = row[1] # 根据政党代码选择对应的子字典 if party_code == 'R': target_dict = party_tweets['Republicans'] elif party_code == 'D': target_dict = party_tweets['Democrats'] else: continue # 跳过非目标政党的数据 # 用dict.get简化统计逻辑,不存在则设为0再加1 target_dict[username] = target_dict.get(username, 0) + 1 return party_tweets
修改说明
- 初始化嵌套结构:直接创建包含
Republicans和Democrats的顶层字典,每个键对应空的用户统计子字典,确保最终结构符合要求。 - 文件资源管理:改用
with语句打开文件,无需手动调用close(),自动处理文件关闭,避免资源泄漏。 - 简化统计逻辑:使用
dict.get()方法替代原有的if-else判断,一行代码完成用户发推数的累加(不存在则默认0,再加1)。 - 明确分支处理:通过
party_code判断后直接指向对应政党的子字典,避免重复代码,逻辑更清晰。
内容的提问来源于stack exchange,提问作者jkDAbOntHeH8rz
相关产品推荐
相关产品推荐

