多文件字符串统计(信息流活动):求Top n浏览推文最多用户
实现统计浏览推文最多的前n位用户
别担心,我来帮你从头到尾搞定这个需求!先把所有信息理清楚,再给出完整的可运行代码。
核心需求
我们需要编写程序,输出浏览推文最多的前n位用户及其浏览数量,其中「已浏览推文」的定义包含以下四类:
- 关注用户发布的推文
- 被提及(@user)的推文
- 收到的私信(DM)
- 转发推文对应的原作者发布的内容
测试文件与参考代码
1. 关注关系文件:follows.txt
这个文件每行首词是用户,后续是该用户的关注对象:
andrew fred fred judy andrew fred george judy andrew john george
2. 参考代码片段(统计关注人数)
你提供的这段代码用于统计每个用户的关注数量,我们可以基于它扩展出完整的关注关系字典:
for line in lines: names = line.split() follow_dict[names[0]] = len(names)-1 if max_follower < len(names)-1: max_follower = len(names)-1
3. 信息流文件:stream.txt
这个文件包含所有的推文活动,包括普通推文、提及、转发、私信:
andrew I hate mondays. fred Python is cool. fred Ko Ko Bop Ko Ko Bop Ko Ko Bop for ever andrew @fred no it isn't, what do you think @john??? judy @fred enough with the k-pop judy RT @fred Python is cool. andrew RT @judy @fred enough with the k pop george RT @fred Python is cool. andrew DM @john Oops john DM @andrew Who are you go away! Do you know him, @judy?
预期示例输出
当输入n=10时,程序输出如下(相同浏览量的用户换行对齐):
Enter n: 10 6 judy 5 fred george 3 andrew john
完整实现代码
下面是满足所有需求的Python代码,每一步都有清晰的注释:
def load_follows(file_path): # 构建用户-关注列表字典,key是用户名,value是该用户关注的所有用户集合 follow_dict = {} with open(file_path, 'r') as f: for line in f: names = line.strip().split() if not names: continue user = names[0] # 用集合存储关注对象,方便快速查找 following = set(names[1:]) follow_dict[user] = following return follow_dict def parse_stream(file_path): # 解析每条推文,提取关键信息:发布者、提及用户、是否私信、转发原作者 stream_data = [] with open(file_path, 'r') as f: for line in f: line = line.strip() if not line: continue # 拆分发布者和内容,maxsplit=1避免内容中的空格拆分发布者 parts = line.split(maxsplit=1) author = parts[0] content = parts[1] if len(parts) > 1 else "" # 提取所有被提及的用户(去掉@符号) mentioned_users = set() for word in content.split(): if word.startswith('@') and len(word) > 1: mentioned_users.add(word[1:]) # 判断是否是私信 is_dm = content.startswith('DM') # 提取转发的原作者(格式是RT @xxx) rt_original_author = None if content.startswith('RT @'): # 拆分出@后面的原作者 rt_part = content.split(maxsplit=2)[1] rt_original_author = rt_part[1:] stream_data.append({ 'author': author, 'mentioned': mentioned_users, 'is_dm': is_dm, 'content': content, 'rt_author': rt_original_author }) return stream_data def calculate_view_counts(follow_dict, stream_data): # 初始化所有用户的浏览量为0 view_counts = {user: 0 for user in follow_dict.keys()} for tweet in stream_data: author = tweet['author'] mentioned = tweet['mentioned'] is_dm = tweet['is_dm'] rt_author = tweet['rt_author'] content = tweet['content'] # 规则1:关注该推文作者的用户,浏览量+1 for user, following in follow_dict.items(): if author in following: view_counts[user] += 1 # 规则2:被提及的用户,浏览量+1 for user in mentioned: if user in view_counts: view_counts[user] += 1 # 规则3:私信接收者,浏览量+1 if is_dm: # 提取私信接收者(格式DM @xxx) dm_recipient = content.split(maxsplit=2)[1][1:] if dm_recipient in view_counts: view_counts[dm_recipient] += 1 # 规则4:转发推文的原作者发布的内容,会被转发者的关注者浏览 if rt_author is not None: # 转发者的关注者能看到这条转发,相当于浏览了原作者的内容 for user, following in follow_dict.items(): if author in following: view_counts[user] += 1 return view_counts def main(): # 加载关注关系和信息流数据 follow_dict = load_follows('follows.txt') stream_data = parse_stream('stream.txt') # 计算每个用户的浏览量 view_counts = calculate_view_counts(follow_dict, stream_data) # 获取用户输入的n值 n = int(input("Enter n: ")) # 按浏览量降序排序,浏览量相同则按用户名升序排序 sorted_users = sorted(view_counts.items(), key=lambda x: (-x[1], x[0])) # 按要求格式输出前n位用户 current_count = None for user, count in sorted_users[:n]: if count != current_count: print(f"{count} {user}") current_count = count else: # 相同浏览量的用户缩进对齐 print(f" {user}") if __name__ == "__main__": main()
内容的提问来源于stack exchange,提问作者Soshi
相关产品推荐
相关产品推荐

