You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多文件字符串统计(信息流活动):求Top n浏览推文最多用户

实现统计浏览推文最多的前n位用户

别担心,我来帮你从头到尾搞定这个需求!先把所有信息理清楚,再给出完整的可运行代码。

核心需求

我们需要编写程序,输出浏览推文最多的前n位用户及其浏览数量,其中「已浏览推文」的定义包含以下四类:

  • 关注用户发布的推文
  • 被提及(@user)的推文
  • 收到的私信(DM)
  • 转发推文对应的原作者发布的内容

测试文件与参考代码

1. 关注关系文件:follows.txt

这个文件每行首词是用户,后续是该用户的关注对象:

andrew fred fred judy andrew fred george judy andrew john george

2. 参考代码片段(统计关注人数)

你提供的这段代码用于统计每个用户的关注数量,我们可以基于它扩展出完整的关注关系字典:

for line in lines:
    names = line.split()
    follow_dict[names[0]] = len(names)-1
    if max_follower < len(names)-1:
        max_follower = len(names)-1

3. 信息流文件:stream.txt

这个文件包含所有的推文活动,包括普通推文、提及、转发、私信:

andrew I hate mondays.
fred Python is cool.
fred Ko Ko Bop Ko Ko Bop Ko Ko Bop for ever
andrew @fred no it isn't, what do you think @john???
judy @fred enough with the k-pop
judy RT @fred Python is cool.
andrew RT @judy @fred enough with the k pop
george RT @fred Python is cool.
andrew DM @john Oops
john DM @andrew Who are you go away! Do you know him, @judy?

预期示例输出

当输入n=10时,程序输出如下(相同浏览量的用户换行对齐):

Enter n: 10
6 judy
5 fred
george
3 andrew
john

完整实现代码

下面是满足所有需求的Python代码,每一步都有清晰的注释:

def load_follows(file_path):
    # 构建用户-关注列表字典,key是用户名,value是该用户关注的所有用户集合
    follow_dict = {}
    with open(file_path, 'r') as f:
        for line in f:
            names = line.strip().split()
            if not names:
                continue
            user = names[0]
            # 用集合存储关注对象,方便快速查找
            following = set(names[1:])
            follow_dict[user] = following
    return follow_dict

def parse_stream(file_path):
    # 解析每条推文,提取关键信息:发布者、提及用户、是否私信、转发原作者
    stream_data = []
    with open(file_path, 'r') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            # 拆分发布者和内容,maxsplit=1避免内容中的空格拆分发布者
            parts = line.split(maxsplit=1)
            author = parts[0]
            content = parts[1] if len(parts) > 1 else ""
            
            # 提取所有被提及的用户(去掉@符号)
            mentioned_users = set()
            for word in content.split():
                if word.startswith('@') and len(word) > 1:
                    mentioned_users.add(word[1:])
            
            # 判断是否是私信
            is_dm = content.startswith('DM')
            
            # 提取转发的原作者(格式是RT @xxx)
            rt_original_author = None
            if content.startswith('RT @'):
                # 拆分出@后面的原作者
                rt_part = content.split(maxsplit=2)[1]
                rt_original_author = rt_part[1:]
            
            stream_data.append({
                'author': author,
                'mentioned': mentioned_users,
                'is_dm': is_dm,
                'content': content,
                'rt_author': rt_original_author
            })
    return stream_data

def calculate_view_counts(follow_dict, stream_data):
    # 初始化所有用户的浏览量为0
    view_counts = {user: 0 for user in follow_dict.keys()}
    
    for tweet in stream_data:
        author = tweet['author']
        mentioned = tweet['mentioned']
        is_dm = tweet['is_dm']
        rt_author = tweet['rt_author']
        content = tweet['content']
        
        # 规则1:关注该推文作者的用户,浏览量+1
        for user, following in follow_dict.items():
            if author in following:
                view_counts[user] += 1
        
        # 规则2:被提及的用户,浏览量+1
        for user in mentioned:
            if user in view_counts:
                view_counts[user] += 1
        
        # 规则3:私信接收者,浏览量+1
        if is_dm:
            # 提取私信接收者(格式DM @xxx)
            dm_recipient = content.split(maxsplit=2)[1][1:]
            if dm_recipient in view_counts:
                view_counts[dm_recipient] += 1
        
        # 规则4:转发推文的原作者发布的内容,会被转发者的关注者浏览
        if rt_author is not None:
            # 转发者的关注者能看到这条转发,相当于浏览了原作者的内容
            for user, following in follow_dict.items():
                if author in following:
                    view_counts[user] += 1
    
    return view_counts

def main():
    # 加载关注关系和信息流数据
    follow_dict = load_follows('follows.txt')
    stream_data = parse_stream('stream.txt')
    
    # 计算每个用户的浏览量
    view_counts = calculate_view_counts(follow_dict, stream_data)
    
    # 获取用户输入的n值
    n = int(input("Enter n: "))
    
    # 按浏览量降序排序,浏览量相同则按用户名升序排序
    sorted_users = sorted(view_counts.items(), key=lambda x: (-x[1], x[0]))
    
    # 按要求格式输出前n位用户
    current_count = None
    for user, count in sorted_users[:n]:
        if count != current_count:
            print(f"{count} {user}")
            current_count = count
        else:
            # 相同浏览量的用户缩进对齐
            print(f"     {user}")

if __name__ == "__main__":
    main()

内容的提问来源于stack exchange,提问作者Soshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:48:28