如何修正代码以从Tweet类实例列表中提取全部话题标签?
修正提取所有推文话题标签的代码
原代码的问题在于for循环内直接使用return语句,这会让程序在遍历第一条推文后就终止循环,因此只能获取单条推文的标签。
修正后的代码如下:
import re # 假设Tweet类已正确定义 tweet1 = Tweet("@realDonaldTrump", "Despite the negative press covfefe #bigsmart", 1249, 54303) tweet2 = Tweet("@elonmusk", "Technically, alcohol is a solution #bigsmart", 366.4, 166500) tweet3 = Tweet("@CIA", "We can neither confirm nor deny that this is our first tweet. #heart", 2192, 284200) tweets = [tweet1, tweet2, tweet3] # 初始化空列表存储所有话题标签 all_hashtags = [] for tweet in tweets: # 提取当前推文的话题标签 current_tags = re.findall(r'#\w+', tweet.content) # 将当前推文的标签添加到总列表(用extend避免嵌套列表) all_hashtags.extend(current_tags) # 输出所有话题标签 print(all_hashtags) # 输出结果: ['#bigsmart', '#bigsmart', '#heart']
如果需要去除重复的话题标签,可以在最后添加一行:
all_hashtags = list(set(all_hashtags))
内容的提问来源于stack exchange,提问作者babygroot
相关产品推荐
相关产品推荐

