如何用Python筛选Twitter用户Timeline中含特定文本的Tweet?
解决Tweepy筛选用户时间线含特定关键词推文的问题
嘿,我发现你的问题所在了——你在调用api.user_timeline的时候用了q="#Gempa"参数,但这个方法根本不支持q过滤条件!q参数是给搜索API(比如api.search_tweets)用的,用户时间线接口本身不会根据关键词筛选内容,所以你之前的代码相当于忽略了q参数,直接拉取了用户的最新推文。
下面给你两种可行的解决方案:
方案1:本地过滤用户时间线推文
先拉取用户的时间线推文,然后在代码里检查每条推文的文本是否包含你要的关键词(比如#Gempa)。这种方法适合需要先获取用户所有推文再做筛选的场景:
修改你的循环部分代码如下:
n = 0 # 先拉取用户的时间线,无需q参数 for tweet in tweepy.Cursor(api.user_timeline, id=108543358, lang="id", result_type="recent", since_id=last_id).items(30): # 检查推文文本是否包含#Gempa(转小写判断更严谨,避免大小写遗漏) if "#gempa" in tweet.text.lower(): print(f"*****{n+1}*****") print(f"ID: {tweet.id_str}") print(f"Text: {tweet.text}") # Python3里无需手动encode,直接用tweet.text即可 print(f"Retweet Count: {tweet.retweet_count}") print(f"Favorite Count: {tweet.favorite_count}") print(f"Date Time: {tweet.created_at}") # 获取地理位置的正确方式: if tweet.coordinates: print(f"Coordinates: {tweet.coordinates}") if tweet.place: print(f"Place: {tweet.place.full_name}") print("************") n += 1 # 插入数据库时直接用tweet.text,你的数据库已配置utf8mb4,支持所有特殊字符 cur.execute("INSERT INTO tweet (no, id, text, retweet_count, favourite_count, date_time) VALUES (%s, %s,%s,%s,%s,%s)", (str(n), tweet.id_str, tweet.text, str(tweet.retweet_count), str(tweet.favorite_count), str(tweet.created_at)))
方案2:用搜索API直接过滤(更高效)
如果你只需要含特定关键词的用户推文,直接用Twitter的搜索API会更高效——让Twitter服务器帮你过滤,不用拉取多余的推文。只需要构造包含用户ID/用户名和关键词的搜索query:
n = 0 # 构造搜索query:from:用户ID/用户名 + 关键词 search_query = "from:108543358 #Gempa" for tweet in tweepy.Cursor(api.search_tweets, q=search_query, lang="id", result_type="recent", since_id=last_id).items(3): print(f"*****{n+1}*****") print(f"ID: {tweet.id_str}") print(f"Text: {tweet.text}") print(f"Retweet Count: {tweet.retweet_count}") print(f"Favorite Count: {tweet.favorite_count}") print(f"Date Time: {tweet.created_at}") # 地理位置获取同方案1 if tweet.coordinates: print(f"Coordinates: {tweet.coordinates}") if tweet.place: print(f"Place: {tweet.place.full_name}") print("************") n += 1 cur.execute("INSERT INTO tweet (no, id, text, retweet_count, favourite_count, date_time) VALUES (%s, %s,%s,%s,%s,%s)", (str(n), tweet.id_str, tweet.text, str(tweet.retweet_count), str(tweet.favorite_count), str(tweet.created_at)))
额外小提示
- 你代码里循环用了
str(i)但没定义i,应该改成str(n+1)或者直接用自增后的n值; pymysql.connect里的port参数应该是整数(比如port=3306),如果是默认端口可以省略这个参数;- 关于地理位置:只有用户发推文时主动开启定位,
tweet.coordinates或tweet.place才会有值,否则为None,建议先判断再打印/存储。
内容的提问来源于stack exchange,提问作者Nanda
相关产品推荐
相关产品推荐

