如何在NewsAPI Python代码中去除重复打印值?
解决NewsAPI获取重复数据的问题
使用NewsAPI的Python代码获取数据时,拿到了近800条结果,但其中仅约4条是非重复数据(数量随关键词变化),需要实现去重逻辑阻止重复数据生成。
原Python代码
news_sources = newsapi.get_sources() for source in news_sources['sources']: #print(source['name']) all_articles = newsapi.get_everything( q='shooting', language='en', #from_param='2023-01-22', #to='2023-01-22' ) for article in all_articles['articles']: #print('Source : ', article['source']['name']) #print('Title : ', article['title']) #print('Date : ', article['publishedAt']) print('Url : ', article['url'], '\n\n')
参考的JS去重代码
你看到过一段JavaScript去重逻辑,但不确定是否适用于当前Python代码:
var isDuplicated = false; for (var new in news) { if (new.title == articleModel.title) { isDuplicated = true; } } if (!isDuplicated) { // Now you can add it news.add(articleModel); }
Python版本的去重解决方案
这段JS的核心逻辑是遍历已存数据对比标题去重,在Python里可以用更高效的方式实现:利用集合(set)存储唯一标识(比如文章的url,因为每个新闻的url是唯一的),集合会自动忽略重复值,比循环遍历列表效率高很多。
修改后的代码如下:
news_sources = newsapi.get_sources() # 用集合存储已处理过的文章url,实现去重 processed_urls = set() for source in news_sources['sources']: all_articles = newsapi.get_everything( q='shooting', language='en', #from_param='2023-01-22', #to='2023-01-22' ) for article in all_articles['articles']: article_url = article['url'] # 检查当前url是否已处理过 if article_url not in processed_urls: processed_urls.add(article_url) # 这里执行你需要的操作,比如打印或存储数据 print('Url : ', article_url, '\n\n') # print('Source : ', article['source']['name']) # print('Title : ', article['title']) # print('Date : ', article['publishedAt'])
逻辑说明
- 初始化空集合
processed_urls,用来记录已经处理过的文章url - 遍历每篇文章时,先提取
url,判断是否在集合中 - 如果不在集合里,就把url加入集合,同时执行后续的打印/存储操作;如果已经存在,直接跳过该文章
内容的提问来源于stack exchange,提问作者TestingPython
相关产品推荐
相关产品推荐

