NY Times API无法返回评论,代码返回文章首段求技术排查
问题根源与修复方案
你的代码返回文章摘要而非评论分析,核心原因是评论获取逻辑完全失效,导致comments始终为None,触发了返回摘要的分支。以下是具体问题点和修复方案:
核心问题点
1. get_comments函数URL完全无效
函数内的url = f"url"是占位符,未使用NY Times评论API的正确格式,请求无效URL必然返回非200状态码,导致comments为None。
2. 文章搜索API请求构造错误
articles_url中的fq参数语法错误:fq=source:(%{api_key}写法完全错误,正确的来源过滤需指定媒体名称,API密钥要放在单独的api-key参数中。- 直接用
article.title作为q参数,未处理空格等特殊字符,会导致URL格式错误。 - 循环匹配文章时,错误将
article_id设为文章URL,NY Times评论API实际可直接通过文章URL查询评论,无需额外提取ID。
3. 异常处理过于宽泛
try...except:捕获所有异常,掩盖了API请求失败、JSON解析错误等问题,无法定位故障原因。
4. 未实际使用bs4
你提到要用bs4抓取评论,但代码完全未引入或使用该库,缺少API失效后的备用抓取逻辑。
修复后的完整代码
from flask import Flask, render_template, request from newspaper import Article from textblob import TextBlob import requests import json import nltk from urllib.parse import quote from bs4 import BeautifulSoup # 引入bs4 nltk.download('punkt') app = Flask(__name__) def get_comments(api_key, article_url): # 使用NY Times评论API正确格式:通过文章URL获取评论 url = f"https://api.nytimes.com/svc/community/v3/user-content/url.json?url={quote(article_url)}&api-key={api_key}" try: response = requests.get(url) response.raise_for_status() # 主动抛出HTTP错误 data = response.json() # 检查API返回是否包含有效评论数据 if 'results' in data and 'comments' in data['results']: return data['results']['comments'] return None except Exception as e: print(f"API获取评论失败: {str(e)}") # API失效时,用bs4抓取网页评论(示例逻辑,需根据NY Times页面结构调整) return scrape_comments_from_page(article_url) def scrape_comments_from_page(article_url): # bs4抓取评论基础逻辑(NY Times评论区可能动态加载,需调整选择器) try: response = requests.get(article_url) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') comments = [] # 注意:需用浏览器开发者工具定位真实评论选择器,此处为示例 for comment_div in soup.select('div.comment-body'): comments.append({ 'commentBody': comment_div.get_text(strip=True), 'commentTitle': '' # 网页无评论标题时设为空 }) return comments if comments else None except Exception as e: print(f"网页抓取评论失败: {str(e)}") return None def process_article(article_url): article = Article(article_url) article.download() article.parse() api_key = '你的API密钥' # 替换为实际NY Times API密钥 # 正确构造文章搜索API请求:编码标题、设置来源过滤、携带API密钥 encoded_title = quote(article.title) articles_url = f"https://api.nytimes.com/svc/search/v2/articlesearch.json?q={encoded_title}&fq=source:(\"The New York Times\")&api-key={api_key}" try: response = requests.get(articles_url) response.raise_for_status() data = response.json() articles = data['response']['docs'] # 匹配目标文章 target_article = None for a in articles: if article_url == a['web_url']: target_article = a break if not target_article: return "未在NY Times数据库中找到对应文章。" # 获取评论 comments = get_comments(api_key, article_url) if not comments: article.nlp() return article.summary # 评论情感分析与主题统计 sentiment_polarity = 0.0 sentiment_subjectivity = 0.0 topics = {} for comment in comments: comment_body = comment.get('commentBody', '') if not comment_body: continue sentiment = TextBlob(comment_body).sentiment sentiment_polarity += sentiment.polarity sentiment_subjectivity += sentiment.subjectivity sentiment_label = 'positive' if sentiment.polarity > 0 else 'neutral' if sentiment.polarity ==0 else 'negative' # 处理评论标题(为空则跳过) comment_title = comment.get('commentTitle', '') if comment_title: for topic in comment_title.split(): topic = topic.lower().strip() # 统一小写去重 if topic not in topics: topics[topic] = {'positive':0, 'neutral':0, 'negative':0} topics[topic][sentiment_label] +=1 num_comments = len(comments) avg_sentiment_polarity = sentiment_polarity / num_comments if num_comments >0 else 0 avg_sentiment_subjectivity = sentiment_subjectivity / num_comments if num_comments >0 else 0 return render_template('results.html', article_title=article.title, article_text=article.text, num_comments=num_comments, avg_sentiment_polarity=round(avg_sentiment_polarity, 2), avg_sentiment_subjectivity=round(avg_sentiment_subjectivity, 2), topics=topics) except Exception as e: print(f"处理文章出错: {str(e)}") article.nlp() return article.summary @app.route('/', methods=['GET', 'POST']) def index(): if request.method == 'POST': article_url = request.form['article_url'] return process_article(article_url) else: return render_template('index.html') if __name__ == '__main__': app.run(debug=True)
额外注意事项
- API密钥替换:将
api_key = '你的API密钥'替换为你在NY Times开发者平台申请的有效密钥。 - bs4抓取逻辑优化:NY Times评论区可能采用动态加载,示例中的选择器需用浏览器开发者工具定位真实元素,或改用
selenium处理动态内容。 - 调试日志:添加的
print语句用于输出错误信息,生产环境可替换为专业日志库。 - 统计优化:对情感结果做了四舍五入,主题统计统一转为小写避免重复计数。
内容的提问来源于stack exchange,提问作者0004
相关产品推荐
相关产品推荐

