You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NY Times API无法返回评论,代码返回文章首段求技术排查

问题根源与修复方案

你的代码返回文章摘要而非评论分析,核心原因是评论获取逻辑完全失效,导致comments始终为None,触发了返回摘要的分支。以下是具体问题点和修复方案:


核心问题点

1. get_comments函数URL完全无效

函数内的url = f"url"是占位符,未使用NY Times评论API的正确格式,请求无效URL必然返回非200状态码,导致comments为None。

2. 文章搜索API请求构造错误

  • articles_url中的fq参数语法错误:fq=source:(%{api_key}写法完全错误,正确的来源过滤需指定媒体名称,API密钥要放在单独的api-key参数中。
  • 直接用article.title作为q参数,未处理空格等特殊字符,会导致URL格式错误。
  • 循环匹配文章时,错误将article_id设为文章URL,NY Times评论API实际可直接通过文章URL查询评论,无需额外提取ID。

3. 异常处理过于宽泛

try...except:捕获所有异常,掩盖了API请求失败、JSON解析错误等问题,无法定位故障原因。

4. 未实际使用bs4

你提到要用bs4抓取评论,但代码完全未引入或使用该库,缺少API失效后的备用抓取逻辑。


修复后的完整代码

from flask import Flask, render_template, request
from newspaper import Article
from textblob import TextBlob
import requests
import json
import nltk
from urllib.parse import quote
from bs4 import BeautifulSoup  # 引入bs4

nltk.download('punkt')

app = Flask(__name__)

def get_comments(api_key, article_url):
    # 使用NY Times评论API正确格式:通过文章URL获取评论
    url = f"https://api.nytimes.com/svc/community/v3/user-content/url.json?url={quote(article_url)}&api-key={api_key}"
    try:
        response = requests.get(url)
        response.raise_for_status()  # 主动抛出HTTP错误
        data = response.json()
        # 检查API返回是否包含有效评论数据
        if 'results' in data and 'comments' in data['results']:
            return data['results']['comments']
        return None
    except Exception as e:
        print(f"API获取评论失败: {str(e)}")
        # API失效时,用bs4抓取网页评论(示例逻辑,需根据NY Times页面结构调整)
        return scrape_comments_from_page(article_url)

def scrape_comments_from_page(article_url):
    # bs4抓取评论基础逻辑(NY Times评论区可能动态加载,需调整选择器)
    try:
        response = requests.get(article_url)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        comments = []
        # 注意:需用浏览器开发者工具定位真实评论选择器,此处为示例
        for comment_div in soup.select('div.comment-body'):
            comments.append({
                'commentBody': comment_div.get_text(strip=True),
                'commentTitle': ''  # 网页无评论标题时设为空
            })
        return comments if comments else None
    except Exception as e:
        print(f"网页抓取评论失败: {str(e)}")
        return None

def process_article(article_url):
    article = Article(article_url)
    article.download()
    article.parse()

    api_key = '你的API密钥'  # 替换为实际NY Times API密钥
    # 正确构造文章搜索API请求:编码标题、设置来源过滤、携带API密钥
    encoded_title = quote(article.title)
    articles_url = f"https://api.nytimes.com/svc/search/v2/articlesearch.json?q={encoded_title}&fq=source:(\"The New York Times\")&api-key={api_key}"
    
    try:
        response = requests.get(articles_url)
        response.raise_for_status()
        data = response.json()
        articles = data['response']['docs']
        # 匹配目标文章
        target_article = None
        for a in articles:
            if article_url == a['web_url']:
                target_article = a
                break
        if not target_article:
            return "未在NY Times数据库中找到对应文章。"

        # 获取评论
        comments = get_comments(api_key, article_url)

        if not comments:
            article.nlp()
            return article.summary

        # 评论情感分析与主题统计
        sentiment_polarity = 0.0
        sentiment_subjectivity = 0.0
        topics = {}
        for comment in comments:
            comment_body = comment.get('commentBody', '')
            if not comment_body:
                continue
            sentiment = TextBlob(comment_body).sentiment
            sentiment_polarity += sentiment.polarity
            sentiment_subjectivity += sentiment.subjectivity
            
            sentiment_label = 'positive' if sentiment.polarity > 0 else 'neutral' if sentiment.polarity ==0 else 'negative'
            
            # 处理评论标题(为空则跳过)
            comment_title = comment.get('commentTitle', '')
            if comment_title:
                for topic in comment_title.split():
                    topic = topic.lower().strip()  # 统一小写去重
                    if topic not in topics:
                        topics[topic] = {'positive':0, 'neutral':0, 'negative':0}
                    topics[topic][sentiment_label] +=1

        num_comments = len(comments)
        avg_sentiment_polarity = sentiment_polarity / num_comments if num_comments >0 else 0
        avg_sentiment_subjectivity = sentiment_subjectivity / num_comments if num_comments >0 else 0

        return render_template('results.html', 
                               article_title=article.title,
                               article_text=article.text,
                               num_comments=num_comments,
                               avg_sentiment_polarity=round(avg_sentiment_polarity, 2),
                               avg_sentiment_subjectivity=round(avg_sentiment_subjectivity, 2),
                               topics=topics)
    except Exception as e:
        print(f"处理文章出错: {str(e)}")
        article.nlp()
        return article.summary

@app.route('/', methods=['GET', 'POST'])
def index():
    if request.method == 'POST':
        article_url = request.form['article_url']
        return process_article(article_url)
    else:
        return render_template('index.html')

if __name__ == '__main__':
    app.run(debug=True)

额外注意事项

  1. API密钥替换:将api_key = '你的API密钥'替换为你在NY Times开发者平台申请的有效密钥。
  2. bs4抓取逻辑优化:NY Times评论区可能采用动态加载,示例中的选择器需用浏览器开发者工具定位真实元素,或改用selenium处理动态内容。
  3. 调试日志:添加的print语句用于输出错误信息,生产环境可替换为专业日志库。
  4. 统计优化:对情感结果做了四舍五入,主题统计统一转为小写避免重复计数。

内容的提问来源于stack exchange,提问作者0004

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 18:24:54