You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Newspaper模块无法提取非英文URL文章问题求助

嘿,我之前也踩过newspaper模块解析非英文站点的坑!你遇到的问题大概率是语言自动检测失效或者模块对非英文站点的HTML适配不好导致的,咱们一步步来解决:

核心原因分析

  • newspaper默认会自动检测文章语言,但像Prothomalo这类孟加拉语站点,自动检测的准确率不高,模块会用错误的语言规则去解析,最终拿不到内容。
  • 部分非英文站点的HTML结构和英文站点差异较大,newspaper的默认解析逻辑没法准确识别正文区域。

解决方案

1. 手动指定目标语言

这是最直接的修复方式,在初始化Article对象时明确指定站点的语言代码(比如孟加拉语是bn),让模块用对应语言的规则解析:

import feedparser
from newspaper import Article

# Prothomalo的订阅源URL
feed_url = "https://www.prothomalo.com/rss/topnews"
feed = feedparser.parse(feed_url)

# 遍历feed中的文章
for entry in feed.entries[:3]:
    article_url = entry.link
    # 手动指定语言为孟加拉语(bn)
    article = Article(article_url, language='bn')
    
    # 可选:添加自定义请求头,避免被反爬拦截
    article.headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    
    article.download()
    article.parse()
    
    print("标题:", article.title)
    print("发布时间:", article.publish_date)
    print("正文预览:", article.text[:200])
    print("---")

2. 结合BeautifulSoup手动解析(如果指定语言后仍失效)

要是指定语言后还是提取不到内容,说明站点的HTML结构比较特殊,可以直接用BeautifulSoup手动定位正文区域:

import feedparser
from newspaper import Article
from bs4 import BeautifulSoup

feed_url = "https://www.prothomalo.com/rss/topnews"
feed = feedparser.parse(feed_url)

for entry in feed.entries[:3]:
    article_url = entry.link
    article = Article(article_url, language='bn')
    article.headers = {'User-Agent': '你的浏览器UA'}
    article.download()
    
    # 用BeautifulSoup解析HTML
    soup = BeautifulSoup(article.html, 'html.parser')
    # 查看Prothomalo网页源码,找到正文对应的容器(比如class为story-content的div)
    content_container = soup.find('div', class_='story-content')
    if content_container:
        # 提取正文文本
        raw_text = content_container.get_text(strip=True, separator='\n')
        print("手动提取的标题:", entry.title)
        print("手动提取的正文:", raw_text[:200])
    print("---")

3. 确保依赖库完整

newspaper依赖nltk的语言数据包,记得提前下载:

import nltk
nltk.download('punkt')

最后提个小建议

如果遇到反爬严格的站点,可以试试更换IP或者使用代理,不过大部分新闻站点只要加个正常的User-Agent就能解决问题。

内容的提问来源于stack exchange,提问作者Istiaque Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:32:24