You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用newspaper3k采集路透社搜索页新闻失败的问题求助

解决newspaper3k无法提取路透社搜索页文章的问题

问题原因

  • 动态内容渲染:路透社搜索页的部分文章链接通过JavaScript动态加载,newspaper3k默认仅抓取静态HTML,不会执行JS,因此无法获取动态生成的内容。
  • 页面结构不匹配:路透社搜索结果的链接DOM结构不符合newspaper3k内置的文章链接识别规则,导致build方法无法识别出有效文章链接。

解决方法

方法1:直接解析静态页面提取链接(适合静态加载的搜索结果)

绕过newspaper3k的build方法,用BeautifulSoup直接提取页面中的文章链接,再逐个用newspaper3k处理单篇文章:

import requests
from bs4 import BeautifulSoup
import newspaper

# 目标搜索URL
url = "https://www.reuters.com/site-search/?query=wheat"
# 设置请求头模拟浏览器,避免反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 获取页面静态源码
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 提取路透社搜索结果的文章链接(根据当前页面结构调整选择器)
article_links = soup.find_all('a', {'data-testid': 'search-result-story-link'})

print(f"Number of articles found: {len(article_links)}")
for link in article_links:
    # 补全相对链接为绝对URL
    article_url = f"https://www.reuters.com{link['href']}" if link['href'].startswith('/') else link['href']
    try:
        # 用newspaper3k处理单篇文章
        article = newspaper.Article(article_url)
        article.download()
        article.parse()
        print(f"URL: {article_url}")
        print(f"Title: {article.title}\n")
    except Exception as e:
        print(f"处理{article_url}失败: {str(e)}")

方法2:用Selenium渲染动态内容(适合需要滚动加载的搜索结果)

如果路透社搜索页需要滚动加载更多文章,用Selenium模拟浏览器执行JS,获取完整渲染后的页面:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import newspaper
import time

url = "https://www.reuters.com/site-search/?query=wheat"

# 配置Chrome浏览器选项(可选,无头模式不显示浏览器窗口)
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

# 启动浏览器并访问页面
driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# 模拟滚动加载3次(可根据需求调整次数)
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 等待内容加载

# 获取完整页面源码
page_source = driver.page_source
driver.quit()

# 解析链接并处理文章
soup = BeautifulSoup(page_source, 'html.parser')
article_links = soup.find_all('a', {'data-testid': 'search-result-story-link'})

print(f"Number of articles found: {len(article_links)}")
for link in article_links:
    article_url = f"https://www.reuters.com{link['href']}" if link['href'].startswith('/') else link['href']
    try:
        article = newspaper.Article(article_url)
        article.download()
        article.parse()
        print(f"URL: {article_url}")
        print(f"Title: {article.title}\n")
    except Exception as e:
        print(f"处理{article_url}失败: {str(e)}")

注意事项

  • 需提前安装依赖包:pip install requests beautifulsoup4 newspaper3k selenium
  • 若使用Selenium,需对应安装ChromeDriver(或其他浏览器驱动)并配置环境变量
  • 注意调整请求头和滚动逻辑,避免触发网站反爬机制

内容的提问来源于stack exchange,提问作者toyop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 06:05:10