使用newspaper3k采集路透社搜索页新闻失败的问题求助
解决newspaper3k无法提取路透社搜索页文章的问题
问题原因
- 动态内容渲染:路透社搜索页的部分文章链接通过JavaScript动态加载,newspaper3k默认仅抓取静态HTML,不会执行JS,因此无法获取动态生成的内容。
- 页面结构不匹配:路透社搜索结果的链接DOM结构不符合newspaper3k内置的文章链接识别规则,导致
build方法无法识别出有效文章链接。
解决方法
方法1:直接解析静态页面提取链接(适合静态加载的搜索结果)
绕过newspaper3k的build方法,用BeautifulSoup直接提取页面中的文章链接,再逐个用newspaper3k处理单篇文章:
import requests from bs4 import BeautifulSoup import newspaper # 目标搜索URL url = "https://www.reuters.com/site-search/?query=wheat" # 设置请求头模拟浏览器,避免反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 获取页面静态源码 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 提取路透社搜索结果的文章链接(根据当前页面结构调整选择器) article_links = soup.find_all('a', {'data-testid': 'search-result-story-link'}) print(f"Number of articles found: {len(article_links)}") for link in article_links: # 补全相对链接为绝对URL article_url = f"https://www.reuters.com{link['href']}" if link['href'].startswith('/') else link['href'] try: # 用newspaper3k处理单篇文章 article = newspaper.Article(article_url) article.download() article.parse() print(f"URL: {article_url}") print(f"Title: {article.title}\n") except Exception as e: print(f"处理{article_url}失败: {str(e)}")
方法2:用Selenium渲染动态内容(适合需要滚动加载的搜索结果)
如果路透社搜索页需要滚动加载更多文章,用Selenium模拟浏览器执行JS,获取完整渲染后的页面:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import newspaper import time url = "https://www.reuters.com/site-search/?query=wheat" # 配置Chrome浏览器选项(可选,无头模式不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # 启动浏览器并访问页面 driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 模拟滚动加载3次(可根据需求调整次数) for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待内容加载 # 获取完整页面源码 page_source = driver.page_source driver.quit() # 解析链接并处理文章 soup = BeautifulSoup(page_source, 'html.parser') article_links = soup.find_all('a', {'data-testid': 'search-result-story-link'}) print(f"Number of articles found: {len(article_links)}") for link in article_links: article_url = f"https://www.reuters.com{link['href']}" if link['href'].startswith('/') else link['href'] try: article = newspaper.Article(article_url) article.download() article.parse() print(f"URL: {article_url}") print(f"Title: {article.title}\n") except Exception as e: print(f"处理{article_url}失败: {str(e)}")
注意事项
- 需提前安装依赖包:
pip install requests beautifulsoup4 newspaper3k selenium - 若使用Selenium,需对应安装ChromeDriver(或其他浏览器驱动)并配置环境变量
- 注意调整请求头和滚动逻辑,避免触发网站反爬机制
内容的提问来源于stack exchange,提问作者toyop
相关产品推荐
相关产品推荐

