使用Python Scrapy与BeautifulSoup爬取网页遇阻:无法获取目标页面所有文章链接
解决Scrapy和BeautifulSoup爬取CNN印度商业页面无输出的问题
嘿,我来帮你排查下问题!你遇到的没输出数据的情况,主要是几个小细节没处理好,咱们分别针对Scrapy和BeautifulSoup来修复:
一、Scrapy版本的问题与修复
你的Scrapy代码有三个关键问题:
- 属性拼写错误:Scrapy Spider类需要的是
start_urls(复数形式),而你写的是start_url,这会导致Scrapy根本不会发起请求。 - 反爬拦截:CNN会检测请求的User-Agent,默认的Scrapy User-Agent容易被拦截,返回的页面没有目标内容。
- 提取逻辑不对:你当前只是打印span元素,而文章链接其实在包含该span的
<a>标签里。
修正后的代码:
import scrapy class CnnSpider(scrapy.Spider): name = "cnn" # 修正为复数形式的start_urls start_urls = ['https://edition.cnn.com/business/india/index.html'] # 设置自定义请求头,避免被反爬拦截 custom_settings = { 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } def parse(self, response): # 找到包含标题span的a标签,提取链接和标题文本 for article in response.css('a.cd__headline'): article_link = article.css('::attr(href)').get() # 拼接完整URL,因为CNN的链接是相对路径 full_link = response.urljoin(article_link) article_title = article.css('span.cd__headline-text::text').get() # 输出结果 print(f"标题: {article_title}, 链接: {full_link}") # 如果要保存到Item,可以返回字典或Item对象 # yield {'title': article_title, 'link': full_link}
二、BeautifulSoup版本的问题与修复
你的BeautifulSoup代码主要问题是:
- 反爬拦截:
requests.get没有设置User-Agent,CNN会返回无内容的页面,导致找不到目标元素。 - 提取逻辑偏差:链接在包含
span.cd__headline-text的<a>标签中,直接找span无法获取链接。
修正后的代码:
import requests from bs4 import BeautifulSoup def get_cnn_links(): url = "https://edition.cnn.com/business/india/index.html" # 设置请求头,模拟浏览器请求 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 带headers发起请求 page = requests.get(url, headers=headers) # 检查请求是否成功 if page.status_code != 200: print(f"请求失败,状态码: {page.status_code}") return soup = BeautifulSoup(page.content, "html.parser") # 找到所有包含标题span的a标签 articles = soup.find_all('a', class_="cd__headline") for article in articles: article_link = article.get('href') # 拼接完整URL full_link = f"https://edition.cnn.com{article_link}" if article_link.startswith('/') else article_link article_title = article.find('span', class_="cd__headline-text").get_text(strip=True) print(f"标题: {article_title}, 链接: {full_link}") # 调用函数 get_cnn_links()
额外说明
如果还是没有数据,可能是CNN启用了更严格的反爬(比如需要Cookie或JS渲染),这时候可以考虑:
- Scrapy中添加
DOWNLOAD_DELAY设置,降低请求频率 - 使用
selenium或playwright处理动态渲染的页面
内容的提问来源于stack exchange,提问作者Marco
相关产品推荐
相关产品推荐

