为何抓取印度经济时报2020年头条的Python代码无结果?
印度经济时报头条抓取CSV无数据的排查与解决方法
可能的原因及对应解决步骤
1. 反爬机制拦截请求
- 直接用
requests发起的请求可能被识别为爬虫,返回空页面或非目标内容。给请求添加浏览器UA头:headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} response = requests.get(url, headers=headers) - 若网站需要Cookie验证,可从浏览器复制对应Cookie加入请求头,或用
requests.Session()维持会话。
2. 日期URL构造错误
- 手动访问你生成的每日URL,确认是否能正常打开并显示当日头条。比如印度经济时报历史存档的URL格式可能为
https://economictimes.indiatimes.com/archivelist/year-2020,month-1,starttime-1200489600.cms,需确保日期转换后的参数完全匹配。 - 检查日期循环逻辑:2020年是闰年,2月有29天,避免生成无效日期的URL导致请求失败。
3. BeautifulSoup解析逻辑错误
- 打印
response.text查看是否抓取到完整页面内容,若内容为空,优先排查反爬问题;若内容正常,调整标签选择器:
比如原选择器可能无法定位头条,可通过浏览器开发者工具查看头条元素的标签和class,替换为soup.find_all('h2', class_='story__heading')这类精准选择器。 - 确认解析器使用正确,推荐用
html.parser或lxml,避免因解析器不兼容导致元素查找失败。
4. CSV写入逻辑问题
- 检查是否在数据为空时就打开了CSV文件,或未调用
writerow()/writerows()写入数据。 - 确保文件打开方式正确,避免编码或换行符问题:
with open('headlines.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['日期', '头条']) - 添加异常捕获,排查请求或解析过程中的隐性错误,避免程序无报错但无数据输出。
修正后的示例代码片段
import requests from bs4 import BeautifulSoup import csv from datetime import datetime, timedelta start_date = datetime(2020, 1, 1) end_date = datetime(2020, 10, 31) headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} with open('et_headlines.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) writer.writerow(['日期', '头条标题']) current_date = start_date while current_date <= end_date: # 替换为印度经济时报历史头条的真实URL格式 date_str = current_date.strftime('%Y/%m/%d') url = f'https://economictimes.indiatimes.com/date/{date_str}' try: response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 替换为页面实际的头条元素选择器 headlines = soup.find_all('h2', class_='story__heading') if headlines: for headline in headlines: writer.writerow([current_date.strftime('%Y-%m-%d'), headline.get_text(strip=True)]) else: print(f"{current_date.strftime('%Y-%m-%d')} 未抓取到头条") except Exception as e: print(f"{current_date.strftime('%Y-%m-%d')} 抓取失败: {str(e)}") current_date += timedelta(days=1)
内容的提问来源于stack exchange,提问作者Doubts
相关产品推荐
相关产品推荐

