如何用BeautifulSoup提取指定网页文章内容?代码返回No Data解决
解决网页文章内容提取失败的方案
针对你提取印度经济时报文章内容时输出“No Data”的问题,可按以下步骤排查修复:
1. 解决请求被反爬拦截问题
很多网站会通过检测User-Agent识别非浏览器请求,直接调用requests.get()会被拦截,返回空白或反爬页面,导致无法解析内容。需添加模拟浏览器的请求头:
import requests from bs4 import BeautifulSoup import pandas as pd # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = 'https://economictimes.indiatimes.com/industry/cons-products/food/heinz-braces-up-for-aggressive-marketing/articleshow/5417995.cms' response = requests.get(url, headers=headers) # 检查请求是否成功,失败则直接抛出异常 response.raise_for_status()
2. 修正文章内容的定位选择器
网页结构可能已变更,原代码的选择器无法匹配目标内容。通过浏览器开发者工具查看该页面源码,可知文章正文位于class="artText"的div标签下,包含多个p标签。更新解析逻辑:
soup = BeautifulSoup(response.text, 'html.parser') # 定位文章内容容器 article_container = soup.find('div', class_='artText') if article_container: # 提取所有有效段落并拼接成完整文本 paragraphs = article_container.find_all('p') full_article = '\n\n'.join([p.get_text(strip=True) for p in paragraphs if p.get_text(strip=True)]) # 用pandas存储或输出结果 df = pd.DataFrame({'文章内容': [full_article]}) print(df) else: print("No Data")
3. 实用调试技巧
- 打印
response.text查看返回内容是否为正常网页,若显示反爬提示或空白,需补充更多请求头字段(如Referer、Accept-Language)。 - 用浏览器开发者工具的“检查元素”功能,右键文章内容选择“复制CSS选择器”,确保定位节点的准确性。
内容的提问来源于stack exchange,提问作者airmon
相关产品推荐
相关产品推荐

