如何使用Python网络爬虫打开URL并获取页面正文内容
Python新闻爬虫获取正文内容实现方案
你现有代码存在几个可调整的点:
- 未导入
re模块,运行时会抛出模块未找到错误 - 抓取到的href大概率是相对路径,需要拼接站点域名生成可直接请求的绝对URL
- 缺少请求头伪装,容易被站点反爬策略拦截
实现步骤
- 首先导入缺失的
re模块,添加请求头伪装浏览器请求 - 封装独立的正文解析函数,请求单个新闻URL后解析页面正文内容
- 遍历抓取到的新闻URL,逐个请求获取正文
完整修改后代码
import requests import re import time from bs4 import BeautifulSoup # 基础配置 BASE_URL = "https://www.aaa.com" # 伪装浏览器请求头,可根据自己的浏览器信息修改 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } def get_news_content(news_url): """获取单个新闻页的正文内容""" try: # 请求新闻详情页 resp = requests.get(news_url, headers=HEADERS, timeout=10) resp.encoding = resp.apparent_encoding # 自动匹配编码避免乱码 soup = BeautifulSoup(resp.text, 'html.parser') # 正文解析规则,需要根据你爬的站点实际结构修改 # 常见正文容器:<article>标签、class为content/article/main的div标签 # 示例:假设正文在class为news-content的div里 content_box = soup.find('div', class_='news-content') if not content_box: # 备选规则,找article标签 content_box = soup.find('article') if content_box: # 提取纯文本,去掉多余换行和空格 content = content_box.get_text(strip=True, separator='\n') return content else: return "未匹配到正文内容,需调整选择器" except Exception as e: return f"请求出错:{str(e)}" if __name__ == "__main__": # 请求列表页 list_resp = requests.get(BASE_URL, headers=HEADERS, timeout=10) list_resp.encoding = list_resp.apparent_encoding soup = BeautifulSoup(list_resp.text, 'html.parser') # 遍历获取所有符合规则的新闻链接 for a_tag in soup.findAll('a', href=True): if re.search(r"\d+$", a_tag['href']): # 拼接绝对URL if a_tag['href'].startswith('http'): full_url = a_tag['href'] else: full_url = BASE_URL + a_tag['href'] if a_tag['href'].startswith('/') else BASE_URL + '/' + a_tag['href'] news_title = a_tag.get_text(strip=True) print(f"新闻标题:{news_title}") print(f"新闻链接:{full_url}") # 获取正文 news_content = get_news_content(full_url) print(f"新闻正文:\n{news_content}") print("-"*80) # 加1秒延迟,避免请求太频繁被封 time.sleep(1)
注意事项
- 代码里的正文选择器
find('div', class_='news-content')需要根据你实际爬取的站点DOM结构调整,打开对应新闻详情页按F12查看正文所在的标签和类名修改即可 - 如果站点有反爬机制,可适当增加请求间隔时间、添加代理IP配置
- 爬取内容请遵守站点的robots协议规定,避免违规爬取
内容的提问来源于stack exchange,提问作者geekPinkFlower
相关产品推荐
相关产品推荐

