You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python网络爬虫打开URL并获取页面正文内容

Python新闻爬虫获取正文内容实现方案

你现有代码存在几个可调整的点:

  • 未导入re模块,运行时会抛出模块未找到错误
  • 抓取到的href大概率是相对路径,需要拼接站点域名生成可直接请求的绝对URL
  • 缺少请求头伪装,容易被站点反爬策略拦截

实现步骤

  • 首先导入缺失的re模块,添加请求头伪装浏览器请求
  • 封装独立的正文解析函数,请求单个新闻URL后解析页面正文内容
  • 遍历抓取到的新闻URL,逐个请求获取正文

完整修改后代码

import requests
import re
import time
from bs4 import BeautifulSoup

# 基础配置
BASE_URL = "https://www.aaa.com"
# 伪装浏览器请求头,可根据自己的浏览器信息修改
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

def get_news_content(news_url):
    """获取单个新闻页的正文内容"""
    try:
        # 请求新闻详情页
        resp = requests.get(news_url, headers=HEADERS, timeout=10)
        resp.encoding = resp.apparent_encoding # 自动匹配编码避免乱码
        soup = BeautifulSoup(resp.text, 'html.parser')
        
        # 正文解析规则,需要根据你爬的站点实际结构修改
        # 常见正文容器:<article>标签、class为content/article/main的div标签
        # 示例:假设正文在class为news-content的div里
        content_box = soup.find('div', class_='news-content')
        if not content_box:
            # 备选规则,找article标签
            content_box = soup.find('article')
        
        if content_box:
            # 提取纯文本,去掉多余换行和空格
            content = content_box.get_text(strip=True, separator='\n')
            return content
        else:
            return "未匹配到正文内容,需调整选择器"
    except Exception as e:
        return f"请求出错:{str(e)}"

if __name__ == "__main__":
    # 请求列表页
    list_resp = requests.get(BASE_URL, headers=HEADERS, timeout=10)
    list_resp.encoding = list_resp.apparent_encoding
    soup = BeautifulSoup(list_resp.text, 'html.parser')
    
    # 遍历获取所有符合规则的新闻链接
    for a_tag in soup.findAll('a', href=True):
        if re.search(r"\d+$", a_tag['href']):
            # 拼接绝对URL
            if a_tag['href'].startswith('http'):
                full_url = a_tag['href']
            else:
                full_url = BASE_URL + a_tag['href'] if a_tag['href'].startswith('/') else BASE_URL + '/' + a_tag['href']
            news_title = a_tag.get_text(strip=True)
            print(f"新闻标题:{news_title}")
            print(f"新闻链接:{full_url}")
            # 获取正文
            news_content = get_news_content(full_url)
            print(f"新闻正文:\n{news_content}")
            print("-"*80)
            # 加1秒延迟,避免请求太频繁被封
            time.sleep(1)

注意事项

  • 代码里的正文选择器find('div', class_='news-content')需要根据你实际爬取的站点DOM结构调整,打开对应新闻详情页按F12查看正文所在的标签和类名修改即可
  • 如果站点有反爬机制,可适当增加请求间隔时间、添加代理IP配置
  • 爬取内容请遵守站点的robots协议规定,避免违规爬取

内容的提问来源于stack exchange,提问作者geekPinkFlower

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 02:36:06