You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:使用BeautifulSoup爬取The961黎巴嫩新闻详情页时p标签返回None无法获取完整文章

问题:爬取The961详情页文章内容返回None的解决方法

我是一名爬虫新手,尝试爬取The961网站(https://www.the961.com/latest-news/lebanon-news/)的黎巴嫩新闻内容,编写的代码如下:

import bs4
import requests
import re
r = requests.get('https://www.the961.com/latest-news/lebanon-news/').text
soup = bs4.BeautifulSoup(r, 'lxml')
for article in soup.find_all('article'):
    title = article.h3.text
    print(title)
    date = article.find('span', class_='byline-part date')
    if date:
        print('Date:', date.text)
    author = article.find('span', class_="byline-part author")
    if author:
        print('Author:', author.text)
    link = article.find('h3', class_='title').a['href']
    link_r = requests.get(link).text
    soup_link = bs4.BeautifulSoup(link_r, 'lxml')
    // scraping link from title, then opening that link and trying to scrape the whole article, very new to this so I don't know what to do!
    for article in soup_link.find_all('article'):
        paragraph = article.find('p')
        print(paragraph)
    print()

当前列表页的标题、日期、作者及文章详情页链接均可正常获取,但进入详情页后,通过article.find('p')获取段落内容时返回None,无法爬取完整的文章内容,希望得到对应的解决方法。


解决方法

嘿,刚看了你的问题和代码,你在列表页的爬取逻辑完全没问题,问题出在详情页的HTML结构定位上!

你现在用soup_link.find_all('article')再找<p>,但这个网站的详情页里,文章正文其实是嵌套在div.entry-content这个容器里的,不是直接在<article>标签下的第一层<p>。咱们调整一下详情页的爬取逻辑就行:

修改后的完整代码

import bs4
import requests
import re

# 列表页请求
r = requests.get('https://www.the961.com/latest-news/lebanon-news/').text
soup = bs4.BeautifulSoup(r, 'lxml')

for article in soup.find_all('article'):
    # 提取列表页信息
    title = article.h3.text.strip()
    print(f"标题:{title}")
    
    date = article.find('span', class_='byline-part date')
    if date:
        print(f"日期:{date.text.strip()}")
    
    author = article.find('span', class_="byline-part author")
    if author:
        print(f"作者:{author.text.strip()}")
    
    link = article.find('h3', class_='title').a['href']
    print(f"详情链接:{link}")
    
    # 处理详情页内容
    try:
        link_r = requests.get(link).text
        soup_link = bs4.BeautifulSoup(link_r, 'lxml')
        
        # 定位文章正文的容器(这是关键!)
        content_box = soup_link.find('div', class_='entry-content')
        if content_box:
            print("文章内容:")
            # 提取所有段落,跳过空内容
            all_paragraphs = content_box.find_all('p')
            for p in all_paragraphs:
                p_text = p.text.strip()
                if p_text:
                    print(p_text)
        else:
            print("未找到文章内容容器")
    
    except Exception as e:
        print(f"爬取该文章时出错:{str(e)}")
    
    # 用分隔线区分不同文章,输出更清晰
    print("-" * 60)

关键调整点说明:

  1. 定位正确的内容容器:该网站详情页的所有正文段落都放在class="entry-content"的div里,所以先找到这个容器再提取段落,就不会返回None了
  2. 获取所有段落:用find_all('p')代替find('p'),这样能拿到完整的文章内容,而不是只取第一个段落
  3. 异常处理:加入try-except块,避免单个文章爬取失败(比如链接失效、网络问题)导致整个程序中断
  4. 文本清洗:用strip()去除文本前后的多余空白字符,让输出更整洁

这样修改后,你就能正常获取到详情页的完整文章内容啦!

内容的提问来源于stack exchange,提问作者Isolated Gamer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 14:47:33