新手求助:使用BeautifulSoup爬取The961黎巴嫩新闻详情页时p标签返回None无法获取完整文章
问题:爬取The961详情页文章内容返回None的解决方法
我是一名爬虫新手,尝试爬取The961网站(https://www.the961.com/latest-news/lebanon-news/)的黎巴嫩新闻内容,编写的代码如下:
import bs4 import requests import re r = requests.get('https://www.the961.com/latest-news/lebanon-news/').text soup = bs4.BeautifulSoup(r, 'lxml') for article in soup.find_all('article'): title = article.h3.text print(title) date = article.find('span', class_='byline-part date') if date: print('Date:', date.text) author = article.find('span', class_="byline-part author") if author: print('Author:', author.text) link = article.find('h3', class_='title').a['href'] link_r = requests.get(link).text soup_link = bs4.BeautifulSoup(link_r, 'lxml') // scraping link from title, then opening that link and trying to scrape the whole article, very new to this so I don't know what to do! for article in soup_link.find_all('article'): paragraph = article.find('p') print(paragraph) print()
当前列表页的标题、日期、作者及文章详情页链接均可正常获取,但进入详情页后,通过article.find('p')获取段落内容时返回None,无法爬取完整的文章内容,希望得到对应的解决方法。
解决方法
嘿,刚看了你的问题和代码,你在列表页的爬取逻辑完全没问题,问题出在详情页的HTML结构定位上!
你现在用soup_link.find_all('article')再找<p>,但这个网站的详情页里,文章正文其实是嵌套在div.entry-content这个容器里的,不是直接在<article>标签下的第一层<p>。咱们调整一下详情页的爬取逻辑就行:
修改后的完整代码
import bs4 import requests import re # 列表页请求 r = requests.get('https://www.the961.com/latest-news/lebanon-news/').text soup = bs4.BeautifulSoup(r, 'lxml') for article in soup.find_all('article'): # 提取列表页信息 title = article.h3.text.strip() print(f"标题:{title}") date = article.find('span', class_='byline-part date') if date: print(f"日期:{date.text.strip()}") author = article.find('span', class_="byline-part author") if author: print(f"作者:{author.text.strip()}") link = article.find('h3', class_='title').a['href'] print(f"详情链接:{link}") # 处理详情页内容 try: link_r = requests.get(link).text soup_link = bs4.BeautifulSoup(link_r, 'lxml') # 定位文章正文的容器(这是关键!) content_box = soup_link.find('div', class_='entry-content') if content_box: print("文章内容:") # 提取所有段落,跳过空内容 all_paragraphs = content_box.find_all('p') for p in all_paragraphs: p_text = p.text.strip() if p_text: print(p_text) else: print("未找到文章内容容器") except Exception as e: print(f"爬取该文章时出错:{str(e)}") # 用分隔线区分不同文章,输出更清晰 print("-" * 60)
关键调整点说明:
- 定位正确的内容容器:该网站详情页的所有正文段落都放在
class="entry-content"的div里,所以先找到这个容器再提取段落,就不会返回None了 - 获取所有段落:用
find_all('p')代替find('p'),这样能拿到完整的文章内容,而不是只取第一个段落 - 异常处理:加入
try-except块,避免单个文章爬取失败(比如链接失效、网络问题)导致整个程序中断 - 文本清洗:用
strip()去除文本前后的多余空白字符,让输出更整洁
这样修改后,你就能正常获取到详情页的完整文章内容啦!
内容的提问来源于stack exchange,提问作者Isolated Gamer
相关产品推荐
相关产品推荐

