You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+BeautifulSoup无法抓取新闻URL正文,求解决方案

问题分析与修复方案

核心问题

  1. 内层循环位置错误:把遍历链接的循环嵌套在文章遍历逻辑里,导致每次新增一个链接就会重复处理所有历史链接,造成冗余抓取甚至数据混乱。
  2. 反爬拦截:requests.get()请求缺少浏览器标识的请求头,被网站识别为非浏览器请求,返回内容异常,导致BeautifulSoup无法定位到正文元素。
  3. 空值处理缺失:如果find()找不到目标元素,直接调用.text会抛出AttributeError,导致程序崩溃。

修复后的代码

from selenium.webdriver import ActionChains, Keys
import time
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException
from bs4 import BeautifulSoup
import requests
import pandas as pd

# 初始化浏览器
driver = webdriver.Chrome()
driver.maximize_window()
url = 'https://www.iol.co.za/news/south-africa/eastern-cape'
wait = WebDriverWait(driver, 5)
driver.get(url)
time.sleep(3)

# 获取文章列表
articles = wait.until(EC.presence_of_all_elements_located(
    (By.XPATH, "//article//*[(name()='h1' or name()='h2' or name()='h3' or name()='h4' or name()='h5' or name()='h6') and string-length(text())>0]/ancestor::article")
))

article_link = []
full_text = []

# 第一步:先批量收集所有文章链接
for article in articles:
    link = article.find_element(By.XPATH, ".//a")
    news_link = link.get_attribute('href')
    article_link.append(news_link)

# 第二步:遍历链接抓取正文,添加浏览器请求头绕过反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

for j in article_link:
    try:
        news_response = requests.get(j, headers=headers)
        news_response.raise_for_status()  # 检查请求是否成功返回
        news_soup = BeautifulSoup(news_response.content, 'html.parser')
        # 定位正文元素,兼容备选选择器
        art_cont = news_soup.find('div', class_='Article__StyledArticleContent-sc-uw4nkg-0')
        if art_cont:
            full_text.append(art_cont.get_text(strip=True))
        else:
            art_cont = news_soup.find('div', class_='article-content')
            full_text.append(art_cont.get_text(strip=True) if art_cont else '正文未找到')
    except Exception as e:
        print(f"处理链接 {j} 时出错: {str(e)}")
        full_text.append('抓取失败')

# 输出结果
print("文章链接列表:")
for link in article_link:
    print(link)

print("\n正文内容列表:")
for text in full_text:
    print(text)

driver.quit()

关键优化点

  • 分离收集与抓取逻辑:先完成所有链接的收集,再统一遍历抓取,避免重复操作。
  • 添加请求头:模拟浏览器请求,绕过网站基础反爬检测。
  • 异常捕获:处理请求失败、元素定位失败等情况,避免程序直接崩溃。
  • 备选定位方案:提供备用选择器,应对网站类名变更的情况,提升代码鲁棒性。

备选方案(纯Selenium抓取)

如果requests仍被反爬拦截,可直接用Selenium模拟浏览器打开每个链接抓取,完全复刻用户浏览行为:

# 替换第二步的抓取逻辑为:
for j in article_link:
    try:
        driver.get(j)
        wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'Article__StyledArticleContent-sc-uw4nkg-0')))
        art_cont = driver.find_element(By.CLASS_NAME, 'Article__StyledArticleContent-sc-uw4nkg-0')
        full_text.append(art_cont.text.strip())
    except Exception as e:
        print(f"处理链接 {j} 时出错: {str(e)}")
        full_text.append('抓取失败')

内容的提问来源于stack exchange,提问作者TG_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 18:31:43