使用Selenium+BeautifulSoup无法抓取新闻URL正文,求解决方案
问题分析与修复方案
核心问题
- 内层循环位置错误:把遍历链接的循环嵌套在文章遍历逻辑里,导致每次新增一个链接就会重复处理所有历史链接,造成冗余抓取甚至数据混乱。
- 反爬拦截:
requests.get()请求缺少浏览器标识的请求头,被网站识别为非浏览器请求,返回内容异常,导致BeautifulSoup无法定位到正文元素。 - 空值处理缺失:如果
find()找不到目标元素,直接调用.text会抛出AttributeError,导致程序崩溃。
修复后的代码
from selenium.webdriver import ActionChains, Keys import time from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium import webdriver from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException from bs4 import BeautifulSoup import requests import pandas as pd # 初始化浏览器 driver = webdriver.Chrome() driver.maximize_window() url = 'https://www.iol.co.za/news/south-africa/eastern-cape' wait = WebDriverWait(driver, 5) driver.get(url) time.sleep(3) # 获取文章列表 articles = wait.until(EC.presence_of_all_elements_located( (By.XPATH, "//article//*[(name()='h1' or name()='h2' or name()='h3' or name()='h4' or name()='h5' or name()='h6') and string-length(text())>0]/ancestor::article") )) article_link = [] full_text = [] # 第一步:先批量收集所有文章链接 for article in articles: link = article.find_element(By.XPATH, ".//a") news_link = link.get_attribute('href') article_link.append(news_link) # 第二步:遍历链接抓取正文,添加浏览器请求头绕过反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } for j in article_link: try: news_response = requests.get(j, headers=headers) news_response.raise_for_status() # 检查请求是否成功返回 news_soup = BeautifulSoup(news_response.content, 'html.parser') # 定位正文元素,兼容备选选择器 art_cont = news_soup.find('div', class_='Article__StyledArticleContent-sc-uw4nkg-0') if art_cont: full_text.append(art_cont.get_text(strip=True)) else: art_cont = news_soup.find('div', class_='article-content') full_text.append(art_cont.get_text(strip=True) if art_cont else '正文未找到') except Exception as e: print(f"处理链接 {j} 时出错: {str(e)}") full_text.append('抓取失败') # 输出结果 print("文章链接列表:") for link in article_link: print(link) print("\n正文内容列表:") for text in full_text: print(text) driver.quit()
关键优化点
- 分离收集与抓取逻辑:先完成所有链接的收集,再统一遍历抓取,避免重复操作。
- 添加请求头:模拟浏览器请求,绕过网站基础反爬检测。
- 异常捕获:处理请求失败、元素定位失败等情况,避免程序直接崩溃。
- 备选定位方案:提供备用选择器,应对网站类名变更的情况,提升代码鲁棒性。
备选方案(纯Selenium抓取)
如果requests仍被反爬拦截,可直接用Selenium模拟浏览器打开每个链接抓取,完全复刻用户浏览行为:
# 替换第二步的抓取逻辑为: for j in article_link: try: driver.get(j) wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'Article__StyledArticleContent-sc-uw4nkg-0'))) art_cont = driver.find_element(By.CLASS_NAME, 'Article__StyledArticleContent-sc-uw4nkg-0') full_text.append(art_cont.text.strip()) except Exception as e: print(f"处理链接 {j} 时出错: {str(e)}") full_text.append('抓取失败')
内容的提问来源于stack exchange,提问作者TG_
相关产品推荐
相关产品推荐

