网页爬取SCMP新闻文章时无法获取正文内容求助
问题描述
我正在从SCMP网站爬取新闻文章,能够获取文章的标题与作者名称,但无法获取正文内容。尝试了两种方法均未成功,代码如下:
第一种方法
options = webdriver.ChromeOptions() lists = ['disable-popup-blocking'] caps = DesiredCapabilities().CHROME caps["pageLoadStrategy"] = "normal" driver.get('https://www.scmp.com/news/asia/east-asia/article/3199400/japan-asean-hold-summit-tokyo-around-december-2023-japanese-official') driver.implicitly_wait(5) bsObj = BeautifulSoup(driver.page_source, 'html.parser') text_res = bsObj.select('div[class="details__body body"]') text = "" for item in text_res: if item.get_text() == "": continue text = text + item.get_text().strip() + "\n"
第二种方法
options = webdriver.ChromeOptions() driver = webdriver.Chrome(executable_path= r"E:\chromedriver\chromedriver.exe", options=options) #add your chrome path driver.get('https://www.scmp.com/news/asia/east-asia/article/3199400/japan-asean-hold-summit-tokyo-around-december-2023-japanese-official') driver.implicitly_wait(5) a = driver.find_element_by_class_name("details__body body").text print(a)
问题分析
- 语法错误:第二种方法中
find_element_by_class_name仅支持单个类名,不能用空格传递多个类名(details__body body是两个类),直接导致元素查找失败。 - 等待策略不足:
implicitly_wait(5)的固定等待时间可能不足以让动态加载的正文完全渲染,尤其是网站有反爬延迟加载的情况。 - 选择器与内容结构不匹配:即使定位到父容器,正文可能嵌套在更下层的子元素(如
<p>标签)中,直接提取父容器文本可能无法获取有效内容;另外网站可能更新了页面结构,原选择器已失效。 - 反爬拦截:SCMP可能通过浏览器特征检测、Cookie验证等方式限制爬虫,未处理的情况下页面可能无法正常加载正文。
解决方案
1. 修复元素定位语法
将多类选择改为css_selector,这是支持多类组合定位的正确方式:
a = driver.find_element_by_css_selector("div.details__body.body").text print(a)
2. 改用显式等待确保内容加载
使用WebDriverWait等待正文元素可见或包含文本,比隐式等待更可靠:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 初始化driver代码... driver.get('目标文章URL') # 最长等待10秒,直到正文容器可见 wait = WebDriverWait(driver, 10) body_container = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "div.details__body.body"))) # 提取容器内所有p标签的文本(SCMP正文通常嵌套在p标签中) paragraphs = body_container.find_elements_by_tag_name("p") full_text = "\n".join([p.text.strip() for p in paragraphs if p.text.strip()]) print(full_text)
3. 添加反爬规避参数
在ChromeOptions中加入模拟正常浏览器的配置,避免被识别为爬虫:
options = webdriver.ChromeOptions() # 禁用自动化识别特征 options.add_argument("--disable-blink-features=AutomationControlled") # 设置正常用户代理 options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 移除自动化提示 options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False)
4. 验证页面结构
右键目标文章页面的正文区域,检查元素确认当前的类名和层级结构。如果SCMP更新了页面布局,需要同步调整选择器(例如正文可能迁移到article-body类容器下)。
内容的提问来源于stack exchange,提问作者Starlord22
相关产品推荐
相关产品推荐

