You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫循环仅返回首个迭代结果问题求助

问题解决:爬虫循环重复抓取首个文章内容

你的问题根源很明确:循环内始终从全局的BeautifulSoup对象中查找元素,每次都取第一个article的内容,而不是基于当前循环迭代的article对象去查找它的子元素。

修改后的代码如下:

import requests
import bs4

url = 'https://coreyms.com'

response = requests.get(url)
response.raise_for_status()

schafer = bs4.BeautifulSoup(response.text, 'html.parser')

# 遍历每一篇文章
for article in schafer.find_all('article'):
    # 基于当前article对象查找标题
    header = article.select_one('article a').getText()
    print(header)

    # 基于当前article对象查找段落
    paragraph = article.select_one('div > p').getText()
    print(paragraph)
    
    # 基于当前article对象查找iframe,先判断是否存在
    iframe = article.select_one('iframe')
    if iframe:
        link = iframe.get('src')
        vidID = link.split('/')[4].split('?')[0]
        ytLink = f'https://youtube.com/watch?v={vidID}'
        print(ytLink)
    print()

关键修改点:

  • 将所有schafer.select()替换为article.select_one():select_one()会返回当前元素下匹配的第一个子元素,正好对应单篇文章里的标题、段落和视频框架
  • 添加了iframe存在性判断:避免部分文章没有视频时出现索引错误
  • 去掉了不必要的索引[0]:select_one()直接返回单个元素,无需再取列表第一个项

内容的提问来源于stack exchange,提问作者Slatercj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 23:48:23