You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何精准使用BeautifulSoup的find_all或select爬取格斗选手链接?

原始代码

Fighter1Main = []
for i in range(1,3):
            url = Request(f"https://www.sherdog.com/events/a-{page}", headers={'User-Agent': 'Mozilla/5.0'})
            response = urlopen(url).read()
            soup = BeautifulSoup(response, "html.parser")
            for test2 in soup.find_all(class_="fighter left_side"):
                test3 = test2.find_all(itemprop="url")
                Fighter1Main.append(test3)
            page = page + 1

当前输出(嵌套列表)

[[<a href="/fighter/Todd-Medina-61" itemprop="url">
<img alt="Todd 'El Tiburon' Medina" itemprop="image" src="/image_crop/200/300/_images/fighter/20140801074225_IMG_5098.JPG" title="Todd 'El Tiburon' Medina">
</img></a>], [<a href="/fighter/Ricco-Rodriguez-8" itemprop="url">
<img alt="Ricco 'Suave' Rodriguez" itemprop="image" src="/image_crop/200/300/_images/fighter/20141225125221_1MG_9472.JPG" title="Ricco 'Suave' Rodriguez">
</img></a>]]

期望输出(独立标签列表)

<a href="/fighter/Todd-Medina-61" itemprop="url">
<a href="/fighter/Ricco-Rodriguez-8" itemprop="url">

已尝试方案

使用CSS选择器后仍得到嵌套列表:

for test2 in soup.select('.fighter.left_side [itemprop="url"]'):
                Fighter1Main.append(test2)

解决方法

问题根源是find_all返回的是列表,直接append会将整个子列表加入主列表,形成嵌套。以下是两种可行修正方式:

方式1:用extend替代append

extend会将子列表中的元素逐个添加到主列表,避免嵌套:

Fighter1Main = []
page = 1  # 需先初始化page变量
for i in range(1,3):
    url = Request(f"https://www.sherdog.com/events/a-{page}", headers={'User-Agent': 'Mozilla/5.0'})
    response = urlopen(url).read()
    soup = BeautifulSoup(response, "html.parser")
    for test2 in soup.find_all(class_="fighter left_side"):
        test3 = test2.find_all(itemprop="url")
        Fighter1Main.extend(test3)  # 替换append为extend
    page = page + 1

方式2:直接遍历目标标签并逐个添加

跳过中间嵌套循环,直接定位所有符合条件的<a>标签,逐个append:

Fighter1Main = []
page = 1
for i in range(1,3):
    url = Request(f"https://www.sherdog.com/events/a-{page}", headers={'User-Agent': 'Mozilla/5.0'})
    response = urlopen(url).read()
    soup = BeautifulSoup(response, "html.parser")
    # 直接遍历所有目标a标签
    for a_tag in soup.select('.fighter.left_side [itemprop="url"]'):
        Fighter1Main.append(a_tag)
    page = page + 1

内容的提问来源于stack exchange,提问作者yfr code

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 19:50:25