You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup遍历HTML元素时循环无法迭代的问题求助

问题分析与解决方案

你的代码问题出在**find_all('div', {'id':'retailers'})**——HTML规范里id是唯一标识,整个页面只有一个id为retailers的div,所以find_all返回的列表只有一个元素。你遍历这个列表时,自然只能拿到这个div里的第一个h2文本。

修正后的代码实现

要提取所有国家名称、对应的市场名称及URL,需要先定位到唯一的retailers容器,再遍历容器内的所有国家区块:

from bs4 import BeautifulSoup
import requests

html_text = requests.get('https://www.freshplaza.com/europe/content/retailers/').text
soup = BeautifulSoup(html_text, 'lxml')

# 获取唯一的retailers容器
retailers_container = soup.find('div', id='retailers')

# 遍历容器内所有的国家标题(h2)及其后续的市场列表
for country_section in retailers_container.find_all(['h2', 'ul'], recursive=False):
    if country_section.name == 'h2':
        current_country = country_section.text.strip()
        print(f"\n国家: {current_country}")
    elif country_section.name == 'ul':
        # 遍历当前国家下的所有市场链接
        for market_item in country_section.find_all('li'):
            market_name = market_item.a.text.strip()
            market_url = market_item.a['href']
            print(f"- 市场: {market_name} | URL: {market_url}")

关键说明

  • 用find('div', id='retailers')替代find_all,直接获取唯一的容器元素。
  • 用find_all(['h2', 'ul'], recursive=False)遍历容器的直接子元素,确保按顺序获取国家标题和对应的市场列表。
  • 区分h2(国家名称)和ul(市场列表),逐个提取对应信息。

内容的提问来源于stack exchange,提问作者harderfasterstronger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 16:50:18