You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup按顺序抓取重复折叠面板内容?

Scrape Accordion Items in Order with BeautifulSoup

Got it, let's break down how to solve this problem exactly as you need it—grabbing every accordion item's content in the exact order it appears, whether it has paragraphs, lists, or a mix.

Step 1: Setup & Target All Accordion Items

First, get your environment set up with BeautifulSoup and parse your HTML (either from a URL or local file). Then, target all div elements that include the accordion-item class—this catches both active and non-active items since you mentioned they repeat with different data.

from bs4 import BeautifulSoup
# Uncomment below if fetching from a URL
# import requests

# Example setup (adjust to your source):
# response = requests.get("your_target_page_url")
# soup = BeautifulSoup(response.text, "html.parser")

# Or for a local HTML file:
# with open("your_html_file.html", "r") as f:
#     soup = BeautifulSoup(f.read(), "html.parser")

# Grab all accordion items (matches any div with "accordion-item" in its class)
accordion_items = soup.find_all('div', class_=lambda cls: cls and 'accordion-item' in cls)

Step 2: Traverse Each Item & Scrape Content in Order

For every accordion item, we'll first extract the title, then work through the content section element by element to preserve the original order. This handles cases where items only have paragraphs, or mix paragraphs with lists.

# Loop through each item and print scraped content
for item_num, item in enumerate(accordion_items, start=1):
    print(f"--- Accordion Item {item_num} ---")
    
    # Extract and print the title
    title_text = item.find('p', class_='accordion-title').get_text(strip=True)
    print(f"**Title:** {title_text}")
    
    # Get the content container
    content_container = item.find('div', class_='accordion-content')
    
    # Iterate through paragraphs and lists in the order they appear
    for element in content_container.find_all(['p', 'ul'], recursive=False):
        if element.name == 'p':
            # Handle paragraph content
            print(f"- {element.get_text(strip=True)}")
        elif element.name == 'ul':
            # Handle list items
            print("  - List items:")
            for li in element.find_all('li'):
                print(f"    * {li.get_text(strip=True)}")
    
    print("\n")

Step 3: Store as Structured Data (Optional)

If you want to save the scraped data instead of just printing it, structure it into a list of dictionaries for easy reuse (like saving to a JSON file later):

scraped_data = []

for item in accordion_items:
    item_details = {}
    
    # Add title to the dictionary
    item_details['title'] = item.find('p', class_='accordion-title').get_text(strip=True)
    
    # Add ordered content
    content_list = []
    content_container = item.find('div', class_='accordion-content')
    
    for element in content_container.find_all(['p', 'ul'], recursive=False):
        if element.name == 'p':
            content_list.append({
                'type': 'paragraph',
                'text': element.get_text(strip=True)
            })
        elif element.name == 'ul':
            content_list.append({
                'type': 'list',
                'items': [li.get_text(strip=True) for li in element.find_all('li')]
            })
    
    item_details['content'] = content_list
    scraped_data.append(item_details)

# Use scraped_data for analysis, saving to a file, etc.

Key Details

  • recursive=False ensures we only grab direct children of the content container—this keeps the content order exactly as it appears in the HTML.
  • The lambda in find_all for accordion items lets us catch both accordion-item and accordion-item-active classes without listing them separately.
  • Items without a <ul> tag are handled automatically—the code just skips the list section and processes only paragraphs.

内容的提问来源于stack exchange,提问作者Apurv Anand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:57:38