如何用Python的BeautifulSoup按顺序抓取重复折叠面板内容?
Got it, let's break down how to solve this problem exactly as you need it—grabbing every accordion item's content in the exact order it appears, whether it has paragraphs, lists, or a mix.
Step 1: Setup & Target All Accordion Items
First, get your environment set up with BeautifulSoup and parse your HTML (either from a URL or local file). Then, target all div elements that include the accordion-item class—this catches both active and non-active items since you mentioned they repeat with different data.
from bs4 import BeautifulSoup # Uncomment below if fetching from a URL # import requests # Example setup (adjust to your source): # response = requests.get("your_target_page_url") # soup = BeautifulSoup(response.text, "html.parser") # Or for a local HTML file: # with open("your_html_file.html", "r") as f: # soup = BeautifulSoup(f.read(), "html.parser") # Grab all accordion items (matches any div with "accordion-item" in its class) accordion_items = soup.find_all('div', class_=lambda cls: cls and 'accordion-item' in cls)
Step 2: Traverse Each Item & Scrape Content in Order
For every accordion item, we'll first extract the title, then work through the content section element by element to preserve the original order. This handles cases where items only have paragraphs, or mix paragraphs with lists.
# Loop through each item and print scraped content for item_num, item in enumerate(accordion_items, start=1): print(f"--- Accordion Item {item_num} ---") # Extract and print the title title_text = item.find('p', class_='accordion-title').get_text(strip=True) print(f"**Title:** {title_text}") # Get the content container content_container = item.find('div', class_='accordion-content') # Iterate through paragraphs and lists in the order they appear for element in content_container.find_all(['p', 'ul'], recursive=False): if element.name == 'p': # Handle paragraph content print(f"- {element.get_text(strip=True)}") elif element.name == 'ul': # Handle list items print(" - List items:") for li in element.find_all('li'): print(f" * {li.get_text(strip=True)}") print("\n")
Step 3: Store as Structured Data (Optional)
If you want to save the scraped data instead of just printing it, structure it into a list of dictionaries for easy reuse (like saving to a JSON file later):
scraped_data = [] for item in accordion_items: item_details = {} # Add title to the dictionary item_details['title'] = item.find('p', class_='accordion-title').get_text(strip=True) # Add ordered content content_list = [] content_container = item.find('div', class_='accordion-content') for element in content_container.find_all(['p', 'ul'], recursive=False): if element.name == 'p': content_list.append({ 'type': 'paragraph', 'text': element.get_text(strip=True) }) elif element.name == 'ul': content_list.append({ 'type': 'list', 'items': [li.get_text(strip=True) for li in element.find_all('li')] }) item_details['content'] = content_list scraped_data.append(item_details) # Use scraped_data for analysis, saving to a file, etc.
Key Details
recursive=Falseensures we only grab direct children of the content container—this keeps the content order exactly as it appears in the HTML.- The lambda in
find_allfor accordion items lets us catch bothaccordion-itemandaccordion-item-activeclasses without listing them separately. - Items without a
<ul>tag are handled automatically—the code just skips the list section and processes only paragraphs.
内容的提问来源于stack exchange,提问作者Apurv Anand

