如何将多个div class中的数据提取至pandas DataFrame?
map-item div data attributes to ordered Pandas DataFrame Got it, let's walk through this step by step—you're already past the tricky scraping part, so converting to a DataFrame is straightforward once you know how to extract those data attributes properly.
Step 1: Extract Data Attributes from Each Div
Each map-item div has prefixed attributes like data-companyname or data-country. Instead of hardcoding each attribute name (which is brittle if the dashboard adds more fields later), we can dynamically extract all attributes starting with data- and clean up their keys by stripping the data- prefix.
Step 2: Build a List of Dictionaries
We'll loop through each div in your list, convert its relevant attributes into a dictionary, and collect all these dictionaries into a list—this is the easiest format for Pandas to turn into a DataFrame.
Step 3: Convert to DataFrame & Add Serial Number
Once we have our list of dictionaries, we can pass it directly to pd.DataFrame(), then insert a serial number column at the front (starting from 1, since that's what most structured tables use).
Full Working Code Example
import pandas as pd from bs4 import BeautifulSoup # Assume you've already fetched the page and parsed it with BeautifulSoup # Example: soup = BeautifulSoup(page_content, 'html.parser') map_items = soup.find_all('div', class_='map-item') # Initialize empty list to store parsed data parsed_data = [] for item in map_items: # Extract all data-* attributes, strip the 'data-' prefix from keys item_dict = { key.replace('data-', ''): value for key, value in item.attrs.items() if key.startswith('data-') } parsed_data.append(item_dict) # Convert list of dicts to DataFrame df = pd.DataFrame(parsed_data) # Add serial number column at position 0 (first column) df.insert(0, '序号', range(1, len(df) + 1)) # Optional: Fill empty values if some divs are missing attributes df = df.fillna('N/A') # Preview the result print(df.head())
Key Notes & Troubleshooting
- Handling Missing Attributes: If some
map-itemdivs lack certain data attributes, Pandas will automatically fill those cells withNaN. Thefillna('N/A')line replaces those with a user-friendly placeholder—adjust this to your needs. - Renaming Columns: If you want to clean up column names (e.g.,
data-company-namebecomescompany_name), usedf.rename(columns={'company-name': 'company_name'}, inplace=True)after creating the DataFrame. - Verifying Extraction: To make sure you're getting all the right attributes, print
item.attrsfor one div first—this will show you every attribute attached to the element, so you can confirm which ones are prefixed withdata-.
内容的提问来源于stack exchange,提问作者Funkeh-Monkeh

