如何用Python BeautifulSoup提取嵌套元素与指定数据并分栏导出Excel
1. Finding <a> Tags in Nested div/p Structures with BeautifulSoup
No need to overcomplicate this—BeautifulSoup handles nested tags seamlessly right out of the box. The core tool here is the find_all() method, which recursively searches through all child elements by default.
Example Code:
Let’s say your HTML structure looks like this:
<div class="listing"> <p>View <a href="/car1">this car</a> for sale.</p> <div class="details"> <p>More info <a href="/car1/details">here</a>.</p> </div> </div>
Here’s how to extract all <a> tags, even the nested ones:
from bs4 import BeautifulSoup # Load your raw HTML content (from a file or web request) with open('page.html', 'r') as f: html_content = f.read() soup = BeautifulSoup(html_content, 'html.parser') # Get every <a> tag in the entire document all_links = soup.find_all('a') # If you only want links inside a specific parent (e.g., div with class "listing") target_links = soup.find('div', class_='listing').find_all('a') # Extract clean text and URLs from each link for link in all_links: link_text = link.get_text(strip=True) link_url = link.get('href') print(f"Link: {link_text} | URL: {link_url}")
Quick Tips:
find_all('a')works recursively, so it’ll dig through every nesteddiv/pautomatically.- Use
find()first to target a specific parent element if you don’t want all links on the page. get_text(strip=True)removes extra whitespace from link text, andget('href')grabs the URL attribute reliably.
2. Splitting année (Year) and kilométrage (Mileage) into Separate Excel Columns
Assuming your combined data looks like 2018 - 95 000 km or 2020 | 110000 km, we can split these values using regex or string splitting before exporting to Excel. Pandas is the easiest tool for this job—it simplifies both data manipulation and Excel exports.
Example Code:
import pandas as pd import re # Sample data where year and mileage are bundled in one column raw_data = { 'combined_info': [ '2017 - 80 000 km', '2019 | 105 000 km', '2021 120000 km' ] } df = pd.DataFrame(raw_data) # Use regex to extract year and mileage df['année'] = df['combined_info'].apply(lambda x: re.search(r'(\d{4})', x).group(1)) df['kilométrage'] = df['combined_info'].apply(lambda x: re.search(r'(\d[\d\s]+)km', x).group(1).replace(' ', '')) # Convert mileage to integer for easier analysis (optional but recommended) df['kilométrage'] = df['kilométrage'].astype(int) # Export to Excel without the original combined column df.drop('combined_info', axis=1).to_excel('vehicle_data.xlsx', index=False)
Alternative for Consistent Separators:
If your data uses a fixed separator like -, you can skip regex and use simple string splitting:
df[['année', 'kilométrage']] = df['combined_info'].str.split(' - ', expand=True) # Clean up the mileage column df['kilométrage'] = df['kilométrage'].str.replace(' km', '').str.replace(' ', '').astype(int)
Notes:
- Adjust the regex or split separator to match your actual data format (e.g., commas, pipes, or just spaces).
- Test the regex on a few sample entries to make sure it captures all edge cases (like missing spaces or different separators).
内容的提问来源于stack exchange,提问作者Jiess

