Python处理列表内容并导出至CSV多列的技术需求
Let's break down how to achieve your three goals step by step. The key is to extract each component cleanly using your fixed tags as delimiters, then format everything into a structured CSV file.
Step 1: Define Your Raw Data
First, let's assume your input list looks like this (swap in your actual gender/country values where needed):
raw_data = [ 'ahref="/name<sdfjnsflnsdlkfnsdlf>male<sdlfknsldfnwlkfn>USA<dfnsdflnsdfsldn>This is a graduate student<yogm>', 'ahref="/name<sdfjnsflnsdlkfnsdlf>female<sdlfknsldfnwlkfn>Canada<dfnsdflnsdfsldn>This is a graduate student<yogm>', 'ahref="/name<sdfjnsflnsdlkfnsdlf>male<sdlfknsldfnwlkfn>UK<dfnsdflnsdfsldn>This is a graduate student<yogm>' ]
Step 2: Process Each Entry to Extract Components
Instead of just deleting unwanted substrings (which would merge all data into one mess), we'll split each string using your fixed tags as separators to pull out each piece of data:
def process_entry(entry): # Extract the link path and add the abc.com prefix link_part = entry.split('ahref="')[1] link_path, remaining = link_part.split('<sdfjnsflnsdlkfnsdlf>') full_link = f"abc.com{link_path}" # Pull out the gender value gender, remaining = remaining.split('<sdlfknsldfnwlkfn>') # Pull out the country value country, remaining = remaining.split('<dfnsdflnsdfsldn>') # Extract the description and clean up extra whitespace description = remaining.split('<yogm>')[0].strip() return full_link, gender, country, description # Process all entries in your list processed_entries = [process_entry(item) for item in raw_data]
Step 3: Write to CSV
Use Python's built-in csv module to save the processed data into a CSV file with clear column headers:
import csv # Define your CSV column headers headers = ["Link", "Gender", "Country", "Description"] # Write the data to a CSV file with open("student_records.csv", "w", newline="", encoding="utf-8") as csv_file: writer = csv.writer(csv_file) writer.writerow(headers) # Add headers first writer.writerows(processed_entries) # Write all processed data rows
Integrating with Your Existing BeautifulSoup Code
You mentioned you already have code handling the link prefix. If your data comes from parsing HTML with BeautifulSoup, here's how to adapt this solution to fit your workflow:
prefix = "abc.com" processed_entries = [] for element in letters: # Reuse your existing link prefix logic link_path = element.a["href"] full_link = f"{prefix}{link_path}" # Extract the text content containing gender/country/description text_content = element.get_text() # Parse the text content the same way as before remaining = text_content.split('<sdfjnsflnsdlkfnsdlf>')[1] gender, remaining = remaining.split('<sdlfknsldfnwlkfn>') country, remaining = remaining.split('<dfnsdflnsdfsldn>') description = remaining.split('<yogm>')[0].strip() processed_entries.append( (full_link, gender, country, description) ) # Write to CSV using the same code as above
Quick Notes
- This code works assuming your tags (
<sdfjnsflnsdlkfnsdlf>, etc.) are consistent across all entries. If there are variations, you could switch to regular expressions for more flexible parsing. - The final CSV will be saved in your script's directory, with separate columns for each piece of data you need.
内容的提问来源于stack exchange,提问作者Lko

