Python3过滤含/Thumbnails的图片链接列表及CSV批量处理问题
Hey there! Let's break down how to solve your two main problems: filtering out links with /Thumbnails from a list, and handling that at scale for your 25,000-column CSV file.
1. Filtering a Single List of Image Links
The cleanest and most efficient way to do this is using a list comprehension—it avoids the pitfalls of remove() (like only deleting the first match, or index errors when modifying the list while iterating).
Here's how it works with your example:
pictures = ['Media/Shop/922180cruv.jpg', 'Media/Shop/922180cruvdet.jpg', 'Media/Shop/Thumbnails/922180cruvx320x240.jpg', 'Media/Shop/Thumbnails/922180cruvdetx320x240.jpg'] # Keep only links that DON'T contain "/Thumbnails" filtered_pictures = [pic for pic in pictures if '/Thumbnails' not in pic] print(filtered_pictures) # Output: ['Media/Shop/922180cruv.jpg', 'Media/Shop/922180cruvdet.jpg']
Why this is better than your previous attempts:
- It creates a new list directly, so you don't have to worry about modifying the original list and breaking iteration.
- It works regardless of where the
/Thumbnailsstring appears in the link, no need to rely on fixed indexes. - It's faster and more readable for large lists.
2. Processing Your 25,000-Column CSV File
Since your CSV has so many columns (each containing a list of links), we'll use tools that handle tabular data efficiently. Below are two approaches:
Option 1: Using Pandas (Recommended for Large CSVs)
Pandas is perfect for bulk processing of CSV data. We'll read the file, apply our filter to every column, then save the cleaned data.
Note: If your CSV stores lists as strings (like "['link1', 'link2']"), we'll use ast.literal_eval to convert them back to actual lists first.
import pandas as pd import ast # 1. Read the CSV file df = pd.read_csv("your_input_file.csv") # 2. Define a function to filter a single list of links def filter_thumbnail_links(list_str): # Convert string representation of list to a real list link_list = ast.literal_eval(list_str) # Apply our filter return [link for link in link_list if "/Thumbnails" not in link] # 3. Apply the filter to every cell in the DataFrame df = df.applymap(filter_thumbnail_links) # 4. Save the cleaned CSV df.to_csv("your_filtered_output.csv", index=False)
Option 2: Using Python's Built-in csv Module
If you prefer not to use pandas, you can use the standard library's csv module. This is a bit more verbose but works just as well.
import csv import ast # Open input and output files with open("your_input_file.csv", "r") as infile, open("your_filtered_output.csv", "w", newline="") as outfile: reader = csv.reader(infile) writer = csv.writer(outfile) # Process each row in the CSV for row in reader: cleaned_row = [] # Process each column (cell) in the row for cell in row: # Convert string list to actual list link_list = ast.literal_eval(cell) # Filter out thumbnail links cleaned_links = [link for link in link_list if "/Thumbnails" not in link] # Convert back to string to store in CSV cleaned_row.append(str(cleaned_links)) # Write the cleaned row to the output file writer.writerow(cleaned_row)
Why Your Previous Methods Failed
remove(): Only deletes the first occurrence of a matching string. If you loop to delete all matches, you'll run into index errors because the list length changes as you remove items.- Fixed indexes: Since each list has a different number of thumbnail links (or none at all), you can't rely on static positions to remove elements. List comprehensions bypass this entirely by building a new list from scratch.
内容的提问来源于stack exchange,提问作者kekw

