批量URL重定向匹配:Bash脚本失效后的最优方案咨询
Hey Kevin, let's tackle this URL matching problem step by step. Here's what you need to know:
Looking at your URL examples, the core issue is that old and new URLs have consistent content but differ in formatting (extra ID path, case, separators like --- vs -, trailing slashes).
In this scenario, Python with string normalization + dictionary mapping is the most accurate approach. Here's why:
- Python lets you flexibly handle all those formatting differences (case conversion, replacing special separators, stripping extra paths) with custom logic.
- Unlike your Bash awk script, it can easily clean up edge cases that break matching (like the
---vs-in your example, which is why your script produced an empty file). - MySQL SQL is less ideal here because string manipulation in SQL is clunky and less flexible for these kinds of text normalization tasks.
Below is a practical Python script that handles URL cleaning, matching, and even generates your .htaccess rules directly.
Core Approach
- First, we'll read the new site CSV, clean each URL's title to a standardized format, and store a map of cleaned titles to original new URLs.
- Then, we'll process the old site CSV, clean each old URL's title the same way, and look up the matching new URL from our map.
- Finally, we'll save the matched pairs to
combined.csvand generate ready-to-use 301 redirect rules for.htaccess.
Full Python Code
import csv import re def clean_url_title(url): """Normalize the title part of a URL for matching""" # Remove trailing slashes url = url.rstrip('/') # Grab the last segment of the URL (the title part) title_part = url.split('/')[-1] # Convert to all lowercase title_part = title_part.lower() # Replace triple dashes with single dashes first title_part = re.sub(r'---+', '-', title_part) # Replace any non-alphanumeric/non-dash characters with dashes title_part = re.sub(r'[^a-z0-9-]', '-', title_part) # Collapse multiple consecutive dashes into one title_part = re.sub(r'-+', '-', title_part) # Remove leading/trailing dashes title_part = title_part.strip('-') return title_part def main(): # Build a map from cleaned titles to new URLs new_url_map = {} with open('newsite.csv', 'r', encoding='utf-8') as f: reader = csv.reader(f) # Skip header row (comment this out if your CSV has no header) next(reader) for row in reader: if not row: continue # Remove quotes around the URL (matching your CSV example) new_url = row[0].strip('"') cleaned_title = clean_url_title(new_url) new_url_map[cleaned_title] = new_url # Process old URLs and match to new ones combined_results = [["old-url", "new-url"]] htaccess_rules = [] with open('oldsite.csv', 'r', encoding='utf-8') as f: reader = csv.reader(f) next(reader) # Skip header row for row in reader: if not row: continue old_url = row[0].strip('"') cleaned_title = clean_url_title(old_url) # Check if we have a match if cleaned_title in new_url_map: matched_new_url = new_url_map[cleaned_title] combined_results.append([old_url, matched_new_url]) # Generate 301 redirect rule for .htaccess htaccess_rules.append(f"Redirect 301 {old_url} {matched_new_url}") # Save the combined CSV with open('combined.csv', 'w', encoding='utf-8', newline='') as f: writer = csv.writer(f) writer.writerows(combined_results) print(f"Success! Generated combined.csv with {len(combined_results)-1} matched pairs.") # Save the .htaccess file with open('.htaccess', 'w', encoding='utf-8') as f: f.write('\n'.join(htaccess_rules)) print("Generated .htaccess with 301 redirect rules.") if __name__ == "__main__": main()
Code Notes
- The
clean_url_titlefunction is the key: it standardizes both old and new URL titles so they can be matched perfectly, even with formatting differences. - The script automatically removes quotes around URLs (since your example CSV has quoted URLs).
- If your CSVs don't have header rows, just comment out the
next(reader)lines. - If you run into cases where URLs don't match, you can tweak the
clean_url_titlefunction—for example, adjusting the regex to handle additional special characters or formatting quirks.
Your awk script uses / as the field separator, which works for grabbing the title segment, but it only handles lowercase conversion and trailing slash removal. It doesn't account for the --- in old URLs vs - in new URLs, which means the standardized titles don't match. Python's flexible string handling fixes this easily.
内容的提问来源于stack exchange,提问作者Kevin

