You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量URL重定向匹配:Bash脚本失效后的最优方案咨询

Hey Kevin, let's tackle this URL matching problem step by step. Here's what you need to know:

Which Method Gives the Highest Matching Accuracy?

Looking at your URL examples, the core issue is that old and new URLs have consistent content but differ in formatting (extra ID path, case, separators like --- vs -, trailing slashes).

In this scenario, Python with string normalization + dictionary mapping is the most accurate approach. Here's why:

  • Python lets you flexibly handle all those formatting differences (case conversion, replacing special separators, stripping extra paths) with custom logic.
  • Unlike your Bash awk script, it can easily clean up edge cases that break matching (like the --- vs - in your example, which is why your script produced an empty file).
  • MySQL SQL is less ideal here because string manipulation in SQL is clunky and less flexible for these kinds of text normalization tasks.
How to Implement This in Python

Below is a practical Python script that handles URL cleaning, matching, and even generates your .htaccess rules directly.

Core Approach

  1. First, we'll read the new site CSV, clean each URL's title to a standardized format, and store a map of cleaned titles to original new URLs.
  2. Then, we'll process the old site CSV, clean each old URL's title the same way, and look up the matching new URL from our map.
  3. Finally, we'll save the matched pairs to combined.csv and generate ready-to-use 301 redirect rules for .htaccess.

Full Python Code

import csv
import re

def clean_url_title(url):
    """Normalize the title part of a URL for matching"""
    # Remove trailing slashes
    url = url.rstrip('/')
    # Grab the last segment of the URL (the title part)
    title_part = url.split('/')[-1]
    # Convert to all lowercase
    title_part = title_part.lower()
    # Replace triple dashes with single dashes first
    title_part = re.sub(r'---+', '-', title_part)
    # Replace any non-alphanumeric/non-dash characters with dashes
    title_part = re.sub(r'[^a-z0-9-]', '-', title_part)
    # Collapse multiple consecutive dashes into one
    title_part = re.sub(r'-+', '-', title_part)
    # Remove leading/trailing dashes
    title_part = title_part.strip('-')
    return title_part

def main():
    # Build a map from cleaned titles to new URLs
    new_url_map = {}
    with open('newsite.csv', 'r', encoding='utf-8') as f:
        reader = csv.reader(f)
        # Skip header row (comment this out if your CSV has no header)
        next(reader)
        for row in reader:
            if not row:
                continue
            # Remove quotes around the URL (matching your CSV example)
            new_url = row[0].strip('"')
            cleaned_title = clean_url_title(new_url)
            new_url_map[cleaned_title] = new_url

    # Process old URLs and match to new ones
    combined_results = [["old-url", "new-url"]]
    htaccess_rules = []

    with open('oldsite.csv', 'r', encoding='utf-8') as f:
        reader = csv.reader(f)
        next(reader)  # Skip header row
        for row in reader:
            if not row:
                continue
            old_url = row[0].strip('"')
            cleaned_title = clean_url_title(old_url)
            # Check if we have a match
            if cleaned_title in new_url_map:
                matched_new_url = new_url_map[cleaned_title]
                combined_results.append([old_url, matched_new_url])
                # Generate 301 redirect rule for .htaccess
                htaccess_rules.append(f"Redirect 301 {old_url} {matched_new_url}")

    # Save the combined CSV
    with open('combined.csv', 'w', encoding='utf-8', newline='') as f:
        writer = csv.writer(f)
        writer.writerows(combined_results)
    print(f"Success! Generated combined.csv with {len(combined_results)-1} matched pairs.")

    # Save the .htaccess file
    with open('.htaccess', 'w', encoding='utf-8') as f:
        f.write('\n'.join(htaccess_rules))
    print("Generated .htaccess with 301 redirect rules.")

if __name__ == "__main__":
    main()

Code Notes

  • The clean_url_title function is the key: it standardizes both old and new URL titles so they can be matched perfectly, even with formatting differences.
  • The script automatically removes quotes around URLs (since your example CSV has quoted URLs).
  • If your CSVs don't have header rows, just comment out the next(reader) lines.
  • If you run into cases where URLs don't match, you can tweak the clean_url_title function—for example, adjusting the regex to handle additional special characters or formatting quirks.
Why Your Bash Script Failed

Your awk script uses / as the field separator, which works for grabbing the title segment, but it only handles lowercase conversion and trailing slash removal. It doesn't account for the --- in old URLs vs - in new URLs, which means the standardized titles don't match. Python's flexible string handling fixes this easily.

内容的提问来源于stack exchange,提问作者Kevin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:50:02