DataFrame中模糊匹配:文章标题与对应URL错位匹配问题求助
Got it, you're dealing with a frustrating issue where your article titles and their corresponding URLs are cross-mismatched in your pandas DataFrame—each title is paired with the URL from the opposite adjacent row. Let's get this sorted out step by step.
First, Let's Confirm the Problem Structure
Based on your example, your DataFrame looks like this (titles and URLs are swapped between consecutive rows):
import pandas as pd # Recreate your sample DataFrame df = pd.DataFrame({ "title": [ "Who will be the next president?", "5 ways to make a cocktail", "2 millions raised by this startup", "How did you find your house" ], "urls": [ "https://website/5-ways-to-make-a-cocktail.com", "https://website/who-will-be-the-next-president.com", "https://website/how-did-you-find-your-house.com", "https://website/2-millions-raised-by-this-startup.com" ] })
Step 1: Split & Swap the Mismatched Data
Since the misalignment follows a consistent pattern (row 0 ↔ row 1, row 2 ↔ row 3), we can split the DataFrame into even-indexed and odd-indexed groups, swap their URL columns, then recombine:
# Split into even and odd rows (reset indexes to avoid alignment issues) even_rows = df.iloc[::2].reset_index(drop=True) odd_rows = df.iloc[1::2].reset_index(drop=True) # Swap the 'urls' values between the two groups even_rows["urls"], odd_rows["urls"] = odd_rows["urls"], even_rows["urls"] # Combine back into a fixed DataFrame and re-sort to original order fixed_df = pd.concat([even_rows, odd_rows]).sort_index().reset_index(drop=True)
Step 2: Verify the Fix
Let's print the result to make sure titles now match their correct URLs:
print(fixed_df)
You'll get this corrected output:
title urls 0 Who will be the next president? https://website/who-will-be-the-next-president.com 1 5 ways to make a cocktail https://website/5-ways-to-make-a-cocktail.com 2 2 millions raised by this startup https://website/2-millions-raised-by-this-startup.com 3 How did you find your house https://website/how-did-you-find-your-house.com
Notes for Edge Cases
- If your dataset has an odd number of rows, add a check to handle the final unpaired row (it likely has the correct URL already, so you can append it directly).
- This method works for any even-sized dataset where the misalignment is consistent across all consecutive row pairs.
内容的提问来源于stack exchange,提问作者ML_Enthousiast

