如何遍历DataFrame列并移除网址前缀'www.'及通用域名后缀(.com/.org/.net等)
www. Prefix and Common Suffixes Hey there! Let's refine your URL cleaning logic into a robust, efficient script for your pandas DataFrame. Your initial loop idea is a solid starting point, but we can make this way better using pandas' built-in vectorized string operations—they’re faster for large datasets and cleaner to write.
Step-by-Step Solution
First, let's set up a sample DataFrame to test our code:
import pandas as pd # Sample DataFrame with various URL formats df = pd.DataFrame({ 'url': [ 'www.example.com', 'test.org', 'www.another-site.net', 'subdomain.example.com', 'plain-url', 'www.blog.co.uk' # Bonus: handles URLs with non-target suffixes ] })
Method 1: Vectorized String Operations (Recommended)
Use pandas' str.replace() with regular expressions to handle both prefix and suffix removal in one go. This is the most efficient approach for large datasets.
# Remove www. prefix and common suffixes (.com, .org, .net) df['cleaned_url'] = df['url'] \ .str.replace(r'^www\.', '', regex=True) # Strip leading www. .str.replace(r'\.(com|org|net)$', '', regex=True) # Strip trailing suffixes
Let’s break down the regex:
^www\.: Matches the exact stringwww.at the start of the URL (^denotes the start,\.escapes the dot since it’s a special regex character)\.(com|org|net)$: Matches.com,.org, or.netat the end of the URL ($denotes the end)
Method 2: Loop-Based Approach (For Reference)
If you prefer a loop (not ideal for large DataFrames), here’s a fixed version of your initial logic. We’ll fix variable overwriting and handle multiple suffixes properly:
cleaned_urls = [] suffixes = ('.com', '.org', '.net') for url in df['url']: temp_url = url # Remove www. prefix if temp_url.startswith('www.'): temp_url = temp_url[4:] # Remove common suffixes if temp_url.endswith(suffixes): # Get the length of the matching suffix to remove suffix_length = next(len(s) for s in suffixes if temp_url.endswith(s)) temp_url = temp_url[:-suffix_length] cleaned_urls.append(temp_url) df['cleaned_url_loop'] = cleaned_urls
Check the Results
Run print(df) to see the cleaned URLs:
url cleaned_url cleaned_url_loop 0 www.example.com example example 1 test.org test test 2 www.another-site.net another-site another-site 3 subdomain.example.com subdomain.example subdomain.example 4 plain-url plain-url plain-url 5 www.blog.co.uk blog.co.uk blog.co.uk
Extend to More Suffixes
If you need to handle additional suffixes (like .io, .co.uk), just update the regex in Method 1:
df['cleaned_url'] = df['url'] \ .str.replace(r'^www\.', '', regex=True) \ .str.replace(r'\.(com|org|net|io|co\.uk)$', '', regex=True)
内容的提问来源于stack exchange,提问作者Andrew Hicks

