You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历DataFrame列并移除网址前缀'www.'及通用域名后缀(.com/.org/.net等)

Clean URLs in a Pandas DataFrame: Remove www. Prefix and Common Suffixes

Hey there! Let's refine your URL cleaning logic into a robust, efficient script for your pandas DataFrame. Your initial loop idea is a solid starting point, but we can make this way better using pandas' built-in vectorized string operations—they’re faster for large datasets and cleaner to write.

Step-by-Step Solution

First, let's set up a sample DataFrame to test our code:

import pandas as pd

# Sample DataFrame with various URL formats
df = pd.DataFrame({
    'url': [
        'www.example.com',
        'test.org',
        'www.another-site.net',
        'subdomain.example.com',
        'plain-url',
        'www.blog.co.uk'  # Bonus: handles URLs with non-target suffixes
    ]
})

Method 1: Vectorized String Operations (Recommended)

Use pandas' str.replace() with regular expressions to handle both prefix and suffix removal in one go. This is the most efficient approach for large datasets.

# Remove www. prefix and common suffixes (.com, .org, .net)
df['cleaned_url'] = df['url'] \
    .str.replace(r'^www\.', '', regex=True)  # Strip leading www.
    .str.replace(r'\.(com|org|net)$', '', regex=True)  # Strip trailing suffixes

Let’s break down the regex:

  • ^www\.: Matches the exact string www. at the start of the URL (^ denotes the start, \. escapes the dot since it’s a special regex character)
  • \.(com|org|net)$: Matches .com, .org, or .net at the end of the URL ($ denotes the end)

Method 2: Loop-Based Approach (For Reference)

If you prefer a loop (not ideal for large DataFrames), here’s a fixed version of your initial logic. We’ll fix variable overwriting and handle multiple suffixes properly:

cleaned_urls = []
suffixes = ('.com', '.org', '.net')

for url in df['url']:
    temp_url = url
    # Remove www. prefix
    if temp_url.startswith('www.'):
        temp_url = temp_url[4:]
    # Remove common suffixes
    if temp_url.endswith(suffixes):
        # Get the length of the matching suffix to remove
        suffix_length = next(len(s) for s in suffixes if temp_url.endswith(s))
        temp_url = temp_url[:-suffix_length]
    cleaned_urls.append(temp_url)

df['cleaned_url_loop'] = cleaned_urls

Check the Results

Run print(df) to see the cleaned URLs:

url       cleaned_url cleaned_url_loop
0       www.example.com           example           example
1             test.org              test              test
2  www.another-site.net  another-site  another-site
3    subdomain.example.com    subdomain.example    subdomain.example
4             plain-url        plain-url        plain-url
5        www.blog.co.uk      blog.co.uk      blog.co.uk

Extend to More Suffixes

If you need to handle additional suffixes (like .io, .co.uk), just update the regex in Method 1:

df['cleaned_url'] = df['url'] \
    .str.replace(r'^www\.', '', regex=True) \
    .str.replace(r'\.(com|org|net|io|co\.uk)$', '', regex=True)

内容的提问来源于stack exchange,提问作者Andrew Hicks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:04:10