求助:在Spyder与Pandas中实现移动端网址m.替换的正则表达式方案——避免域名误匹配
Hey there! Let's sort out this URL conversion issue you're dealing with in Spyder using Pandas. The key problem here is targeting only the m. prefix at the start of the domain (right after http:// or https://) instead of any random m in the URL. Your initial regex tried an exclusion approach, which got messy—let's use a precise matching strategy instead.
The Right Regex Approach
Instead of trying to exclude certain domains, we can directly match the exact pattern we want to replace: the m. that comes immediately after the HTTP/HTTPS protocol. Here's the practical code for your Pandas workflow:
import pandas as pd # Sample DataFrame with your test URLs df = pd.DataFrame({ 'mobile_url': [ 'https://m.gsmarena.com/sony_xperia_5_iv-price-11838.php', 'https://redfm.ie/', 'http://m.example.org/page.html', 'https://www.m.fake.com' # Bonus: handles edge case with www. before m. ] }) # Convert mobile URLs to desktop versions df['desktop_url'] = df['mobile_url'].str.replace(r'^(https?://)m\.', r'\1', regex=True)
Breakdown of the Regex:
^: Anchors the match to the start of the string, ensuring we only modify the beginning of the URL.(https?://): Captures the HTTP/HTTPS protocol (the?makes thesoptional, so it works for bothhttpandhttps). We save this as group\1to reuse it in the replacement.m\.: Matches the literalm.(we escape the dot with\because dots act as wildcards in regex).
Test Results:
| Original Mobile URL | Converted Desktop URL |
|---|---|
https://m.gsmarena.com/sony_xperia_5_iv-price-11838.php | https://gsmarena.com/sony_xperia_5_iv-price-11838.php |
https://redfm.ie/ | https://redfm.ie/ (no change, which is exactly what we want!) |
http://m.example.org/page.html | http://example.org/page.html |
Why Your Initial Regex Didn't Work
Your original regex m\.^((?!(\.ie|.com)).)*$ had a couple of critical issues:
- The
^anchor was placed afterm\., which inverted the logic—you were trying to matchm.at the end of the string instead of the start. - The exclusion-based approach was overcomplicating the problem. We don't need to block
.ieor.comdomains; we just need to target the specificm.prefix we want to remove.
Bonus: Handling Edge Cases
If you also need to account for URLs like https://www.m.example.com (where m. comes after www.), adjust the regex to preserve the www. part:
df['desktop_url'] = df['mobile_url'].str.replace(r'^(https?://)(www\.)?m\.', r'\1\2', regex=True)
This will turn https://www.m.example.com into https://www.example.com.
内容的提问来源于stack exchange,提问作者Aki

