移除Pandas文本序列中的URL并转换特殊字符,适配langdetect检测
嘿,我碰到过一模一样的问题!爬取的数据总是带着各种脏东西,把langdetect搞崩太正常了。给你一套我亲测有效的清洗流程,一步步来解决:
第一步:先清掉HTML标签和冗余URL
首先得把那些乱七八糟的HTML标签、无关URL去掉,同时得保住NLP需要的标点。我一般用BeautifulSoup提取纯文本,再用正则干掉URL,最后过滤掉离谱的特殊字符但留着常用标点:
from bs4 import BeautifulSoup import re import pandas as pd def clean_html_and_urls(text): # 先处理空值,避免报错 if pd.isna(text): return "" # 扒掉HTML标签,提取纯文本 soup = BeautifulSoup(text, "html.parser") clean_text = soup.get_text(strip=True) # 把所有http/https开头的链接删掉 clean_text = re.sub(r'https?://\S+|www\.\S+', '', clean_text) # 只保留字母、数字、空格和常用标点,其他奇怪字符全去掉 clean_text = re.sub(r'[^\w\s.,!?;:"\'’“”()\-——]', '', clean_text) return clean_text # 把这个函数应用到你的Pandas序列上 df['clean_text'] = df['raw_text'].apply(clean_html_and_urls)
第二步:解码那些烦人的ASCII编码
爬取数据里常见的两种ASCII编码坑:HTML实体(比如'对应单引号)和Unicode转义(比如\u0027),分开处理就行:
处理HTML实体编码
直接用Python自带的html.unescape()就能搞定,把实体转成正常字符:
import html def decode_html_entities(text): if pd.isna(text): return "" return html.unescape(text) # 把解码步骤加到清洗流程里 df['clean_text'] = df['clean_text'].apply(decode_html_entities)
处理Unicode转义的ASCII编码
如果数据里有\uXXXX这种转义字符,用encode再decode的方式转回来:
def decode_unicode_escape(text): if pd.isna(text): return "" try: # 先转成字节再解码转义 return text.encode('utf-8').decode('unicode-escape') except: # 万一解码失败就返回原文本,别搞崩整个流程 return text df['clean_text'] = df['clean_text'].apply(decode_unicode_escape)
第三步:给langdetect排雷,避免报错
langdetect特别怕空文本、超短文本和残留的奇怪字符,所以清洗后再加一层过滤:
from langdetect import detect, LangDetectException def detect_language(text): # 太短的文本直接标记为未知,别让langdetect瞎猜报错 if len(text.strip()) < 3: return "unknown" try: return detect(text) except LangDetectException: # 遇到识别不了的也标记未知 return "unknown" # 给每条数据打语言标签 df['language'] = df['clean_text'].apply(detect_language) # 现在就能按语言拆分数据了,比如提取英文数据集 english_df = df[df['language'] == 'en']
额外小技巧
- 如果数据还有编码乱码(比如GBK转UTF-8出错那种),可以用
chardet自动检测编码再修复:
import chardet def fix_encoding(text): if pd.isna(text): return "" # 检测文本编码 result = chardet.detect(text.encode('utf-8', errors='ignore')) # 按检测到的编码解码 return text.encode(result['encoding']).decode('utf-8', errors='replace')
这个按需用就行,不是所有情况都需要。
- 可以把所有清洗步骤整合到一个函数里,跑起来更快:
def full_clean(text): if pd.isna(text): return "" # 先解码Unicode转义 text = decode_unicode_escape(text) # 再解码HTML实体 text = html.unescape(text) # 清理HTML和URL soup = BeautifulSoup(text, "html.parser") text = soup.get_text(strip=True) text = re.sub(r'https?://\S+|www\.\S+', '', text) text = re.sub(r'[^\w\s.,!?;:"\'’“”()\-——]', '', text) return text.strip()
这样处理完,数据应该能被langdetect正常识别,同时保住了NLP需要的标点,后续任务也能顺利开展啦!
内容的提问来源于stack exchange,提问作者Nick Duddy
相关产品推荐
相关产品推荐

