面向Python的巴西葡萄牙语与欧洲葡萄牙语文本分类方法咨询
葡萄牙语文本(巴西/欧洲)分类的Python实现方案
一、基于专属词汇的规则分类法
两类葡语的词汇差异比语法差异更易识别,适合快速过滤明确文本:
- 整理专属词汇集:收集仅在巴西葡语(BR)或欧洲葡语(PT)中高频使用的词汇,例如:
- BR专属:
celular,ônibus,sanduíche,pão de queijo,xícara,óleo diesel - PT专属:
telemóvel,autocarro,bocadillo,chávena,gasóleo,carrinho de compras
- BR专属:
- 实现逻辑:统计文本中两类词汇的出现频次,高频者对应类别
- 代码示例:
# 定义专属词汇集合 br_vocab = {'celular', 'ônibus', 'sanduíche', 'pão de queijo', 'xícara', 'óleo diesel'} pt_vocab = {'telemóvel', 'autocarro', 'bocadillo', 'chávena', 'gasóleo', 'carrinho de compras'} def classify_by_vocab(text): text_lower = text.lower() br_count = sum(1 for word in br_vocab if word in text_lower) pt_count = sum(1 for word in pt_vocab if word in text_lower) if br_count > pt_count: return '巴西葡萄牙语' elif pt_count > br_count: return '欧洲葡萄牙语' else: return '无法确定' # 测试示例 test_br = 'Comprei um celular novo e fui de ônibus para o shopping.' test_pt = 'Comprei um telemóvel novo e fui de autocarro para o centro comercial.' print(classify_by_vocab(test_br)) # 输出:巴西葡萄牙语 print(classify_by_vocab(test_pt)) # 输出:欧洲葡萄牙语
二、机器学习分类方案
针对模糊文本,机器学习的效果远优于规则方法,以下是两种可行实现:
1. 预训练葡语BERT模型微调
利用Hugging Face的transformers库,基于葡语预训练BERT模型(neuralmind/bert-base-portuguese-cased)微调分类任务,适合高精度需求:
from transformers import BertTokenizer, BertForSequenceClassification from torch.utils.data import Dataset, DataLoader import torch # 加载预训练模型与分词器 tokenizer = BertTokenizer.from_pretrained('neuralmind/bert-base-portuguese-cased') model = BertForSequenceClassification.from_pretrained('neuralmind/bert-base-portuguese-cased', num_labels=2) # 自定义数据集类 class PortugueseTextDataset(Dataset): def __init__(self, texts, labels, tokenizer, max_len=128): self.texts = texts self.labels = labels self.tokenizer = tokenizer self.max_len = max_len def __len__(self): return len(self.texts) def __getitem__(self, idx): text = str(self.texts[idx]) label = self.labels[idx] encoding = self.tokenizer.encode_plus( text, add_special_tokens=True, max_length=self.max_len, return_token_type_ids=False, padding='max_length', truncation=True, return_attention_mask=True, return_tensors='pt', ) return { 'input_ids': encoding['input_ids'].flatten(), 'attention_mask': encoding['attention_mask'].flatten(), 'labels': torch.tensor(label, dtype=torch.long) } # 假设已准备好标注数据:train_texts(文本列表)、train_labels(0=巴西,1=欧洲) train_dataset = PortugueseTextDataset(train_texts, train_labels, tokenizer) train_loader = DataLoader(train_dataset, batch_size=8, shuffle=True) # 训练配置 optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5) device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') model.to(device) # 训练循环(简化版) model.train() for epoch in range(3): for batch in train_loader: input_ids = batch['input_ids'].to(device) attention_mask = batch['attention_mask'].to(device) labels = batch['labels'].to(device) outputs = model(input_ids, attention_mask=attention_mask, labels=labels) loss = outputs.loss loss.backward() optimizer.step() optimizer.zero_grad() # 预测函数 def predict_bert(text): encoding = tokenizer.encode_plus( text, add_special_tokens=True, max_length=128, return_token_type_ids=False, padding='max_length', truncation=True, return_attention_mask=True, return_tensors='pt', ) input_ids = encoding['input_ids'].to(device) attention_mask = encoding['attention_mask'].to(device) with torch.no_grad(): outputs = model(input_ids, attention_mask=attention_mask) predicted_class = torch.argmax(outputs.logits, dim=1).item() return '巴西葡萄牙语' if predicted_class == 0 else '欧洲葡萄牙语'
2. 传统机器学习管道(TF-IDF + SVM)
无GPU时可选用,基于TF-IDF提取特征结合SVM分类,兼顾效率与准确率:
from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.svm import SVC from sklearn.pipeline import Pipeline from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score # 假设已准备好标注数据:texts(文本列表)、labels(0=巴西,1=欧洲) X_train, X_test, y_train, y_test = train_test_split(texts, labels, test_size=0.2, random_state=42) # 构建分类管道 pipeline = Pipeline([ ('tfidf', TfidfVectorizer(ngram_range=(1,2), max_features=10000)), ('svm', SVC(kernel='linear')) ]) # 训练与评估 pipeline.fit(X_train, y_train) y_pred = pipeline.predict(X_test) print(f"分类准确率: {accuracy_score(y_test, y_pred):.2f}") # 单文本预测 def predict_tfidf(text): return '巴西葡萄牙语' if pipeline.predict([text])[0] == 0 else '欧洲葡萄牙语'
三、混合优化方案
先通过词汇规则过滤明确分类的文本,再用机器学习处理规则无法判断的模糊文本,平衡效率与准确率:
- 用规则方法标记所有文本,分离出“无法确定”的子集
- 用机器学习模型预测该子集的类别
内容的提问来源于stack exchange,提问作者Milosh Bankovic
相关产品推荐
相关产品推荐

