You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

面向Python的巴西葡萄牙语与欧洲葡萄牙语文本分类方法咨询

葡萄牙语文本(巴西/欧洲)分类的Python实现方案

一、基于专属词汇的规则分类法

两类葡语的词汇差异比语法差异更易识别,适合快速过滤明确文本:

  • 整理专属词汇集:收集仅在巴西葡语(BR)或欧洲葡语(PT)中高频使用的词汇,例如:
    • BR专属:celular, ônibus, sanduíche, pão de queijo, xícara, óleo diesel
    • PT专属:telemóvel, autocarro, bocadillo, chávena, gasóleo, carrinho de compras
  • 实现逻辑:统计文本中两类词汇的出现频次,高频者对应类别
  • 代码示例:
# 定义专属词汇集合
br_vocab = {'celular', 'ônibus', 'sanduíche', 'pão de queijo', 'xícara', 'óleo diesel'}
pt_vocab = {'telemóvel', 'autocarro', 'bocadillo', 'chávena', 'gasóleo', 'carrinho de compras'}

def classify_by_vocab(text):
    text_lower = text.lower()
    br_count = sum(1 for word in br_vocab if word in text_lower)
    pt_count = sum(1 for word in pt_vocab if word in text_lower)
    
    if br_count > pt_count:
        return '巴西葡萄牙语'
    elif pt_count > br_count:
        return '欧洲葡萄牙语'
    else:
        return '无法确定'

# 测试示例
test_br = 'Comprei um celular novo e fui de ônibus para o shopping.'
test_pt = 'Comprei um telemóvel novo e fui de autocarro para o centro comercial.'
print(classify_by_vocab(test_br))  # 输出:巴西葡萄牙语
print(classify_by_vocab(test_pt))  # 输出:欧洲葡萄牙语

二、机器学习分类方案

针对模糊文本,机器学习的效果远优于规则方法,以下是两种可行实现:

1. 预训练葡语BERT模型微调

利用Hugging Face的transformers库,基于葡语预训练BERT模型(neuralmind/bert-base-portuguese-cased)微调分类任务,适合高精度需求:

from transformers import BertTokenizer, BertForSequenceClassification
from torch.utils.data import Dataset, DataLoader
import torch

# 加载预训练模型与分词器
tokenizer = BertTokenizer.from_pretrained('neuralmind/bert-base-portuguese-cased')
model = BertForSequenceClassification.from_pretrained('neuralmind/bert-base-portuguese-cased', num_labels=2)

# 自定义数据集类
class PortugueseTextDataset(Dataset):
    def __init__(self, texts, labels, tokenizer, max_len=128):
        self.texts = texts
        self.labels = labels
        self.tokenizer = tokenizer
        self.max_len = max_len

    def __len__(self):
        return len(self.texts)

    def __getitem__(self, idx):
        text = str(self.texts[idx])
        label = self.labels[idx]
        
        encoding = self.tokenizer.encode_plus(
            text,
            add_special_tokens=True,
            max_length=self.max_len,
            return_token_type_ids=False,
            padding='max_length',
            truncation=True,
            return_attention_mask=True,
            return_tensors='pt',
        )
        
        return {
            'input_ids': encoding['input_ids'].flatten(),
            'attention_mask': encoding['attention_mask'].flatten(),
            'labels': torch.tensor(label, dtype=torch.long)
        }

# 假设已准备好标注数据:train_texts(文本列表)、train_labels(0=巴西,1=欧洲)
train_dataset = PortugueseTextDataset(train_texts, train_labels, tokenizer)
train_loader = DataLoader(train_dataset, batch_size=8, shuffle=True)

# 训练配置
optimizer = torch.optim.AdamW(model.parameters(), lr=2e-5)
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model.to(device)

# 训练循环(简化版)
model.train()
for epoch in range(3):
    for batch in train_loader:
        input_ids = batch['input_ids'].to(device)
        attention_mask = batch['attention_mask'].to(device)
        labels = batch['labels'].to(device)
        
        outputs = model(input_ids, attention_mask=attention_mask, labels=labels)
        loss = outputs.loss
        
        loss.backward()
        optimizer.step()
        optimizer.zero_grad()

# 预测函数
def predict_bert(text):
    encoding = tokenizer.encode_plus(
        text,
        add_special_tokens=True,
        max_length=128,
        return_token_type_ids=False,
        padding='max_length',
        truncation=True,
        return_attention_mask=True,
        return_tensors='pt',
    )
    
    input_ids = encoding['input_ids'].to(device)
    attention_mask = encoding['attention_mask'].to(device)
    
    with torch.no_grad():
        outputs = model(input_ids, attention_mask=attention_mask)
        predicted_class = torch.argmax(outputs.logits, dim=1).item()
    
    return '巴西葡萄牙语' if predicted_class == 0 else '欧洲葡萄牙语'

2. 传统机器学习管道(TF-IDF + SVM)

无GPU时可选用,基于TF-IDF提取特征结合SVM分类,兼顾效率与准确率:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import SVC
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

# 假设已准备好标注数据:texts(文本列表)、labels(0=巴西,1=欧洲)
X_train, X_test, y_train, y_test = train_test_split(texts, labels, test_size=0.2, random_state=42)

# 构建分类管道
pipeline = Pipeline([
    ('tfidf', TfidfVectorizer(ngram_range=(1,2), max_features=10000)),
    ('svm', SVC(kernel='linear'))
])

# 训练与评估
pipeline.fit(X_train, y_train)
y_pred = pipeline.predict(X_test)
print(f"分类准确率: {accuracy_score(y_test, y_pred):.2f}")

# 单文本预测
def predict_tfidf(text):
    return '巴西葡萄牙语' if pipeline.predict([text])[0] == 0 else '欧洲葡萄牙语'

三、混合优化方案

先通过词汇规则过滤明确分类的文本,再用机器学习处理规则无法判断的模糊文本,平衡效率与准确率:

  1. 用规则方法标记所有文本,分离出“无法确定”的子集
  2. 用机器学习模型预测该子集的类别

内容的提问来源于stack exchange,提问作者Milosh Bankovic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 12:30:55