如何从社交媒体文本中提取音乐艺术家姓名?NLP新手求助
Hey there! 作为NLP新手,从社交媒体这种格式随意的文本里精准提取音乐艺术家确实挺头疼的——毕竟这类内容里有全大写艺名、复合姓名,通用NER模型很容易漏判或误判。我给你几个亲测好用的方案,帮你搞定这个需求:
1. 结合规则匹配+spaCy自定义优化
spaCy默认的NER模型对音乐领域的实体覆盖有限,你可以给它加一层规则匹配,再结合音乐艺术家词典过滤结果,效果会好很多:
import spacy from spacy.matcher import Matcher # 加载spaCy基础模型 nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) # 定义匹配模式:匹配全大写的单词或多词艺名 # 单词艺名(比如CHANGE) pattern_single = [{"IS_UPPER": True, "IS_ALPHA": True, "IS_STOP": False}] # 多词艺名(比如TAYLOR SWIFT) pattern_multi = [{"IS_UPPER": True, "IS_ALPHA": True, "IS_STOP": False}, {"IS_SPACE": True}, {"IS_UPPER": True, "IS_ALPHA": True, "IS_STOP": False}] matcher.add("ARTIST_PATTERN", [pattern_single, pattern_multi]) # 处理示例文本 doc = nlp("Today bandcamp is waiving fees again! CHANGE, TAYLOR SWIFT and POP SMOKE will be using all funds collected through bandcamp to donate to Anti Repression Committee. No Justice No Peace.") matches = matcher(doc) # 提取匹配结果,再用艺术家词典过滤 artist_dict = {"CHANGE", "TAYLOR SWIFT", "POP SMOKE"} # 可以扩展成更全的词典 found_artists = [] for match_id, start, end in matches: span_text = doc[start:end].text if span_text in artist_dict: found_artists.append(span_text) print(found_artists) # 输出: ['CHANGE', 'TAYLOR SWIFT', 'POP SMOKE']
2. 用领域预训练的NER模型
通用NER模型对娱乐领域的实体识别精度不够,你可以试试用Hugging Face Transformers加载针对人物/娱乐领域微调的模型,再结合音乐艺术家词典做二次过滤:
from transformers import pipeline # 加载针对人物识别优化的预训练模型 ner_pipeline = pipeline("ner", model="dbmdz/bert-large-cased-finetuned-conll03-english", aggregation_strategy="simple") text = "Today bandcamp is waiving fees again! CHANGE, TAYLOR SWIFT and POP SMOKE will be using all funds collected through bandcamp to donate to Anti Repression Committee. No Justice No Peace." # 识别文本中的人物实体 ner_results = ner_pipeline(text) candidate_artists = [res["word"] for res in ner_results if res["entity_group"] == "PER"] # 用艺术家词典过滤,得到最终结果 artist_dict = {"CHANGE", "TAYLOR SWIFT", "POP SMOKE"} final_artists = [artist for artist in candidate_artists if artist in artist_dict] print(final_artists) # 输出: ['CHANGE', 'TAYLOR SWIFT', 'POP SMOKE']
这里的模型会把所有人物都识别出来,所以必须用词典过滤,确保只保留音乐艺术家。
3. 规则+词典的轻量方案(新手友好)
如果不想折腾复杂模型,这个方案最快见效:核心是用一个足够全的音乐艺术家词典,通过字符串匹配直接提取:
def extract_music_artists(input_text, artist_dictionary): # 统一转成大写,避免大小写问题 text_upper = input_text.upper() found_artists = [] # 优先匹配多词艺名(避免被拆成单词) for artist in sorted(artist_dictionary, key=lambda x: len(x.split()), reverse=True): artist_upper = artist.upper() if artist_upper in text_upper: found_artists.append(artist) # 替换已匹配内容,防止重复识别 text_upper = text_upper.replace(artist_upper, "") # 再匹配单词艺名 for artist in artist_dictionary: if len(artist.split()) == 1 and artist.upper() in text_upper.split(): found_artists.append(artist) # 去重后返回 return list(set(found_artists)) # 示例词典(可以扩展成包含上万条艺术家的列表) artist_dict = {"CHANGE", "TAYLOR SWIFT", "POP SMOKE"} sample_text = "Today bandcamp is waiving fees again! CHANGE, TAYLOR SWIFT and POP SMOKE will be using all funds collected through bandcamp to donate to Anti Repression Committee. No Justice No Peace." print(extract_music_artists(sample_text, artist_dict)) # 输出: ['CHANGE', 'TAYLOR SWIFT', 'POP SMOKE']
这个方案的关键是不断扩充你的艺术家词典——你可以从维基百科音乐分类、Spotify公开数据或者音乐平台API导出名单,慢慢完善。
内容的提问来源于stack exchange,提问作者masahiro okamura
相关产品推荐
相关产品推荐

