使用PYABSA批量情感分析自定义文件时遇RuntimeError求助
问题:自定义数据集批量情感推理触发RuntimeError
我按照指定教程完成情感推理,单句子列表的推理代码可正常运行,但修改批量推理代码使用自定义文件test.dat.apc(共2行,每行含2个句子)时,触发以下错误:
修改后的代码
# inference_sets = ABSADatasetList.Phone # original code inference_sets = 'test.dat.apc' # this is my own file that I want to infer sentiment for each sentence results = sent_classifier.batch_infer(target_file=inference_sets, print_result=True, save_result=True, ignore_error=False, )
报错信息
RuntimeError Traceback (most recent call last) Input In [56], in <cell line: 2>() 1 test = 'test.dat.apc' ----> 2 results = sent_classifier.batch_infer(target_file=test, 3 print_result=False, 4 save_result=True, 5 ignore_error=False, 6 ) File ~\Anaconda3\envs\spacy\lib\site-packages\pyabsa-1.16.15-py3.9.egg\pyabsa\core\apc\prediction\sentiment_classifier.py:197, in SentimentClassifier.batch_infer(self, target_file, print_result, save_result, ignore_error, clear_input_samples) 193 self.clear_input_samples() 195 save_path = os.path.join(os.getcwd(), 'apc_inference.result.json') --> 197 target_file = detect_infer_dataset(target_file, task='apc') 198 if not target_file: 199 raise FileNotFoundError('Can not find inference datasets!') File ~\Anaconda3\envs\spacy\lib\site-packages\pyabsa-1.16.15-py3.9.egg\pyabsa\functional\dataset\dataset_manager.py:302, in detect_infer_dataset(dataset_path, task) 300 if os.path.isdir(dataset_path.dataset_name): 301 print('No inference set found from: {}, unrecognized files: {}'.format(dataset_path, ', '.join(os.listdir(dataset_path.dataset_name)))) --> 302 raise RuntimeError( 303 'Fail to locate dataset: {}. If you are using your own dataset, you may need rename your dataset according to {}'.format( 304 dataset_path, 305 'https://github.com/yangheng95/ABSADatasets#important-rename-your-dataset-filename-before-use-it-in-pyabsa') 306 ) 307 if len(dataset_path) > 1: 308 print(colored('Please DO NOT mix datasets with different sentiment labels for training & inference !', 'yellow')) RuntimeError: Fail to locate dataset: ['test.dat.apc']. If you are using your own dataset, you may need rename your dataset according to https://github.com/yangheng95/ABSADatasets#important-rename-your-dataset-filename-before-use-it-in-pyabsa
问题原因与解决办法
1. 文件名不符合PyABSA识别规则
PyABSA的detect_infer_dataset函数会严格校验数据集文件名,test.dat.apc不符合它的命名规范,导致无法识别为有效的APC任务数据集。
- 解决:将文件名改为
test.apc,去掉多余的.dat后缀,让系统能识别这是测试集文件。
2. 文件内容格式不符合APC任务要求
APC(基于方面的情感分类)要求每行输入为「句子 + 目标方面词」的组合,用[SEP]分隔,示例格式:
The camera quality is excellent [SEP] camera quality This phone gets hot quickly [SEP] heat issue
你当前文件每行包含2个句子,完全不符合输入格式,即使文件名正确,后续也会报错。
- 解决:调整文件内容,每行对应一个待推理的「句子-方面」对,用
[SEP]分隔开。
3. 确保文件路径正确
如果文件不在当前工作目录,需要传入完整的绝对路径,比如r"C:\your_path\test.apc",避免系统找不到文件。
内容的提问来源于stack exchange,提问作者Nemo
相关产品推荐
相关产品推荐

