从Excel读取标签字符串赋值给OCI Data Labeling的labels参数时类型不匹配报错问题
从Excel读取标签字符串赋值给OCI Data Labeling的labels参数时类型不匹配报错问题
看起来你遇到的问题很典型——OCI Data Labeling的labels参数要求的是Label对象的列表,但你直接把Excel里读出来的字符串赋值给它了,自然会触发类型不匹配的错误。我来给你两种解决思路,看哪种更适合你的情况:
方案一:简化Excel的标签存储格式(推荐)
如果可以修改Excel文件,建议把单元格里的内容改成纯标签名称的逗号分隔形式,比如把原来的:
oci.data_labeling_service_dataplane.models.Label(label="scar"), oci.data_labeling_service_dataplane.models.Label(label="nevus"), oci.data_labeling_service_dataplane.models.Label(label="increased_cup_disc")
改成:
scar,nevus,increased_cup_disc
这样代码里处理起来就非常简单,只需要拆分字符串,然后逐个生成Label对象即可:
import oci import pandas as pd config = oci.config.from_file() data_labeling_service_dataplane_client = oci.data_labeling_service_dataplane.DataLabelingClient(config) excel_file_path = r'C:\SMALLFILEEXCEL.xlsx' df = pd.read_excel(excel_file_path, header=None) for index, row in df.iterrows(): if len(row) < 2: # 确保有record_id和labels两列 print(f"Skipping row due to insufficient data: {row}") continue record_id2 = row[0].strip() label_names = row[1].strip().split(',') # 拆分逗号分隔的标签名称 print(f"Using record_id: {record_id2}") try: # 生成Label对象列表 label_objects = [oci.data_labeling_service_dataplane.models.Label(label=name.strip()) for name in label_names] create_annotation_response = data_labeling_service_dataplane_client.create_annotation( create_annotation_details=oci.data_labeling_service_dataplane.models.CreateAnnotationDetails( record_id=record_id2, compartment_id="compartmentid", entities=[ oci.data_labeling_service_dataplane.models.GenericEntity( entity_type="GENERIC", labels=label_objects # 这里传入Label对象列表 ) ], freeform_tags={'example_key_2': 'example_value_2'}, defined_tags={'example_key_3': {'example_nested_key': 'example_nested_value'}} ), opc_retry_token="a12354123", opc_request_id="example_opc_request_id" ) print(create_annotation_response.data) except oci.exceptions.ServiceError as e: print(f"ServiceError: {e.message}") print(f"Request ID: {e.request_id}") print(f"Status Code: {e.status}") print(f"Error Code: {e.code}") print(f"Using record_id: {record_id2}")
方案二:直接解析Excel中的现有字符串(不修改Excel)
如果不想改动Excel,那需要解析单元格里的字符串,提取出每个标签的值,再生成Label对象。这里不建议用eval(因为如果Excel里的内容不可控,会有安全风险),而是用正则表达式来提取标签值:
import oci import pandas as pd import re config = oci.config.from_file() data_labeling_service_dataplane_client = oci.data_labeling_service_dataplane.DataLabelingClient(config) excel_file_path = r'C:\SMALLFILEEXCEL.xlsx' df = pd.read_excel(excel_file_path, header=None) # 正则表达式:匹配label="xxx"里的xxx内容 label_pattern = re.compile(r'label="([^"]+)"') for index, row in df.iterrows(): if len(row) < 2: print(f"Skipping row due to insufficient data: {row}") continue record_id2 = row[0].strip() label_str = row[1].strip() print(f"Using record_id: {record_id2}") try: # 提取所有标签名称 label_names = label_pattern.findall(label_str) # 生成Label对象列表 label_objects = [oci.data_labeling_service_dataplane.models.Label(label=name) for name in label_names] create_annotation_response = data_labeling_service_dataplane_client.create_annotation( create_annotation_details=oci.data_labeling_service_dataplane.models.CreateAnnotationDetails( record_id=record_id2, compartment_id="compartmentid", entities=[ oci.data_labeling_service_dataplane.models.GenericEntity( entity_type="GENERIC", labels=label_objects ) ], freeform_tags={'example_key_2': 'example_value_2'}, defined_tags={'example_key_3': {'example_nested_key': 'example_nested_value'}} ), opc_retry_token="a12354123", opc_request_id="example_opc_request_id" ) print(create_annotation_response.data) except oci.exceptions.ServiceError as e: print(f"ServiceError: {e.message}") print(f"Request ID: {e.request_id}") print(f"Status Code: {e.status}") print(f"Error Code: {e.code}") print(f"Using record_id: {record_id2}")
这两种方案都能解决你遇到的类型不匹配问题,其中方案一因为格式更简洁,后续维护也更方便,是优先推荐的。
备注:内容来源于stack exchange,提问作者Robert Giljohann
相关产品推荐
相关产品推荐

