使用TAPAS流水线进行表格问答时遇TypeError报错求助
问题原因
报错的核心原因是表格中的数值为整数类型,TAPAS流水线在处理表格内容时会执行正则替换操作,而正则函数只接受字符串/字节类型的输入,整数类型无法被处理,因此抛出TypeError。
另外代码里还有一个小问题:TAPAS流水线返回的是单个字典对象,不是列表,result[0]['query']的写法会引发索引错误,需要调整。
解决方法
- 将表格中所有数值类型的元素转换为字符串
- 修正结果提取的方式
修改后的完整代码:
from transformers import pipeline # Load the TAPAS pipeline nlp = pipeline(task="table-question-answering", model="google/tapas-large-finetuned-wtq") # Sample tabular data - 所有数值转成字符串 table = [ ["Name", "Department", "Salary"], ["Alice", "HR", "50000"], ["Bob", "Finance", "60000"], ["Charlie", "IT", "75000"], ] # Sample natural language query query = "What is the average salary in the IT department?" # Perform table-based question answering result = nlp(table=table, query=query) # Extract the SQL query from the result - 直接取字典的query键 sql_query = result['query'] # Print the generated SQL query print("Generated SQL query:", sql_query)
验证说明
运行修改后的代码,会输出类似如下的SQL查询:
SELECT AVG(Salary) FROM table WHERE Department = 'IT'
这说明流水线已经正常工作,能够根据自然语言查询生成对应的SQL语句。
内容的提问来源于stack exchange,提问作者Milan Manoj
相关产品推荐
相关产品推荐

