You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ValueError: 传入值shape为(20,1)与indices隐含(20,6)不匹配问题排查

问题排查
  • 核心报错原因:循环内对data执行了两次追加操作,导致列表内元素长度不统一
    第一次data.append(text)追加的是单个字符串(长度1),第二次data.append([text, dru, asker, datr, thema, frage])追加的是长度为6的列表。最终构造DataFrame时指定了6列,但传入的data中部分元素只有1个值,形状不匹配直接触发报错。删除第一次data.append(text)即可解决该问题,因为第二次追加的列表中已经包含text作为raw_text字段的取值。
  • 隐藏逻辑错误:正则变量名不匹配
    代码中定义主题匹配正则的变量名为th1,但后续匹配时调用的是thema1.search(text),变量名不一致会导致匹配逻辑失效,虽然被异常捕获后thema会默认赋值为None,但无法拿到正确的匹配结果,需要将变量名统一。
  • 可优化点:正则编译操作可以移到循环外部
    目前所有re.compile逻辑都写在for循环内部,每处理一个tif文件都会重复执行编译,没有必要,移到循环外可提升运行效率。
修正后参考代码
import glob
import re
import pandas as pd
import pytesseract
from PIL import Image

# 正则编译移到循环外
duck1 = re.compile(r'(CHE)(.*)\n', flags = re.DOTALL | re.MULTILINE)
asker1 = re.compile(r'(my|the)\s+zzzz(.*)office', flags = re.DOTALL | re.MULTILINE)
date1 = re.compile(r'\s+dfg(\d{2}\.\d{2}\.\d{4})', flags = re.DOTALL | re.MULTILINE)
# 变量名统一为thema1
thema1 = re.compile(r'(gh|gh)\s+fg\s+sdf(.*)(rrr,\s+rtr)', flags = re.DOTALL | re.MULTILINE)
frage1 = re.compile(r'(\neee)(.*)(we|wzz)\s+drte\s+Srrr:', flags = re.DOTALL | re.MULTILINE)

data = []
listOfPages = glob.glob(r"C:/Users/name/*.tif")
for entry in listOfPages:
    text = pytesseract.image_to_string(
            Image.open(entry), lang="en"
        )
    try:
        d2 = duck1.search(text)
        dru = d2.group(1) if d2 else None
    except:
        dru = None
    try:
        asker2 = asker1.search(text)
        asker = asker2.group(1) if asker2 else None
    except:
        asker = None
    try:
        date2 = date1.search(text)
        datr = date2.group(0) if date2 else None
    except:
        datr = None
    try:
        thema2 = thema1.search(text)
        thema = thema2.group(1) if thema2 else None
    except:
        thema = None
    try:
        frage2 = frage1.search(text)
        frage = frage2.group(1) if frage2 else None
    except:
        frage = None
    data.append([text, dru, asker, datr, thema, frage])
    
df0 = pd.DataFrame(data, columns =['raw_text', 'wer', 'asker', 'date', 'area', 'que_text'])
print(df0)

内容的提问来源于stack exchange,提问作者id345678

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 19:54:07