You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Python爬虫+Tesseract OCR写入CSV时ValueError: 值长度与索引长度不匹配问题

解决Pandas写入CSV时ValueError:列表长度不匹配问题(保留完整OCR文本)

问题描述

我从CSV读取URL列表,爬取指定图片后用Tesseract OCR识别文本,尝试把URL、图片URL和OCR结果写入新CSV时,df1['image text'] = img_text_list触发报错:

ValueError: Length of values (20) does not match length of index (10)

url和img url列能成功插入,但img_text_list存的是完整OCR长文本,截断后长度匹配就能写入,但我需要保留完整文本,该怎么解决?

期望的CSV结构示例:

urlimg urlimage text
abc.com/imageabc.com/image/2.jpgHow the Sangh Parivar Business
Understands the fundamentals.
Goyal Confederation of Indian Industry
(CII) in a meeting
Said that the Tata group new e-
abc.com/imageabc.com/image/3.jpgGST implemented in a hurry, increasing
Tariff barriers and now countries
amid accusations of being hostileIndus
If the business world
confused about relationships-
As a human always said
You are welcome. But

我使用的代码如下:

from PIL import Image
import pytesseract as pt
from pytesseract import image_to_string
import io
import requests
import pandas as pd
from bs4 import BeautifulSoup
import pathlib

img_text_list = []
urls_list = []
img_list = []
df1 = pd.DataFrame(columns=['url', 'img url', 'image text'])

img_formats = [".jpg", ".jpeg", ".gif", ".png", ".webp"]
df = pd.read_csv("urls.csv")
urls = df["urls"].tolist()

for y in urls:
    response = requests.get(y)
    soup = BeautifulSoup(response.text, 'html.parser')
    img_tags = soup.find_all('img', class_='pick')
    img_srcs = ["https://abc.in/" + img['src'].replace( '\\', '/') if img.has_attr('src') else '-' for img in img_tags]
    img_alts = [img['alt'] if img.has_attr('alt') else '-' for img in img_tags]
    
    for count, x in enumerate(img_srcs):
        img_list.append(x)
        urls_list.append(y)
        if x != '-':
            print(pathlib.Path(x).suffix)
            if pathlib.Path(x).suffix in img_formats:
                response = requests.get(x)
                print(response)
                img = Image.open(io.BytesIO(response.content))
                text = pt.image_to_string(img, lang="hin")
                img_text_list.append(text)
            else:
                img_text_list.append("not supported")
        else:
            img_text_list.append("no src")

df1['url'] = urls_list
df1['img url'] = img_list
df1['image text'] = img_text_list
df1.to_csv('data.csv')

原因分析

这个报错的核心是三个列表的长度不一致:urls_list和img_list的长度对应DataFrame的行数(10行),但img_text_list的长度是20,导致无法一一对应赋值。

大概率是你在处理OCR文本时,不小心把每个换行的文本单独添加到了列表中(比如用split('\n')后循环append,或者用extend代替append),而不是把完整的OCR结果作为一个字符串添加进去。

解决方案

1. 确保列表长度严格一致

检查OCR处理逻辑,每个图片只对应img_text_list中的一个元素——把完整的OCR文本(包含所有换行)直接append到列表,不要分割后添加:

# 正确操作:直接添加完整的OCR结果字符串
text = pt.image_to_string(img, lang="hin")
img_text_list.append(text)  # 不要对text做split或extend操作

2. 保留换行写入CSV

Pandas在写入包含换行的字符串到CSV时,会自动用双引号包裹该单元格,确保换行被识别为单元格内的内容,而不是新的行。你只需要正常调用to_csv即可,建议加上index=False去掉默认的索引列:

df1.to_csv('data.csv', index=False)

3. 调试长度定位问题

如果还是不确定哪里出了问题,可以在赋值前添加调试代码,确认三个列表的长度是否一致:

print(f"urls_list长度: {len(urls_list)}")
print(f"img_list长度: {len(img_list)}")
print(f"img_text_list长度: {len(img_text_list)}")

如果长度不一致,就顺着代码排查哪个步骤多添加/少添加了元素。

修正后的完整代码

我给你的代码加了异常处理(避免请求失败导致流程中断)和一些细节优化,确保稳定性:

from PIL import Image
import pytesseract as pt
from pytesseract import image_to_string
import io
import requests
import pandas as pd
from bs4 import BeautifulSoup
import pathlib

img_text_list = []
urls_list = []
img_list = []
df1 = pd.DataFrame(columns=['url', 'img url', 'image text'])

img_formats = [".jpg", ".jpeg", ".gif", ".png", ".webp"]
df = pd.read_csv("urls.csv")
urls = df["urls"].tolist()

for y in urls:
    try:
        # 检查URL请求是否成功
        response = requests.get(y)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        img_tags = soup.find_all('img', class_='pick')
        img_srcs = [
            "https://abc.in/" + img['src'].replace('\\', '/') 
            if img.has_attr('src') else '-' 
            for img in img_tags
        ]
        
        for x in img_srcs:
            img_list.append(x)
            urls_list.append(y)
            if x != '-':
                # 转小写避免大小写格式问题(比如.JPG和.jpg)
                suffix = pathlib.Path(x).suffix.lower()
                if suffix in img_formats:
                    try:
                        img_response = requests.get(x)
                        img_response.raise_for_status()
                        img = Image.open(io.BytesIO(img_response.content))
                        text = pt.image_to_string(img, lang="hin")
                        img_text_list.append(text)
                    except Exception as e:
                        print(f"处理图片{x}出错: {str(e)}")
                        img_text_list.append("图片处理失败")
                else:
                    img_text_list.append("不支持的图片格式")
            else:
                img_text_list.append("无图片src属性")
    except Exception as e:
        print(f"处理URL{y}出错: {str(e)}")
        # 如果URL请求失败,可根据需求添加对应处理逻辑

# 调试长度,确认一致
print(f"urls_list长度: {len(urls_list)}")
print(f"img_list长度: {len(img_list)}")
print(f"img_text_list长度: {len(img_text_list)}")

df1['url'] = urls_list
df1['img url'] = img_list
df1['image text'] = img_text_list
df1.to_csv('data.csv', index=False)

内容的提问来源于stack exchange,提问作者jeniffer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:38:11