解决Python爬虫+Tesseract OCR写入CSV时ValueError: 值长度与索引长度不匹配问题
问题描述
我从CSV读取URL列表,爬取指定图片后用Tesseract OCR识别文本,尝试把URL、图片URL和OCR结果写入新CSV时,df1['image text'] = img_text_list触发报错:
ValueError: Length of values (20) does not match length of index (10)
url和img url列能成功插入,但img_text_list存的是完整OCR长文本,截断后长度匹配就能写入,但我需要保留完整文本,该怎么解决?
期望的CSV结构示例:
| url | img url | image text |
|---|---|---|
| abc.com/image | abc.com/image/2.jpg | How the Sangh Parivar Business Understands the fundamentals. Goyal Confederation of Indian Industry (CII) in a meeting Said that the Tata group new e- |
| abc.com/image | abc.com/image/3.jpg | GST implemented in a hurry, increasing Tariff barriers and now countries amid accusations of being hostileIndus If the business world confused about relationships- As a human always said You are welcome. But |
我使用的代码如下:
from PIL import Image import pytesseract as pt from pytesseract import image_to_string import io import requests import pandas as pd from bs4 import BeautifulSoup import pathlib img_text_list = [] urls_list = [] img_list = [] df1 = pd.DataFrame(columns=['url', 'img url', 'image text']) img_formats = [".jpg", ".jpeg", ".gif", ".png", ".webp"] df = pd.read_csv("urls.csv") urls = df["urls"].tolist() for y in urls: response = requests.get(y) soup = BeautifulSoup(response.text, 'html.parser') img_tags = soup.find_all('img', class_='pick') img_srcs = ["https://abc.in/" + img['src'].replace( '\\', '/') if img.has_attr('src') else '-' for img in img_tags] img_alts = [img['alt'] if img.has_attr('alt') else '-' for img in img_tags] for count, x in enumerate(img_srcs): img_list.append(x) urls_list.append(y) if x != '-': print(pathlib.Path(x).suffix) if pathlib.Path(x).suffix in img_formats: response = requests.get(x) print(response) img = Image.open(io.BytesIO(response.content)) text = pt.image_to_string(img, lang="hin") img_text_list.append(text) else: img_text_list.append("not supported") else: img_text_list.append("no src") df1['url'] = urls_list df1['img url'] = img_list df1['image text'] = img_text_list df1.to_csv('data.csv')
原因分析
这个报错的核心是三个列表的长度不一致:urls_list和img_list的长度对应DataFrame的行数(10行),但img_text_list的长度是20,导致无法一一对应赋值。
大概率是你在处理OCR文本时,不小心把每个换行的文本单独添加到了列表中(比如用split('\n')后循环append,或者用extend代替append),而不是把完整的OCR结果作为一个字符串添加进去。
解决方案
1. 确保列表长度严格一致
检查OCR处理逻辑,每个图片只对应img_text_list中的一个元素——把完整的OCR文本(包含所有换行)直接append到列表,不要分割后添加:
# 正确操作:直接添加完整的OCR结果字符串 text = pt.image_to_string(img, lang="hin") img_text_list.append(text) # 不要对text做split或extend操作
2. 保留换行写入CSV
Pandas在写入包含换行的字符串到CSV时,会自动用双引号包裹该单元格,确保换行被识别为单元格内的内容,而不是新的行。你只需要正常调用to_csv即可,建议加上index=False去掉默认的索引列:
df1.to_csv('data.csv', index=False)
3. 调试长度定位问题
如果还是不确定哪里出了问题,可以在赋值前添加调试代码,确认三个列表的长度是否一致:
print(f"urls_list长度: {len(urls_list)}") print(f"img_list长度: {len(img_list)}") print(f"img_text_list长度: {len(img_text_list)}")
如果长度不一致,就顺着代码排查哪个步骤多添加/少添加了元素。
修正后的完整代码
我给你的代码加了异常处理(避免请求失败导致流程中断)和一些细节优化,确保稳定性:
from PIL import Image import pytesseract as pt from pytesseract import image_to_string import io import requests import pandas as pd from bs4 import BeautifulSoup import pathlib img_text_list = [] urls_list = [] img_list = [] df1 = pd.DataFrame(columns=['url', 'img url', 'image text']) img_formats = [".jpg", ".jpeg", ".gif", ".png", ".webp"] df = pd.read_csv("urls.csv") urls = df["urls"].tolist() for y in urls: try: # 检查URL请求是否成功 response = requests.get(y) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') img_tags = soup.find_all('img', class_='pick') img_srcs = [ "https://abc.in/" + img['src'].replace('\\', '/') if img.has_attr('src') else '-' for img in img_tags ] for x in img_srcs: img_list.append(x) urls_list.append(y) if x != '-': # 转小写避免大小写格式问题(比如.JPG和.jpg) suffix = pathlib.Path(x).suffix.lower() if suffix in img_formats: try: img_response = requests.get(x) img_response.raise_for_status() img = Image.open(io.BytesIO(img_response.content)) text = pt.image_to_string(img, lang="hin") img_text_list.append(text) except Exception as e: print(f"处理图片{x}出错: {str(e)}") img_text_list.append("图片处理失败") else: img_text_list.append("不支持的图片格式") else: img_text_list.append("无图片src属性") except Exception as e: print(f"处理URL{y}出错: {str(e)}") # 如果URL请求失败,可根据需求添加对应处理逻辑 # 调试长度,确认一致 print(f"urls_list长度: {len(urls_list)}") print(f"img_list长度: {len(img_list)}") print(f"img_text_list长度: {len(img_text_list)}") df1['url'] = urls_list df1['img url'] = img_list df1['image text'] = img_text_list df1.to_csv('data.csv', index=False)
内容的提问来源于stack exchange,提问作者jeniffer

