如何从市场传单PDF提取产品信息至CSV?多款工具尝试未果
可行解决方案
针对你提取市场传单产品数据的需求,结合你之前遇到的OCR识别错误、顺序混乱、提取不全问题,提供以下3种可行方案:
方案1:开源OCR+布局分析(Tesseract + LayoutParser)
这个组合能精准识别页面布局,避免文本顺序混乱,同时提升OCR准确率:
- 步骤1:获取目标页面的高清资源
因为目标是网页版传单,用Selenium渲染页面并截取第10页的高清截图:from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() driver.get("目标网页URL") # 切换到第10页(根据网页翻页按钮定位,示例为点击文本为"10"的按钮) page_10_btn = driver.find_element(By.XPATH, "//button[text()='10']") page_10_btn.click() time.sleep(2) # 等待页面加载完成 driver.save_screenshot("page_10.png") driver.quit() - 步骤2:用LayoutParser识别产品区块
LayoutParser能自动检测页面中的独立产品单元,避免跨产品的文本混杂:import layoutparser as lp import cv2 image = cv2.imread("page_10.png") # 加载预训练的布局检测模型(适配传单类文档) model = lp.Detectron2LayoutModel( config_path="lp://PubLayNet/mask_rcnn_X_101_32x8d_FPN_3x/config", label_map={0: "Text", 1: "Title", 2: "List", 3: "Table", 4: "Figure"}, extra_config=["MODEL.ROI_HEADS.SCORE_THRESH_TEST", 0.8] ) layout = model.detect(image) # 过滤出产品相关的文本区块(可根据坐标和大小调整筛选条件) product_blocks = [block for block in layout if block.width > 100 and block.height > 50] - 步骤3:逐个区块识别文本并提取字段
对每个产品区块单独用Tesseract识别,再拆分名称、原价、折扣价:import pytesseract pytesseract.pytesseract.tesseract_cmd = r'你的Tesseract安装路径' for block in product_blocks: x1, y1, x2, y2 = block.coordinates product_region = image[int(y1):int(y2), int(x1):int(x2)] text = pytesseract.image_to_string(product_region, lang='ita') # 意大利语模型适配目标传单 # 根据文本格式拆分字段(示例按换行和关键词分割) lines = [line.strip() for line in text.split('\n') if line.strip()] product_name = lines[0] original_price = [line for line in lines if '€' in line and 'sconto' not in line][0] discount_price = [line for line in lines if 'sconto' in line or '€' in line][-1] print(f"产品名称:{product_name},原价:{original_price},折扣价:{discount_price}")
方案2:商用高精度OCR API(Google Cloud Vision/AWS Textract)
商用API自带成熟的布局分析和实体识别能力,能大幅降低识别错误率:
- 以Google Cloud Vision为例:
- 上传第10页的图片或PDF到Google Cloud Storage,调用
DOCUMENT_TEXT_DETECTION接口,获取带坐标的文本结果。 - 根据文本的坐标位置分组,同一产品的名称、价格会处于相邻区域,通过y轴坐标判断归属关系。
- 从分组后的文本中提取对应字段,API会返回文本的置信度,可过滤低置信度结果。
- 上传第10页的图片或PDF到Google Cloud Storage,调用
方案3:直接抓取网页DOM数据(最优,避免OCR)
目标传单是网页形式,直接分析HTML结构抓取数据是最精准的方式:
- 步骤1:用浏览器开发者工具查看第10页的产品DOM结构,找到产品容器的类名或标签(比如每个产品对应
<div class="product-item">)。 - 步骤2:用Python的BeautifulSoup抓取数据:
注意:如果网页是动态加载的,需要用Selenium代替requests渲染页面后再解析。import requests from bs4 import BeautifulSoup url = "目标网页URL" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 定位第10页的产品列表(需根据实际DOM结构调整选择器) products = soup.select("div.page-10 .product-item") for product in products: product_name = product.select_one(".product-name").text.strip() original_price = product.select_one(".original-price").text.strip() discount_price = product.select_one(".discount-price").text.strip() print(f"产品名称:{product_name},原价:{original_price},折扣价:{discount_price}")
内容的提问来源于stack exchange,提问作者user22028791
相关产品推荐
相关产品推荐

