You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从市场传单PDF提取产品信息至CSV?多款工具尝试未果

可行解决方案

针对你提取市场传单产品数据的需求,结合你之前遇到的OCR识别错误、顺序混乱、提取不全问题,提供以下3种可行方案:

方案1:开源OCR+布局分析(Tesseract + LayoutParser)

这个组合能精准识别页面布局,避免文本顺序混乱,同时提升OCR准确率:

  • 步骤1:获取目标页面的高清资源
    因为目标是网页版传单,用Selenium渲染页面并截取第10页的高清截图:
    from selenium import webdriver
    from selenium.webdriver.common.by import By
    import time
    
    driver = webdriver.Chrome()
    driver.get("目标网页URL")
    # 切换到第10页(根据网页翻页按钮定位,示例为点击文本为"10"的按钮)
    page_10_btn = driver.find_element(By.XPATH, "//button[text()='10']")
    page_10_btn.click()
    time.sleep(2)  # 等待页面加载完成
    driver.save_screenshot("page_10.png")
    driver.quit()
    
  • 步骤2:用LayoutParser识别产品区块
    LayoutParser能自动检测页面中的独立产品单元,避免跨产品的文本混杂:
    import layoutparser as lp
    import cv2
    
    image = cv2.imread("page_10.png")
    # 加载预训练的布局检测模型(适配传单类文档)
    model = lp.Detectron2LayoutModel(
        config_path="lp://PubLayNet/mask_rcnn_X_101_32x8d_FPN_3x/config",
        label_map={0: "Text", 1: "Title", 2: "List", 3: "Table", 4: "Figure"},
        extra_config=["MODEL.ROI_HEADS.SCORE_THRESH_TEST", 0.8]
    )
    layout = model.detect(image)
    # 过滤出产品相关的文本区块(可根据坐标和大小调整筛选条件)
    product_blocks = [block for block in layout if block.width > 100 and block.height > 50]
    
  • 步骤3:逐个区块识别文本并提取字段
    对每个产品区块单独用Tesseract识别,再拆分名称、原价、折扣价:
    import pytesseract
    
    pytesseract.pytesseract.tesseract_cmd = r'你的Tesseract安装路径'
    for block in product_blocks:
        x1, y1, x2, y2 = block.coordinates
        product_region = image[int(y1):int(y2), int(x1):int(x2)]
        text = pytesseract.image_to_string(product_region, lang='ita')  # 意大利语模型适配目标传单
        # 根据文本格式拆分字段(示例按换行和关键词分割)
        lines = [line.strip() for line in text.split('\n') if line.strip()]
        product_name = lines[0]
        original_price = [line for line in lines if '€' in line and 'sconto' not in line][0]
        discount_price = [line for line in lines if 'sconto' in line or '€' in line][-1]
        print(f"产品名称:{product_name},原价:{original_price},折扣价:{discount_price}")
    

方案2:商用高精度OCR API(Google Cloud Vision/AWS Textract)

商用API自带成熟的布局分析和实体识别能力,能大幅降低识别错误率:

  • 以Google Cloud Vision为例:
    1. 上传第10页的图片或PDF到Google Cloud Storage,调用DOCUMENT_TEXT_DETECTION接口,获取带坐标的文本结果。
    2. 根据文本的坐标位置分组,同一产品的名称、价格会处于相邻区域,通过y轴坐标判断归属关系。
    3. 从分组后的文本中提取对应字段,API会返回文本的置信度,可过滤低置信度结果。

方案3:直接抓取网页DOM数据(最优,避免OCR)

目标传单是网页形式,直接分析HTML结构抓取数据是最精准的方式:

  • 步骤1:用浏览器开发者工具查看第10页的产品DOM结构,找到产品容器的类名或标签(比如每个产品对应<div class="product-item">)。
  • 步骤2:用Python的BeautifulSoup抓取数据:
    import requests
    from bs4 import BeautifulSoup
    
    url = "目标网页URL"
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    # 定位第10页的产品列表(需根据实际DOM结构调整选择器)
    products = soup.select("div.page-10 .product-item")
    for product in products:
        product_name = product.select_one(".product-name").text.strip()
        original_price = product.select_one(".original-price").text.strip()
        discount_price = product.select_one(".discount-price").text.strip()
        print(f"产品名称:{product_name},原价:{original_price},折扣价:{discount_price}")
    
    注意:如果网页是动态加载的,需要用Selenium代替requests渲染页面后再解析。

内容的提问来源于stack exchange,提问作者user22028791

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 16:45:15