You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过PDF Miner提取并规整PDF中含contrat n°的指定格式文本?

How to Reformat Your Extracted PDF Text to the Desired Structure

Absolutely, this is totally doable! The core idea is to parse the messy extracted text, pull out the key pieces (Client number, Contrat number, Produit value), then rearrange them into the clean format you want. Let's walk through how to implement this, using Python since you're already using PDF Miner.

First: Understand the Pattern

Your extracted text looks like this:
Client Contrat Produit n°XXXXXX n°XXXXX XXXXX
And you want to turn it into:
Client n°XXXX Contrat n°XXXX Produit XXXXXXX

From what I can see, the extraction has dumped all the labels first, then the corresponding values. So we just need to map each label to its value and reorder.

Option 1: Simple String Splitting (For Consistent Structure)

If every line you extract follows exactly the same order (labels first, then Client n°, Contrat n°, Produit value), this straightforward method works:

extracted_line = "Client Contrat Produit n°XXXXXX n°XXXXX XXXXX"

# Split the line into individual words/tokens
parts = extracted_line.split()

# Grab the labels and values separately
labels = parts[:3]  # ["Client", "Contrat", "Produit"]
values = parts[3:]  # ["n°XXXXXX", "n°XXXXX", "XXXXX"]

# Put it all together in your desired format
formatted_line = f"{labels[0]} {values[0]} {labels[1]} {values[1]} {labels[2]} {values[2]}"
print(formatted_line)  # Output: Client n°XXXXXX Contrat n°XXXXX Produit XXXXX

Option 2: Regular Expressions (More Flexible)

If the extraction order is sometimes inconsistent (like values showing up in different places), regex is more robust. It lets you capture all the n° numbers and the Produit value regardless of their position:

import re

extracted_line = "Client Contrat Produit n°XXXXXX n°XXXXX XXXXX"

# Capture all n° followed by digits
contract_numbers = re.findall(r'n°\d+', extracted_line)
# Extract the Produit value by removing labels and n° entries
produit_value = re.sub(r'(Client|Contrat|Produit|n°\d+)', '', extracted_line).strip()

# Build the formatted line (adjust indices if your n° order is swapped)
formatted_line = f"Client {contract_numbers[0]} Contrat {contract_numbers[1]} Produit {produit_value}"
print(formatted_line)

Integrate This With Your PDF Miner Workflow

You can slot this logic right into your existing code that extracts text from the PDF. Here's a full example:

from pdfminer.high_level import extract_text

def format_pdf_line(line):
    # Skip lines that don't have the target pattern
    if "contrat n°" not in line.lower():
        return None
    
    # Use regex to extract components
    import re
    contract_numbers = re.findall(r'n°\d+', line)
    produit_value = re.sub(r'(Client|Contrat|Produit|n°\d+)', '', line).strip()
    
    # Make sure we have all the parts we need
    if len(contract_numbers) >= 2:
        return f"Client {contract_numbers[0]} Contrat {contract_numbers[1]} Produit {produit_value}"
    return line  # Fallback to original if something's missing

# Extract text from your PDF
pdf_content = extract_text("your_pdf_file.pdf")

# Process each line
for line in pdf_content.split('\n'):
    result = format_pdf_line(line)
    if result:
        print(result)

Quick Tips:

  • If the n° values are swapped (e.g., first n° is Contrat instead of Client), just swap the indices in contract_numbers (use contract_numbers[1] for Client and contract_numbers[0] for Contrat).
  • If the Produit value has special characters, tweak the regex to include them (e.g., re.sub(r'(Client|Contrat|Produit|n°\d+)', '', line).strip() should handle most cases).

This should give you the clean, structured lines you're looking for!


内容的提问来源于stack exchange,提问作者Nathan Cheval

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:16:26