使用Apache Tika 1.15提取PDF内容:能否获取<h1>、<strong>等HTML标签?
Great question! Let's tackle this from two angles: getting Apache Tika 1.15 to preserve richer HTML structure like <h1> or <strong>, and recommending proven alternative tools if that doesn't work out.
First off, Tika 1.15 is quite old (released in 2017), so its PDF parsing capabilities are limited compared to modern versions. But there are still tweaks you can try to pull out more structural tags:
Adjust parsing parameters via custom headers: When calling Tika from Python, you can pass headers to enable bookmark parsing (which often maps to headings) and extract styled marked content. Using the
tikaPython library, here's how:from tika import parser headers = { "X-Tika-PDFenableAutoBookmarkParsing": "true", "X-Tika-PDFextractMarkedContent": "true" } parsed_result = parser.from_file("your_document.pdf", headers=headers)These flags tell Tika to prioritize structural elements and styled text over plain paragraphs.
Request HTML output directly: Instead of raw XML, ask Tika to return HTML content, which is more likely to include semantic tags like
<h1>or<strong>:headers = {"Content-Type": "text/html"} parsed_html = parser.from_file("your_document.pdf", headers=headers) print(parsed_html["content"])That said, if the PDF was generated with poor structural metadata, even these tweaks might only yield
<p>tags. If possible, upgrading Tika to the latest 2.x version is highly recommended — its rewritten PDF parser does a far better job of detecting and preserving document structure.
If Tika 1.15 can't meet your needs, here are three battle-tested tools that consistently deliver better structured output, with minimal setup:
1. PyMuPDF (fitz)
My go-to for most PDF extraction tasks — it's fast, lightweight, and excels at detecting text styling (bold, italic) and structural hierarchy. You can map font properties to HTML tags easily:
import fitz doc = fitz.open("your_document.pdf") html_output = "<html><body>" for page in doc: blocks = page.get_text("dict")["blocks"] for block in blocks: if block["type"] == 0: # Target text blocks only for line in block["lines"]: for span in line["spans"]: text = span["text"] style_tags = [] # Detect bold/italic from font flags if span["flags"] & 1: # Bold flag style_tags.append("strong") if span["flags"] & 2: # Italic flag style_tags.append("em") # Map font size to heading tags (adjust thresholds for your PDFs) if span["size"] > 16: html_output += f"<h1>{text}</h1>" elif span["size"] > 14: html_output += f"<h2>{text}</h2>" else: # Wrap text in style tags if needed if style_tags: open_tags = "".join([f"<{tag}>" for tag in style_tags]) close_tags = "".join([f"</{tag}>" for tag in reversed(style_tags)]) html_output += f"<p>{open_tags}{text}{close_tags}</p>" else: html_output += f"<p>{text}</p>" html_output += "</body></html>" print(html_output)
2. pdfplumber
Perfect for PDFs with consistent layouts — it retains detailed text attributes like font size, weight, and position, making it easy to identify headings and styled text:
import pdfplumber with pdfplumber.open("your_document.pdf") as pdf: html_output = "<html><body>" for page in pdf.pages: text_objects = page.extract_words(extra_attrs=["fontname", "size"]) current_paragraph = [] for obj in text_objects: text = obj["text"] is_bold = "Bold" in obj["fontname"] # Check for headings first if obj["size"] > 16: if current_paragraph: html_output += f"<p>{' '.join(current_paragraph)}</p>" current_paragraph = [] html_output += f"<h1>{text}</h1>" elif obj["size"] > 14: if current_paragraph: html_output += f"<p>{' '.join(current_paragraph)}</p>" current_paragraph = [] html_output += f"<h2>{text}</h2>" else: # Add bold styling if detected current_paragraph.append(f"<strong>{text}</strong>" if is_bold else text) # Add remaining paragraph content if current_paragraph: html_output += f"<p>{' '.join(current_paragraph)}</p>" html_output += "</body></html>" print(html_output)
3. pdfminer.six
A low-level tool that gives you full control over parsing — ideal for complex PDFs, though it has a steeper learning curve:
from pdfminer.high_level import extract_pages from pdfminer.layout import LTTextContainer, LTChar html_output = "<html><body>" for page_layout in extract_pages("your_document.pdf"): for element in page_layout: if isinstance(element, LTTextContainer): for text_line in element: line_content = "" active_styles = [] for char in text_line: if isinstance(char, LTChar): # Handle bold styling if "Bold" in char.fontname and "strong" not in active_styles: line_content += "<strong>" active_styles.append("strong") elif "strong" in active_styles and "Bold" not in char.fontname: line_content += "</strong>" active_styles.remove("strong") # Handle italic styling if "Italic" in char.fontname and "em" not in active_styles: line_content += "<em>" active_styles.append("em") elif "em" in active_styles and "Italic" not in char.fontname: line_content += "</em>" active_styles.remove("em") line_content += char.get_text() # Close any remaining style tags for tag in reversed(active_styles): line_content += f"</{tag}>" # Map font size to headings if char.size > 16: html_output += f"<h1>{line_content}</h1>" elif char.size > 14: html_output += f"<h2>{line_content}</h2>" else: html_output += f"<p>{line_content}</p>" html_output += "</body></html>" print(html_output)
- If you can stick with Tika, upgrading to the latest 2.x version will drastically improve structural extraction.
- PyMuPDF is the best balance of ease-of-use and performance for most use cases.
- Always test with your specific PDFs — structure quality varies wildly depending on how the PDF was generated (e.g., native vs. scanned, different authoring tools).
内容的提问来源于stack exchange,提问作者Filip Górczyński

