Python2的file()与Python3的open()对象不匹配,求对应替代方法
file() in Python3 (And How to Fix Your PDFminer Issue) First off, the straight answer: there’s no built-in object in Python3 that’s identical to Python2’s file() type. Python3 completely overhauled its IO system, so open() returns modern classes like _io.BufferedReader (for binary mode) instead of the old file object. The good news is that this doesn’t have to stop you from processing PDFs with pdfminer—here’s how to fix this:
1. Use the Python3-Compatible pdf2txt.py from pdfminer.six
Since you already installed pdfminer.six, you don’t need to mess with the old Python2 code from the repository. The package comes with a fully updated pdf2txt.py that works natively in Python3. Just run this command in your terminal to generate XML output directly:
pdf2txt.py -o your_output.xml -t xml path/to/your/file.pdf
This will handle the IO correctly out of the box, no code changes required.
2. Modify the Old Code to Work with Python3 IO Objects
If you really need to use the original repository code, you’ll need to adjust parts that rely on the file object’s specific behavior:
- Replace any type checks like
isinstance(fp, file)with checks against Python3’s IO base class:from io import IOBase if isinstance(fp, IOBase): # Your existing processing logic here - Swap out Python2-specific methods like
xreadlines()with Python3-compatible alternatives—for example, just iterate over the file object directly (for line in fp:) instead of callingfp.xreadlines(). - Double-check that all file openings use
open()with the correct mode (alwaysrbfor PDF files, since they’re binary).
The key point here is that Python3’s IO classes implement the same core methods (read(), readline(), close()) as Python2’s file—the issue is usually hardcoded type checks or deprecated methods in the old code, not a lack of functionality.
内容的提问来源于stack exchange,提问作者Bengi Koseoglu

