使用Python的pdfx模块读取PDF正常,命令行转JSON报错求助
Hey there, let's figure out why you're hitting that SyntaxError with your pdfx command and get it sorted out quickly.
The Root Cause
Looking at your command:
pdfx -d D:/Output/ -j -o output.json pdf
The issue is right at the end: you're using pdf as a placeholder, but pdfx expects an actual PDF file path/name here (like the 1951.pdf you used in your Python code). That's what's triggering the syntax error—pdfx can't parse "pdf" as a valid input file.
Fixes to Try
Here's how to adjust your command to work correctly:
Replace the placeholder with your actual PDF file
Swap outpdffor the name or full path of your target PDF. For example, if your file is in the current working directory:pdfx -d D:/Output/ -j -o output.json 1951.pdfIf it's in another directory, use the full path:
pdfx -d D:/Output/ -j -o output.json "C:/Documents/1951.pdf"Double-check your output directory
Make sure the directory you specified with-d D:/Output/already exists. pdfx might fail if it can't find the directory to download referenced PDFs into—create it manually if needed before running the command.Handle spaces in file paths (if applicable)
If your PDF file or output path has spaces, wrap the path in quotes to avoid parsing issues:pdfx -d "D:/My Output Folder/" -j -o output.json "My Research Paper.pdf"Update pdfx (optional but recommended)
Outdated versions might have parameter bugs. Check your version with:pdfx --versionIf it's not the latest, upgrade using pip:
pip install --upgrade pdfx
Once you make these adjustments, the command should run without the SyntaxError and generate your JSON output as expected.
内容的提问来源于stack exchange,提问作者shikha singh

