求助:用Python3+tex2py提取arXiv天体物理TeX文件纯净正文
I feel your pain—arXiv TeX files are notoriously inconsistent, with authors using custom macros, varying section tags, and text split across commands that break naive parsers like tex2py. Here's a robust Python-based approach tailored to astrophysics papers that will get you the pure body text you need:
Step 1: Install Required Packages
First, grab the tools we'll use for parsing and pattern matching:
pip install pylatexenc regex
Step 2: Implement the Cleanup Pipeline
This function combines regex-based removal of non-body sections with a robust TeX-to-text parser to handle messy, split text. It targets the common non-body elements you want to exclude:
import regex from pylatexenc.latex2text import LatexNodes2Text def extract_clean_body(tex_file_path): # Read the raw TeX content with open(tex_file_path, 'r', encoding='utf-8', errors='ignore') as f: tex_content = f.read() # -------------------------- # Phase 1: Remove non-body sections # -------------------------- # Remove title, author blocks, and maketitle command tex_content = regex.sub(r'\\title\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\author\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\maketitle', '', tex_content) # Remove abstract (handles both environment and command forms) tex_content = regex.sub(r'\\begin\{abstract\}.*?\\end\{abstract\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\abstract\{.*?\}', '', tex_content, flags=regex.DOTALL) # Remove references and bibliography sections tex_content = regex.sub(r'\\begin\{thebibliography\}.*?\\end\{thebibliography\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\bibliography\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\bibitem\{.*?\}', '', tex_content, flags=regex.DOTALL) # Remove acknowledgments (handles starred sections and custom commands) tex_content = regex.sub(r'\\section\*?\{Acknowledgments\}.*?(?=\\section|\\subsection|$)', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\acknowledgments\{.*?\}', '', tex_content, flags=regex.DOTALL) # Remove tables, figures, and image commands tex_content = regex.sub(r'\\begin\{tabular\}.*?\\end\{tabular\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\begin\{figure\}.*?\\end\{figure\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\table\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\figure\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\includegraphics\{.*?\}', '', tex_content, flags=regex.DOTALL) # Remove footnotes tex_content = regex.sub(r'\\footnote\{.*?\}', '', tex_content, flags=regex.DOTALL) tex_content = regex.sub(r'\\footnotetext\{.*?\}', '', tex_content, flags=regex.DOTALL) # -------------------------- # Phase 2: Convert TeX to plain text # -------------------------- parser = LatexNodes2Text() plain_text = parser.latex_to_text(tex_content) # -------------------------- # Phase 3: Clean up whitespace and artifacts # -------------------------- # Fix line breaks and excess spaces plain_text = regex.sub(r'\n\s*\n', '\n\n', plain_text) plain_text = regex.sub(r' +', ' ', plain_text).strip() return plain_text # Example usage clean_body_text = extract_clean_body("sample_astro_paper.tex") print(clean_body_text)
Step 3: Handle Astrophysics-Specific Edge Cases
Astrophysics papers often use custom macros (like \Msun for solar mass). Add this step right after reading the TeX content to replace them with readable text:
# Replace common astro macros with plain text macro_map = { r'\\Msun': 'solar mass', r'\\Lsun': 'solar luminosity', r'\\kpc': 'kiloparsec', r'\\Mpc': 'megaparsec', r'\\keV': 'kiloelectronvolt', # Add more macros from your corpus here } for macro, replacement in macro_map.items(): tex_content = regex.sub(macro, replacement, tex_content)
Step 4: Iterate and Refine
Since arXiv papers vary widely, test this on your sample files and adjust regex patterns as needed. For example:
- If some papers use
\section{Appendix}that you want to exclude, add a regex to remove that section. - If inline equations are causing issues, you can either keep them as LaTeX syntax (useful for some NLP tasks) or add a regex to remove them entirely.
内容的提问来源于stack exchange,提问作者brienna

