You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:用Python3+tex2py提取arXiv天体物理TeX文件纯净正文

Solution for Extracting Clean Body Text from arXiv Astrophysics TeX Files

I feel your pain—arXiv TeX files are notoriously inconsistent, with authors using custom macros, varying section tags, and text split across commands that break naive parsers like tex2py. Here's a robust Python-based approach tailored to astrophysics papers that will get you the pure body text you need:

Step 1: Install Required Packages

First, grab the tools we'll use for parsing and pattern matching:

pip install pylatexenc regex

Step 2: Implement the Cleanup Pipeline

This function combines regex-based removal of non-body sections with a robust TeX-to-text parser to handle messy, split text. It targets the common non-body elements you want to exclude:

import regex
from pylatexenc.latex2text import LatexNodes2Text

def extract_clean_body(tex_file_path):
    # Read the raw TeX content
    with open(tex_file_path, 'r', encoding='utf-8', errors='ignore') as f:
        tex_content = f.read()
    
    # --------------------------
    # Phase 1: Remove non-body sections
    # --------------------------
    # Remove title, author blocks, and maketitle command
    tex_content = regex.sub(r'\\title\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\author\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\maketitle', '', tex_content)
    
    # Remove abstract (handles both environment and command forms)
    tex_content = regex.sub(r'\\begin\{abstract\}.*?\\end\{abstract\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\abstract\{.*?\}', '', tex_content, flags=regex.DOTALL)
    
    # Remove references and bibliography sections
    tex_content = regex.sub(r'\\begin\{thebibliography\}.*?\\end\{thebibliography\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\bibliography\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\bibitem\{.*?\}', '', tex_content, flags=regex.DOTALL)
    
    # Remove acknowledgments (handles starred sections and custom commands)
    tex_content = regex.sub(r'\\section\*?\{Acknowledgments\}.*?(?=\\section|\\subsection|$)', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\acknowledgments\{.*?\}', '', tex_content, flags=regex.DOTALL)
    
    # Remove tables, figures, and image commands
    tex_content = regex.sub(r'\\begin\{tabular\}.*?\\end\{tabular\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\begin\{figure\}.*?\\end\{figure\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\table\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\figure\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\includegraphics\{.*?\}', '', tex_content, flags=regex.DOTALL)
    
    # Remove footnotes
    tex_content = regex.sub(r'\\footnote\{.*?\}', '', tex_content, flags=regex.DOTALL)
    tex_content = regex.sub(r'\\footnotetext\{.*?\}', '', tex_content, flags=regex.DOTALL)
    
    # --------------------------
    # Phase 2: Convert TeX to plain text
    # --------------------------
    parser = LatexNodes2Text()
    plain_text = parser.latex_to_text(tex_content)
    
    # --------------------------
    # Phase 3: Clean up whitespace and artifacts
    # --------------------------
    # Fix line breaks and excess spaces
    plain_text = regex.sub(r'\n\s*\n', '\n\n', plain_text)
    plain_text = regex.sub(r' +', ' ', plain_text).strip()
    
    return plain_text

# Example usage
clean_body_text = extract_clean_body("sample_astro_paper.tex")
print(clean_body_text)

Step 3: Handle Astrophysics-Specific Edge Cases

Astrophysics papers often use custom macros (like \Msun for solar mass). Add this step right after reading the TeX content to replace them with readable text:

# Replace common astro macros with plain text
macro_map = {
    r'\\Msun': 'solar mass',
    r'\\Lsun': 'solar luminosity',
    r'\\kpc': 'kiloparsec',
    r'\\Mpc': 'megaparsec',
    r'\\keV': 'kiloelectronvolt',
    # Add more macros from your corpus here
}
for macro, replacement in macro_map.items():
    tex_content = regex.sub(macro, replacement, tex_content)

Step 4: Iterate and Refine

Since arXiv papers vary widely, test this on your sample files and adjust regex patterns as needed. For example:

  • If some papers use \section{Appendix} that you want to exclude, add a regex to remove that section.
  • If inline equations are causing issues, you can either keep them as LaTeX syntax (useful for some NLP tasks) or add a regex to remove them entirely.

内容的提问来源于stack exchange,提问作者brienna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:48:50