You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让python-bibtexparser输出纯文本而非LaTeX格式结果?

问题描述

我通过自定义的DOI辅助类获取BibTeX元数据,再用bibtexparser解析后,得到的是LaTeX格式的文本。比如针对DOI 10.1145/800001.811672,当前解析结果为:

The structure of the {\textquotedblleft}the{\textquotedblright}-multiprogramming system

但我需要的是纯文本格式:

The structure of the "the"-multiprogramming system

想确认当前版本的bibtexparser是否支持将LaTeX格式转换为纯文本的功能,还是需要提交功能请求?

核心实现代码:

doi=DOI(self.doi)
meta_bibtex=doi.fetchBibtexMeta()
bd=bibtexparser.loads(meta_bibtex)
btex=bd.entries[0]

自定义DOI辅助类代码(doi.py):

'''
Created on 2023-02-12

@author: wf
'''
import urllib.request
import json
from dataclasses import dataclass

@dataclass
class DOI:
    """
    获取DOI数据
    """
    doi:str
    
    def fetchMeta(self,headers:dict)->dict:
        """
        获取该DOI的元数据
        
        参数:
            headers(dict): 请求头
            
        返回:
            dict: 对应请求头的元数据
        """
        url=f"https://doi.org/{self.doi}"
        req=urllib.request.Request(url,headers=headers)
        response=urllib.request.urlopen(req)
        encoding = response.headers.get_content_charset('utf-8')
        content = response.read()
        text = content.decode(encoding)
        return text
        
    def fetchBibtexMeta(self)->dict:
        """
        通过获取BibTeX格式数据来获取DOI元数据
         
        返回:
            dict: 元数据
            
        """
        headers= {
            'Accept': 'application/x-bibtex; charset=utf-8'
        }
        text=self.fetchMeta(headers)
        return text
    
    def fetchCiteprocMeta(self)->dict:
        """
        通过获取Citeproc JSON格式数据来获取DOI元数据
            
        返回:
            dict: 元数据
        """
        headers= {
            'Accept': 'application/vnd.citationstyles.csl+json; charset=utf-8'
        }
        text=self.fetchMeta(headers)
        json_data=json.loads(text)
        return json_data   
解决方案

当前版本的bibtexparser本身没有内置LaTeX转纯文本的功能,可通过以下两种方式解决:

  • 手动替换常用LaTeX命令
    针对常见的LaTeX特殊字符编写替换逻辑,示例代码:

    def latex_to_plain(latex_content):
        # 可根据需求扩展更多转换规则
        replace_map = {
            r'{\textquotedblleft}': '"',
            r'{\textquotedblright}': '"',
            r'{\textit{': '',
            r'}}': '',
            r'{\textbf{': ''
        }
        for latex_expr, plain_text in replace_map.items():
            latex_content = latex_content.replace(latex_expr, plain_text)
        return latex_content
    
    # 使用示例
    plain_title = latex_to_plain(btex['title'])
    
  • 借助专业LaTeX转文本库
    使用pylatexenc这类专门处理LaTeX转换的库,能覆盖更复杂的格式场景,示例:

    from pylatexenc.latex2text import LatexNodes2Text
    
    converter = LatexNodes2Text()
    plain_title = converter.latex_to_text(btex['title'])
    

如果希望bibtexparser原生支持该功能,可向项目提交功能请求。


内容的提问来源于Stack Exchange,提问作者Wolfgang Fahl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 13:40:33