You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python从文档构建biword index并生成相邻双词组合列表?

双词索引(Biword Index)构建Python实现

功能说明

实现输入文本内容自动拆分生成相邻双词索引项,输出结果为字符串列表。

实现代码

def build_biword_index(doc_content):
    # 按空格分割为单个单词列表
    word_list = doc_content.split()
    # 单词数量不足2时返回空列表
    if len(word_list) < 2:
        return []
    # 生成相邻双词组合
    return [f"{word_list[i]} {word_list[i+1]}" for i in range(len(word_list) - 1)]

# 测试样例
if __name__ == "__main__":
    sample_text = "There have been biographies of Dewey that briefly describe his system, but this is the first attempt to provide a detailed history of the work that more than any other has spurred the growth of librarianship in this country and abroad."
    biword_result = build_biword_index(sample_text)
    print(biword_result)

输出验证

运行上述代码后,输出的前4项和需求示例完全匹配:
['There have', 'have been', 'been biographies', 'biographies of', ...]

拓展说明

  • 如果需要去除单词附带的标点符号,可在分割前添加清洗逻辑:
    import re
    doc_content = re.sub(r'[^\w\s]', '', doc_content)
    
  • 中文场景使用时,可先通过jieba等分词工具完成分词,将分词后的列表替换代码中的word_list即可生成中文双词索引。

内容的提问来源于stack exchange,提问作者user3237451

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 18:42:01