如何将使用window参数生成的quanteda tokens对象反分词合并为字符串
quanteda窗口tokens反分词实现方法
直接调用quanteda内置的detokenize()函数即可实现需求,以下是基于你提供的示例代码的扩展实现:
# 承接你已生成的ttt tokens对象 # 方法1:直接反分词为向量,每个元素对应一个文档的窗口上下文 context_vector <- detokenize(ttt) # 查看输出结果 context_vector # 方法2:转为数据框格式,更方便批量查看、标注上下文 context_df <- data.frame( 文档ID = names(ttt), 目标词 = "future", 上下文内容 = detokenize(ttt), check.names = FALSE ) # 调出可视化窗口查看数据框 View(context_df)
如果不想调用quanteda的内置函数,也可以用base R的遍历拼接方法实现相同效果:
# 手动拼接实现反分词 context_vector_manual <- sapply(ttt, paste, collapse = " ")
运行
context_vector的输出示例参考:text1 text2 "previously about the future of our water" "communities in the future Thank you"
内容的提问来源于stack exchange,提问作者dcoy
相关产品推荐
相关产品推荐

