You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Botpress机器人对接Word文档文本读取API开发技术问询

Hey there! Let's work through your two Botpress + Word document integration questions to get everything running smoothly:

1. API Structure: Where to put your functions & how to connect them

You absolutely can keep everything in a single app.js file—this works great for simpler projects where you don't have tons of complex logic to manage. That said, here's how to structure it effectively:

  • Single file approach: Wrap your textract document-reading logic into a reusable function directly in app.js, then create an Express (or whatever framework you're using) API route that calls this function. Botpress can then hit this route via HTTP requests from its action files.
  • Scalable approach (for future growth): If you think you'll add more document types, authentication, or other features later, split your code into modules. For example:
    • Create a services/documentHandler.js file to hold all your textract logic (reading files, splitting paragraphs, etc.)
    • Import this module into your main app.js and hook it up to your API routes. This keeps your code organized and easier to maintain.

2. Fetching & Returning Specific Word Document Paragraphs

The key here is to first convert the raw text from your Word doc into a structured array of paragraphs, then pull the specific one you need. Here's how to do it:

Step 1: Process the Word Document into Paragraphs

Use textract to pull the full text, then split it into paragraphs (adjust the splitting rule based on how your doc is formatted—empty lines are a common separator):

const getParagraphsFromWordDoc = (filePath) => {
  return new Promise((resolve, reject) => {
    textract.fromFileWithPath(filePath, (err, fullText) => {
      if (err) reject(err);
      // Split by empty lines to get individual paragraphs, filter out empty entries
      const paragraphs = fullText.split(/\n\s*\n/).filter(para => para.trim() !== '');
      resolve(paragraphs);
    });
  });
};

Step 2: Build an API Endpoint to Fetch Specific Paragraphs

Add this route to your app.js to let Botpress request a specific paragraph by index:

const express = require('express');
const textract = require('textract');
const app = express();

// ... include the getParagraphsFromWordDoc function here ...

app.get('/api/get-paragraph', async (req, res) => {
  try {
    const { filePath, paraIndex } = req.query;
    if (!filePath || paraIndex === undefined) {
      return res.status(400).json({ error: 'Missing filePath or paraIndex parameter' });
    }
    const paragraphs = await getParagraphsFromWordDoc(filePath);
    const index = parseInt(paraIndex);
    if (index < 0 || index >= paragraphs.length) {
      return res.status(404).json({ error: 'Paragraph index out of range' });
    }
    res.json({ content: paragraphs[index] });
  } catch (err) {
    res.status(500).json({ error: err.message });
  }
});

app.listen(3000, () => console.log('Document API running on port 3000'));

Step 3: Call the API from Botpress Action

Create a Botpress action file (e.g., fetchWordParagraph.js) to call your API and store the result in the session:

module.exports = async (bp, event, args, { session }) => {
  try {
    const apiResponse = await bp.http.get('http://localhost:3000/api/get-paragraph', {
      params: {
        filePath: './your-document-path.docx',
        paraIndex: '1' // Replace with the index of the paragraph you want
      }
    });
    // Store the paragraph content in session to use in your bot's reply
    session.targetParagraph = apiResponse.data.content;
  } catch (err) {
    session.targetParagraph = 'Sorry, I couldn\'t retrieve that content right now.';
    console.error('Error fetching paragraph:', err);
  }
};

Step 4: Use the Content in Botpress Replies

In your bot's response template, just reference the stored value:

{{session.targetParagraph}}

Bonus: Skip the Separate API (Lightweight Option)

If you don't want to run a separate API, you can call textract directly inside a Botpress action—no extra app.js needed:

const textract = require('textract');

module.exports = (bp, event, args, { session }) => {
  textract.fromFileWithPath('./your-document-path.docx', (err, fullText) => {
    if (err) {
      session.response = 'Oops, I had trouble reading the document.';
      return;
    }
    const paragraphs = fullText.split(/\n\s*\n/).filter(para => para.trim() !== '');
    // Return the 3rd paragraph (index 2)
    session.response = paragraphs[2] || 'That paragraph doesn\'t exist.';
  });
};

Then use {{session.response}} in your bot's reply template like you already do.


内容的提问来源于stack exchange,提问作者AVu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:59:53