You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Endpoint Utterance Log生成LUIS应用训练所需的JSON文件?

Can I generate the JSON dataset needed to train a LUIS app using the Endpoint Utterance Log?

Absolutely! You can absolutely leverage the Endpoint Utterance Log to create the labeled JSON dataset required for training your LUIS app. Here’s a practical, step-by-step breakdown of how to pull this off:

  • Export your endpoint utterance log first
    Head over to your LUIS portal, open your target app, and navigate to the Manage tab. From there, select Endpoint utterances—this is where all the real user queries sent to your app’s endpoint are stored. Click the Export button, and you can choose to export either all utterances or just the ones that were misclassified/unlabeled. The latter is especially handy if you’re looking to fix specific gaps in your model’s performance.

  • Convert the exported data to LUIS’s training JSON format
    The exported log will come in either CSV or JSON format (your call). If you go with CSV, you’ll need to map it to LUIS’s standard training structure. Here’s what that core structure looks like:

    {
      "utterances": [
        {
          "text": "Reserve a table for two at downtown Italian spot",
          "intent": "BookRestaurant",
          "entities": [
            {
              "entity": "PartySize",
              "startPos": 18,
              "endPos": 19
            },
            {
              "entity": "Cuisine",
              "startPos": 33,
              "endPos": 39
            }
          ]
        }
      ]
    }
    

    For each entry in the log:

    • Plug the user’s actual query into the text field
    • Assign the correct intent—you can use the log’s predicted intent as a starting point, but always double-check it (real users might phrase things in unexpected ways!)
    • Add any relevant entities with their exact start/end positions in the text. If the log includes entity predictions, use those as a guide, but verify accuracy to avoid training your model on bad data.
  • Clean and validate the dataset
    This step can’t be overstated! The endpoint log might have duplicate utterances, misclassified intents, or missing entity labels. Take the time to review each entry, fix errors, and remove any irrelevant queries. High-quality training data is the key to a better-performing LUIS app—cutting corners here will hurt your model’s accuracy later.

  • Import the JSON back into LUIS
    Once your formatted JSON is ready, jump back to the Build tab of your LUIS app, go to Utterances, and click Import. Upload your JSON file, and the utterances will be added to your training set. From there, you can retrain your model and test it to see how the new data improves its performance.

A quick pro tip: If you’re working with a huge log, use a simple script (Python, PowerShell, whatever you’re comfortable with) to automate the conversion process. It’ll save you tons of manual work and reduce the chance of human error. Also, remember that this log is made up of real user interactions—this kind of diverse, real-world data is way more valuable than synthetic test utterances for making your app robust.

内容的提问来源于stack exchange,提问作者Daniel Edwards

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:50:37