You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Watson Knowledge Studio和Watson Discovery提取指定单位的值(基于指定报告的Matplot方法)

Alright, let’s walk through how to extract numerical values tied to specific units using Watson Knowledge Studio (WKS) and Watson Discovery, aligning with the Matplotlib-based approach from that empirical literature assessment report. Here’s a practical, step-by-step breakdown:

1. Align Your Approach with the Original Report’s Logic

First, let’s map the report’s Matplotlib-focused workflow to Watson’s toolchain: the original study used pattern matching and validation to pull units like Cohen’s d, η², or p-values from literature. We’ll replicate that pattern recognition (plus add machine learning for edge cases) using WKS to define what counts as a "value + target unit" entity, then use Discovery to scale extraction across your document set.

2. Define Custom Entities in Watson Knowledge Studio

Start here to teach the system to recognize your target numerical values and their associated units:

  • Create a new WKS project tailored to academic literature (PDFs, text files from psychology/cognitive neuroscience papers work best).
  • Define custom entity types (e.g., EffectSizeWithUnit, StatisticalValueWithUnit). For specificity, you can even use sub-types like CohensDValue or PValue, but a flexible parent entity works too.
    • Manual annotation: Label 50-100 sample documents first—mark instances like d = 0.85, η² = 0.12, or p < 0.05, linking the numerical value to its unit. This trains the ML model to pick up non-standard formatting (like .7 instead of 0.7 or Cohen’s d vs d).
    • Add rule-based extractors: Write regex patterns to cover consistent formats, just like the report’s pattern matching. Example patterns:
      • For Cohen’s d: (-?\d+\.?\d*) (Cohen's d|d)
      • For p-values: p (?:=|<|>) (\d+\.?\d*|1e-\d+)
      • Wrap these in backticks in WKS’s rule editor to ensure proper parsing.
  • Train and validate the combined ML+rule model: Test it against unannotated documents to adjust rules or add more labels if you miss edge cases.
3. Integrate the WKS Model with Watson Discovery

Once your model is solid, move it to Discovery for bulk processing:

  • Export the trained entity recognition model from WKS (look for the "Export model" option in the project settings).
  • Import this model into your Watson Discovery instance, then enable it in your collection’s document processing pipeline. Make sure Discovery is set to parse PDFs/text correctly (enable OCR if you have scanned documents).
4. Extract and Filter Target Values at Scale

Now run the extraction to match the report’s bulk processing step:

  • Upload your full set of literature documents to the Discovery collection.
  • Use Discovery’s query interface (or API) to pull the entities you care about:
    • Basic query to get all target entities: enriched_text.entities.type:EffectSizeWithUnit
    • Filter by specific units: enriched_text.entities.type:EffectSizeWithUnit AND enriched_text.entities.text:"Cohen's d"
  • Export the results as CSV or JSON—this gives you a clean list of values and their units, ready for analysis with Matplotlib (just like the original report did for statistical assessment and visualization).
Pro Tips for Better Accuracy
  • Cover edge cases in regex: Account for scientific notation (e.g., 1.2e-5 for p-values) and inconsistent spacing (e.g., p<.001 vs p < .001).
  • Add context filters: Use Discovery’s passages feature to pull the paragraph around each entity—this lets you verify the value is actually an effect size/statistic, not a random number in the paper.
  • Iterate on annotations: If you notice the model missing values, go back to WKS and add those cases to your training set—this improves ML accuracy over time.

内容的提问来源于stack exchange,提问作者Blue_IoT_Mix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:16:11