You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:从Python转用Perl实现术语相关性计算

How to Compute Term Relatedness in Python (or Adapt Your Perl HSO Script)

Hey there, let's break down how to solve this—whether you want a pure Python implementation or to reuse that Perl HSO script you found. First, a quick note: in NLP, "relatedness" and "similarity" often overlap, but HSO (the algorithm from your Perl script) is a hybrid method that combines WordNet's hierarchical structure and information content to calculate semantic relatedness.

Option 1: Pure Python Implementation of HSO

You can replicate the HSO logic using Python's NLTK library and WordNet. Here's a step-by-step guide:

1. Setup Dependencies

First, install NLTK and download the required WordNet data:

pip install nltk

Then run this in Python to grab the necessary datasets:

import nltk
nltk.download('wordnet')
nltk.download('omw-1.4')  # Multilingual WordNet support
nltk.download('wordnet_ic')  # Precomputed information content data

2. Implement HSO Relatedness

HSO's core idea is to find shared parent terms (hypernyms) of your target words, then calculate a score based on the path length to those hypernyms and their information content (IC). Here's a working implementation that matches the Perl script's logic:

from nltk.corpus import wordnet as wn
from nltk.corpus import wordnet_ic

# Load precomputed information content from the Brown corpus
brown_ic = wordnet_ic.ic('ic-brown.dat')

def hso_relatedness(synset_a, synset_b, ic_data):
    # Get all shared hypernyms between the two word senses
    common_hyps = synset_a.common_hypernyms(synset_b)
    if not common_hyps:
        return 0.0  # No shared semantic meaning, score 0
    
    max_score = 0.0
    for hyp in common_hyps:
        # Calculate total path length from both synsets to the shared hypernym
        path_length = synset_a.shortest_path_distance(hyp) + synset_b.shortest_path_distance(hyp)
        # Fetch the hypernym's information content
        hyp_ic = hyp.ic(ic_data)
        # HSO's core scoring formula (aligns with the Perl implementation)
        current_score = (2 * hyp_ic) / (path_length + 2 * hyp_ic)
        if current_score > max_score:
            max_score = current_score
    return max_score

# Example usage (matches your Perl script's terms: car#n#1 and bus#n#2)
car_syn = wn.synset('car.n.01')  # "car" as an automobile
bus_syn = wn.synset('bus.n.02')  # "bus" as a public transit vehicle
score = hso_relatedness(car_syn, bus_syn, brown_ic)
print(f"car (sense 1) <-> bus (sense 2) Relatedness Score: {score:.4f}")

This will give you a score comparable to your Perl script's output. You can tweak the formula slightly if you need exact parity with the Perl version (check the WordNet::Similarity HSO docs for edge case handling).

Option 2: Adapt Your Perl Script to Run in Python

If you'd rather reuse the mature Perl implementation instead of redoing it, you can call the Perl script directly from Python using the subprocess module.

1. Prepare the Perl Script

Save this as hso_calculator.pl:

use WordNet::Similarity::hso;
use WordNet::QueryData;

# Initialize WordNet and HSO object
my $wn = WordNet::QueryData->new();
my $hso = WordNet::Similarity::hso->new($wn);

# Get target terms from command line arguments
my ($term1, $term2) = @ARGV;
my $relatedness = $hso->getRelatedness($term1, $term2);

# Handle errors
my ($err_code, $err_msg) = $hso->getError();
die "Error: $err_msg\n" if $err_code;

# Output the calculated score
print $relatedness;

Make sure you have Perl installed, plus the required modules:

cpanm WordNet::Similarity WordNet::QueryData

2. Call the Script from Python

Use subprocess to pass terms and retrieve the score:

import subprocess

def perl_hso_relatedness(term1, term2):
    try:
        result = subprocess.run(
            ['perl', 'hso_calculator.pl', term1, term2],
            capture_output=True,
            text=True,
            check=True
        )
        return float(result.stdout.strip())
    except subprocess.CalledProcessError as e:
        raise Exception(f"Perl script failed: {e.stderr}")

# Test with your example terms
score = perl_hso_relatedness("car#n#1", "bus#n#2")
print(f"Relatedness Score (from Perl): {score:.4f}")

Which Option Should You Choose?

  • Go with the pure Python implementation if you want a self-contained, Python-native solution (easier to integrate with existing Python workflows).
  • Use the Perl script wrapper if you need exact parity with the original HSO implementation from WordNet::Similarity, or if you don't want to reimplement edge cases.

内容的提问来源于stack exchange,提问作者belzebubele

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:11:28