论文部分内容阅读
This paper investigates relations between word semantic den-sity and word frequency.A distributed representations based word av-erage similarity is defined as the measure of word semantic density.We find that the average similarities of low frequency words are always big-ger than that of high frequency words,when the frequency approaches to 400 around,the average similarity tends to stable.The finding keeps cor-rect with changes of the size of training corpus,dimension of distributed representations and number of negative samples in skip-gram model.It also keeps on 17 different languages.Basing on the finding,we propose a pseudo context skip-gram model,which makes use of context words of semantic nearest neighbors of target words.Experiment results show our model achieves significant performance improvements in both word similarity and analogy tasks.